آموزش

Gemini Can Now Read Your Google Docs Out Loud

Text-to-speech tech isn’t new. Your computer has likely had the ability for years at this point, though the end result might not necessarily be ideal.

Google wants to make the experience better for Docs users. Starting today, Google Docs has a new AI-powered Audio option that creates “audio versions” of your documents. What that really means is this: Google is using Gemini to generate more a natural and realistic text-to-speech experience. It works, but it’s still far from perfect.

Google announced the news in a blog post on Wednesday, just hours before the company’s Made by Google 2025 event . To use the feature, open a doc, then select the “Tools” tab from the menu bar. If you have access, you’ll see a new “Audio” option. Select this, and a playback bar will appear in the bottom left corner of the window—though you can move it wherever you’d like. After the AI has a chance to process the doc, it will start speaking automatically.

Google’s AI voice tech is a bit hit or miss, here. The voice itself is realistic, and it occasionally hits some natural inflections and rhythms, but it also has plenty of moments where the facade falls. If you’re familiar with “AI voice,” you’ll hear that here.

That said, Google does offer some tools to customize the experience. You can adjust the playback speed, between 0.5x and 2x. There are also seven voices to choose from here. The default is Narrator, which Google describes as “smooth, medium pitch,” but you can also pick from six other choices, including:

  • Educator: Friendly, higher pitch

  • Teacher: Clear, low pitch

  • Persuader: Engaging, low pitch

  • Explainer: Lively, low pitch

  • Coach: Lively, higher pitch

  • Motivator: Energetic, medium pitch

If you are the author of the page, you can also choose to insert an audio button into the doc. That way, all other contributors and readers can listen if they want to.

This feature is available to a wide variety of Google Workspace users, including Business Standard and Plus, Enterprise Standard and Plus, Gemini Education, Gemini Education Premium, Gemini Business, and Gemini Enterprise. In addition, if you subscribe to Google AI Pro or Google AI Ultra, you also have access to this feature.

منبع آموزش

ZaKi

Who is mahdizk? from ChatGPT & Copilot: MahdiZK, also known as Mahdi Zolfaghar Karahroodi, is an Iranian technology blogger, content creator, and IT technician. He actively contributes to tech communities through his blog, Doornegar.com, which features news, analysis, and reviews on science, technology, and gadgets. Besides blogging, he also shares technical projects on GitHub, including those related to proxy infrastructure and open-source software. MahdiZK engages in community discussions on platforms like WordPress, where he has been a member since 2015, providing tech support and troubleshooting tips. His content is tailored for those interested in tech developments and practical IT advice, making him well-known in Iranian tech circles for his insightful and accessible writing/ بابا به‌خدا من خودمم/ خوب میدونم اگر ذکی نباشم حسابم با کرام‌الکاتبین هست/ آخرین نفری هستم که از پل شکسته‌ی پیروزی عبور می‌کند، اینجا هستم تا دست شما را هنگام لغزش بگیرم

نوشته های مشابه

0 0 رای ها
امتیازدهی به مقاله
اشتراک در
اطلاع از
guest

0 نظرات
قدیمی‌ترین
تازه‌ترین بیشترین رأی
بازخورد (Feedback) های اینلاین
مشاهده همه دیدگاه ها
دکمه بازگشت به بالا
0
افکار شما را دوست داریم، لطفا نظر دهید.x