Vol.01 · No.10 Daily Dispatch September 24, 2026

Latest AI News

AI · PapersDaily CurationOpen Access
AI NewsBusiness
8 min read

Google launches Gemini 3.8 TTS for custom, controllable voices

Flash TTS and Flash‑Lite TTS let teams design voices from scratch, direct delivery line by line, and add watermarked voice replication with consent checks. Rollout starts in Google AI Studio and the Gemini API, with integrations into Google Vids and partner apps.

Reading Mode

One-Line Summary

Google ships expressive text-to-speech across its stack while YouTube adds creator AI in Studio, and open-weights pushes from Xiaomi and BFL expand cheaper, deployable options.

Big Tech

Google rolls out Gemini 3.8 TTS models for custom voices

Google releases two text-to-speech models, Gemini 3.8 Flash TTS and Flash‑Lite TTS, that let you create realistic custom voices and control how each line is spoken across Google AI Studio, the Gemini API, Gemini Notebook, and Google Vids. The models support more than 100 languages and offer a library of 2,000+ production-ready voices alongside from-scratch generative voice design. 1

Creators and teams can direct performances line by line, handle long-form narration with minimal speaker drift, stage native two-speaker scenes, and add realistic backchannel cues. Voice replication works from a 30‑second sample with built-in consent verification, and all audio is watermarked with SynthID and carries C2PA credentials; voice replication in AI Studio is unavailable in Illinois, Texas, the EEA, UK, Switzerland, and India. 1

Google cites leading results: Gemini 3.8 Flash TTS tops Hume AI’s Voice Design Benchmark at 71.4 and leads accent modeling at 60.8, while Flash and Flash‑Lite rank #1 and #2 on Hume AI’s Overall Quality Index. Blind human preference tests on Voice Arena place the models strongly across languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi. 1

Availability starts Sep 23, 2026 in Google AI Studio and the Gemini API, with API access coming to Gemini Enterprise and broader use in Google Vids. Developer platforms like Agora, LiveKit, Pipecat, and Vercel can simplify deployment, and partners including Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang are integrating for dubbing, localization, and voice agents. 1

YouTube brings AI draft feedback, thumbnail tests, and research to Studio

YouTube adds AI to its Studio app to analyze unpublished videos for pacing, structure, and storytelling feedback, expands Ask Studio (its AI Q&A) to iOS and Android, and introduces a research feed showing what content is working so creators can adapt it. The app can also generate channel-matched thumbnails and titles, plus dynamic thumbnails to test different options with audience segments. 2

YouTube says creators have run over 40 million A/B tests on titles and thumbnails, and Studio will soon let them test up to three opening cuts and automatically monitor thumbnail performance to swap in better performers. Analytics will explain not just the numbers but why videos perform and offer advice to increase engagement. 2

Additional updates include an expansion of YouTube Shopping’s affiliate program to 35 countries (with more than 1.3 million creators participating), opt-in AI comment moderation that adapts to creator preferences, enhanced likeness protection combining voice and facial detection, Custom Feeds for U.S. users, and new shopping queries and voice commands in Ask YouTube on TV. These were detailed in reporting carried by Idaho Statesman. 3

Industry & Biz

BFL launches FLUX 3 Action, an open-weights robot model aiming for faster control

Black Forest Labs unveils FLUX 3 Action, a 7‑billion‑parameter open-weights “World Action Model” that converts camera observations, a robot’s state, and natural-language instructions into physical actions. The company says it scores 42.92% on NVIDIA’s RoboLab‑120, outpacing listed models on the public leaderboard at announcement time, and runs 1.43× faster than NVIDIA’s 16B Cosmos3‑Nano‑Policy while using 44% as many parameters. 4

BFL tells VentureBeat it will release the weights, code, fine‑tuning recipe, benchmarks, and reproducible examples, including an implementation using the relatively inexpensive SO‑101 robot with Hugging Face’s LeRobot. Demos include real‑world drone piloting and beating Doom without in‑game deaths, though the company frames these as early experiments; VentureBeat also notes NVIDIA’s leaderboard had not updated before launch. 4

For developers, the draw is speed and deployability: BFL claims distilled variants set a new Pareto frontier on RoboLab success rate vs real‑time factor, but reminds that categories differ (WAMs vs VLAs vs action‑reasoning models), making direct comparisons hard. The practical takeaway is to validate on your own robot and tasks where consistent, low‑latency action really matters. 4

Xiaomi’s MiMo‑V2.6‑Pro tops open‑weights ranking and undercuts on price

VentureBeat reports that Xiaomi’s MiMo‑V2.6‑Pro debuts with an Artificial Analysis Intelligence Index score of 46, ranking it the top open‑weights model as of Sep 21–22, 2026, and tying with Grok 4.7 while surpassing some proprietary peers on that composite. The release also includes the cheaper MiMo‑V2.6‑Flash and a Pro‑UltraSpeed variant. 5

Times of AI notes Xiaomi publishes MIT‑licensed weights on Hugging Face with native multimodal input (text, image, video, audio) and a one‑million‑token context window. Through Xiaomi’s API, Pro is priced at $0.435 per million input tokens and $0.87 per million output tokens, while Flash is $0.14 and $0.28; cache‑hit input is listed at $0.0036 per million tokens. 6

Xiaomi attributes the jump to a large reinforcement learning program: each model reportedly completed 30 RL steps over roughly 750,000 trajectories in under six days, with estimated costs of $2.62 million for Pro and $850,000 for Flash, and the company says it is releasing training environments and RL code. For teams, that combination—open weights and aggressive pricing—makes long‑context agent work more affordable to test. 6

Community Pulse

Hacker News (216↑) — Mixed: commenters say Google competes on price while some dislike exaggerated prosody and prefer restrained “computer” voices. 7

"IMHO its obviously: they cant compete with SotA models in benchmaxxing. They can and do compete on price though." — Hacker News 7

"Yes, I find nearly every "SOTA" voice model I try intolerable to listen to because of the fake exaggerated expression/emotion. It's actively distracting because it pulls focus to emphasize randomly. ChatGPT Voice models are so insufferable to put up with for a conversation longer than 45 seconds. All I want is a clear, technically flawless, even/restrained "computer voice" for pretty much every use case (except audiobooks). But that doesn't make for splashy demos/score well for RLHF raters." — Hacker News 7

What This Means for You

Google’s new TTS is not just another preset voice; it is a voice design studio with safety guardrails. For content teams, customer support, and internal training, it means you can ship brand‑consistent narration and agents faster, with consent checks and invisible watermarks that help with disclosure policies. If your org uses Gemini or Google Workspace, this lowers integration friction. 1

For creators, YouTube Studio’s AI feedback plus thumbnail and opening‑hook tests can raise click‑through and watch time without new headcount. The research feed can shortcut ideation by showing what’s resonating now, and the analytics updates help explain performance, not just measure it. Treat the tools like a product manager: draft, test, and iterate. 2

On budgets, Xiaomi’s MIT‑licensed, open‑weights models point to cheaper million‑token runs for docs, codebases, and long agent workflows. If you route tasks between a premium model and a low‑cost open one, you can cut spend on long contexts while keeping quality where it matters. Validate security posture before production use of a new API. 5

If your team touches robotics or interactive 3D, BFL’s open‑weights policy model signals faster iteration for embodied or simulated tasks. The benchmark claims are promising, but deployability hinges on your hardware, tasks, and latency targets—run a small, well‑instrumented pilot before committing. 4

Action Items

  1. Design a brand voice in Google AI Studio: Use the new audio playground to create one custom voice and a 60‑second script; test line‑by‑line direction, two‑speaker scenes, and enable consent/watermark checks before sharing.
  2. Run a YouTube Studio optimization sprint: Take one unpublished video, get AI draft feedback, and A/B test three thumbnails and two titles; log click‑through and retention changes for a week.
  3. Trial MiMo‑V2.6 for a long‑context task: Use Xiaomi’s API (Pro or cheaper Flash) to summarize a 200k‑token spec or codebase and compare cost/speed versus your current model.
  4. Write a one‑page voice‑replication policy: Define consent recording, watermark retention, and disclosure language so your team can use synthetic voices responsibly.
  5. Scope FLUX 3 Action fit: Read the model materials and RoboLab task list, then pick one high‑latency‑sensitive task to prototype when weights/examples land.

Sources 8

Helpful?

Comments (0)