Vol.01 · No.10 Daily Dispatch July 23, 2026

Latest AI News

AI · PapersDaily CurationOpen Access
AI NewsBusiness
6 min read

OpenAI launches Presence to run enterprise agents in production

Presence bundles policies, guardrails, and evaluations so voice and chat agents can take approved actions at work. Google trims Gemini costs, and China’s Moonshot pushes open‑weight performance, signaling a shift to governed, cost‑aware AI.

Reading Mode

One-Line Summary

OpenAI moves beyond model access with Presence for governed enterprise agents, while Google cuts Gemini costs and China’s Moonshot advances open‑weight performance—pointing to hands‑on, cost‑aware AI deployments.

Big Tech

Google ships Gemini 3.6 Flash, plus 3.5 Flash-Lite and Flash Cyber

Google releases three Gemini updates aimed at speed and cost: 3.6 Flash for efficiency, 3.5 Flash-Lite for ultra-fast responses, and 3.5 Flash Cyber tuned for cybersecurity, while the delayed 3.5 Pro remains in testing and pretraining for Gemini 4 is underway. 1

3.6 Flash reduces output token usage by about 17% versus 3.5 Flash, improves coding (DeepSWE 49% vs. 37%) and computer-use (OSWorld 83% vs. 78.4%), and lowers API pricing to $1.50 per million input tokens and $7.50 per million output tokens. Flash-Lite targets throughput at 350 tokens per second, with pricing at $0.30 per million input and $2.50 per million output. 1

Access is staggered: 3.6 Flash begins rolling out through the API and Gemini app, 3.5 Flash-Lite is available to developers and in the Gemini app and is rolling into Google Search, and Flash Cyber starts as a limited pilot via the CodeMender agent for governments and trusted partners. 2

Industry & Biz

Moonshot’s Kimi K3 debuts as 2.8T open-weight model

Moonshot AI, maker of the Kimi assistant, unveils K3 with 2.8 trillion parameters and a 1 million‑token context window, calling it the world’s largest open‑weight system and claiming performance that closes the gap with leading U.S. models. Reuters frames the release as part of China’s accelerating AI push. 3

Moonshot’s documentation says K3 uses Kimi Delta Attention and Attention Residuals with a sparse Mixture‑of‑Experts (16 of 896 experts active) and plans to release full model weights by July 27, 2026. The docs note native vision, long‑horizon coding, and knowledge‑work scenarios. 4

Independent coverage adds that K3 still trails the very top U.S. models overall, but has posted standout results on some coding and agentic benchmarks; as of July 16, Arena’s Frontend Code Arena ranked it first, according to Tom’s reporting. 5

Glow exits stealth with $180M to secure AI on endpoints

Glow, founded by former leaders from Meta and Snowflake, raises a $180 million Series A at a $1.2 billion valuation to build an endpoint security platform that maps enterprise devices, governs AI agents and developer tools, and enforces policies before risky software lands on laptops and servers. Backers include Sequoia, Cyberstarts, Greenoaks, and Redpoint. 6

The company pitches prevention over after‑the‑fact detection as attackers automate with generative AI; early deployments (undisclosed customers) have reportedly blocked malicious npm packages and flagged missing endpoint protections. Glow says it orchestrates Anthropic and Google models through Amazon Bedrock alongside its own software. 6

New Tools

OpenAI launches Presence for governed enterprise agents

Presence is OpenAI’s enterprise product for running trusted AI agents in real‑time voice and chat across customer support, sales, and high‑risk internal workflows, bundling policies, guardrails, simulations, and evaluations so agents can take approved actions and escalate to humans. It is offered to eligible enterprises via a limited general‑availability program led by OpenAI’s Forward Deployed Engineers and select integrators. 7

Each deployment is scoped to a specific job (for example, billing issues or IT tickets), with the agent granted only the knowledge and system access required. Companies define what the agent can do, when approvals are needed, and when to hand off to a person; after launch, production sessions and escalations feed an improvement loop where Codex proposes updates for teams to test and roll out. 7

OpenAI says Presence already powers its English‑language phone support at 1‑888‑GPT‑0090, resolving 75% of inbound issues without human assistance, and that its Codex‑powered iteration reduced human handoffs by 15 percentage points in 10 days. BBVA, SoftBank, and IAG are exploring deployments on the same foundation. 7

OpenAI positions Presence as a governed layer around its models; a spokesperson tells VentureBeat it uses OpenAI models for the core agent while allowing third‑party models or services via APIs for guardrails, tools, and parts of the workflow. OpenAI has not disclosed pricing or contractual terms, and Presence is not self‑serve. 8

Community Pulse

Hacker News (59↑) — Concern centers on privacy of voice agents and timing amid safety questions. 9

"What privacy guarantees do these deployments make to end users? If you call a bar in SF or a concierge in Vegas, you'll often get a voice chat agent. There's no terms of service you agree to. In aggregate, it's a pretty rich dataset on openai's end. What happens to your voice, your queries, your PII?" — Hacker News 9

"I wonder if this is really the best timing to release a new product. Why would an enterprise customer choose an OpenAI solution when it's all over the news one of their agents "went rogue"?" — Hacker News 9

What This Means for You

Presence is a packaged way to move from pilot demos to governed production agents—if your organization can commit to a hands‑on deployment with OpenAI’s engineers. It’s limited GA and not self‑serve, so treat it like a consultative rollout: pick one high‑value workflow, define policies and escalation rules up front, and plan a measured launch. 7

If you manage teams that spend heavily on AI tokens, Google’s 3.6 Flash and 3.5 Flash‑Lite offer practical levers: about 17% fewer tokens on common tasks, lower output‑token pricing for 3.6, and up to 350 tokens/second speed from Flash‑Lite—useful for high‑volume content, support replies, or agent handoffs. Run side‑by‑side trials to see if quality‑to‑cost fits your workload. 1

If you’re evaluating model options, the open‑weight trend matters: Kimi K3’s weights are slated for release by July 27, enabling in‑house or partner‑hosted tests where policy permits. For regulated orgs, confirm your compliance stance on Chinese open‑weight models before any sandboxing. 4

Security teams should assume AI agents and dev tooling now live on endpoints. Whether or not you trial Glow, the category signal is clear: inventory which agents and tools can run on employee devices, tighten approvals, and verify endpoint coverage before automating more workflows. 6

Action Items

  1. Request an OpenAI Presence briefing: Ask your OpenAI account team to scope one workflow (e.g., billing issues or IT tickets) for a limited GA deployment and list required policies and escalation triggers.
  2. Draft agent SOPs for a pilot: Write a two‑page standard operating procedure the agent must follow, including tool permissions, data sources, and “hand off to human” criteria.
  3. Test Gemini 3.6 Flash vs. 3.5 Flash‑Lite: In the Gemini app or Google AI Studio, run your top five prompts and compare speed, token use, and output quality.
  4. Review Kimi K3 docs with compliance: If permitted, skim the K3 quickstart and architecture notes, then create a small, policy‑approved sandbox plan for post‑weights testing.
  5. Run a 30‑minute endpoint AI audit: With IT/security, list which AI agents, npm packages, and dev tools can install on employee laptops, and note where approvals or EDR coverage are missing.

Sources 12

Helpful?

Comments (0)