Indexed — Issue 003
Monday, July 27, 2026. Google shipped three Gemini models built for agents, skipped the flagship again, and quietly raised its 2026 capex ceiling by $15B. If you run AI automation jobs in production, the cost math just moved.
Speak naturally. Send without fixing.
Wispr Flow turns your voice into clean, professional text you can send the moment you stop talking. Not rough transcription you have to clean up. Actual polished text — ready for email, Slack, or any app.
Speak the way you think. Go on tangents. Change your mind mid-sentence. Flow strips the filler, fixes the grammar, and gives you text that reads like you spent five minutes writing it.
89% of messages sent with zero edits. Millions of professionals use Flow daily, including teams at OpenAI, Vercel, and Clay. Works on Mac, Windows, and iPhone.
The Signal
Google released Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber on July 21, framing the launch around efficiency, latency, and reliability for agent workloads at scale. Gemini 3.5 Pro, last updated in February, is still delayed — DeepMind product lead Logan Kilpatrick says it is now in partner testing and hopes to “land soon.”
Two days later, Alphabet raised its 2026 capex plan to $195B–$205B, up from $180B–$190B last quarter, on the back of an 82% jump in Google Cloud revenue to $24.8B. Shares dropped about 3% after hours as analysts pressed Sundar Pichai on whether Google can stay at the frontier.
Meanwhile Moonshot's Kimi K3 landed at $15 per million output tokens — half of GPT-5.6 Sol and roughly a third of Claude Fable 5, per The Verge.
What it means for you: the Flash tier is where operators live. If you route agent traffic through Gemini or a router, re-run your cost model this week — the cheaper Flash tier plus Chinese open models pressure your incumbent's per-token bill. Do not wait for 3.5 Pro before switching production workloads to Flash 3.6; the delay is now a pattern, not a hiccup.
The Stack
Three tools worth putting on the shortlist while pricing shifts:
Google Gemini — the front door to the new Flash tier if your team already lives inside Docs, Sheets, and Gmail. Deep Research and Workspace grounding are the reasons to keep it in rotation even if Pro slips again.
Claude — Opus 4.6's 1M-token context is the honest answer when you need to feed a full contract, policy set, or research packet into one conversation without chunking.
Perplexity AI — the fastest way to fact-check a claim like “Kimi K3 is half the price of GPT-5.6” before you paste it into a board deck.
Pair one chat model, one long-context reasoner, and one cited search tool. That's the minimum viable stack for any operator running AI automation jobs across research, drafting, and QA.
Prompt of the Day
Use this to pressure-test whether Gemini 3.6 Flash actually fits your workload. Paste your current model, your task, and your monthly token spend into ChatGPT or Claude, then send:
You are a pragmatic AI infrastructure advisor. I currently run [my task] on [my model] and spend roughly [$ amount] per month on inference. Google just released Gemini 3.6 Flash, positioned as cheaper and more token-efficient than the prior Flash tier, and Moonshot Kimi K3 is priced at $15 per million output tokens. Walk me through: (1) which of my workloads are latency-sensitive vs. quality-sensitive, (2) which could plausibly move to a cheaper Flash-class model without quality loss, (3) a 2-week A/B test plan with concrete pass/fail criteria, and (4) the switching cost I am ignoring. Be specific and skeptical.
The trick is forcing the model to name switching costs — data plumbing, prompt rewrites, eval regressions — instead of just quoting per-token savings. If you want a second set of eyes on the output, operators are comparing their A/B results and switching notes inside AI Freedom Lab this week.
Tool of the Day
ChatGPT — still the default entry point for most operators, now on GPT-5 with o3 reasoning available for the multi-step work.
Why it earns the slot today: while Google resets its Flash pricing and Anthropic pushes Opus, ChatGPT is the one workspace where you can draft the client email, debug the ingestion script, analyze the uploaded CSV, and generate the launch visual without leaving the tab. For most solo operators and small teams, the honest bottleneck is context-switching, not model quality.
Concrete use today: upload your last three client reports and ask GPT-5 to build a reusable Custom GPT that produces the fourth in your voice. That is a one-hour investment that pays back on every future report. Verdict: if you are choosing one paid seat this quarter and you do not already have a strong reason to prefer Claude's long context or Gemini's Workspace grounding, ChatGPT is still the safest default.
Quick Hits
Cisco open-sourced Antares-350M and Antares-1B on Hugging Face, small models tuned to hunt vulnerabilities in code repos. Access is gated to verified cyber defenders. Useful if you run security review on customer code and want to stop shipping source to an outside API. Axios has the details. Caveat: benchmarks are Cisco's own.
Substack integrated Pangram so readers can scan posts and comments for AI-written content. If you publish there, assume every post now carries an implicit AI-percentage score in a reader's mind. Full story on TechCrunch.
Moonshot paused new Kimi K3 subscriptions after launch demand overwhelmed capacity, per its own claim to The Verge. Treat availability as unstable for the next two weeks before you plan production dependencies on it.
Alphabet is now booking revenue from direct TPU sales for the first time. If Nvidia's pricing power ever cracks, this is the quarter to circle.
Closing the loop
The pattern this week: incumbents are competing on Flash-tier cost while their flagship models slip. That is good news for anyone running AI automation jobs on thin margins — the useful tier just got cheaper twice in a month. It is bad news if your product's moat depended on being the only team with access to a top-tier model.
Two things to do before Friday. One: pull your last 30 days of token spend and mark which workloads are truly quality-sensitive versus which you kept on the expensive model out of habit. Two: run one real A/B between your current default and Gemini 3.6 Flash on a workload you can measure. Ship the winner, log the loser.
Today's Tool of the Day: ChatGPT on the AI Tools Index
Missed an issue? Browse the Indexed archive


