Category: AI and Machine Learning

Four shapes, one cluster, very different operational headaches.
“Agent” is doing some heroic heavy lifting as a word right now.
Depending on who you ask, an agent might be:
- A python script firing off 10,000 parallel …
Two weeks ago, Model Context Protocol shipped its 2026-07-28 specification revision.
The biggest updates in this release are: no more stateful initialize handshake, Mcp-Session-Id headers have been retired, and every request …
I created a skill to save users 40-60% on token costs when migrating AI inference workloads from a serverless platform (like Cloud Run or Gemini Enterprise Agent Platform) to GKE. So what is it? What isn’t it? Why a …
Introduction to Distributed RL Sandboxing on GKE
Reinforcement Learning (RL) is the cornerstone of modern AI training. Rather than train a model to produce an expected output, we verify if it has achieved a particular outcome. This is …
The Am Dash And Discerning Human Writing from AI
Was This Post Written by Gemini?*
The other day I was listening to one of my favorite podcasts talking about AI’s writing style. And of course, they were talking about the Em dash. I’ll let …
Fine-tuning Gemma 3 on an A4 Slurm Cluster
Note: This post has been adapted into a formal tutorial on the Google Cloud documentation site.
Overview
This post demonstrates how to fine-tune the Gemma 3 language model on a multi-node, …
Fine-tuning Gemma 3 on an A3 Mega Slurm Cluster
Note: This post has been adapted into a formal tutorial on the Google Cloud documentation site.
Overview
This post demonstrates how to fine-tune the Gemma 3 language model on a multi-node, …