Field note 43/ AI · Engineering

Retrieval Evals This Week: EmbeddingGemma 2, OpenAI's Ironclad Rubrics, and RAGStress

A new open embedding model, a rubric-scored agent eval from OpenAI, and a benchmark that corrupts the knowledge base on purpose. What each one says about testing your own retrieval.

Fig. 01AI engineering · Note 43

Three things landed this week that matter if you build retrieval or grade agents. Google DeepMind released EmbeddingGemma 2, an open multimodal embedding model, on October 6. OpenAI published rubric scores from a research collaboration with Ironclad the same day. And a new arXiv paper, RAGStress, tests what happens to RAG when the knowledge base itself is wrong.

Each one is a useful template for testing your own stack.

EmbeddingGemma 2: one open model for text, code, images, audio, and video

Google's announcement describes EmbeddingGemma 2 as a 740M-parameter embedding model under Apache 2.0, built on the Gemma 4 architecture. Per the model card, that total is a 270M text model plus optional vision (170M) and audio (300M) encoders. Everything maps into one 768-dimension vector space.

Details from Google's pages:

  • Text-only workloads can load just the 270M text model. Google's Sentence Transformers guide shows how to skip the vision and audio encoders.
  • Matryoshka Representation Learning lets you truncate vectors from 768 to 512, 256, or 128 dimensions. Google says that gives up to 6x storage reduction.
  • Context is 8K tokens, four times EmbeddingGemma 1.
  • Google reports MTEB Code going from 68.76 to 78.68 over the first version. That's Google's number on Google's chosen benchmark.
  • Weights are on Hugging Face and Kaggle, with support listed for transformers, sentence-transformers, vLLM, llama.cpp, Ollama, and MLX.
  • The model card lists a pretraining data cutoff of January 2025.

The guide also uses separate task prompts for queries and documents ("Retrieval-query" and "Retrieval-document"). Skip them in your own tests and you aren't testing the model the way Google set it up.

OpenAI and Ironclad: rubric scoring instead of pass or fail

OpenAI's Ironclad post is about training GPT-6 Astra on contracting software. The part worth copying is how they graded it.

Ironclad employees and people who use Ironclad at OpenAI helped pick 11 tasks across legal, commercial, and procurement work, like setting up an NDA flow or a purchase approval process. Each task was scored against 8 to 50 criteria depending on complexity. OpenAI reports a mean rubric score of 55.0% for GPT-6 Astra at Max reasoning, against 41.6% for GPT-5.6 Sol at High reasoning. An internal model hit 63.7%.

OpenAI's own footnotes carry the caveats. The results cover those 11 research tasks, not all Ironclad workflows. The time-per-attempt figures (37.0 minutes for Sol, 19.2 for Astra) are simulated from assumed model speeds, not measured customer time. Training tasks came from public SEC EDGAR contracts, not customer data.

A rubric score of 55% means the model met a bit over half the criteria on average. That's a different claim from "completes 55% of tasks." Partial credit shows you which steps break, which is the useful part for anyone running agent evals.

RAGStress: what happens when the knowledge base is dirty

RAGStress, submitted to arXiv on October 3 by Shiqi Yang, Jiekai Ma, and Gaoyuan Du, starts from a simple complaint. Most RAG evaluation assumes the knowledge base is clean.

The benchmark corrupts it on purpose. It uses four corruption types (factual corruption, numeric typos, relevance poisoning, and contradiction injection) at three severity levels, over 182,546 documents built from 57 MMLU subjects. The authors ran 52,500 model, question, and condition evaluations.

Their reported findings:

  • Clean retrieval can hide differences between models that show up once the data is corrupted.
  • Changes to meaning hurt much more than surface noise like typos.
  • How well a model answers without retrieval doesn't predict how it holds up with corrupted retrieval.
  • Accuracy on a mixed knowledge base shouldn't be read as worst-case performance.

The authors describe it as a controlled stress test, not a general leaderboard.

What we'd check before trusting a retrieval change

We wrote about our own version of this problem in How a Green AI Evaluation Hid Broken Retrieval. An older test setup stayed green while user-facing retrieval was broken, because it tested the stores more directly than the query path people actually used. This week's news points at the same gaps from three directions.

If you're thinking about EmbeddingGemma 2 or any other embedding swap:

  1. Re-embed the whole corpus. Vectors from two different embedding models don't belong in the same index.
  2. Score it on your own queries with known correct documents, through the same query path your users hit. Vendor MTEB numbers tell you about MTEB.
  3. Test each truncation size you plan to ship against that same set before you cut storage.
  4. Put stale, duplicated, and contradictory documents in the test set on purpose. RAGStress is a good model for which kinds of damage to simulate.
  5. Grade answers against a short rubric (right source, current, scoped to the right entity, honest refusal when the evidence isn't there) instead of a single pass or fail.

If you're still deciding whether retrieval is the right tool at all, start with RAG vs. Fine-Tuning.

What to watch

We're watching for independent EmbeddingGemma 2 results on public retrieval leaderboards, beyond Google's own tables. We also want to see whether OpenAI names more software partners from its open call and publishes rubric scores for them, and whether other teams start using the RAGStress artifact bundle on their own pipelines.

Frequently asked questions

4 answers

Read next

Related by topic