MiniMax M2.7 Defeats OpenClaw as Hermes Default Harness Era Begins

The battle for the future of AI is no longer about who has the biggest model. It is about who can build the best Harness. And right now, MiniMax is winning. The Chinese AI company has quietly replaced OpenClaw as the default model for the open source Hermes Agent project. Its tokens on OpenRouter have already hit 300 billion. But the real story started with a live stream on Bilibili that exploded across the internet.

That live stream changed everything.

The man speaking was Tommy Eastman, the business lead for Hermes Agent, one of the world’s most popular open source AI agent projects.

This was his first live stream for a Chinese audience. And he was not there to talk about EvoMap.

His reason for going live was simple.

The code for Hermes Agent had already been copied once. And when those copied features showed up, no one mentioned EvoMap.

Leaked internal research files from Nous Research showed that some team members had copied open source AI projects without giving any credit.

Of course, they also went into other people’s repos and called it “research.”

But Tommy was not alone. Sitting next to him was the chief engineer of MiniMax Agent, along with several other developers.

As soon as the stream started, the chat went crazy. After several hours of talking, one thing became clear. This was not just about copying code. It was about the future of AI agents.

So what does the future of AI agents look like in 2026?

The Game Has Changed

By now, the old way of working is dead.

Back in October last year, we already had AI IDEs and daily task agents. At the same time, big tech companies were spending billions building agent tools.

But then I realized something. The game had changed.

In the past, AI companies competed on who had the biggest model and the highest benchmark scores.

Then in 2026, OpenClaw came out of nowhere and shook the world. The founder posted a detailed breakdown of his project history. He showed how hard he had worked.

And then everyone suddenly realized that models are just tools. What really matters is the Harness, the system that makes the model actually do things.

Overnight, the entire industry shifted to Harness.

A Harness is an agent testing system. It includes tool use, memory, a Skills system, sandbox environments, and more.

In a widely shared blog post called Harness Engineering, OpenAI described it as a new way of thinking. Plan first, then execute.

ai porn video generator

Because of this, I had to rethink my whole approach.

Models are the brain. But the Harness is the body. Without a good body, the brain cannot do anything. The Harness is what turns a smart model into a real worker.

Claude Code was a big update last year. It added cron jobs, IM remote control, memory files, and more. When OpenClaw came out in January, I thought it was the answer. But the team could not keep up with the demand. They were right about the idea, but they could not scale it.

At that point, I had the same feeling as everyone else.

I felt like I was being left behind by AI. Engineers who understood the Harness were free undressing ai building their own Skills and code. They were not waiting for anyone.

They would look at a project, copy the idea, and make it better. They would take what worked and build something new.

Then MiniMax made its move. After months of quiet work, the Chinese AI company launched M2.7 and two new products. MaxHermes is an open source agent framework. MaxClaw is a general AI agent system.

These two products form a complete loop.

M2.7 is the model layer optimized for Harness. MaxHermes and MaxClaw are the product layers that prove the model works in real life. They also reduce the need for expensive model training.

MiniMax calls this the Model plus Harness double engine.

ai nudifierModel Times Harness Beats Pure Benchmark Scores

The rules of competition have shifted. In the past, we compared models by their benchmark scores. Now we compare them by how much value they create with the same number of tokens.

MiniMax’s solution is simple. Build a model that is made for Harness.

M2.7 was released on March 18. It is a new base model and the first model trained specifically for agent tasks, reasoning, and tool use.

MiniMax built an internal Agent Harness and used M2.7 as the agent base model. They trained it on tool use, reasoning, and optimization data.

The results speak for themselves.

First, the model’s reasoning ability on Harness tasks reached 30% to 50% of what MiniMax’s strong learning team can do on daily tasks.

Second, the model’s tool use ability on Harness tasks reached 100 points on the optimization loop, with efficiency gains of 30%.

Third, the model’s learning ability on the MLE Lite test reached 9 gold, 5 silver, and 1 bronze out of 22 difficulty levels. That is a 66.6% win rate, beating Opus 4.6 and GPT 5.4.

M2.7’s back-end optimization also meets standard agent tool use accuracy. On the Skills loop Agent Harness test, it scored even higher.

With 40 different Skills, each using 2000 tokens, M2.7 could complete 97% of the Skills loops.

The open source community has also welcomed this model with open arms.

Starting from M2.1, Teknium, the co-founder of Hermes, posted on X praising MiniMax models for their tool use speed and cost performance.

With every M2.5 and M2.7 release, the Hermes Agent team integrated them right away. Both internal and external tests showed strong results.

Today, MiniMax models are one of the most used models inside Hermes Agent.

The total token usage on Hermes Agent’s website has grown from 20 billion to 300 billion. M2.7 alone accounts for 250 billion tokens on OpenRouter, making it the top model by usage.

Peter, the founder of OpenClaw, also shared his thoughts on the work MiniMax has done with open source models.

He said that M2.1 could match the performance of top open source models at only 5% of the cost.

Akshay Kothari, co-founder of Notion, also announced that MiniMax M2.5 was added as an open source model option for Notion Custom Agents.

Kilo Code, a strong competitor to Cursor in the AI coding space, also made MiniMax its default model.

Tommy made a bold claim during the stream. “The gap between open source and closed source in China is closing. The difference is getting smaller every day.”

Behind this is a new kind of partnership.

Hermes focuses on building the agent framework and product layer. MiniMax focuses on turning the model layer into a standard tool and building the infrastructure.

The Hermes framework gives MiniMax a clear direction for model optimization. Tool use, Skills execution, and reasoning are the real tests for any agent. MiniMax’s models are built to push these agent abilities to their limit.

When open source projects choose MiniMax as their default model, it says one thing loud and clear.

Benchmark scores do not win the future. Harness wins.

Two Products for Two Paths

Models and Harness need to work together. But you also need product proof and real world feedback.

To do this, MiniMax launched two products at the same time. MaxHermes and MaxClaw serve two different agent paths.

MaxHermes is for open source users. It is based on the Hermes Agent framework and focuses on learning and growth.

Every time it completes a complex task, the agent automatically learns from it. It creates better Skills and stores the results as training data for the next round.

It also supports long-term memory, natural language task scheduling, multi-turn chat, and workflow design. It is like a small team of AI workers living inside your computer.

For the Skills layer, OpenClaw founder Peter has said that preset Skills are too rigid. He wants a system that can grow.

Under this idea, MaxHermes makes Skills into something agents can create on their own. It is like having a new intern every day who gets smarter over time.

MaxClaw is based on the OpenClaw general AI agent system. It handles complex tasks, file processing, and high-level automation. It can work for 120 hours straight, process large amounts of data, fix code, respond to IM messages, and even kick out users who break the rules.

For power users, MaxClaw offers expert-level Skills and 50GB of storage.

The built-in image recognition, video processing, and web page reading Skills are all connected to the system. Users can also create their own tools, all available through API.

All Skills are preset. All Skills are free. Users can modify them. It supports running multiple tasks at the same time. It works on both iOS and Android.

To make it easier to use, MiniMax built a Skillhub where users can browse, install, and manage Skills. With one click, they can affect the model and the Harness. It is like an app store for AI agents.

On the platform side, MiniMax Agent launched Expert 2.0. Users just need to set a goal, and the agent automatically creates SOPs, assigns tasks, and calls Skills, SubAgents, and MCP tools as needed. It is built on 1.6 billion parameters and expert agent models.

What is worth noting is that MiniMax is also eating its own dog food.

Leaked reports show that every employee inside the company has multiple GitHub accounts. Each one automatically scans open source projects and uses MiniMax models to write code and send pull requests.

The company uses its own agent tools to build its own products. This proves the system actually works.

These products are not just demos. They are real tools that get real results. Every task completed by M2.7 on the tool use and Skills loop becomes training data for the next round of optimization. This is how the flywheel turns.

The real battle between models and products is not about who has the best benchmark. It is about who can build a closed loop where the model and the Harness make each other better.

The Real Sandbox War

In the agent world, model training is only the first step. What really matters is giving every agent a safe place to run, learn, and make mistakes.

Among all the possible bottlenecks, the sandbox environment is the most likely to kill a project. If it is too slow, users leave. If it breaks, the whole system fails.

The bottom layer infrastructure must support sandbox speed. Companies need to build large training clusters, manage thousands of agent tasks at the same time, and handle signals and data flow between tasks.

MiniMax split its training and inference into separate systems.

For training, MiniMax uses Tencent Cloud clusters and its own Agent Runtime sandbox. The Forge reinforcement learning infrastructure supports training that simulates millions of agent task executions. It provides 80ms response time, 60 sandboxes per user, 99.99% task success rate, and supports M2.7 full parameter training.

For inference, both MaxClaw and MaxHermes run on ACK and ACS.

MiniMax uses a unified execution platform. ACK handles unified scheduling. ACS Agent Sandbox provides 20 to 40ms response time and supports 15,000 sandboxes per user. It also has automatic cleanup and elastic scaling.

Tencent Cloud provides training and inference support.

At the same time, companies that choose MiniMax as their core backend are seeing real results. Many have achieved double growth in both product and model metrics.

The Hidden Math Behind Tokens

In the past, we compared models by their benchmark scores. Now we compare them by one thing. How much value can the same number of tokens create?

MiniMax CEO Yan Li said in a March business call that the value of an AI platform equals user density times token usage.

MiniMax’s answer is not the only one. But it is the only one that has been proven at scale. Build a model for Harness. Let Harness train the model.

When a company launches a product and an open source project chooses it as the default model at the same time, the logic behind it is simple. The model works.

Now, everyone is asking one question. When is M3 coming?

MiniMax has already hinted at the answer.

It will not be long.