热心市民王先生

X 上的 张小珺 Xiaojun Zhang:“The Second Interview with Kimi's Yang Zhilin: "Standing at the Beginning of Infinity"” / X

[

张小珺 Xiaojun Zhang

](https://x.com/zhang_benita)

[

@zhang_benita

](https://x.com/zhang_benita)

[

图像

](https://x.com/zhang_benita/article/2079758352213762367/media/2079749672290390016)

The Second Interview with Kimi’s Yang Zhilin: “Standing at the Beginning of Infinity”

8

41

279

[

5.2万

](https://x.com/zhang_benita/status/2079758352213762367/analytics)

This was my second interview with Yang Zhilin, conducted shortly after Kimi released K2.

As described in my earlier post, our first interview was published on Kimi’s first anniversary; the second took place a year and a half later (published in August 2025) — making it Yang Zhilin’s most recent in-depth interview to date.

The year between the two interviews was a difficult one for Kimi, swinging back and forth between peaks and troughs. You can read the founder’s changing state of mind between the lines.

Compared with the previous piece, this interview goes much deeper into technical detail — it records Yang Zhilin’s complete account of how AI escapes the “brain in a vat,” moves toward agents, and ultimately toward AI self-evolution. Many of the technical judgments made here are already visible today, on the day K3 is released.

For the first interview, we only had text and an audio podcast. For the second, we also have a video version, which you can watch on YouTube (with English subtitles).

Here are the links for you to choose from:

The first interview with Yang Zhilin (published March 2024)

· The original Chinese text:

https://mp.weixin.qq.com/s/w2NdX-9O2VVP9qQIa-Wcmg

· The audio podcast:

https://open.spotify.com/episode/6DyY8CH0TIeaHmdCQV53Se?si=6KcYPziBRhCH7aHlWr6Kfg

The second interview with Yang Zhilin (published August 2025)

· The original Chinese text:

https://mp.weixin.qq.com/s/uqUGwJLO30mRKXAtOauJGA

· The audio podcast:

https://open.spotify.com/episode/4bZkVrLi3O0hDQksf4cIOq?si=K3LplonOTfSKeMGTL6sw0g

· The YouTube video podcast (with English subtitles):

https://youtu.be/91fmhAnECVc?si=BgYEodr_CgQ2SBsr

Yang Zhilin: “Maybe one day we’ll discover that this snow mountain has no summit — I don’t know — I hope it never does. That is what The Beginning of Infinity means: It is an infinite mountain.”

By Zhang Xiaojun

Yang Zhilin arrived a little late, wearing a wrinkled black T-shirt, and grabbed a McDonald’s burger, taking a few hurried bites. It was late July 2025, shortly after the release of the Kimi K2 model.

A year and a half had passed since our last interview. In January 2024, I first met Yang Zhilin in a dim conference room with the hum of air conditioning. That March, I published “The Second Interview with Kimi’s Yang Zhilin: Marching Toward an Endless, Unknown Snow Mountain.” He was 32 at the time, and his one-year-old foundation-model company had drawn widespread attention.

As a founder whose academic background fits large language models perfectly, he quickly rose to fame — and controversy soon followed.

For the year after that, he almost stopped giving interviews or attending events, as if he had sealed himself off. In early 2025, DeepSeek reached the peak of global renown. Inside the company, employees could hardly detect any emotional fluctuation in him; they only noticed that he had added the suffix “a friend of time” to his Feishu display name.

In July 2025, the Kimi K2 model was released, and the company returned to the public eye. It is an open-source coding and agentic large language model based on a Mixture-of-Experts (MoE) architecture. Figuratively speaking, through its coding ability the model escaped the enclosed “brain in a vat,” grew “hands,” and began to manipulate the external digital world. Nature even described it as “another DeepSeek moment.”

Throughout the past year’s storms of public opinion and entrepreneurial ups and downs, Yang Zhilin repeatedly read The Beginning of Infinity by the British physicist David Deutsch. Two sentences from the book influenced him deeply — “Problems are inevitable,” and “Problems are soluble.”

In his view, human civilization is a process of continually conquering problems and expanding the boundaries of knowledge. Solving one problem gives rise to more new problems; the space for research thereby opens up, and the frontier of knowledge extends.

“That is exactly what makes it interesting — you always have new problems to solve, and every time you solve one, technology climbs a few hundred meters higher. Maybe one day we’ll find that this snow mountain has no summit — I don’t know — I hope it never does. That is what The Beginning of Infinity means: it is an infinite mountain.”

This is a story of mutual conquest between “humans and AI,” between “the climber and the mountain,” between “the human individual and the machine system.”

Only now, Yang Zhilin is still on the mountainside, and the storm has not stopped.

The following is an edited excerpt of the interview with Yang Zhilin (the author has polished the language).

Chapter One: An Infinite Mountain

01 The Beginning of Infinity

Zhang Xiaojun: At the end of your first year of entrepreneurship, our 2024 interview was titled “Marching Toward an Endless, Unknown Snow Mountain.” Another year has passed. Standing here now, in July 2025, how do you feel?

Yang Zhilin: The phrase you just mentioned… it feels like ages ago. One day in AI is a year in the human world. I don’t even know how many human days make up one AI year. Many things have indeed changed, but the “snow mountain feeling” you describe is pretty much the same.

Toward the summit, we have walked another stretch of the way.

Zhang Xiaojun: Where have you gotten to now?

Yang Zhilin: From where we stand today, the model’s progress has been substantial — two years ago it couldn’t even write a coherent article; now it can not only write very good articles, but also work continuously for hours to help you complete a highly complex coding task. That was hard to imagine two years ago.

In the process of climbing the snow mountain, we have unlocked some new scenarios and roughly know what the path in between looks like. But at the same time, as you climb higher, you still observe a similar landscape — ahead, there are still many unknown technical problems waiting to be solved.

Zhang Xiaojun: On a peak surrounded by heavy snow, have you become clearer, or more lost?

Yang Zhilin: Many things have definitely become clearer. Two years ago, it wasn’t clear how the various reinforcement learning (RL) paradigms should be done, or how to give models stronger reasoning ability or stronger agentic capability. Back then the focus was more on how to do pre-training better and how to use RLHF (Reinforcement Learning from Human Feedback) to improve the conversational experience.

But now some of those questions have been answered — and those answers have unfolded into new questions of their own.

Today, although we can do reinforcement learning, it ultimately depends on good evaluation or verification. If you ask a model to solve a math problem, or a coding task with test cases, it can do quite well. But if you ask it to perform a more complex end-to-end task, it is sometimes hard to find a way to evaluate or measure it. So the system generates new problems.

This is a bit like a book I’ve been reading recently, called The Beginning of Infinity. I’ve read it several times. Deutsch says there are two sentences that could be carved in stone: one is “problems are inevitable”; the other is “problems are soluble.”

You could say that before the Enlightenment, society was static. People didn’t pursue innovation; you would explain the phenomena you saw with mysticism, but those explanations were not good explanations. For example, when you saw thunder in the sky, you would think the thunder god was angry; when it snowed in winter, you would say some deity was in a bad mood — the entire social structure was static, and only a tiny number of people were truly doing scientific research or creating knowledge.

But after the Enlightenment, society became dynamic, and new knowledge is constantly being created. Every time you solve a problem, it brings new problems. Problems are inexhaustible, because as the boundary of your knowledge expands, you encounter new ones.

AI research and development happens to be in exactly this state right now. You solve some problems in reinforcement learning, and then you run into problems of evaluation, measurement, and verification, which require us to find new answers.

But that is exactly what makes it interesting — you always have new problems to solve, and every time you solve one, technology climbs a few hundred meters higher.

Maybe one day we’ll discover that this snow mountain has no summit — I don’t know — I hope it never does.

That is what The Beginning of Infinity means:

It is an infinite mountain.

02 It’s Still a “Brain in a Vat”

Zhang Xiaojun: Let’s look back. Over the past year, what were the most important things in global foundation models in your mind? Which of them were paradigm-level changes in artificial intelligence?

Yang Zhilin: One is the long-thinking reasoning model, represented by o1 as the first of its kind. Essentially, it works by letting the model make many attempts and reflections along the way — reflection being the key.

Reflection consists of two abilities: one is proposing new conjectures; the other is verifying them.

You can understand it this way: in the process of solving a problem, the model continually proposes new conjectures, and those conjectures get self-verified. For instance, after it proposes a conjecture, judging whether it is right or wrong requires a certain verification capability. Although you are not explicitly training a verification model, verification happens implicitly during the reasoning process. It makes multiple conjectures and verifications for a problem, and finally arrives at an answer.

This greatly improves the model’s capability. Originally you could only take one shot and give an answer directly — that answer might be right or wrong, and there was no such process. But now you can keep proposing conjectures and verifying them, which is equivalent to having tried several times.

You can turn Pass@k into Pass@1 — that is essentially the principle.

This is very similar to how humans do scientific research or solve problems: continually proposing new conjectures and then verifying them.

Zhang Xiaojun: It’s a process of free exploration rather than a linear process.

Yang Zhilin: The way it works effectively is, in many cases, still fairly linear. If you don’t consider parallel sampling — suppose you’re doing serial sampling — then it is a linear process.

Each time you propose a new conjecture, that conjecture may be based on previous conjectures, or even on conjectures you have already rejected, and you propose a new one from there. It approximates a more linear process.

Of course, you can now combine the linear process with parallel strategies — for example, sampling many candidates at once, blending parallel and serial approaches. Some recent papers also argue that serial sampling has a higher ceiling, which is consistent with our experimental conclusions.

What I described above is one paradigm, but it is still a “brain in a vat” — it doesn’t need to interact with the outside world.

Zhang Xiaojun: A brain in a vat?

Yang Zhilin: Imagine a fish tank. You put a brain inside it, with no connection to the outside world. It just thinks inside its own head, keeps thinking, and can solve a problem without any interaction with the outside world.

But there is another very important paradigm: multi-turn agent reinforcement learning, or agentic models trained with RL techniques. Its defining feature is that it interacts with the outside world a lot.

For example, I think while performing actions — possibly many rounds of actions: calling a search engine one moment, using a browser the next, writing a few lines of code after that — solving a problem over multiple turns.

It is no longer a “brain in a vat”; it interacts with the outside world — my next action depends on the feedback I get from those interactions and on the updated state the outside world gives me.

But both of these things point to the same thing: test-time scaling. It means you can achieve better scaling at test time, or at inference time.

Take Chat as it used to be: it was mostly about outputting a result in a single turn. I ask you to write an article, and you write an article; I ask you to polish it, and you output another few hundred tokens — a very small number of tokens. But both long-thinking reinforcement learning and agent reinforcement learning are essentially ways of scaling up tokens at prediction time.

Whether you run more turns or generate more thinking tokens within each turn, both are methods of scaling tokens, enabling you to complete more complex tasks.

This also comes with longer completion times. You can now spend hours doing one complex thing without any human involvement in the process. For example: clone a code repository, translate it into another programming language, debug it, test it, fix all the bugs, and get it running properly. This kind of work can be done end to end, thanks to the scaling of test-time compute.

Another interesting trend is that more model companies are now building “first-party agent products.”

At the beginning — say, half a year or a year ago — many products were built on top of foundation models: you would erect some scaffolding on top, or design tools for the model to use better, and thereby assemble a product.

Zhang Xiaojun: Enjoying the capabilities that spill over from the model.

Yang Zhilin: Right. What it essentially does is reverse-engineer the model’s training process. Because the model’s training process also relies on all kinds of means — you can imagine that Anthropic trained such a model using its in-house environments, tools, and scaffolding, but it didn’t open those directly to you.

Through reverse engineering, you get closer to fitting its distribution — which tools work best? Which system prompt works best? What kind of context engineering works best? It is a process of reverse engineering.

But you’ll find that when a model company builds a “first-party product,” the logic is completely different.

You no longer need that reverse-engineering process; it becomes a forward approach. I design the tools first, design my context-engineering methods first, and then I train the model inside this very environment — so the model naturally performs better in your environment.

These are two different lines of thinking, but the second one probably has a higher ceiling.

You can integrate tools and models much better. If the model handles something poorly, you can adjust the tool design, make it better, and at the same time train end to end. This is also a fairly large variable in the way development is done.

Zhang Xiaojun: To make it easier for everyone to understand: the first approach you described, the scaffolding kind, is the development style of something like Manus; the second is your approach — training the model end to end.

Yang Zhilin: Right. Of course, our investment in “first-party products” isn’t particularly large yet; the model remains our main line. But Claude Code or ChatGPT Agent are “first-party products,” and that should be a major trend as well.

What comes next depends on how “first-party” and “third-party” products fit together, and what that looks like in the ecosystem.

03 L1 to L5 Are Not Necessarily a Serial Relationship

Zhang Xiaojun: Speaking of the main line — OpenAI laid out a five-level framework from L1 to L5, and I’ve always been curious about the internal logic behind it.

L1 is Chatbot, L2 is Reasoner, L3 is Agent, L4 is Innovator, and L5 is Organizer.

Why does Agent come only after Chatbot and Reasoner? And why do Innovator and Organizer come next? How does this sequence progress in terms of capability structure?

Yang Zhilin: It is a step-by-step dependency of capabilities. The ceiling of Agent (L3) depends on having strong Reasoning (L2) capability — but it is not a strict prerequisite that Reasoning must come first.

Suppose we slightly rearranged the order of technological development: you first develop Agent capability, and then develop Reasoning in the narrow sense — that is, long-CoT (long chain-of-thought) reasoning. That could work too.

You could say Claude’s route is a bet on exactly this: it hasn’t done particularly much on Reasoning, but it does extremely well on Agent. Behind this are bets on different technical paths. But ultimately you can’t get around it — if you want to climb a few more steps toward the summit, you need both capabilities; it’s only a matter of time.

So they are not necessarily mutually dependent. It’s not that you must do Reasoning first and then Agent. But to build the best Agent, you must also make Reasoning the best. And once you have Agent capability, you can do some of the things that come after.

Why is the next stage (L4) Innovation? The most critical point here is: when can a model participate in the development of the model itself? Only when the model participates in the development process can you unlock the true Innovator stage.

We hope K2 can participate in the development of K3. Without agentic capability, that would be very hard. But once you have agentic capability, the model can propose new ideas, run the corresponding experiments, analyze the results, draw conclusions, iterate on the next version of the idea, or optimize the performance of some piece of infrastructure. All of these depend on strong agentic capability.

Innovation (L4) and Organization (L5) are not necessarily a completely linear relationship either; some of it is parallel. We can already see this trend —

For example, once you have one agent, you can expand it into a multi-agent system. You can fork many different agents from one agent and let them do different things. Some work serially, some in parallel; then you merge, then split into different tasks again. Maybe some write tests, some write documentation, some design the software architecture — each with its own division of labor.

So it’s not necessarily a linear relationship; the two (L4 and L5) may happen simultaneously.

But Reasoning (L2) and Agent (L3) are probably the prerequisites for Innovation (L4) and Organization (L5).

Zhang Xiaojun: The hallmark of Innovation (L4) is the model iterating on itself. What about Organization (L5)?

Yang Zhilin: A relatively simple way to think about it is that it will be a multi-agent system.

Of course, how to train a multi-agent system well end to end — without overfitting to a few specific agent types, so that it generalizes better — is quite challenging.

Zhang Xiaojun: Is Organization the “summit of the snow mountain” for models?

Yang Zhilin: Not really. Maybe there truly is no summit.

Zhang Xiaojun: So what kind of yardstick do you take the L1-to-L5 framework to be?

Yang Zhilin: It marks several important technical milestones, but they are not necessarily in a serial relationship. It’s not that we expect one capability to be fully solved before moving on to the next problem.

Take Reasoning. If you really want to solve open-ended reasoning problems, you need strong Innovation — proposing new model architectures — which in turn places higher demands on reasoning capability. Essentially, you use L4 methods to solve L2 problems and make L2 better.

These capabilities will keep getting better over time.

That said, the different technical bets you make along the way will cause some divergence in your short-term path. And those short-term divergences matter — because you are facing a dynamic market.

Zhang Xiaojun: We used to take it for granted that the destination was AGI. If the destination today is no longer AGI, then what is AGI?

Yang Zhilin: AGI is not a particular step on the staircase, where you climb onto that step and suddenly, overnight, achieve AGI. Rather, it is a direction.

In many domains today, you could argue it’s already AGI — better than 99% of humans. In many math or programming competitions, at the current rate of improvement, one can expect many problems to be fully solved soon.

AGI has two dimensions: on one hand, the technology keeps improving; on the other, there’s the technology’s impact on human society. The latter is a much longer-cycle matter, and it is also part of AGI.

It’s a bit like the steam engine: after its invention, society needed decades, even centuries, to digest the change. Some jobs become unnecessary, but new jobs are created. Everyone becomes a “superhuman” who can do more. The way society works and its operating efficiency will change dramatically.

Although we named our company “Moonshot,” it differs from a moon landing. With a moon landing, the moment you stand on the moon you can proclaim “I made it.” With AI, it’s hard to suddenly shout a slogan at some point in time and say “we have achieved AGI right now.”

You just keep climbing.

Even, after a while, it may not be you doing the climbing — it’s you using AI to climb. Right now we have K2 do data processing, model analysis, and model training — things that used to require humans — and we’ll gradually hand more of that over to models. It’s like an amplifier that helps you climb the mountain better.

Zhang Xiaojun: If the snow mountain is endless, what are you pursuing?

Yang Zhilin: The process of climbing itself.

You used to be at the foot of the mountain; now you’ve climbed a bit higher, and the view you can see is different.

It is a process of dynamic evolution.

Chapter Two: K2 Is Mount K2

04 Fed the Same Amount of Data, the “Brain” Grows More

Zhang Xiaojun: Let’s review the key decisions of your two years of entrepreneurship. From 2023 to 2024, your key decisions were: deciding to start the company in February 2023, raising funds, and building the team; then in the second half of the year, Kimi went live and you bet on long context.

From 2024 to 2025, what were your key decisions this year?

Yang Zhilin: One very important point is that, technically, we shifted from an R&D paradigm centered on pre-training and SFT (Supervised Fine-Tuning) to one centered on pre-training and reinforcement learning. That required doing many things — both building up talent and changing the way we do R&D.

Another point is that the shift from conversation to Agent is an important paradigm change, and it has greatly affected the way we actually work.

Zhang Xiaojun: Over the past half year you released K1.5 and K2. What does each of them mean for Kimi?

Yang Zhilin: K1.5 was more about validating reinforcement learning techniques.

Zhang Xiaojun: Catching up with o1?

Yang Zhilin: Right. We invested in this technical route relatively early, got some results, and figured out how the underlying technology actually works.

At the time, we found that you don’t really need much process reward or a value function — they even had some side effects during training.

We found that you can train extremely well using end-to-end reward signals directly. That wasn’t entirely clear in the early days. Through this process, we accumulated reinforcement learning infrastructure and some algorithmic know-how.

K2 has several priorities. First, we wanted it to be a very good base model. If you want a better base model, you have to look at where the current pre-training bottleneck lies for the whole field.

We found that the growth of high-quality data is indeed very slow, and multimodal data cannot effectively improve the “IQ” of text itself — you can think of high-quality data as approaching a constant. Under these circumstances, we wanted to maximize the use of every piece of data — the so-called token efficiency.

You want the brain to grow more while consuming the same amount of data; you want to extract more intelligence.

This line of thinking differs from before. Suppose you do a lot of performance optimization on the training system to make training faster — that’s certainly valuable. But faster training by itself doesn’t raise the ceiling of intelligence, because the number of tokens is still the same. Training faster just means finishing the training in less time; the model doesn’t necessarily get better. That is an optimization of training efficiency, or compute efficiency.

People have worked on that before. What we want more now is to improve token efficiency — to make one piece of data do the work of several.

We pay close attention to things like the Muon optimizer. It’s very interesting and improves token efficiency significantly. The Adam optimizer has been used for ten years — most model training uses Adam — but its token efficiency isn’t good enough.

The Muon optimizer doesn’t treat each element independently; instead, it considers the parameters of an entire matrix together, along with the dependencies among them. In this way, it achieves better learning efficiency — from the same piece of data, you learn more intelligence.

In our early experiments, under compute-optimal conditions, there was basically a twofold improvement. In other words, learning from one piece of data was equivalent to learning from two pieces with Adam. If you have 30T high-quality tokens, it’s equivalent to now having 60T high-quality tokens.

Zhang Xiaojun: But the actual amount of data is still the same.

Yang Zhilin: After it learns, the brain grows faster. Because the learning efficiency is higher and the optimizer is better, it absorbs faster.

You feed it the same amount of data, but it absorbs it better — the compression rate rises faster and the loss drops faster.

Zhang Xiaojun: Is the Muon optimizer your original invention?

Yang Zhilin: The Muon optimizer was proposed by Keller Jordan (a computer scientist and machine learning engineer who joined OpenAI in December 2024). We made many optimizations on top of his work so that it could be adapted to train very large-scale language models.

We previously had a project called Moonlight that enabled it, for the first time, to train language models at a certain scale. Later, as we scaled it further, we discovered many new pitfalls — for example, the max logit (an indicator used to observe whether training is proceeding normally) could explode.

This problem is hard to detect in small-scale experiments, but you encounter it in large-scale training. So we proposed some new methods, such as clipping techniques, that allow it to train well even at very large scale.

This is extremely important — because the number of tokens is limited, you want every token to generate more value.

Zhang Xiaojun: I read your technical report. You tried using existing models to rewrite existing data and generate new corpora. What exactly is the rewriting strategy? — The report didn’t mention this.

Yang Zhilin: We perform a lot of rephrasing on the data. Say you have 30T tokens, but the high-quality portion is much smaller — maybe only tens or hundreds of billions of tokens. You want this high-quality data to be well utilized. We did some rewriting of this data so that it can be better absorbed by the model and generalize better.

The main idea is that if you learn the same piece of data many times, its generalization may not be as good — there are overfitting issues. Through rephrasing, we hope to give it a degree of generalization.

There are many specific ways to rewrite, and we found one that works relatively well in experiments.

Zhang Xiaojun: Which one?

Yang Zhilin: That space is also very large; there are many research opportunities.

Zhang Xiaojun: What do you think of this view — “Rewriting and expansion are actually useless. If knowledge can be written out, that means the knowledge was already in there; there’s no new knowledge, unless you use other methods while rewriting.”

Yang Zhilin: That’s a great question, and it does depend on how you rewrite. Theoretically, it still comes down to whether there is an input of new entropy. That places certain requirements on the rewriting method. But we’re not necessarily using the best rewriting method right now; there’s a lot of room to explore.

Back to the earlier point: with the K2 model, on the one hand we wanted it to be a good base model. We very much wanted to improve its token efficiency, and these are our corresponding designs — including adding more parameters through greater sparsity.

That also raises its token efficiency, because with more parameters, even though you learn from the same amount of data, you absorb it better. In any case, experiments verify that there is indeed better token efficiency.

Second, we wanted it to have strong agentic capability. Through various kinds of reinforcement learning, or simulations of tools and environments, you give it relatively good generalization.

For an agentic model, the biggest challenge right now is generalization.

Because current RL techniques are limited in that both the training tasks and the evaluation metrics are, in many cases, single points. For example, if you train on data from the same distribution as SWE-bench, it improves SWE-bench — that’s a certainty. But after the metric goes up, it doesn’t mean the model’s generalization has become better.

We are also trying to solve part of the generalization problem. We don’t want to overfit to certain tools, or to certain environments, or to certain specific tasks. Those tasks may be good observations, but we don’t want to overfit to them.

This problem is more severe in agent training. Compared with conversational models, generalization for agents is a bigger challenge.

Zhang Xiaojun: Agentic capability is currently trained mostly in the post-training stage. Why not train it in the pre-training stage?

Yang Zhilin: That’s also something we want to explore next.

Zhang Xiaojun: Could that improve generalization?

Yang Zhilin: It depends on how you do it — for example, whether the data distribution is broad enough, and whether there are good evaluation methods.

Evaluation as a whole is now an important bottleneck preventing agent models from becoming more general. You will gradually notice that there aren’t many benchmarks available for agents. When you observe a score on those benchmarks, it often doesn’t reflect the actual capability — it’s rather one-sided. This is a problem everyone needs to figure out.

One potential line of thinking: we need to train AI in a more AI-native way. We want models to participate in more of the training process. For example, if your AI can do good alignment research, it will theoretically generalize better, rather than merely optimizing a few single-point tasks.

Today, agents don’t yet generalize as well as conversation does — the next few hundred steps up the snow mountain may well be this.

Zhang Xiaojun: It sounds like K1.5 was following OpenAI, while K2 is racing ahead.

Yang Zhilin: We’ve drawn on many technical directions, but we also hope to have some innovations of our own. At least in the public record, we are the first to use a non-Adam optimizer — one based on matrix orthogonalization — to train a model at such a large scale. That is one innovation.

Some of our approaches to agent data were also, at least among publicly available materials, done relatively early.

What’s very interesting is that as you climb higher up the snow mountain, you find the space actually gets bigger. Because the number of tokens used to accomplish the same task is increasing, and the problems are becoming more complex.

Just as I said earlier: problems are inevitable, but problems can always be solved.

These inevitable problems may look more numerous than before, but your space for research grows broader along with them.

05 When You Train with Muon, It Blows Up

Zhang Xiaojun: Let’s talk specifically about the K2 project. How was it initiated? How long was the preparation?

Yang Zhilin: The preparation took a fairly long time; many of the technologies involved have been under research since last year.

A technology like Muon requires a relatively long research cycle. At the beginning you run early experiments and find that the idea has potential. We run small experiments to validate how much potential an idea has.

Once you have the idea, getting to the point where you can apply it to train a trillion-parameter model requires validating its effectiveness through scaling experiments at different scales. Some problems only reveal themselves after you scale to a certain size — so the cycle is fairly long.

Of course, if you only count the model training itself — from pressing the training button to the end of training — the time isn’t that long. But R&D requires doing many things further upstream to ensure the final training goes relatively smoothly.

Zhang Xiaojun: When did you make the bet on building an agentic LLM?

Yang Zhilin: That also required a lot of accumulation; it’s just that the approach differs at different points in time.

At first you don’t necessarily do it end to end, but you accumulate environments and data; later on, you do reinforcement learning more end to end. In between you need a lot of infrastructure and data accumulation. It’s hard to do it very well in just one or two months.

Overall, I feel that foundation models and the related technologies really require time to accumulate. It’s still about “being a friend of time.” The technology curve is somewhat steep; you can’t just decide to do it today and have it done.

Zhang Xiaojun: So when was the project initiated? Why initiate this project?

Yang Zhilin: We spent a year accumulating various technologies, but K2 was definitely a decision made in the last few months — we decided to train such a model and chose which technologies to put into it. That’s roughly how the decision went. But it’s not that I decided today to train this model and started from zero.

We’re always training the next-generation model. It’s simply a matter of deciding: which technologies should the next-generation model incorporate? What kind of model do you expect it to be? Just as we’re now considering what the model after K2 should look like — this is something you continually think about and decide.

Each time, you look at what new things have been added to the toolbox and decide which ones to take out and use. That’s the process.

Zhang Xiaojun: Are your research team and training team separate? If you started researching these technologies a year ago, is it one team doing the whole thing when it comes to formal training?

Yang Zhilin: It’s one team. These things are hard to separate. You run into problems in actual training, and if you didn’t understand them beforehand, there’s no way to solve them.

Zhang Xiaojun: Did you encounter any challenges during K2’s development?

Yang Zhilin: When you train with Muon, it blows up.

We have some figures in the paper: your max logit climbs very high — into the hundreds or even higher. We believe this affects training stability. If you train for a long time, many of the so-called internal metrics become abnormal, which harms the model’s ceiling.

So we went back and revisited it, and fixed it. Because this is something you cannot predict in small-scale experiments — the explosion doesn’t happen at small scale.

Everything else was basically fine. We did many small-scale experiments, and the results transferred; the problems weren’t big. The only issue was this one, which you can’t validate at small scale and have to solve on the fly during scaling.

Zhang Xiaojun: K2 has become a hit recently. Has your mood swung at all?

Yang Zhilin: Not really, I’m fine — this is a long journey.

You have to keep building the next-generation model. It still comes back to those two sentences carved in stone: new problems will arise, and then you go solve them. That’s also the most interesting part.

Zhang Xiaojun: I heard that in an internal group chat, you described K2 as meaning Mount K2.

Yang Zhilin: K2 was already one of the hardest mountains in the world to climb; the names happen to coincide.

It’s not the destination, because it’s not the highest mountain — but it may be the hardest. That’s because there are many paradigm shifts right now: from conversation to agent, and the base model getting bigger, which is inherently difficult.

Zhang Xiaojun: Did K2’s release results exceed your expectations?

Yang Zhilin: About as expected. You already know how well the model trained while it’s in progress. There were no surprises, pleasant or otherwise.

Zhang Xiaojun: What were the most important pieces of know-how you gained from training K2?

Yang Zhilin: We wrote them all in the paper.

We’re very open — we still want to share more with the community.

06 From a “Brain in a Vat” to a System That Interacts with the World

Zhang Xiaojun: Since K2 is an agentic large language model, how would you define an agent, and how would you classify agents?

Yang Zhilin: It’s perhaps a transformation from a “brain in a vat” into something that can interact with the world, because the most important characteristic of an agent is that it can use tools over multiple turns.

There are two key points: one is multi-turn; the other is tools.

Multi-turn means you can act many times — it’s a form of test-time scaling. Tools are how this “brain” connects to the external world.

For example, with a search engine, you can connect the model to the entire internet; by writing code, you can connect the “brain” to the digital world, because almost all automation in the digital world can be described in code — it gives the model this capacity for automation.

These two are the characteristics of an agent in my mind, and there will be more and more tools going forward. Of course, tools follow a long-tail distribution. If the model generalizes well, it won’t just use common tools — it will be able to use highly personalized ones.

For example, the model can access a company’s internal databases, personal documents, or even custom APIs to complete business operations like refunding a ticket or placing an order. It should be able to generalize to tools it has never seen. I’ve always felt that what agents lack most is generalization ability.

If generalization is strong, the various vertical agents people discuss become less necessary. Because a general agent that generalizes to long-tail tools can solve many domain-specific problems simply by plugging into different tools. Give it a custom database, custom APIs, or custom document interfaces, and you have a very vertical agent. Its universality becomes much stronger.

Multi-turn is mainly about achieving test-time scaling and being able to do complex tasks. Unlike a conversational model that outputs one turn at a time, this can do different things — just like a human. A human’s daily work, you could say, is a sequence of using tools over many turns. You want to fit the human sequence into the model, but since you can’t collect such digital data, you can use reinforcement learning to construct it.

It is essentially simulating human behavior — though you can’t simply call it simulating humans; “simulating human behavior” isn’t quite accurate. It’s actually general.

Zhang Xiaojun: What do you mean by “you can’t simply say it’s simulating humans”? Humans are very general too.

Yang Zhilin: Right, humans are general. Humans are the so-called universal constructors.

But its main purpose is not to simulate humans; its main purpose is generality — that’s the design goal. Its resemblance to the way humans do things is merely a coincidental outcome, not the purpose of designing the system.

Zhang Xiaojun: This is like designing an airplane: the goal is to make a means of transportation, not to fly like a bird.

Yang Zhilin: When we build agent systems, it’s more about building a general intelligence, aligned with that goal; it just happens to be similar to humans.

Zhang Xiaojun: How do you improve an agent’s generality? Have you explored any methods?

Yang Zhilin: That’s a very hard problem. Today there’s a risk that agent generalization falls into overfitting to certain benchmarks, yet we also lack good benchmarks. That’s the challenge ahead.

But there may be some solutions. I still believe that using more AI to train AI can alleviate this problem to a certain extent.

Zhang Xiaojun: When will it be possible to use AI to train AI? What’s the bottleneck now?

Yang Zhilin: Part of it is already happening, but you want it to do more. A lot still depends on human design right now.

Zhang Xiaojun: That takes us to the next stage — the Innovator stage (L4).

Yang Zhilin: This is very interesting. You have to use Innovation-style approaches to solve Agent problems. Because agents don’t generalize well enough, you have to use innovation to solve it — using L4 techniques to solve L3 problems.

So the L1-to-L5 framework really may not be linear. Without good innovation — without using AI to train AI, or using AI to align AI — it’s hard for agents to achieve good generalization.

You manually define some tasks and only fit those tasks, but performance suffers on other tasks you can’t see. You pump up the scores on a few tasks, but users don’t feel it in the many more out-of-distribution (OOD) scenarios.

The field is currently in a stage where benchmarks are insufficient or failing, and agent generalization is a problem.

Zhang Xiaojun: Why are math and code relatively easy domains to generalize in?

Yang Zhilin: Actually, they’re not that easy either. If you do reinforcement learning, similar problems exist now.

Reinforcement learning itself generalizes better than SFT, because the process involves more on-policy samples — the model learns from its own samples — so generalization looks better, and there are negative gradients. These two factors lead to demonstrably better generalization.

But the generalization is limited. For example, if you reach 99 points on a certain type of math competition, other kinds of math problems might improve by five points, but it’s hard to reach 99 directly. Without doing the corresponding RL tasks, it’s hard to achieve that kind of generalization directly.

So doing math problems has similar issues — it’s constrained by the distribution. It’s still “you reap what you sow.”

But we hope that reinforcement learning, or post-training, uses more AI, so that the model can break out of the “you reap what you sow” situation.

Zhang Xiaojun: Is it possible that, in the end, you just can’t break out of it — that generalization can’t be substantially improved?

Yang Zhilin: I’ll go back to what I said earlier: problems are inevitable, but problems are soluble. You push forward every time — generalization will keep getting better. There’s not necessarily an end to it; there will always be better generalization.

Zhang Xiaojun: For agents, tasks and environments are extremely important. How do you define a good task? How do you define a good environment? Have you had any thoughts during your exploration?

Yang Zhilin: One approach is: given a model, you design some environments to reverse-fit that model. Of course, you can also design in the forward direction — suppose you’re a first-party developer, you design tools and environments up front, and let the model improve its capabilities within those environments.

The key is to make that design more general. It should be able to handle many tasks; you shouldn’t design tools and environments specifically for certain tasks. When the design is general enough, the model can learn within it, rather than you fitting the model in reverse. That’s probably the better approach.

Zhang Xiaojun: I’ve noticed one thing: people generally believe that in task design, you should lean toward designing sufficiently challenging tasks, because that spawns more fundamental new methods; but K2 was designed with medium-difficulty tasks. What was the consideration? Does that affect generality?

Yang Zhilin: It’s also a climbing process. You can’t have the model start by trying to prove a math problem no one has ever proved — the sample efficiency would be extremely low.

The better approach right now is: if reinforcement learning is paired with a good sampling strategy, it essentially becomes an implicit curriculum learning mechanism — you want the model to start learning at an appropriate difficulty and gradually ramp up, rather than learning very hard tasks from the outset. Otherwise, sample efficiency is low, you basically learn nothing, and the compute may all be wasted.

But the challenge is that many of today’s tasks are still based on existing human data or manually designed tasks; the AI-native portion is still relatively small, which brings generalization problems.

Zhang Xiaojun: In your eyes, what is the relationship between a coding agent and a general agent?

Yang Zhilin: A coding agent is a subset of tasks — but possibly a very important subset.

In the end, we want to do more than just coding. Even the model we’re training now isn’t meant to only do coding, because coding itself has some limitations.

Zhang Xiaojun: Can I put it this way — coding is like the human hand?

Relatively speaking, coding is an easier task for agents, right?

Yang Zhilin: It’s easier to verify, so it’s easier to learn. But it faces similar challenges — the generalization problem. Even coding agents encounter the same challenge.

A coding agent is a very important subset because it represents the automation of the digital world. Many agent tool sets today are fixed. If you want to create a new tool, you essentially write a piece — or a large chunk — of code to implement it. Or if you want better context management (context engineering), behind it there’s also a corresponding tool, and that tool may also be implemented in code. Code occupies a unique position and plays a unique role here.

But building a coding agent alone is not enough. Many non-programmers also use Claude Code to get things done — lawyers, product managers, designers. They use Claude Code because the model has a degree of generalization ability; it’s not just about writing code.

Zhang Xiaojun: What you want to build is a general agent, not a coding model?

Yang Zhilin: We still hope to build a general model.

Zhang Xiaojun: From writing code to manipulating the entire digital world, what capabilities do agents still lack?

Yang Zhilin: The use of high-frequency tools is still not good enough; there’s a lot of room in terms of capability. This also shows that we lack better benchmarks for observation. SWE-bench may saturate very soon. Many benchmarks are not good enough and don’t truly reflect actual user experience.

There’s room in the high-frequency tools themselves. And for long-tail tools — in situations you haven’t seen, completely OOD — how do you achieve better generalization? That’s also a very important problem that needs solving.

Zhang Xiaojun: For agents, do long context and long-term memory matter?

Yang Zhilin: Long context matters a lot too. Many tasks today simply can’t be solved with a 128K or 256K context; you need a million tokens or even more.

The challenge is that you not only have to handle such a long context, you also have to keep the “brain” sharp — the IQ has to be very high.

That’s a big challenge for model training. On the one hand, you want a high enough compression rate, which means the model has to be big enough; on the other hand, you want it to be long. There’s a natural tension between the two, so you need a better architecture.

But with some architectures you’ll find that they improve results at longer contexts while not necessarily helping at short contexts — sometimes they even hurt. That involves a balancing problem in architecture.

Still, these problems can be solved step by step. I think there are some solutions.

In addition, current RL training methods still have a lot of room for improvement. For example, when training complex multi-agent systems, using only end-to-end rewards is likely not enough. How should intermediate rewards be generated? Can we get rid of some of the manual design?

That’s also a direction well worth exploring.

Chapter Three: A System Both Simple and Complex

07 Open Source vs. Closed Source

Zhang Xiaojun: I went back and read our conversation from last year, and there’s a question I really want to ask you.

Last year you said that open source would lag behind closed source. Because open source works differently now — in the past, everyone could contribute to open source, whereas open-sourcing a foundation model today is essentially centralized, and community contributions haven’t been validated by compute. By contrast, the closed-source camp concentrates talent and capital; it is an integration of market resources.

You said at the time: “The leader won’t open-source; only the laggards will.”

But today, you open-sourced.

Yang Zhilin: Because right now, globally, we’re not fully in the lead yet (laughs).

Some of those judgments still hold in their broad direction: when you release a model, the community can contribute certain things. For example, you can do a lot on the inference side; you can let more people use the model for free.

But when it comes to contributing to the model itself — making the model stronger — currently only the original developer can do that.

Of course, if you look at the base model, that’s indeed the case. But doing extensive post-training on top of an open-source model — especially agentic post-training — may give rise to new opportunities.

Suppose you really want to build a law-related agent, and you’re a startup. You can absolutely train a specialized agent on top of K2, with your own specific set of tools, and it can perform extremely well in the scenarios you care about. That kind of opportunity exists.

It’s more about empowering downstream applications than feeding back into the improvement of the base model. Of course, this question needs to be observed dynamically.

Zhang Xiaojun: Will you choose open source for the long term?

Yang Zhilin: That’s what we hope to do for the long term, but it doesn’t have to be open source only. We want to share technical know-how with the community — that’s an important way to accelerate technological progress.

People don’t have to be purely competitors; there can also be collaboration. All the open-source companies could even form an ecosystem that better advances the technology — so the snow mountain can be climbed better. A race to the top.

But not everything has to be open-sourced either. For example, in collaborations with certain companies, not everything will be opened up.

Zhang Xiaojun: All in all, is open source a belief in a technical system, or a strategy of market maneuvering?

Yang Zhilin: Objectively speaking, it’s both, and both bring benefits. But ultimately, we hope that through this, the technology becomes safer and reaches a better level faster.

Zhang Xiaojun: How will the open- and closed-source ecosystems evolve? In your view, how many players will ultimately remain globally, open and closed combined?

Yang Zhilin: Not many — but there will still be a few. If you look at the past two years, the trend is fairly clear: the market is gradually becoming more concentrated, more convergent, more focused. Maybe it started with several hundred players, then went to several dozen, then to a few.

A few — that’s probably the final stable number. As it looks now, that’s highly likely.

Zhang Xiaojun: Which side do you belong to — open source or closed source?

Yang Zhilin: That has to be observed dynamically. We hope to share more technology over the long term.

Zhang Xiaojun: Why have most Chinese companies gone open source?

Yang Zhilin: Objectively speaking, there’s an element of market maneuvering. But it’s a good thing for the community.

08 Multimodality That Doesn’t Damage the “Brain” Is Already Good Enough

Zhang Xiaojun: How do you view products in the AI era? How is building an AI product different from building a mobile-internet product? — You used to love saying “the model is the product.”

Yang Zhilin: I can only speak about AI products; I’ve never built a mobile-internet product.

(The idea that the model is the product) hasn’t changed. When you build an agent product, you need to combine the model with tools and context. But you’ll find that when you train the model, you basically have to have this entire system built before you can train it.

Once the model is trained, the product is basically done. Making interaction improvements on top of that is certainly valuable, but that’s the icing on the cake.

Your model’s performance is polished during training and fits the tools and environments extremely well — in other words, the product is completed in the training process.

Zhang Xiaojun: Last year, you mentioned that the development paradigm has evolved into building an enormous system — like Google building its search engine system in the early 2000s. Today, do you have more imagination about the enormous systems of the AI era?

Yang Zhilin: The complexity of today’s systems lies in wanting to make the model general. On one hand, it has become simpler; on the other, more complex.

Simpler in that you only need to put everything into a single model — you don’t need to maintain so many models or devise a bunch of routing strategies. Conceptually, and in terms of engineering implementation, it has become simpler.

But at the same time, it has also become more complex. If you want it to be general, you want the model to work in all kinds of scenarios. For example, when you build an agent model, you don’t want it to work only with your own tool set — you want it to work when other people use the model with different tool sets, even tools you’ve never seen, or tools defined and implemented differently. That’s a very high bar.

Within agents today, there may be several different types of tasks — coding agents, search agents, and other kinds of agents. If you put them all into a single general model, conflicts may arise. Perhaps the tool definitions differ, or the data patterns differ.

In other words, the process of making it a general model comes with many technical challenges.

But if you don’t make it a general model, its generalization suffers, and it can only do one thing. Agents in particular need many steps to complete a task. Even a programmer doesn’t just write code; and even writing code isn’t just SWE-bench. To build something truly general and genuinely usable, the more steps there are, the higher the requirement for generality.

Its systemic complexity shows up in the training process: you have to make the model general enough, rather than fitting it to a few single-point capabilities. If you only fit single-point capabilities, your benchmark scores may look great, but the generality won’t be there — that’s the biggest challenge I can observe in this system today.

One example: if you want to add multimodal capability to a model, you need to make sure that multimodal capability doesn’t damage its “brain.”

Zhang Xiaojun: Multimodality can only aim for “not damaging”?

Yang Zhilin: Right. Not damaging it is already quite good.

You want the multimodal mode and the text mode to share one “brain.” You want it, in multimodal mode, to still bring out the IQ of the text part — rather than routing into a separate set of parameters, in which case it might completely lose what it learned from text.

When you build a general model, you face challenges like this. When you have all kinds of modalities, all kinds of task types, plus Agent, Reasoning, and Chat, fusing them all together is challenging.

And now it’s not just SFT — you also have to do RL, which makes the challenge even heavier.

General pre-training is relatively easy: you just put all the text together, and there basically aren’t many problems. But the further you get into post-training, and the more you get into RL, the more severe this problem becomes — that’s the systemic complexity of it.

09 When New Interactions Reduce the Noise in the Signals You Collect

Zhang Xiaojun: You see, the search engine system was built on top of the PC internet, and the recommendation engine system was built on top of the mobile phone — the mobile internet. In the AI era, where will the new super-nodes appear?

Yang Zhilin: They will run in many data centers — the “AI factory” that Jensen (NVIDIA founder and CEO Jensen Huang) often talks about. But there will still be more terminals. Some terminals will reuse today’s; some may be new.

Zhang Xiaojun: Will new forms of interaction be born?

Yang Zhilin: Definitely. Two years ago, Chat was a new form of interaction. Now with agents, there are many new forms of interaction — for example, you can have it execute a task asynchronously and watch the intermediate results.

Look at coding: first there was Copilot, then Cursor, then Claude Code — the interaction changed with each generation. Interaction changes as the model changes.

When you have a new generation of models with much-improved capabilities, you find the interaction can change. You no longer need to click “accept” on edits one by one; instead, you execute an agentic coding task over multiple steps.

Of course, Claude Code’s interaction today isn’t the final form either, because models will keep improving, and as capabilities improve, interaction will keep changing. For instance, if you have a multi-agent system, what will the interaction look like? It will probably keep shifting as the capability boundary moves.

Zhang Xiaojun: Has the scaling law slowed down today?

Yang Zhilin: The scaling law has hit the data wall — that’s an objective fact. To break through the data wall, you need to improve token efficiency. That’s why we’re working on improving token efficiency. The data wall exists; at the same time, you need to scale more compute into all kinds of RL tasks.

But from what we observe now, the speed at which models improve has not decreased — it’s even accelerating.

Zhang Xiaojun: Why haven’t AI products formed a data flywheel by now?

Yang Zhilin: Because compute-based scaling is too powerful.

For example, you first scale pre-training, then scale RL — and RL scales much more efficiently than pre-training. Because it’s on-policy training with negative gradients, its scaling efficiency is higher.

When your scaling efficiency is that high, directly scaling compute — scaling FLOPs — brings enormous gains, while the gains from other means are small by comparison. That’s one aspect.

On the other hand, the so-called data flywheel depends heavily on feedback from the external environment. We don’t want that feedback to contain a lot of noise. But this problem hasn’t been solved very well yet.

Foundation models are quite sensitive to noise when they learn; they’re different from traditional systems like recommender systems. A recommender system may not fear noise as much, but foundation models are sensitive to it.

Right now, FLOPs-based scaling looks like the more effective path. But when will this balance change? — It’s also possible that through new forms of interaction, the noise in the signals you collect can be reduced.

Zhang Xiaojun: That would require creating a new interaction paradigm.

Yang Zhilin: Right. But that interaction also has to fit the development of the model’s capabilities. Your interaction can’t outrun the model’s capabilities; you should design a good interaction within the scope of what the model can currently do.

It’s worth trying. It’s just that, as of today, scaling along the FLOPs dimension — or improving learning efficiency — is a more certain and more effective approach.

Zhang Xiaojun: If we go by what Yan Junjie (founder and CEO of MiniMax) says — that user data cannot improve a model’s intelligence — then is there no point in building consumer-facing (To C) products today? Should you just single-mindedly improve intelligence?

Yang Zhilin: It depends on how you look at it. You may not be able to train directly on user feedback. But the benefit of having a certain user base is that you know what the demand distribution looks like — you know where users have a good or bad experience — and you can abstract those things into evaluations and then optimize the model. If nobody uses your model at all, you don’t know which direction to optimize in.

You also have to consider the commercial value of users. We’ve now reached a new watershed: users can generate commercial value. Look at OpenAI — consumer users generate enormous commercial value, accounting for a relatively large share of its revenue.

Especially now that many agent products can generate value end to end, it also depends on who your users are. If it’s just chit-chat or checking the weather, the commercial value isn’t that big. But professional agent users have real productivity value in themselves.

Zhang Xiaojun: Over the past year, what new thinking do you have about consumer products?

Yang Zhilin: I think more about how to build the model, because once the model is trained well, the product is mostly done. We’ll keep going down this path.

Zhang Xiaojun: Some people say Kimi started out wanting to be “China’s OpenAI” — a characterization you never agreed with — and has shifted to wanting to be “China’s Anthropic.” Is there such a repositioning inside the company?

Yang Zhilin: It’s hard to define it that way. The contexts and the soil in China and the U.S. are different. Today we think about problems more from a global perspective. “Being China’s so-and-so” doesn’t really hold.

To put it simply: we want to keep climbing the mountain, be a friend of time, and accelerate the advance of technology together with the community.

10 Long-Context Architectures Affect “IQ”

Zhang Xiaojun: As a founder, what is the rhythm of your life like now?

Yang Zhilin: I probably sleep fairly late, ha — every day is different.

But it’s fine. I spend a lot of time on how to train the model better.

Zhang Xiaojun: Your time mainly goes into model training?

Yang Zhilin: More or less. But model training is an abstract concept; what matters is technical strategy, which is the most critical part of company strategy — what to do next and what not to do. Because the technical space is huge, you always have to pick certain directions to invest in heavily.

We placed our bets relatively early in many directions, and they paid off. We started doing long-CoT RL very early and reacted quickly; we worked on the optimizer; we did larger-scale pre-training; we built the first open agentic model. These are all key technical decisions. These decisions can determine fifty or sixty percent of the company’s direction.

But to make good decisions, you need a lot of evidence, and you have to run many experiments. You have to understand the specific experimental results very well; you can’t just decide on a whim — you need to know more.

Zhang Xiaojun: Among these decisions, which one did you wrestle with the most?

Yang Zhilin: Not particularly any of them. The key is that it’s a process of gathering data. Run experiments and see whether they’re solid. Then make the judgment combined with your understanding of the technology. In many cases, once the data is sufficient, the judgment is fairly obvious.

Next — at least for now — K2’s performance potential hasn’t been fully squeezed out yet. What we released was closer to a base model. We can add more FLOPs in the post-training stage, and the ceiling should be much higher than it is now.

We’ll also build the next-generation model, but how exactly to do it — we decide through experiments.

Zhang Xiaojun: You’ll add multimodality too, right?

Yang Zhilin: Multimodality is fairly certain.

But doing multimodal capability well is not easy in itself. There’s a lot of work in it: how to make it borrow the brain of the text side, rather than growing a separate brain of its own. For example, suppose your MoE (Mixture of Experts) has 20 experts dedicated to multimodality — you wouldn’t want that to happen. Because then the multimodality you train might be a “dumb multimodality.”

We want it to be a “smart multimodality.”

Zhang Xiaojun: What important technical milestones lie ahead?

Yang Zhilin: Agent generalization is the most important.

We’ll keep working on long-context support. In particular, having a longer context while maintaining very high IQ is also a very important problem.

Many long-context architectures today still affect “IQ.”

Zhang Xiaojun: Why do long-context architectures affect “IQ”?

Yang Zhilin: Pure linear attention may indeed affect IQ, because that architecture carries certain biases, and those biases don’t perform as well in some scenarios. But to a certain extent, this can be solved.

Zhang Xiaojun: What do you think of what Zhang Xiangyu (chief scientist of StepFun) said about the fundamental flaw of next-token prediction?

His point is that as models scale up, conversational ability, knowledge, and emotional intelligence all keep getting stronger, but reasoning ability — especially on math — rises first, then plateaus, and with further scaling actually declines. Using a bigger model for math problems makes it prone to skipping steps and being “dishonest.” He calls this a fundamental flaw of next-token prediction.

Yang Zhilin: That’s why you have to pair it with the scaling of reinforcement learning. Without reinforcement learning today, a model can hardly be called smart. It won’t necessarily do very well on math problems, for instance.

But a bigger base model gives reinforcement learning a higher ceiling. Because there’s more knowledge — RL essentially activates a reasoning paradigm that unlocks that knowledge. Its ceiling is higher, but it needs to be paired with RL to be activated.

Zhang Xiaojun: How do you view world models? — Some people say that building a world model is building the world, while building an agent is building the human.

Yang Zhilin: Using AI to train AI has something of that flavor. If you have a very good world model, you can simulate these things — that’s a way of using AI to train AI.

It may be a path toward better generalization.

11 Boundaries and Reality

Zhang Xiaojun: Now, let’s discuss some practical questions.

Between foundation-model companies and application companies building agent products, where does the boundary lie in the long run?

Yang Zhilin: I don’t have a definite answer. All I can say is that today, “first-party products” have one advantage: vertical integration. You put the model inside and train it; the model and the tools become one, rather than being built separately and then reverse-engineered.

But because the agent space is vast, “first-party products” can’t necessarily cover everything. If you can find certain spaces — for example, where implementing the tools requires a great deal of domain know-how, or where the evaluation is something a “first-party product” can’t fully cover — there are opportunities.

Because open-source models like K2 exist, everyone can fine-tune on top of them, making specialized agents and vertical agents much easier to create.

Zhang Xiaojun: That depends on how general the general model becomes.

Yang Zhilin: No matter how general it gets, there will always be some tools you have to build. You might not build the model at all, but instead build tools extremely well.

If a tool is made too general, it overlaps a lot with the “first-party product” — and in that case, the vertically integrated player has the bigger advantage.

But if the tool is specific to a certain scenario — even something others can’t build — for instance, if you control certain offline service entry points, and nobody else can build your ordering or transaction tools, then you may generate unique value.

There’s also another possibility: when the traffic and business model of “first-party products” or general agents become mature enough, many proprietary, previously monopolized tools will also be willing to plug in, because overall monetization efficiency will be higher. But improving monetization efficiency takes time. Within that window, proprietary agents will have room.

Ultimately, the reason “general” works is that overall monetization efficiency is higher. Today — including many content platforms — it’s possible that once you connect your content to a general agent, monetization efficiency will be higher than it is today. But that may take a very long time.

Zhang Xiaojun: Would a company like Manus be your potential customer or your competitor?

Yang Zhilin: It’s still very early; it’s hard to judge what it will look like, and the product itself will evolve.

In the short term, it’s more cooperation than competition. Today you can see K2 inside Cursor, Perplexity, and Genspark as well.

But the future will evolve — a bit like the relationship between Claude and Cursor. Cursor may also need to dynamically adjust its product strategy. On one hand, it may need a certain level of model capability, because the technology curve is still steep; on the other hand, can it have tools or environments that others can’t build?

There’s no direct answer right now. All I can say is that, for now, the advantage of first-party integration still exists.

Zhang Xiaojun: How do you think about business models today? Is API a good business?

Yang Zhilin: The business models that are clear at present: one is API services; the other is “first-party products.” We’ll try both. Today’s top priority is still making the model better — that remains the primary goal.

In the process of improving the model, if you lead in certain areas, there is indeed room for commercialization. The market is growing very fast today; the leading companies have billions or even tens of billions of dollars in ARR (annual recurring revenue), and may double or triple every one or two quarters. We’ll observe dynamically and make corresponding attempts.

Zhang Xiaojun: One of your users said he loves Kimi, but worries that Kimi can’t make money. Can you make money?

Yang Zhilin: The priority is still investing for now. Whether we can make money depends on how good the model is.

We’re also willing to serve the last mile of the user experience — delivering deliverables for users’ high-value problems.

In a global AI market worth tens of billions of dollars and growing at high speed, focusing on making the technology great actually gives you more certainty in everything else.

Chapter Four: In His Own Story

12 Manage with RL, Not SFT

Zhang Xiaojun: Over the past year, have you had any new thinking about organization?

Yang Zhilin: Good question. Lately I’ve been thinking about one thing — several things are connected.

You see, scientific research — or the process of creating new knowledge — is a lot like a reinforcement learning (RL) process. There used to be an empiricist theory that people acquire knowledge through experience. But many later views hold that this isn’t the case. Humanity existed on Earth for a very long time, yet until a few hundred years ago, no one said “the Earth is a sphere.” You are immersed in experience all along, but you don’t know what the facts are — experience does not directly give you knowledge. Instead, you propose a conjecture — “I believe this sphere is round” — and you think of all kinds of ways to verify it.

The same goes for training neural networks. You may observe that some internal metrics look off; you propose a conjecture about “why it behaves this way,” and you design experiments to verify it. This process is very much like reinforcement learning.

At the same time, you find that managing a team works the same way. This is what Tim (Zhou Xinyu, co-founder of Moonshot AI) tells me every day — manage with RL, not with SFT.

Of course, these days, when you do RL, you also want to add a bit of SFT, because SFT is a good prior that keeps the model from “flying too far.” But you also have to hold your own hand back — you can’t do too much SFT. Too much SFT, and team members lose their initiative and can’t innovate. I’m practicing this myself now, and it seems to have some effect.

The core is striking the balance between SFT and RL.

· SFT is you telling them “this is how this thing should be done, do it this way.”

· RL is you giving them a reward — if the outcome looks like this, it’s good; it’s more about the goal.

You probably want RL to be primary, with some SFT to provide priors as guardrails, or to prevent it from forgetting important things.

RL is something very fundamental. It connects scientific research, model training, and organizational management.

But this also brings a challenge: in the RL process, how do you define the reward? If you simply set a goal — say, “push up all the benchmark scores” — people will overfit the metrics by any means necessary. The scores go up, but the model itself hasn’t actually gotten better.

So the definition of the reward matters a great deal, and you need to understand very well how the specifics actually work. Otherwise you get reward hacking.

Zhang Xiaojun: So within an organization, how do you define the reward?

Yang Zhilin: You build more observational metrics and try as hard as possible not to overfit. That’s effective to a degree. Only then do you get better generalization and avoid being hacked.

The biggest problem with managing a team by RL is that you can easily get hacked. Everything looks great on the surface, but in reality it hasn’t achieved what you ultimately wanted — that’s the risk. The risk of managing a team by SFT is that everyone loses their creativity. In the end, these things have to be balanced to some degree.

Of course, I’m still learning. It’s far from perfect today.

13 AI Is an Amplifier of Human Civilization

Zhang Xiaojun: Over the past year, Kimi swung back and forth between peaks and troughs. Living inside that, what was your state of mind? Did you need to manage it?

Yang Zhilin: My state of mind is: be a friend of time.

Zhang Xiaojun: A real person’s state of mind can’t be that simple.

Yang Zhilin: As you said, there are highs and lows. What’s probably most important is that I still love doing this and want to do it well. So you don’t have to think about anything else — just think about how to do it well. It seems fairly simple, actually; nothing particularly complicated.

A lot of complexity is artificially imposed by people. It’s actually not that complicated.

Zhang Xiaojun: Have you gained a deeper understanding of human nature?

Yang Zhilin: That still needs time to temper. I wouldn’t claim anything that deep.

All I can say is that inside my own story, you continually feel out what kind of person you really are, and why you’re doing this — you keep thinking about these questions.

Zhang Xiaojun: What kind of person are you? Why are you doing this?

Yang Zhilin: I just find it interesting.

Zhang Xiaojun: What’s interesting? Is doing experiments interesting, is doing research interesting, or is doing AI interesting?

Yang Zhilin: The process of searching for truth. The process of continually discovering new problems and solving them.

Zhang Xiaojun: But you could solve other problems too. Why does it have to be this one?

Yang Zhilin: Because this thing matters. AI matters.

I asked Kimi this question too. It said this thing is “an amplifier of human civilization,” and I think that’s very well put.

Back to The Beginning of Infinity: from the Enlightenment to today, humanity has been searching for new ways to break through the boundaries of knowledge. But the next breakthrough of the boundary may come through AI — it’s an enormous lever.

Today, in any frontier discipline, you have to spend twenty or thirty years to learn the most cutting-edge knowledge. But AI can learn it overnight and go on to make new breakthroughs.

AI will become meta-science.

It is an amplifier of human civilization.

Zhang Xiaojun: Could it destroy human civilization?

Yang Zhilin: You can’t say that risk doesn’t exist, but there are many things we can do — whether it’s safer alignment or better social mechanisms.

For example, when AI can do certain things, it may well create new jobs, and we need ways to manage that transition.

I asked Kimi exactly this question. It said: although such a risk exists, we must not give up. Because giving up would mean giving up the ceiling of human civilization — you don’t know how high that ceiling could go. It’s a bit like “giving up eating for fear of choking.”

But I admit we have to do a lot to cope. Because many of the AI capabilities you see today are somewhat shocking — half a year ago, you simply couldn’t have imagined it could do these things.

At the same time, I believe humanity’s unique value will endure through this process. Human experience and emotion cannot be replaced by AI. So there may be different ways of living — hopefully, better ones.

Zhang Xiaojun: What different ways of living?

Yang Zhilin: I used to think a human life has a few sources of meaning: creation, experience, and love. Of course, everyone is different; that’s how it is for me.

A large part of “creation” may be doable by AI. I enjoy that process, but I have to admit that one day much creative work will be done by AI. But the latter two — “experience” and “love” — will remain human-centered.

Zhang Xiaojun: If AI takes away creation, it also takes away productivity.

Yang Zhilin: That’s fine. People can enjoy the fruits of production, if we have good mechanisms.

But it’s a slow process. It won’t be done in a year or two; it will take ten or twenty years of gradual adjustment.

Zhang Xiaojun: Do you chat with Kimi frequently?

Yang Zhilin: Of course. I have to test the model.

Zhang Xiaojun: Do you talk with it about profound topics, or about self-exploration?

Yang Zhilin: Sometimes. It’s fine — some of it is also work-related questions.

Zhang Xiaojun: Over this past year, have you had moments of really low spirits?

Yang Zhilin: I think it’s been fine. It’s mostly this: some things work, some things don’t; you solve some problems, and new problems arise — you’re constantly in that process.

As long as you find the thing interesting, you just want to keep doing it.

14 Any Intermediate State Can Become a Target of Criticism

Zhang Xiaojun: Over the past year, did you take any detours?

Yang Zhilin: Definitely. There are many, many decisions along the way — some technical, some business.

What’s very important is a company’s ability to gradually adjust through the process. The process of creating knowledge is like this too — it’s impossible that everything in the knowledge you create is right. You’ll find that some things are wrong; but they may be right for a period of time, and then wrong again after a while. And once they’re wrong, you have to adjust.

Take Newton: much of what he did was the best theory of its time, but it wasn’t perfect — in some scenarios it’s completely wrong. Universal gravitation needed other explanations; it needed relativistic explanations through the warping of spacetime.

I think the evolution of an organization and the development of a company are the same — it’s a dynamic process. At any intermediate point, you may be right at one moment and wrong at another.

This is also something Kimi told me — any intermediate state can become a target of criticism. You will always have the limitations of your era.

What matters more is how, through this process, you invest on the one hand in things that “don’t change” — talent, technical accumulation — and on the other hand adapt and adjust to environmental changes and feedback signals. Both are important.

Zhang Xiaojun: Internet products expand DAU and market size through marketing. AI products seem different — growth and customer acquisition depend more on big leaps in model capability. Between intelligence and promotion, which is more fundamental?

Yang Zhilin: It still depends on which of the two variables is bigger. In a stage of rapid technological development, it’s very hard to win the war through marketing. It’s more of an auxiliary means.

The question is what the ratio should be between this auxiliary means and your primary means. That needs dynamic adjustment, and it also depends on your current commercialization progress, or how strong your PMF really is. Different points in time call for different strategies.

And maybe in another year or two, you’ll find that some strategy is a good one again. I don’t think anything is set in stone.

We look at it with a more open mind, but at each point in time the most important thing is to grasp — which variable is the biggest one.

Zhang Xiaojun: Another year has passed. Do you think Kimi’s probability of success has increased, or its probability of failure?

Yang Zhilin: I think (the probability of success) has increased. Every step you climb, your probability of success grows, because some people stop climbing.

Zhang Xiaojun: Are you afraid of falling?

Yang Zhilin: There’s definitely fear.

But it’s more important to focus on the step you’re on right now — what can you do? That’s the question that matters more.

Zhang Xiaojun: I keep asking about your emotions, and you always say: ah, fine, fine. What were your most recent one or two moments of pure inner delight?

Yang Zhilin: “Not pleased by external gains, not saddened by personal losses.” Hard as that is to achieve, you have to avoid emotional decision-making.

Zhang Xiaojun: Do you get emotional?

Yang Zhilin: To some degree, definitely — you’re human, after all. But you have to avoid emotional decisions. In the end, when it comes to decision-making and execution, you need to be more rational.

Zhang Xiaojun: Over the past year, what was your biggest growth?

Yang Zhilin: Realizing this: problems are inevitable; they will always exist; continually solving new problems is the most important thing — and probably the most interesting one. That’s a change of mindset, and it changes the way you do many things.

Zhang Xiaojun: That sounds like a kind of mindfulness.

Yang Zhilin: I don’t know how to characterize it, but probably something like that. (Laughs)

Zhang Xiaojun: Let me finish with a few quick-fire questions.

One food you love, globally speaking.

Yang Zhilin: Ramen!

Zhang Xiaojun: Why?

Yang Zhilin: It’s delicious!

Zhang Xiaojun: One piece of knowledge few people know but everyone must know.

Yang Zhilin: I don’t think I’m very good at answering this kind of question.

Zhang Xiaojun: Based on all the books you’ve read, recommend a must-read.

Yang Zhilin: There’s one book I’ve been talking about this whole time — I’ll recommend that one.

Zhang Xiaojun: In your mind, which papers have shaped the course of AI?

Yang Zhilin: The most important papers are backpropagation, Transformer, and GPT-3.

Of course, some are building blocks and also very important — like ResNet (residual networks), which may be the foundation of optimization. And Adam. And now maybe Muon as well.

Zhang Xiaojun: Based on your current understanding, what is the single most critical bet?

Yang Zhilin: A generalizable agent.

Using Innovation — using L4 to do L3.

Zhang Xiaojun: This year, have you had any moments of epiphany?

Yang Zhilin: I don’t know. I feel like my brain has turned to mush.

I’ve already said a whole year’s worth of words.

A third person at the scene: This might just be long context affecting IQ.

Yang Zhilin: Can’t help it.

The limitations of carbon-based life. (Laughs)

(This translation was generated by K3)

想发布自己的文章?

升级为 Premium

上午10:40 · 2026年7月22日

[

5.2万

查看](https://x.com/zhang_benita/status/2079758352213762367/analytics)

8

41

279

234

相关

查看引用