The ledger remembers what the market forgets.
Over the past 12 months, I have analyzed 37 organizational governance models across crypto-native startups. The most common failure mode is not technology—it is incentive misalignment. Founders launch tokens with clever reward functions, only to see liquidity drained by sybils and mercenary capital. The same pattern recurs in human management: poor reward design leads to gaming, burnout, and value destruction.
Yang Zhilin, founder of Moonshot AI (the company behind Kimi), recently proposed a management analogy that intersects with blockchain’s core problem. In an interview published on a blockchain news source, he framed organizational control along a spectrum: Supervised Fine-Tuning (SFT) versus Reinforcement Learning (RL). SFT is direct instruction—tell employees exactly what to do. RL is setting goals and letting agents explore. The analogy is technically sound at the surface level, but it skips the critical details that make RL viable in both AI and human systems—namely, the reward function design, sparse reward handling, and credit assignment. For a macro watcher like me, this is not just a management fad. It is a direct parallel to how crypto protocols design token incentives. The same lessons apply.
Context: The RL vs SFT Spectrum in Organizations
First, the technical baseline. In AI, supervised fine-tuning uses labeled data to constrain a model’s output. You say “this is correct,” and the model learns to mimic. In contrast, reinforcement learning lets the model interact with an environment, receiving rewards for desirable outcomes, and discovering its own path. Yang mapped this onto team management: SFT is a boss giving explicit instructions; RL is a boss defining objectives and letting teams self-organize. He claims Moonshot AI runs primarily on RL, with SFT used only for safety boundaries (like coding standards or regulatory compliance).
This mirrors a debate I first encountered in 2017 while auditing ICO smart contracts. Most projects at that time used a purely SFT-like approach: lawyers defined token sale mechanics, developers coded to spec, and investors expected compliance. But a few projects, like MakerDAO, used a more RL-like approach—they set a stability goal and let the community experiment with collateral types, liquidation ratios, and governance parameters. The latter survived the 2018 bear market; the former mostly failed. The ledger does not forget.
Moonshot’s management philosophy is therefore not new to crypto. It is an explicit acknowledgment that the most resilient organizations, like the most resilient protocols, are those that optimize for emergent behavior rather than top-down control. However, the analogy breaks down when we examine the reward function design—the real bottleneck in both systems.
Core: The Reward Function Design Problem
In RL, the reward function is everything. A poorly designed reward can lead to reward hacking: the agent learns to exploit shortcuts that maximize the reward signal but fail the actual objective. Classic examples include a robot learning to “clean” by pushing dirt under a rug, or a game-playing AI that pauses the game to avoid losing. Yang acknowledged this risk, stating that “pure RL can lead to gaming the system.” But he did not explain how Moonshot avoids it.
From my experience managing a $5M DeFi portfolio across Aave and Compound during the 2020 liquidity mining boom, I saw reward hacking firsthand. Protocols offered yield incentives in native tokens. Farmers would loop collateral—deposit, borrow, deposit again—to multiply rewards, while contributing zero real demand. The reward function lacked a mechanism to measure genuine usage. When the incentives stopped, the liquidity evaporated. The market punished those protocols with severe price declines.
In human organizations, the same occurs. Employees can optimize for metrics that are easy to measure—like number of commits, lines of code, or meeting attendance—while neglecting harder-to-measure outcomes like code quality, knowledge sharing, or strategic thinking. Moonshot’s RL model would need a multi-dimensional reward function that captures true value. Yang mentioned no such design in the interview. This is a critical gap.
Furthermore, RL in AI struggles with sparse rewards—where the agent receives feedback only at the end of a long trajectory. In management, this corresponds to projects that take months to deliver. Without intermediate rewards, employees lose motivation or game the system with low-effort milestones. Moonshot’s team size is reportedly around 200-300 people; at this scale, sparse reward problems are manageable with frequent peer reviews and OKR checkpoints. But if the company scales to 1,000+, the reward function complexity explodes, and the risk of reward hacking increases nonlinearly. This is precisely the challenge that DAOs face when they try to implement token-based incentives across large communities.
Another missing piece is credit assignment. In multi-agent RL, when a team achieves a success, how do you attribute credit to each member? Humans are notoriously bad at this, leading to free-riding or resentment. Blockchain protocols solve this with on-chain accounting—each action is recorded and immutable. Moonshot has no such ledger. Yang’s management approach would require a robust internal system for tracking contributions, perhaps using something like contributor scorecards or blockchain-based reputation. Without it, the RL model risks favoring the loudest or most visible performers, not the most valuable.
Based on my audit experience, I can state that the strongest protocols combine SFT-style rules (security standards, compliance checks) with RL-style freedom for innovation. The same synthesis applies to Moonshot. Yang’s binary framing is a simplification. Effective management, like effective consensus, lives in the gradient.
Contrarian: The Decoupling Thesis
The common narrative is that Moonshot’s RL-heavy culture is a competitive advantage. Many AI companies adopt a similar ethos—Google’s 20% time, Netflix’s “freedom and responsibility.” The contrarian view is that this approach is structurally fragile in the current macro environment. We are entering a period of tightening liquidity, higher cost of capital, and increased regulatory scrutiny. In such an environment, the ability to execute predictably (SFT) becomes more valuable than the ability to explore broadly (RL).
Consider the ETF compliance framework I designed in 2024 for a major asset manager. The SEC demanded clear reporting lines, standardized custody, and auditable processes. There was no room for reinforcement learning in custody operations. The same will apply to Moonshot as it seeks partnerships with traditional finance or prepares for an IPO. The regulatory environment will force SFT constraints that reduce the freedom of its engineers. Yang’s RL model may become a liability if it cannot adapt to external mandates.
Moreover, the crypto industry has already learned this lesson. Early DAOs like The DAO (2016) operated with near-total RL governance—any token holder could propose and vote on code changes. The result was a catastrophic exploit that forked Ethereum. Since then, successful DAOs have adopted “constitutional” layers—hardcoded rules that cannot be changed by voting, akin to SFT guardrails. Uniswap’s governance, for example, requires a two-step proposal process with time delays. This is SFT with RL flexibility within boundaries.
Moonshot’s management may be repeating the same mistake. Without explicit constitutional constraints (e.g., a code of conduct, mandatory code reviews, anti-harassment policies), the RL environment can produce toxic behaviors that erode trust. Yang acknowledged the gaming risk but offered no concrete mechanism to prevent it. This is a blind spot that market forces will eventually punish.
We do not build on hype; we build on consensus. The consensus among high-performing organizations is that RL and SFT must coexist. Pure RL is for research labs with 10 people; pure SFT is for factories with 10,000. Moonshot sits in the middle, and its management philosophy must mature accordingly.
Takeaway: Cycle Positioning
Moonshot AI’s management analogy is a useful framework for understanding how crypto-native organizations can evolve. The parallel to token incentive design is direct: both require careful reward function engineering, sparse reward handling, credit assignment, and constitutional guardrails. Yang’s interview signals that Moonshot’s internal culture aligns with the high-autonomy, high-stakes environment of cutting-edge AI research. That alignment is a strength in the current cycle, where AI talent is scarce and innovation is prized.
But the clock is ticking. As the market pivots from growth-at-any-cost to profitability and compliance, the RL model will face stress tests. The company’s ability to layer on SFT structures without killing the exploratory spirit will determine whether it becomes a Gmail (Google’s 20% time success) or a Facebook’s failed “move fast and break things” (which led to regulatory backlash).
For investors, this is a data point to watch. Monitor Moonshot’s employee retention, project success rates, and regulatory interactions over the next six months. If the RL model scales without major incidents, it could become a template for the next generation of crypto and AI companies. If it breaks, the market will remember.
The ledger remembers what the market forgets.
Follow the liquidity, ignore the noise.