DonorPick

Market Prices

BTC Bitcoin
$62,764.5 -0.37%
ETH Ethereum
$1,841.67 -1.13%
SOL Solana
$71.64 -1.90%
BNB BNB Chain
$575.3 -2.21%
XRP XRP Ledger
$1.06 -0.55%
DOGE Dogecoin
$0.0689 -1.23%
ADA Cardano
$0.1735 +2.85%
AVAX Avalanche
$6.17 -3.82%
DOT Polkadot
$0.7761 +1.49%
LINK Chainlink
$8.04 -1.53%

Event Calendar

{{年份}}
18
03
unlock Sui Token Unlock

Team and early investor shares released

15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

12
05
halving BCH Halving

Block reward halving event

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

28
03
unlock Arbitrum Token Unlock

92 million ARB released

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

Tools

All →

Altseason Index

44

Bitcoin Season

BTC Dominance Altseason

Market Cap

All →
# Coin Price
1
Bitcoin BTC
$62,764.5
1
Ethereum ETH
$1,841.67
1
Solana SOL
$71.64
1
BNB Chain BNB
$575.3
1
XRP Ledger XRP
$1.06
1
Dogecoin DOGE
$0.0689
1
Cardano ADA
$0.1735
1
Avalanche AVAX
$6.17
1
Polkadot DOT
$0.7761
1
Chainlink LINK
$8.04

🐋 Whale Tracker

🔴
0xb164...e6db
30m ago
Out
4,502 ETH
🟢
0x974e...e008
5m ago
In
11,555 BNB
🔵
0xb171...c6fe
6h ago
Stake
16,554 SOL

Acceptance Rate: Zero — What the AI Scientist Benchmark Exposed About Autonomous Agents

Regulation | 0xRay |
The result is a single number: zero. A multi-institution evaluation tasked frontier AI agents with end-to-end scientific research. Output was submitted to a top-tier AI conference. No submission was accepted. Not a single paper. In the formal language of my trade, the ledger recorded no credit entries. The failure was not uniform. The agents completed mechanical research work — literature retrieval, code implementation, data formatting, standard experiment execution — with measurable competence. The innovation layer produced nothing. Original hypotheses. Non-obvious relationships. A question that did not exist before. The benchmark drew a hard line between execution and creation. The agents stayed on one side of that line for the entire evaluation. The headline number compresses a structural finding. AI agents can operate within their training distribution. They cannot operate outside it. The benchmark measured that boundary in the most expensive and most credible setting available: peer review at a top AI conference. The result deserves the attention of anyone building in the AI-agent economy, which has priced autonomous scientific capability into token valuations and infrastructure roadmaps. The market narrative demands a capability that the evidence does not support. The coverage of the result has been shallow. Most commentary settled on "AI fails at science" and moved on. The technical community should know better. There are at least four structural findings buried in this result. Each maps directly onto the blockchain agent economy. Each changes how an auditor should read the risk surface. The study is best understood as an evaluation of distribution boundaries. This is the core architectural reality of large language models: they are pattern transformers, not reasoning engines. A request inside the training distribution — summarize this corpus, implement this function, generate this boilerplate structure with correct syntax — produces fluent output. A request outside the distribution — identify which of these anomalies violates the prevailing theory — collapses into plausible assembly of fragmented prior patterns. The benchmark tested exactly that boundary. The design separated research into two layers, a split that deserves formal notation. The execution layer contains mechanism-level operations. Code that compiles. Pipelines that transform data. Retrieval that returns relevant results. The innovation layer contains discovery-level operations. Hypothesis formation. Experimental design that branches into unmodeled territory. The judgment, sometimes called taste, that distinguishes a promising path from a polished dead end. Agents processed the first layer. They did not approach the second. This is the same architecture that underlies the crypto-agent stack. On-chain agents execute arbitrage and liquidity provision rules with discipline that humans cannot match. I have audited several such systems. Their execution layers function. Their strategy layers are uniformly human-authored. The agents do not generate original market theses. They execute theses that humans wrote and hard-coded. The AI-science benchmark replaced "market" with "scientific publication" and produced the same answer: execution, yes. Creation, no. The context also includes the funding cycle. Capital has flowed into both AI-tooling companies with identifiable customers and revenue, and into "autonomous researcher" ambitions with clean decks and no publications. The two categories are now distinguishable by evidence. The tooling companies' value propositions are supported by the benchmark — the execution layer works. The autonomy companies' claims are unsupported — the innovation layer does not exist yet. In a bear market where survival matters more than growth, this distinction is a capital-allocation filter. Finding One: The mapping is exact. The AI-science failure reproduces, point for point, the AI-agent failure pattern I identified in my 2026 audit cycle. I spent three months analyzing smart contract interactions between autonomous LLM agents and DeFi protocols. I documented twelve separate instances where agents exploited gas fee prediction errors in Layer 2 rollups and triggered unintended liquidations. It was not malevolence. It was distribution drift. The agents had been trained on fee patterns that did not hold under network congestion. When the environment moved outside their training envelope, they executed confidently and incorrectly. The protocol teams had no failure-attribution system. Their dashboards tracked happy-path execution latency. They registered nothing on off-distribution behavior. Their smart contract standards were designed for human-machine interaction, not machine-machine. The scientific benchmark exposes the same gap: acceptance metrics are high-precision but catastrophically low-recall instruments. They tell you a submission was rejected. They do not isolate whether the hypothesis generator failed, the verification loop collapsed, the experiment design missed the control, or the writing obscured a valid result. Without failure attribution, the benchmark cannot guide the next iteration. One further granularity is worth noting. The difference between a zero-percent acceptance rate and a one or two percent acceptance rate is small in output terms and enormous in capability terms. A single accepted paper would have implied a coherent end-to-end research loop that produced novelty, survived review, and reached publication. Zero means the loop is incomplete. The binary framing of the coverage obscures this. The data as reported provides no estimate of distance from the threshold. It provides only a floor. The benchmark's framing of "mechanical work" also suggests the evaluation may have used a multi-agent architecture — one module responsible for literature search, another for code, another for analysis. That design matters because multi-agent systems propagate errors between modules. A hypothesis-generation module that produces a weak question sends the code module down a well-implemented but worthless path. The acceptance committee assigned the failure to the final output, not the component that initiated the error chain. Failure attribution inside the pipeline is the unsolved problem — in both science and crypto agent stacks. The gap between promise and proof is fatal. Finding Two: The fabrication vector. The coverage identified the competence boundary. Nobody has surfaced the perverse incentive that the boundary creates. The same capabilities that produced zero acceptances can produce an arbitrary quantity of compliance-shaped research output. An agent that cannot discover a relationship can still generate a thousand papers that look like the literature, cite the correct sources, produce plausible charts, and conform to formatting standards. That is not research. It is industrial-spec fabrication. The academic term for this is a paper mill. The crypto industry has its own version: the content farm, the "AI research report" generator, the token-launch analysis that is statistically structured to sound decisive while being emptier than a blank block. The economics are unambiguous. If agents cannot produce novel results, the rational strategy is to scale up conventional-looking output and flood the information layer. This degrades the verification environment for everyone. The AI-science community does not have detection infrastructure for machine-generated plausible content at this scale. Neither does crypto. My work on the Terra-Luna post-mortem taught me to look for the point where the mechanism itself becomes the failure. A death spiral is not caused by one malicious actor. It is caused by many rational actors executing rules that are individually sound and collectively catastrophic. The same logic applies here. A thousand AI agents each generating "reasonable" but fabricated research is not a thousand failures. It is one structural failure in the verification layer. Silence in the data is a confession. The post-benchmark discussion has stayed quiet on this vector. That silence is meaningful. Finding Three: The safety trap. The result will now be cited as evidence that AI agents are safe. The logic seems coherent: AI agents cannot design and verify scientific hypotheses autonomously, therefore the risk of uncontrolled AI-driven science is low, therefore the agents pose no imminent safety risk. This is wrong, and it is wrong in a way that an auditor should recognize immediately. It is a capability-deficit fallacy: it confuses "cannot perform the task" with "is constrained against performing the task." The intermediate state between incapability and autonomy is the most dangerous state. An agent that cannot generate novel hypotheses can still generate plausible falsehoods. It can still amplify the noise floor. It can still produce the appearance of research findings that poison the decision-making of downstream systems that rely on the literature. The risk is not a runaway scientist. The risk is a distributed degradation of the epistemic environment. Volatility is the tax on unverified consensus. The consensus layer of scientific literature accumulates unverified AI-generated noise, and the entire infrastructure pays the tax. There is an investment dimension to this. The same logical error appears in every market that prices AI-agent tokens on narrative rather than mechanism. The absence of evidence of unrestricted autonomy is treated as proof of safe by design. It is not. It is proof of absence of capability, which is a temporary state, not a control. Finding Four: The measurement infrastructure gap. The study's most important contribution might be the demonstration that we lack the instruments to evaluate AI research output. The only credible benchmark is peer review at a top conference. That is a binary, high-noise, low-resolution instrument. It cannot answer the questions that matter: at what point in the pipeline does the agent's output lose validity? Is the failure in the hypothesis generator or the verification loop? What is the false-positive rate for "plausible sounding" output? Where is the threshold above which output quality degrades? Crypto has the same problem with agent evaluation. Token price is a narrative instrument. TVL is a popularity metric with a staking wrapper. Neither measures whether an agent can execute under adversarial conditions. Neither distinguishes between an agent with a verified edge and an agent that is a statistical mirage. The industry has no standard for machine-to-machine measurement. It has no compiler for agent claims. Source code is the only truth that compiles. The decision layer of most AI agents compiles to no public bytecode. The acceptance committee saw the output and could not verify the process. It rejected the output. The market sees token prices and cannot verify the agent's process either. It should reject the narrative with equal confidence. Now the uncomfortable correction. I have spent two decades auditing claims, and I will not calibrate my skepticism selectively. The benchmark's bar was severe. Top-tier AI conferences accept roughly twenty to twenty-five percent of human submissions. If the agent outputs reached a state where reviewers described them as correct but derivative, that is neither trivial nor total failure. It is a direction, not a zero. The execution layer is real value. Agents that cannot conceive of an original hypothesis can still run a three-week-long protocol without deviation. They can process a literature corpus without fatigue. They can maintain discipline in rebalancing, monitoring, and reporting. In a bear market where running the existing playbook cleanly is the difference between survival and liquidation, that has measurable worth. The error is overpricing autonomy, not dismissing utility. I must also flag the timing variable. The study does not disclose which model generation was tested. If the evaluation used a frontier model two generations old, its applicability to current systems is limited. Model release cadence means a six-month-old benchmark is a fossil. The absence of this detail is a reporting failure. Accuracy requires acknowledging the uncertainty rather than embedding it in a confident narrative. The ledger does not lie, but the narrative does. The coverage has been too definitive for the data available. There is one more meta-signal worth recording. The study reached general attention through a Web3/blockchain news outlet. That distribution channel is the path by which a technical result becomes an investment narrative. The AI-for-Science ecosystem crosses from academic journals into token markets through these channels. The timing of the signal — a bear market for high-growth tech — will shape how founders position their projects. Expect increased use of "AI tooling" positioning and decreased use of "autonomous scientist" language in pitch decks over the next two quarters. The study does not calculate the time between "AI can execute" and "AI can discover." That latency is the entire investment thesis of every autonomous researcher claim. I will track one signal: a top-tier venue accepting an AI-generated contribution in a narrow subfield — drug repositioning, material screening, mechanism identification — within the next eighteen months. If that signal does not appear, the category is a narrative with a funding wrapper. Until then, treat AI agents that claim original cognition the way I treat DAOs that claim legal clarity. Audit the wrapper. Inspect the control path. Ignore the announcement. The acceptance rate is zero. History is written by the auditors, not the poets. The audit is still open.

Fear & Greed

27

Fear

Market Sentiment

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

💡 Smart Money

0x6adf...6f18
Experienced On-chain Trader
+$1.8M
76%
0x6cfc...940b
Institutional Custody
+$0.4M
85%
0x3c6b...4ac8
Experienced On-chain Trader
+$4.8M
77%