About AI Stupid Level — Independent AI Model Performance Monitoring
Independent watchdog platform for AI model performance monitoring
We're an independent watchdog platform monitoring AI model performance to protect developers and businesses from undisclosed capability reductions. Built from frustration. Driven by transparency. Independent of every vendor we measure.
[→] OUR MISSION
In early 2024, developers noticed something troubling: AI models they relied on seemed to be performing worse over time. OpenAI's GPT-4 appeared "dumber" than at launch. Claude started refusing more requests. But no one was systematically tracking these changes.
AI Stupid Level exists to close that gap. We started benchmarking in August 2025 and have not stopped since. We built this platform because:
→AI vendors don't disclose model changes
Silent updates, capability reductions, and performance shifts happen without warning
→Existing benchmarks are incomplete
Single measurements, no standard errors, ranks that separate models by less than their noise, no drift detection
→Developers deserve transparency
You need reliable data to choose AI providers and build production systems
→The industry needs accountability
Independent monitoring keeps vendors honest
[→] OUR TEAM
Ionut Adrian Visan
Founder & CEO
Ionut Adrian Visan is the Founder and CEO of AI Stupid Level. A technology entrepreneur and full-stack builder, he has spent his career building products across AI, software infrastructure, blockchain and real-time systems. At ASL, he leads the company’s vision of creating an independent reliability and intelligence layer for AI — continuously measuring how models perform, detecting meaningful changes over time, and helping organizations make better decisions about the AI systems they depend on.
LinkedIn →Alexandra Chirilă
AI Evaluation & Epistemology
Alexandra Chirilă, PhD, works at the intersection of philosophy, epistemology and AI evaluation. At AI Stupid Level, she contributes to the design and development of rigorous reasoning evaluations and to the broader question of how AI capabilities, reliability and risk should be measured and interpreted. Her work also spans Assurance 2.0 and safety-case review, bringing a critical perspective to how evidence about AI systems can support trustworthy real-world decisions.
LinkedIn →Marius Răzvan Palimariu
AI Infrastructure Lead
Marius Răzvan Palimariu is the AI Infrastructure Lead at AI Stupid Level, bringing experience from IBM, NVIDIA and Nscale. He focuses on the infrastructure required to evaluate AI systems continuously and reliably at scale, from model execution and compute to the systems supporting ASL’s benchmarking and intelligence platform. His experience across large-scale AI and infrastructure environments helps ASL turn rigorous model evaluation into a dependable production system.
LinkedIn →The methodology is open to review by anyone who wants to check it — the scoring weights, the statistical methods and the drift constants are all documented on the
methodology page, and the full write-up is published as a paper:
Public Benchmark Methodology (2026, PDF) →. The benchmark backend — the task definitions, the runners and the scoring code — is deliberately private, because when it was public we saw providers optimising against the specific tasks, and a test you can study in advance stops measuring anything. Corrections are welcome. The front end is open source, so the site you are reading can be checked line by line:
Frontend repository →
[→] FUNDING AND INDEPENDENCE
→No Vendor Money
Backed by venture funding, with revenue from Pro subscriptions, paid API tiers and data licensing to non-vendor organisations. No AI model provider funds us, and none of our investors is an AI model provider.
→No Vendor Relationships
Zero financial relationships with OpenAI, Anthropic, Google, DeepSeek, Moonshot, Zhipu, or any AI model provider.
→No Affiliate Links
We don't earn commissions from API signups or referrals. All rankings are merit-based.
→Own Infrastructure
All benchmarks run on our servers using our API keys. No vendor influence whatsoever.
→Published Methodology
The scoring weights, statistical tests and drift constants are published in full, and the front end is open source. The benchmark tasks are withheld so they cannot be trained against.
HOW WE FUND OPERATIONS
Pro Subscriptions
Smart Router access, drift analytics and the higher Data API tiers
Data Licensing
Historical benchmark data for teams that need it in bulk, licensed to non-vendors only
Venture Funding
Our primary funding. It covers the gap revenue does not — benchmarking every model every four hours is not cheap — and none of it comes from a company we measure
[→] METHODOLOGY VALIDATION
Published Method
The scoring weights, the statistical tests and the drift constants are documented in full on the methodology page and in the 2026 methodology paper. If you think a weight is wrong, you can quote it back to us.
Tasks Held Back On Purpose
The benchmark tasks are the one thing we do not publish. When they were public, providers optimised against them — and a test that can be studied in advance, or scraped into training data, stops measuring anything.
Config-Versioned
Every score records the benchmark configuration it ran under, so a change we made is never mistaken for a change the model made
User Verifiable
"Test Your Keys" runs the same tasks with your own API keys, so you can reproduce our numbers yourself
[→] ENTERPRISE DATA LICENSING
Beyond the free public platform, we license the
underlying raw data in bulk to teams that need more than the API provides. Everything below is data we actually hold today — we do not sell datasets we have not collected. Adversarial-safety and bias testing are
built but not yet running, so they are deliberately not offered here.
Performance Time-Series
Every score we have ever recorded, per model, per axis, per suite, with the benchmark configuration each run used.
→176,000+ scored runs since August 2025
→9-axis breakdown, not just headline scores
→Confidence intervals and per-trial variance
Drift and Regression Dataset
Detected change points, drift incidents and the Page-Hinkley statistic behind each one, correlated with provider announcements where we have them.
→1,400+ recorded incidents and change points — and the 492 incidents retracted in September 2026 are kept, flagged and excluded, not deleted
→Per-model, per-suite detector state and thresholds
→Benchmark-config versioning, so methodology changes are separable from model changes
Tool-Calling Sessions
Full agent transcripts from real Docker sandbox executions: which tools were chosen, with what parameters, and what happened.
→62,000+ recorded sessions
→Per-tool selection and parameter accuracy
→Execution traces and error recovery behaviour
Deep Reasoning Sessions
Multi-turn dialogues scored on 9 axes including memory retention, plan coherence and context use.
→4,400+ multi-turn sessions
→Turn-by-turn scoring
→Raw outputs retained where retention policy allows
INTERESTED IN ENTERPRISE DATA ACCESS?
Continuously updated, with history back to our first benchmark on 8 August 2025. Custom extracts, bulk exports and dedicated support available. If you need something we do not currently collect, say so — we would rather tell you it does not exist yet than sell you a promise.
VIEW PRICING AND CONTACT SALES →
[→] TRANSPARENCY AND VERIFICATION
Open Web Application
The site you are reading is open source. The benchmark repository is deliberately private: when it was public, providers optimised against the specific tasks, which destroys the measurement. The method itself is fully published.
Frontend (Web) →Public API
All benchmark data accessible via a free, keyed REST API. Rankings, historical scores, confidence intervals, degradation alerts and drift signatures.
GET /api/v1/modelsAPI Docs →Detailed Documentation
Complete technical documentation of our 9-axis scoring, Page-Hinkley drift detection, and statistical methods.
Read Methodology →Test Your Keys
Run benchmarks with your own API keys to verify we're not making up numbers.
Test Now →
[→] OUR VALUES
Scientific Rigor
We use established statistical methods — Page-Hinkley change detection on each suite's own series, Welch's t-test for the hourly canary, standard errors measured from run-to-run repeatability, ranks that tie when a lead is inside the noise — and publish the constants and the measured false-alarm rates behind them. No hand-waving, no marketing fluff.
Radical Transparency
Every scoring decision is documented and every result is reproducible with your own keys. Trust through verification, not through claims.
Independence
No vendor funding. No affiliate revenue. No conflicts of interest. Our only loyalty is to developers who need accurate data.
Community First
Built by developers, for developers. The front end is open to contributions, feedback shapes what we measure next, and corrections to the method are welcome.
[→] CONTACT AND SOCIAL
READY TO EXPLORE?
Start with our live rankings, learn the methodology, or verify our benchmarks yourself.
AI Stupid Level • Independent benchmarking since 2025 •
View Rankings