agent-leaderboard
This app lets you browse interactive leaderboards that compare performance across different categories and methods. Simply pick a category, methodology, and metric to see updated tables and charts....
This app lets you browse interactive leaderboards that compare performance across different categories and methods. Simply pick a category, methodology, and metric to see updated tables and charts....
Observability and devtools for AI agents.
Build, seed, and call twins of the services your software acts on via API, CLI, or MCP.
AI observability platform with open-source Phoenix.
Platform for evaluating and monitoring AI products.
CLIRank scores 420 APIs on agent-friendliness: auth method, CLI tools, pricing transparency, headless operation. Search by capability, compare alterna
Comet's open-source LLM evaluation and tracing platform.
Open-source LLM evaluation framework and platform.
Monitoring platform with LLM observability.
Build, orchestrate, and monitor data pipelines with dlt and dbt. Multi-tenant, multi-language, open-source.
AI agents for web testing
Platform for testing and improving AI products.
Evaluation and observability for AI agents.
Open-source observability for LLM applications.
Evaluation and observability for AI agents.
Platform for evaluating and improving LLM applications.
Integrated AI environment in the terminal. Build, test and instruct agents. - blob42/Instrukt
LLM monitoring and debugging platform.
KushoAI is AI-native infrastructure for software maintenance. Autonomous agents that live in your CI/CD and handle testing, healing, and monitoring co
Open-source LLM engineering and observability platform.
LangChain's platform for debugging, testing, and monitoring agents.
LLM observability and agent evaluation platform.
This web app lets you browse the Agent Memory Leaderboard, where textual and coding memory systems are ranked side‑by‑side. Choose academic or commercial entries, view overall scores, detailed capa...
This app shows interactive leaderboards that rank AI models on the Gaia2 benchmark, displaying scores like pass@1, search, execution, adaptability, and more. Users simply open the page and can refr...
Surface silent failures, pull context across traces, improve your agent before users churn.
Agentic Observability Platform for Mobile
Open-source observability and evaluation for LLM apps.
Evaluation and observability platform for AI agents.
MOVEdot puts AI agents on the manual load of vehicle development, so testing, simulation, and validation keep pace. Runs in your own cloud.
AI agents for end-to-end testing
NOFire connects your stack, builds a live production graph, and gates every agent action at runtime. Ship fast without leaving production exposed.
Browse and filter through detailed leaderboard results of Open Agent performance across various math and multi-modal benchmarks. Select evaluation dimensions, algorithms, datasets, and models to cu...
Orbit Sentinel monitors FCC, ITU, UNOOSA & FAA filings, updated daily. Track spectrum applications, launch permits & regulatory changes across agencie
Evaluation and security for AI agents.
The fastest path from submission to better decisions.
Matched from 47,000+ marketing agencies in 30 seconds. Ranked on verified review data - never on payment. Free, no commission, no paid placements.
Simulation environments for testing AI agents
AI gateway and observability platform for agents.
Prompt management and observability platform.
Open-source evaluation for RAG pipelines and agents.
Monitor your AI Agent the right way. Get alerted when your agent fails in production, trace exactly what went wrong, and prove your fix worked.
All-in-one operations platform for SaaS. Session replay, monitoring, support, code scanning, roadmap, and changelogs in one visual platform.
Find your research gap, match your journal, and get an 8-agent AI review in 15 minutes. 25 free credits to start — pay-as-you-go, no subscription.
Error monitoring with AI agent observability.
This web app shows the current ranking of students in the Agents Course Unit 4 Challenge, displaying each user’s name, score, and submission time. The list refreshes automatically every minute, and...
The @testdriverai in any GitHub repo and TestDriver writes UI tests and catches regressions before they merge. AI-powered end-to-end testing.
Open-source observability (OpenLLMetry) for LLM apps.
Open-source LLM app evaluation library.
Platform for building, evaluating, and deploying AI agents.
Scan SaaS and e-commerce pages for automated WCAG 2.1 AA issues, mapped to ADA, Section 508, or EN 301 549. Then plan manual review and fixes.
Cognee is the open-source agent memory platform for LLM agents. Build persistent memory across sessions with graph, vector, and relational retrieval that runs self-hosted, in Docker, on-prem, or on Co
ML platform with Weave for agent tracing and evals.