Ai-research
evaluating-llms-harness
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic…
evaluating-llms-harness is an Agent Skill — a folder with a SKILL.md and any helper scripts that a coding agent loads on demand. Drop it in your skills directory and the agent picks it up the next time the task matches.
Everything on this page — the summary, the category and the how-to — is written by HowToPrompts; the link below goes to the original project.
| Category | Ai-research |
|---|
/skill add evaluating-llms-harnessHow to use it
- Copy the skill folder into your agent's skills directory, or run the add command above.
- Reload the agent so it re-scans available skills.
- Trigger it by describing the task — the agent loads the skill when it matches.
Compiled and written by HowToPrompts from public sources.
← Back to Skills