We are evaluating AI wrong
Recently, I wanted to test personal AI assistants like Hermes, OpenClaw, and Grokbot. I wasn't particularly interested in finding out which model could score highest on another benchmark. I wanted to answer a much simpler question: How practical can these systems actually be? What purpose can an AI