On the test bench
Reviews measured against real tasks, not leaderboards. First up:
-
Next
Harness comparisons - how different agent runtimes handle tools, context, and long sessions.
-
Queued
The hardware each local setup actually needs, from Mac to DGX Spark.
-
Elsewhere
Qwen models get dedicated coverage in the writing archive.