Local AI, benchmarked in the garage.

Tensor Garage runs open-weight models on real hardware — consumer GPUs, Apple Silicon, mini PCs — and reports what actually happens: tokens per second, memory at real context lengths, what fits, what breaks, and what it costs. Every number is sourced and linked in the description.

New video every two days: quantization explained with measurements, buying guides per memory tier, runtime comparisons (llama.cpp, MLX, vLLM, Transformers), and the model releases that matter for local users.