
Last time I wrote about GLM-4.7-Flash, I was annoyed. It ran at a third of the speed I expected on my mini PC, fell off a cliff as context grew, and the fix turned out to be a runtime that actually implements Multi-head Latent Attention…
View original source — Hacker Noon ↗



