Loading blogs...
Posts tagged Performance.
Workstation WSL Proxy: compile OpenResty once into bwalia/wslproxy-openresty, warm Buildx/GHA layers, extract on bare metal, prefer DEPLOY_MODE=code — and why AI-assisted delivery made this cross-cutting work finishable.
Workstation on turbocharging LLMs: PagedAttention paging for KV cache, vLLM serving, Self-Debugging for agents, PowerInfer token rates, and EG-MLA memory cuts — with production trade-offs.