Insights and hands-on guides on frontend, backend, databases, IT infrastructure, and AI solutions.
Running an LLM on CPU Only - Real Numbers from an Old Xeon with 256GB RAM
Our Servers Catch Burglars Every Day - TREEWALL Field Test Report
NVMe Storage Benchmark: Optane to HDD - 20 Devices Compared with fio
Local AI Video Generation with ComfyUI + Wan 2.2 - An RTX PRO 6000 96GB Build Log
Structuring Unstructured Documents with an LLM - Measured Accuracy on a 19-Item Menu
A spare server generously stuffed with RAM but no GPU? We measured what CPU-only LLM inference really delivers: 5 models from 120B to 4B, thread scaling, power and thermals.
NVIDIA driver 590→595 upgrade benchmarks on RTX PRO 6000 Blackwell — SGLang inference speed, CUDA compute, P-State changes, and operational warnings.
Qwen3-32B is 2x slower than 14B. We tested both with 60 identical questions to measure exactly how much quality you gain for the speed trade-off.
AWQ INT4 delivers 1.88x at 4B scaling to 2.94x at 32B. MoE Qwen3-30B-A3B runs faster than Dense 14B-AWQ. 16 models benchmarked with full data on RTX PRO 6000.
vLLM looks faster at median latency. But at P95, SGLang wins by 6.5x. With 3x throughput and zero errors at 200 users, we explain why P95 matters more than median for production LLM serving.
Comparing 6 local LLMs on Korean fluency — honorifics, business tone, language contamination (Chinese/English mixing), and natural expression across 10 practical questions.
Guarding the Front Door of Your Network - TREEWALL, an AI Security Gateway
Build a CCTV Recording Server Without an NVR - RTSP + ffmpeg + systemd, 180-Day Retention
Our MoE Model Started Spewing Garbage - What Auditing All 18,432 Expert Weights Revealed
What We Learned Building and Auditing 990 AI-Generated English Learning Scenes
BGE-M3 vs mE5-Large - Comparing Korean RAG Search Quality Head to Head