treeru.com

TreeRU Tech Blog

Insights and hands-on guides on frontend, backend, databases, IT infrastructure, and AI solutions.

posts.tsx — 28 posts
Load More18 remaining
  1. Local LLM Benchmark: 6 Models Tested Across 60 Questions and 7 Business Scenarios
  2. Building a Local RAG Pipeline — From Embedding Model Selection to Hallucination Elimination
  3. MoE vs Dense: Why Qwen3-30B-A3B Is Slower Than 14B — And Why Hybrid Inference Failed
  4. Qwen3-32B vs 14B — Is 2x Slower Speed Worth the Quality Gain?
  5. Text2SQL Real-World Test — When LLM Writes SQL Directly
  6. AWQ Quantization Speed Benchmark: 16 Models, INT4 vs BF16, and the MoE Reversal
  7. SGLang 23-Model Serving Guide — Optimal Configuration for Every Model
  8. Qwen3-14B Deep Review — Why It Is Our Top-Ranked Local LLM
  9. Cross-Server AI Inference — Boosting Throughput 70% with a $450 Secondary GPU
  10. Serving 7 Companies on 1 GPU — Multi-Tenant Isolation Testing in Practice
  11. LoRA Fine-Tuning for Custom AI Chatbots — From 10 Training Pairs to Multi-Tenant Serving
  12. SGLang vs vLLM — The Secret Behind the 3x Throughput Gap
  13. Local LLM Concurrent User Load Test — How Many Users Can an RTX PRO 6000 Handle?
  14. Local LLM Business Test (Part 2) — Shopping, Legal & Automation
  15. Local LLM Business Test (Part 1) — Manufacturing, SaaS, Healthcare
  16. LLM Hallucination Test — Which Local Models Fabricate Information?
  17. Local LLM Korean Language Comparison — 6 Models Tested with 10 Real Questions
  18. 8B vs 14B vs 32B LLM: Concurrent User Benchmark on a Single GPU