KubesimplifyKubesimplify
ProductsLearnBlogWorkshopsPartnershipsAboutNewsletter
← All posts

Topic

quantization

2 articles

Cover for The Local LLM Glossary: Every Term, Flag, and Number in Plain English
local-aillmAug 18, 2026

The Local LLM Glossary: Every Term, Flag, and Number in Plain English

Plain-English definitions for every term you hit in local LLM posts: prefill and decode, tokens per second, FP8 and NVFP4, Q4_K_M, KV cache, YaRN, Gated DeltaNet, speculative decoding, and every vLLM, llama.cpp, and Ollama flag worth knowing.

Saiyam Pathak
Saiyam Pathak · 19 min
Read →
Cover for Day 4: Quantization Demystified. BF16, FP8, NVFP4, MXFP4, INT4, GGUF, and Why It All Matters
nvidiadgxsparkJun 10, 2026

Day 4: Quantization Demystified. BF16, FP8, NVFP4, MXFP4, INT4, GGUF, and Why It All Matters

A practical, beginner-friendly guide to BF16, FP8, NVFP4, MXFP4, INT4, and GGUF Q4_K_M on NVIDIA DGX Spark. Bytes per parameter, quality vs size, and which format to pick when.

Saiyam Pathak
Saiyam Pathak · 28 min
Read →

Help Us Do More

All funds go toward providing free cloud native & AI education to everyone.

Kubesimplify© 2026 Kubesimplify
AboutBlogWatch & LearnResourcesContact