Research

Research

How we build and why — methods, evaluations, and engineering notes in the open.

Mixed-precision compression of DeepSeek-V4-Flash

Mixed-precision compression of DeepSeek-V4-Flash

Assigning precision per tensor family compresses DeepSeek-V4-Flash-0731's 284B weights to 80.76 GiB with all 256 routed experts intact, running a 262,144-token physical context on a single 128 GB machine. This report gives the recipe, the two places the engine overruled it, the measured capacity, and the half that compression does not solve: the speculative decoding brought in to close the remaining speed gap is not lossless in this implementation. Same prompt, temperature 0, three verification depths, three different stable outputs.