unbug@unbug9月12日MoE models use 30–50% less power and run quieter than dense models. Qwen3.8-27B vs Qwen3.8-Flash.105
unbug@unbug9月12日Holy wild, 16G RAM run DeepSeek v4.1 flash at 23s/token, it's so not real.FP4 Brain@thefp4brain9月11日got deepseek V4.1 flash running locally on a 16GB m1 mac mini original FP4/FP8 weights, ssd streaming + custom mlx runner 108s ttft and about 23s/token (not to be confused with tok/s)00:001116
FP4 Brain@thefp4brain9月11日got deepseek V4.1 flash running locally on a 16GB m1 mac mini original FP4/FP8 weights, ssd streaming + custom mlx runner 108s ttft and about 23s/token (not to be confused with tok/s)00:00