FreeToken - first look and test

FreeToken is a new project that is making some bold claims about running large or very large models on constrained hardware. Let's take a look at it, and see what we can do with only 12GB of VRAM at our disposal! FreeToken: https://github.com/FlashML-org/FreeToken Chapters: 00:00 Intro 00:21 llama-server baseline 03:34 FreeToken's turn 09:56 OpenCode 11:19 Verdict 12:05 Caveats 14:21 Outro Join our discord!   / discord   Become a certified localhoster!    / @noplacelikelocalhost   Or, support the channel via Patreon:   / noplacelikelocalhost   Thanks for watching!