ollama
๐๐ฅ๐จ๐ฎ๐ ๐๐๐๐ฌ ๐๐ซ๐ ๐๐ฑ๐ฉ๐๐ง๐ฌ๐ข๐ฏ๐. ๐๐๐ญ๐๐ซ๐ฆ๐๐ฅ๐จ๐ง ๐ข๐ฌ ๐ง๐จ๐ญ. ๐
Vision models grow in their capabilities while the execution demands way fewer resources.
Means: my sales slip scanner-project from two years ago needed an upgrade. Before I was using GPT-4 Vision via the @OpenAI API, which costs you control over your data and a few cents for every analysis.
Now, after intensive benchmarking across a range of 15 models (with ๐๐ฅ๐ฅ๐๐ฆ๐) and some surprising regressions (newer harness doesn’t mean the models execute faster; it can also mean they suddenly run into model load errors…), I put my money on ๐๐ฐ๐๐ง๐.๐:๐๐. A casual language model with a ๐ฏ๐ข๐ฌ๐ข๐จ๐ง ๐๐ง๐๐จ๐๐๐ซ. My own benchmark over the selected models against the huge sample size of three receipts: all three were successfully evaluated. ๐
Then I built a ๐ก๐จ๐ญ-๐๐จ๐ฅ๐๐๐ซ-๐๐จ๐ง๐๐๐ฉ๐ญ around the script, and it works (thank you @Mr Zahorsky โ without whom I would have never come across that idea). For instance: in less than ten seconds, 13 samples were processed, and you get the result sum plus a fancy HTML overview for comparison.
๐๐๐ญ๐ ๐๐ง๐ญ๐ซ๐ฒ can be done for cheap: give the kids the smartphone, snap all receipts, throw out some watermelon slices ๐๐
If you want to run the tool as well, check my GitHub: https://github.com/marcelpetrick/sales-slip-scanner-ng. As said: a local GPU is the only thing you need; ๐ง๐จ ๐๐ฅ๐จ๐ฎ๐, ๐ง๐จ ๐๐๐ ๐ค๐๐ฒ๐ฌ, ๐ง๐จ ๐๐ซ๐๐๐ข๐ญ ๐๐๐ซ๐. Everything is now self-hosted. You’ll also find the benchmark of the vision models there, including the results.
๐๐ก๐ ๐๐ข๐ ๐ช๐ฎ๐๐ฌ๐ญ๐ข๐จ๐ง ๐ข๐ฌ: ๐ฐ๐ก๐๐ญ ๐ฌ๐ก๐จ๐ฎ๐ฅ๐ ๐ ๐๐ฎ๐ญ๐จ๐ฆ๐๐ญ๐ ๐ง๐๐ฑ๐ญ ๐ฐ๐ข๐ญ๐ก ๐ ๐ฅ๐จ๐๐๐ฅ ๐ฏ๐ข๐ฌ๐ข๐จ๐ง ๐ฆ๐จ๐๐๐ฅ? ๐๐ก๐๐ญ ๐ฐ๐จ๐ฎ๐ฅ๐ ๐ฒ๐จ๐ฎ ๐๐จ?
๐ ๐ฅ๐ฎ๐ป ๐ฎ ๐ฎ๐ณ๐ ๐ ๐ผ๐ฑ๐ฒ๐น ๐ผ๐ป ๐ฎ๐ป ๐ด๐๐ ๐๐ฃ๐จ.
๐ฐ๐ฌ ๐บ๐ถ๐ป๐๐๐ฒ๐. That’s all it took me to get Bonsai 27B running locally on an RTX A2000 Laptop GPU with just 8GB of VRAM.
Bonsai 27B is based on Qwen 3.6, but uses PrismML’s custom native ๐-๐๐ข๐ญ ๐๐จ๐ง๐ฌ๐๐ข format, reducing the model to just 3.9GB. A specialized ๐ฅ๐ฅ๐๐ฆ๐.๐๐ฉ๐ฉ fork implements custom CUDA kernels for the 1-bit inference path, making it possible to run the model directly on an 8GB GPU. The runtime exposes an OpenAI-compatible API, so existing tools and agents work without modification.
Performance is a different topic: 15 down to 9 tokens/s.
If you want to give it a try: find my notes and setup-scripts at GitHub: https://github.com/marcelpetrick/codingWithGPT/tree/master/bonsaiTestrun




