ollama

๐‚๐ฅ๐จ๐ฎ๐ ๐€๐๐ˆ๐ฌ ๐š๐ซ๐ž ๐ž๐ฑ๐ฉ๐ž๐ง๐ฌ๐ข๐ฏ๐ž. ๐–๐š๐ญ๐ž๐ซ๐ฆ๐ž๐ฅ๐จ๐ง ๐ข๐ฌ ๐ง๐จ๐ญ. ๐Ÿ‰

Written by  on July 15, 2026

Vision models grow in their capabilities while the execution demands way fewer resources.

Means: my sales slip scanner-project from two years ago needed an upgrade. Before I was using GPT-4 Vision via the @OpenAI API, which costs you control over your data and a few cents for every analysis.

Now, after intensive benchmarking across a range of 15 models (with ๐Ž๐ฅ๐ฅ๐š๐ฆ๐š) and some surprising regressions (newer harness doesn’t mean the models execute faster; it can also mean they suddenly run into model load errors…), I put my money on ๐๐ฐ๐ž๐ง๐Ÿ‘.๐Ÿ“:๐Ÿ’๐. A casual language model with a ๐ฏ๐ข๐ฌ๐ข๐จ๐ง ๐ž๐ง๐œ๐จ๐๐ž๐ซ. My own benchmark over the selected models against the huge sample size of three receipts: all three were successfully evaluated. ๐Ÿ˜‰

Then I built a ๐ก๐จ๐ญ-๐Ÿ๐จ๐ฅ๐๐ž๐ซ-๐œ๐จ๐ง๐œ๐ž๐ฉ๐ญ around the script, and it works (thank you @Mr Zahorsky โ€” without whom I would have never come across that idea). For instance: in less than ten seconds, 13 samples were processed, and you get the result sum plus a fancy HTML overview for comparison.

๐ƒ๐š๐ญ๐š ๐ž๐ง๐ญ๐ซ๐ฒ can be done for cheap: give the kids the smartphone, snap all receipts, throw out some watermelon slices ๐Ÿ‰๐Ÿ˜‰

If you want to run the tool as well, check my GitHub: https://github.com/marcelpetrick/sales-slip-scanner-ng. As said: a local GPU is the only thing you need; ๐ง๐จ ๐œ๐ฅ๐จ๐ฎ๐, ๐ง๐จ ๐€๐๐ˆ ๐ค๐ž๐ฒ๐ฌ, ๐ง๐จ ๐œ๐ซ๐ž๐๐ข๐ญ ๐œ๐š๐ซ๐. Everything is now self-hosted. You’ll also find the benchmark of the vision models there, including the results.

๐“๐ก๐ž ๐›๐ข๐  ๐ช๐ฎ๐ž๐ฌ๐ญ๐ข๐จ๐ง ๐ข๐ฌ: ๐ฐ๐ก๐š๐ญ ๐ฌ๐ก๐จ๐ฎ๐ฅ๐ ๐ˆ ๐š๐ฎ๐ญ๐จ๐ฆ๐š๐ญ๐ž ๐ง๐ž๐ฑ๐ญ ๐ฐ๐ข๐ญ๐ก ๐š ๐ฅ๐จ๐œ๐š๐ฅ ๐ฏ๐ข๐ฌ๐ข๐จ๐ง ๐ฆ๐จ๐๐ž๐ฅ? ๐–๐ก๐š๐ญ ๐ฐ๐จ๐ฎ๐ฅ๐ ๐ฒ๐จ๐ฎ ๐๐จ?

#neverstoplearning

๐—œ ๐—ฅ๐—ฎ๐—ป ๐—ฎ ๐Ÿฎ๐Ÿณ๐—• ๐— ๐—ผ๐—ฑ๐—ฒ๐—น ๐—ผ๐—ป ๐—ฎ๐—ป ๐Ÿด๐—š๐—• ๐—š๐—ฃ๐—จ.

Written by  on July 15, 2026

๐Ÿฐ๐Ÿฌ ๐—บ๐—ถ๐—ป๐˜‚๐˜๐—ฒ๐˜€. That’s all it took me to get Bonsai 27B running locally on an RTX A2000 Laptop GPU with just 8GB of VRAM.

Bonsai 27B is based on Qwen 3.6, but uses PrismML’s custom native ๐Ÿ-๐›๐ข๐ญ ๐๐จ๐ง๐ฌ๐š๐ข format, reducing the model to just 3.9GB. A specialized ๐ฅ๐ฅ๐š๐ฆ๐š.๐œ๐ฉ๐ฉ fork implements custom CUDA kernels for the 1-bit inference path, making it possible to run the model directly on an 8GB GPU. The runtime exposes an OpenAI-compatible API, so existing tools and agents work without modification.
Performance is a different topic: 15 down to 9 tokens/s.

If you want to give it a try: find my notes and setup-scripts at GitHub: https://github.com/marcelpetrick/codingWithGPT/tree/master/bonsaiTestrun