Claude vision seems like a massively overpriced VLM choice given Gemini has great VLM tower since 2.4 and my gut feeling is this is a task that local Qwen’s would handle easily.
Disclaimer: my daily workflows include millions of Claude and Gemini tokens spent on dev and video vision intelligence.
This was what I had available, but adding support for other models/providers should be trivial. Adding support for local models is probably a bit harder and I don't have any experience with it. I doubt anything good enough will run on my macbook m1 though.
Disclaimer: my daily workflows include millions of Claude and Gemini tokens spent on dev and video vision intelligence.