Projects

Here are some of the open source projects I’ve been working on. Everything is free to use and it’s all available on GitHub.

Qbort

github.com/balevine/qbort

When you need a large set of support conversations, Qbort can create them for you. You can specify the types of conversations, the distribution of topics and sentiments, the product you’re (pretending to be) supporting, and more. This is great for creating demo data to show customers or for testing automations without passing customer data around.

Qval

github.com/balevine/qval

Created to test LLM-based automations, Qval is a Claude Code skill for comparing your LLM output to a set of expected outputs. You create a set of rules and an output schema, you import a set of support conversations (using Qbort’s data format), you manually evaluate the conversations according to your rules and schema, and then you see how well the LLM does with that set of rules. This works well for testing and iterating on your LLM prompts, especially for automations running on platforms like n8n.

Qbench

github.com/balevine/qbench

When Qval doesn’t give you enough information, reach for Qbench. This takes the basic premise of comparing LLM outputs to an expected set of results and amplifies it. Compare outputs across multiple models, all routed through OpenRouter. Qbench is a desktop app built for macOS, which makes it a little more accessible to folks who aren’t in the terminal or in Claude Code all the time. It’s a great tool for when you’re building LLM- or classifier-based automations and want to find the right combination of prompts and models for your use cases.