The research behind Qelly

Intelligence that earns its conclusions.

We're training financial agents to investigate a business, discover feasible alternatives, and defend a decision with evidence.

Read the research thesis
Conceptual decision search: many proposals converge on verified feasible alternatives
Candidate proposalsFinancial verificationFeasible alternatives
A conceptual view of decision search. Only proposals that satisfy the stated evidence and constraints become alternatives.

Our thesis

Better answers begin with better decisions.

A useful financial agent must know what to investigate, which tools to use, and when the evidence should change its recommendation.

We are developing a specialized Qwen-based agent around an explicit model of a business: cash, dates, obligations, available actions, and constraints. Financial records and computed outcomes remain authoritative.

Learned representations may help discover useful combinations. Each proposal must decode into an action we can check, calculate, and explain. A fluent answer alone does not establish feasibility.

See the product example

From evidence to a reviewable decision.

Each part of the system has a specific job.

  1. 1

    Ground the state

    Documents, accounts, and timestamps become structured financial facts with provenance and material gaps.

  2. 2

    Propose and investigate

    The agent asks targeted questions and proposes typed actions. Search explores combinations and stress scenarios.

  3. 3

    Compute and verify

    Independent tools calculate cash outcomes and check evidence, permissions, obligations, and hard constraints.

  4. 4

    Explain and review

    A decision brief preserves alternatives and uncertainty for human review. Accepted cases can improve training.

One product, an independent research program

Learn from evidence.
Improve the same workflow.

Qelly launches through familiar assistant interfaces. Specialized models can later improve selected Qelly tools without asking customers to change how they work.

Our training pipeline excludes host-platform requests and generated answers by default. Research material follows a separate provenance and rights process: original verified cases, eligible public data, independently collected permissioned records, and expert annotations.

A reviewer's approval is feedback, not proof of correctness. Real outcomes inform the record, but do not reveal the outcomes of alternatives that were never taken.

Deployment requires better quality on defined workflows, or comparable quality at materially lower cost, with reliable constraints and uncertainty handling. Quantum methods are evaluated where they help; they are not a prerequisite for launch.

Training program

Teach the procedure. Measure the judgment.

Begin with the unmodified model using the same tools and evidence. Then change the training objective only when the data and evaluation support it.

Supervised fine-tuning

Initial experiments

LoRA adapters on a quantized Qwen base learn verified demonstrations: request information, call a tool, compare feasible actions, and produce a supported conclusion. User content and tool observations are context, not fabricated output targets.

Multi-turn reinforcement learning

Planned

GRPO will compare fresh trajectories through a controlled financial environment. Independent checks score numerical correctness, material gaps, constraint satisfaction, and decision usefulness.

Expert preference learning

Conditional

DPO is reserved for trustworthy preferred/rejected examples, such as explaining a covenant condition clearly. An unverified model score is insufficient to establish the preferred answer.

What would count as progress

A financial decision has to survive the test.

We evaluate whole workflows, with the same available evidence, tools, and inference budget. General models and unmodified Qwen remain active comparisons.

Evaluation commitments
Question Evidence we require
Are the numbers right? Executable calculations with checked amounts, units, and dates.
Is the recommendation feasible? Hard constraints hold. Unconfirmed financing or consent stays conditional.
Did it ask what matters? Material missing facts are identified; irrelevant questions do not earn credit.
Does it generalize? Business, source-document, period, and scenario-family holdouts; related examples stay together.
Did training help? Base, prompted, and trained models compared on frozen cases, with cost and latency recorded.

Classical and quantum-assisted search

Can better search reveal a better next move?

We are investigating whether learned representations and D-Wave annealing can discover useful action combinations or stress scenarios that other methods miss.

Direct annealing requires an appropriate binary quadratic formulation. Hybrid solvers are evaluated separately. Every returned candidate undergoes the same independent financial checks.

The test: better feasible decisions, broader useful scenario coverage, or more efficient learning against strong classical search at comparable total cost.

Research hypothesis · advantage not established

Research record

The evidence comes first.

Our first training cycle is a bounded local pilot. Public benchmark wins, customer performance, and quantum advantage have not been established. Measured results belong here only after the run, split, and evaluation can be traced.

FoundationDeterministic cash-decision exampleInspect the calculation
In developmentCFO training and comparative evaluation

Baseline → supervised adapter → frozen holdout. Independent numerical checks accompany qualitative model review.

Next evidenceDesign-partner workflows

Permissioned inputs and adviser-reviewed conclusions, followed by real operating feedback.

Technical foundations

Our research builds on established training methods and independently checkable financial tasks.

For researchers, partners, and investors

Help build financial intelligence we can put to the test.

Start a conversation