Text, Voice, and Video: Why Financial AI Needs All Three Input Modalities
Multimodal Interaction Design for Enterprise Financial Intelligence Systems
---
Your CFO dictates a supplier query during their morning commute. A bookkeeper snaps a stack of paper receipts at a client site and asks Stralevo to categorize each one. A partner reviews a physical contract during a site visit and asks a question about payment terms without taking the document back to the office. By the time any of them reach a desktop keyboard, the answers are already waiting.
The financial question that matters most often arrives when you're furthest from a keyboard. Stralevo answers it by voice, by camera, or by text — the same intelligence, whichever input is available.
---
Where Financial Questions Actually Arise
Finance AI has been text-first since its inception — open a browser, type a query, wait for a response. That design reflects where enterprise AI was built: for the office-based accountant at a desk in 2015. Finance professionals in 2026 spend a significant portion of their working day away from that desk — in client visits, in meetings, on the road.
Think about the last five financial questions you asked informally: "Is that invoice paid?" "What's our balance with that supplier?" "Does this receipt exceed the expense limit?" How many of those questions did you actually get answered the same day? The rest became "I'll check that when I'm back at my desk" — notes that, for most finance teams, have an uncomfortably low follow-through rate.
Each unanswered informal financial question compounds. The receipt never categorized becomes a month-end reconciliation problem. The supplier query deferred from Tuesday becomes an unresolved anomaly that surfaces in the quarterly review. The question asked during a client visit and forgotten two hours later becomes a decision made on incomplete information. None of those gaps is visible in isolation. Accumulated across a finance team over a quarter, they represent the operational cost of intelligence that was only accessible from a specific device in a specific room.
---
Voice: The Query That Doesn't Wait for a Keyboard
A CFO driving to a board meeting asks "what's our total spend with Supplier X this year?" through Stralevo Companion, the iOS and Android mobile app. The query runs against the full invoice history already indexed in the Knowledge Graph — supplier transactions, purchase orders, payment records — and returns a sourced answer before the CFO arrives. Nothing was typed. No browser was opened. The answer was verified against original documents and includes the source reference.
Voice accuracy depends on audio quality, background noise, and the specificity of financial terminology spoken. Queries with complex supplier names or account codes may require a confirmation prompt. Stralevo surfaces a suggested interpretation before acting on a voice query for high-stakes financial questions — the same verification step a careful analyst would apply before delivering a board-meeting figure.
Practical voice use cases are clear: quick balance checks during travel, supplier queries between meetings, approval-chain questions during client calls. Complex multi-variable analysis — reconcile all invoices against purchase orders across 18 months for all suppliers above 50,000 euros — is still better approached from a desktop where the query can be reviewed before submission. Voice is not a replacement for structured query input. It is access to intelligence for all the moments when structured query input is unavailable.
---
Camera: The Document That Never Needed to Be Typed
For a bookkeeper processing 20 paper receipts from a client site visit, each image photographed through Stralevo Companion's camera input runs through SightCapture™ ICR — the same document processing engine that handles PDFs — which reads every field: amount, vendor, date, VAT, purchase category, project code, serial numbers where present. The bookkeeper asks "which budget category does each of these go under?" and receives categorization recommendations for all 20 documents before walking back to the car.
Camera input handles 50-plus document formats — printed invoices, photographed receipts, handwritten delivery notes, stamped contracts. Physical paper that would previously have required scanning at the office, uploading to the accounting system, and manual categorization is processed at the moment of capture. SightCapture™ doesn't require perfect lighting or exact framing — it applies the same field-extraction logic to photographed documents that it applies to PDF uploads.
One use case nobody discusses: the partner at a client site reviewing a physical supplier contract who photographs the payment terms page and asks "does this conflict with our standard terms?" The answer doesn't require the document to travel to the office. It doesn't require the client to email a PDF. The camera processes it where it exists.
---
Text: Still the Right Tool for Complex Queries
Desktop text input remains the best choice for complex, multi-variable financial queries where precision matters: "which of our top 20 suppliers raised prices more than 10% in the last two quarters, broken down by product category?" That query benefits from the ability to review it before submitting and to refine it if the initial phrasing is ambiguous.
Text input through Stralevo's desktop interface also supports the longer analytical workflows: building supplier concentration reports, preparing FEC audit documentation — the standardised accounting export French tax authorities can demand on 15 days' notice — running scenario analysis on payment terms. The full ContextUX™ interface — with source citation review, document drill-down, and report formatting — is designed for desktop use.
None of the three modalities compete. Text handles precision at a desk. Voice handles immediacy in motion. Camera handles physical documents in the field. All three route through the same intelligence — the same document index, the same Knowledge Graph, the same zero-hallucination answer verification. The modality is the entry point. The intelligence behind it is identical.
---
The Same Intelligence, Three Entry Points
ContextUX™ — Stralevo's interaction layer — was designed from the start for three input types. This is not text-first with voice added as an afterthought. The intelligence layer doesn't know or care whether the query arrived by voice, camera, or keyboard. It processes structured intent and returns a sourced answer regardless of how the intent was submitted.
Phones replaced desktops for most consumer information tasks not because desktop information was inferior — the question arose when the desktop wasn't available. Navigation apps, banking apps, and messaging became standard on mobile because the use cases — directions while moving, balances at the point of purchase — happen away from desks. Financial intelligence has the same mobility pattern. The CFO's question about supplier spend doesn't arise exclusively at a keyboard. The answer should be available wherever the question does.
---
What Becomes Possible
First-order effect: finance professionals can ask questions and capture documents wherever they are. Second-order effect: the volume of financial questions that actually get answered increases — because the barrier to asking drops from "I need to be at my desk with the right browser tab open" to "I can ask this right now."
Partners equipping their field-based finance teams with Stralevo Companion are setting a different standard for client service. Processing 14 paper receipts at a client site visit and providing categorized output before leaving is not a feature demonstration. It is a workflow change that eliminates a manual step that has existed in accounting for 40 years.
Download Stralevo Companion on iOS or Android. Dictate a financial question you would normally defer until you're back at your desk. The intelligence that usually waits for you is already there.