AI that has to survive an audit, not a demo

LLM integration · retrieval over internal documents · agents · evaluation

We build AI into the systems a company already runs — with the sources attached, the failure cases measured, and a description of the data flow that your security and compliance people can actually read. If a task is better solved without a model, we say so in the framing stage.

Six kinds of work, described by the task and not by the promise

LLM integration into systems you already run

A model is the smallest part of the work. The rest is the seam: your ERP, CRM, ticketing or document store, the permissions people already have, and the places where a wrong answer costs money.

Typical tasks

  • Answer drafting inside a case-management tool, with the sources attached
  • Classification and routing of inbound tickets or claims
  • Summaries of long records, with the page each statement came from

Retrieval over internal documents (RAG)

Answers grounded in your own corpus — policies, contracts, manuals, tickets. We build the ingestion, the chunking, the retrieval and the evaluation set, and we show retrieval quality separately from generation quality.

Typical tasks

  • Policy assistant for staff, restricted by department permissions
  • Contract review support with clause-level references
  • Technical manual search across drawings, PDFs and spreadsheets

Agents and process automation

Multi-step work with tools: reading a queue, calling internal APIs, filling forms, escalating to a person. Every step is logged and reversible; nothing irreversible happens without a human decision.

Typical tasks

  • First-line triage that drafts, then hands over with full context
  • Data reconciliation between two systems with a review queue
  • Back-office reporting assembled from several sources on a schedule

Evaluation and hallucination control

A model that answers fluently is not a model that answers correctly. We build the test set from your own hard cases, measure retrieval and answer quality separately, and keep the numbers visible after every change.

Typical tasks

  • Regression suite over a labelled set of real questions
  • Refusal and escalation rules for questions outside scope
  • Per-answer citation checks before anything reaches a user

Data privacy and regulatory requirements

Where the data lives, what leaves the perimeter, what is retained and what is logged. On-premise or in your own cloud account where the requirement demands it, with a data-flow description your compliance team can review.

Typical tasks

  • Processing inside your cloud tenant, no vendor-side retention
  • Redaction and pseudonymisation before any external call
  • Audit trail of prompts, retrievals and outputs per request

Support after launch

Models and documents change. We hand over with documentation and monitoring, or stay on for a fixed monthly engagement that includes evaluation re-runs and incident handling.

Typical tasks

  • Monitoring for retrieval drift and refusal spikes
  • Re-indexing when the document set changes
  • Quarterly evaluation report with the failure cases named

Stages, what each one produces, and what we need from you

Lengths are the usual range for a single process with a cooperative data owner. They are not a quote: the framing stage exists to replace them with a real plan for your case.

  1. 01

    Framing · 1–2 weeks

    One process, written down as it happens today, with the volumes, the exceptions and the cost of a mistake. Ends with a written scope and a decision on whether AI is the right tool at all.

    Needed from you

    Process owner, two or three real examples, a list of systems involved.

  2. 02

    Data and constraints · 1–2 weeks

    Where the documents live, who may see what, what may leave the perimeter, which model providers are acceptable and what the auditors will ask for.

    Needed from you

    Access to a representative sample, security contact, any existing policy.

  3. 03

    Pilot · 3–6 weeks

    A narrow, real slice in production conditions: retrieval, prompting, evaluation set, UI in the tool people already use. Ends with measured quality on your own cases.

    Needed from you

    A named user group, test environment, one person who can approve access.

  4. 04

    Evaluation and hardening · 2–4 weeks

    Failure cases from the pilot become the regression suite. Permissions, logging, rate limits, refusal paths and hand-over to a human are tightened.

    Needed from you

    Feedback from the user group, security review, decision on hosting.

  5. 05

    Rollout · 2–6 weeks

    Wider release, training for the people using it, monitoring and alerting, documentation for whoever runs it after us.

    Needed from you

    Rollout plan, support contact, agreement on what is measured monthly.

  6. 06

    Support · ongoing

    Fixed monthly engagement: monitoring, re-evaluation after model or document changes, and a named engineer who answers.

    Needed from you

    Nothing beyond the retainer decision.

Chosen per constraint, not per fashion

Models

Models from the main providers, and open-weight models hosted in your own environment when data may not leave it. The choice follows the constraint, not the other way round.

Retrieval

Vector and keyword search over your documents, hybrid ranking, permission-aware filtering, citation back to the source page.

Orchestration

Explicit pipelines and tool-calling agents with step logging, retries and a human hand-over point.

Evaluation

Labelled question sets from your own operation, retrieval and answer metrics tracked per release, failure cases kept as tests.

Integration

APIs and web services against your existing systems; no rip-and-replace, no second copy of your data where it can be avoided.

Delivery

Containerised services in your cloud account or ours, infrastructure as code, environment separation for test and production.

Data handling in one paragraph

Your documents and records stay in the environment you choose: your cloud tenant, your own servers, or ours if you prefer and the classification allows it. We describe what leaves the perimeter, keep the audit trail of prompts, retrievals and outputs, and do not use your data to train anything. Where a requirement forbids external model calls, we run open-weight models inside your perimeter instead — with the accuracy trade-off stated in writing.

Fixed stages, so the cost of the next step is known before you take it

Discovery sprint

Two to four weeks to turn an idea into a written scope, a data inventory and a go/no-go decision. Fixed price, no commitment afterwards.

fixed fee, quoted per scope

Pilot

One process, one user group, measured quality on your own cases. Ends with a working slice and the numbers to decide about production.

fixed fee per stage

Production build

Hardening, permissions, monitoring, rollout and documentation, delivered in stages with a demo at the end of each one.

fixed fee per stage

Support

Monthly engagement after launch: monitoring, re-evaluation, incident handling and small changes.

monthly retainer

The scope and the price of a stage are agreed in writing before it starts. Prices depend on the number of systems, the state of the documents and the hosting constraint — a discovery sprint is what turns that into a number we can stand behind.

Start with the process, not with the model

Tell us what has to change, which systems are involved and who would use the result. If an NDA is needed before details are shared, say so in the form and we will sign yours or send ours first.

Hours
Mon–Fri 9:00–18:00 Central, calls by appointment
Office
500 Example Street, Suite 12, Austin, TX 78701

Discuss a project

Tell us what has to change in the process, what systems are involved and who would use the result. We reply within one business day; an NDA can be signed before any detail is shared.

Prefer email? hello@northboundai.com · or call +1 (555) 010-0600.