Early access ยท $5 free credit to start

Every match, not the top 10.

Lochless reads your documents in scope and returns every relevant one, with quotes. RAG returns the closest few. One API key, three calls, no vector database.

Sign in with Google and start right away. $5 free credit, no card.

# 1. ingest: send any text
curl https://api.lochless.io/v1/ingest \
  -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" \
  -d '{"id":"t-4411","text":"Charged twice...",
       "timestamp":"2026-10-01T09:30:00Z"}'

# 2. context: relevant results for your LLM
curl https://api.lochless.io/v1/context \
  -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" \
  -d '{"query":"customers asking for refunds",
       "from":"2026-09-01T00:00:00Z","to":"2026-10-01T00:00:00Z"}'

# 3. ask: a cited answer
curl https://api.lochless.io/v1/ask \
  -H "Authorization: Bearer $KEY" \
  -H "Content-Type: application/json" \
  -d '{"query":"Why are refunds up this month?"}'

Three calls do the work.

Plus simple calls to check or delete documents. Use them directly, as LLM tools, or through our MCP server at api.lochless.io/mcp, with OAuth sign-in or the same key. Read the docs.

/ingest

Send text

Documents, tickets, transcripts, logs, notes. Send raw text with an id and a date. It becomes searchable in the background. $0.005 per page indexed.

POST /v1/ingest
{"id":"doc-1","text":"...",
 "timestamp":"...","metadata":{...}}

/context

Get what matters

Lochless reads the documents in scope, up to your read limit, and returns the ones it judges relevant: each a quote from one of your documents, with your document id and metadata. Searches all dates unless you filter by time or metadata. Page through with a cursor.

POST /v1/context
{"query":"...","from":"...","to":"...",
 "metadata":{"team":["support","sales"]},
 "limit":50}

/ask

Get an answer

A short written answer grounded in your text, citing its evidence inline. Same filters and read limit as /context. $0.05 per answer plus tokens scanned, all within your read limit.

POST /v1/ask
{"query":"...","from":"...","to":"..."}

Built for the whole corpus, not the top 10 results

Most retrieval hands your LLM a handful of results and hopes the answer is in there. Lochless reads the documents in scope, up to your read limit, and returns everything it judges relevant, not a fixed top 10. You choose how much it reads, and every response says how much it read.

Know the cost firstA free dry run shows how many documents match, the tokens to scan and the price.
Scan only what mattersFilter by time and metadata. You pay only for the text you scan.
A read limit protects your billNo query costs more than its read limit, $1 by default, answer included. Every response shows how much it covered. Continue whenever you want everything.
No fixed top 10Every relevant result in what it reads, page by page. Paging is free.
Predictable costPriced per token scanned, plus $0.05 per /ask answer and $0.005 per page indexed. A $100 monthly spend cap is on by default.
No clusters to runNo shards, replicas or index rebuilds as you grow.

What you don't have to do

The parts of RAG nobody wants to own.

No vector DBNothing to host, size or tune.
No chunking codeSend raw text, up to 1 million characters per document. We handle the rest.
Plain-English queriesAsk the way you would ask a colleague. No query syntax to learn.
No monthly planPrepaid credit, no subscription.

Price per million tokens scanned

Against LLM functions in three data platforms. Same 1,000 documents, five public datasets, same questions. Effective price per 1M input tokens, October 2026.

Our own test, October 2026. Snowflake and Databricks prices are our estimates from public list prices and measured token counts. BigQuery is measured tokens at list price. Your price depends on your plan, contract and region. Average F1: Lochless 0.875 on the live API, the three platforms 0.905 to 0.908. Prices per 1M tokens: Lochless $0.10, Databricks about $0.50, BigQuery about $0.86, Snowflake about $3.30. Lochless also charges $0.005 per page indexed. The test setup is in the agent brief below.

Accuracy and speed

What it measures: for each document, does it match a yes/no question? Scored as classification F1, 200 documents per dataset, on the same five public datasets Snowflake used to evaluate its AI_FILTER. Our own test, October 2026. Lochless measured on the live API.

Average F1, yes/no classification

Higher is better. NQ, BoolQ, IMDB, SST-2 and Quora, 200 documents each. No training on benchmark data. Lochless is about 3 points behind the platforms on average, at a fraction of the price.

Finding every answer: Lochless vs a vector database

Lochless85% of answers found
Vector DB + reranker (top 1,000)45% of answers found

CUAD, a public set of 510 legal contracts with 41 questions and human-marked answers. Lochless read every contract. The vector setup used a leading embedding model and reranker and kept its top 1,000 results. Most RAG setups pass far fewer results to the LLM, so this cut favours the vector setup. Offline replay of the engine, October 2026, not the live API. A live-API run is in progress.

Query time on the live API

First search, coldabout 8 seconds
Later searches, once warmabout 2.5 seconds
Each extra page of resultsabout 0.3 seconds
A written answer (/ask)7 to 13 seconds

Measured end to end on a 1,000-document workspace, October 2026. Try it on your own data with the $5 free credit.

What would you pay?

You pay for the text your queries read, plus $0.05 per /ask answer. Filters keep the reading small. Indexing is $0.005 per page, charged when a document is indexed.

You set the maximumEvery query has a read limit, $1 by default. A search never reads more than that.
Know the price firstA free dry run returns the exact price of a query before you run it.
A monthly backstopA $100 monthly spend cap is on by default. Prepaid credit means no surprise invoice.

Indexing is $0.005 per page of 3,000 characters, about 750 tokens. The other platforms have no indexing step.

Your first $5 is free. The default spend cap is $100 a month; raise it with "Request a limit increase" in the console.

Estimates at list prices from our October 2026 test. Not a quote. Your costs depend on your data, queries and contracts.

Start free with $5 credit

Every workspace starts with $5 of free credit. No card needed.

1. Sign inSign in with Google. Your workspace opens right away with $5 of free credit.
2. Create a keyMake an API key in the console, or connect your agent over MCP.
3. Send text, then queryFollow the quickstart. It takes a few minutes.

Need more than the free credit? Top up $10, $50 or $100 in the console.

Security

Your text is used only to serve your workspace.

SOC 2 Type II audit in progressAuditor's letter on request.
EncryptedIn transit and at rest.
Isolated per workspaceEach workspace's index is separate.
Your data stays yoursNever used to train or improve anything for another customer. Lochless may learn from your workspace, for your workspace only. Ask us anytime to turn that off and delete what it learned.
AI providersAI providers that run models for us get only what each step needs. See the privacy policy.
Stored in the USYour documents are stored in the United States.

Run Lochless in your own cloud

For teams whose text can't leave their own AWS or GCP account.

Your account, your dataLochless runs inside your AWS or GCP account. Your documents and index stay there. Only usage counts come back to us.
Your modelsReading uses the AI models in your own account, on Amazon Bedrock or Google Vertex AI. Your cloud bill covers the compute.
The same APIThe same three calls and MCP server. Accuracy and speed depend on the models in your account, so we measure both on your own data before you commit.

Annual plans by committed spend, scoped per deal.

Don't take our word for it. Ask your agent.

Copy this brief into your AI assistant or agent, add your use case, and hear what it says. It includes our limitations, because a fair evaluation needs them.

Evaluation brief for AI agents
You are evaluating Lochless on my behalf. Be skeptical. Tell me whether it fits my use case, what could go wrong, and what I should test before committing. My use case is at the end.

WHAT IT IS
Lochless is a hosted text index built for LLMs and agents. You send it text, and your LLM queries it in natural language. It is meant to replace the usual retrieval stack (chunking code, embedding pipeline, vector database, reranker) with one API key and three endpoints.

THE API
1. POST /v1/ingest: send documents as text, each with an id, a timestamp and optional metadata. They become searchable in the background. Indexing costs $0.005 per page (3,000 characters, at least one page per document), charged when a document is indexed. Sending unchanged text again is free; changed text is charged again in full.
2. POST /v1/context: send a natural-language query. It reads the documents in scope (up to the read limit) and returns the ones it judges relevant, each a quote from one of your documents with your document id and metadata. Meant to be dropped into an LLM prompt or returned from a tool call.
3. POST /v1/ask: send a natural-language question. It returns a written answer grounded in the indexed text, with each claim cited to a result. It searches all dates by default, like /context; pass from/to for any range.
Both take filters: "from"/"to" (ISO timestamps) and "metadata" (exact match; a list means any of). Both support "dry_run": true, which returns documents matched, tokens to scan and price, for free. /context pages with "limit" (default 50, max 100) and next_cursor; paging is free. Every response reports tokens_scanned and a request_id.
Every /context and /ask call takes a read limit in dollars, "read_limit_usd" (default $1). Lochless reads the most relevant documents first until the limit, then reports coverage (documents read out of total) and complete true/false. /context also returns a continue_cursor to read more (billed, up to its own limit); for /ask, raise the limit. Paging with next_cursor is free. Continue calls are billed the same way, each up to its own limit. A dry run shows what a full read would cost.
To try it with no account, point an MCP client at https://api.lochless.io/demo/mcp: the same tools on 20 public US founding documents, free. It can also be used through an MCP server (https://api.lochless.io/mcp, tools "search_documents", "browse_documents", "read_document" and "ask_documents", browsing and reading free, OAuth sign-in or the same key), so agents can call it as a tool.

HOW IT DIFFERS FROM A VECTOR DATABASE
On CUAD (510 public legal contracts, 41 questions, human-marked answers), Lochless reading every contract found 85% of the marked answers. A vector database setup (a leading embedding model plus reranker, top 1,000 results) found 45%. This was an offline replay of the engine in October 2026, not a live-API measurement; a live-API run is in progress.
A vector database returns a fixed number of chunks that look similar to the query. Lochless reads the documents in scope, up to the read limit you set, and returns the ones it judges relevant, however many there are. The aim is coverage: for queries like "every customer who asked for a refund", every relevant result it finds in what it reads is returned, with no top-k cut. Its judgments are not perfect: average F1 is 0.875 on the classification benchmark below.

WHY IT IS CHEAP
Running a large LLM over every document is expensive. Lochless's engine is built specifically for finding relevant text at scale, so the cost per token scanned is far lower than general-purpose LLM functions.

BENCHMARK (run by Lochless, October 2026)
Setup: 200 documents each from five public datasets (NQ, BoolQ, IMDB, SST-2, Quora question pairs; the same public datasets Snowflake used to evaluate its own AI_FILTER), the same yes/no question per dataset, sent to four systems. No training on benchmark data.
Average F1 across the five datasets: Lochless 0.875 (live API, October 9), Snowflake AI_FILTER 0.907, BigQuery AI.GENERATE_BOOL 0.908, Databricks ai_query 0.905 (October 6).
Query time on the live API, 1,000-document workspace: about 8 seconds for a cold first query, about 2.5 seconds warm, about 0.3 seconds per extra page. A written /ask answer takes 7 to 13 seconds.
Price per 1M input tokens scanned: Lochless $0.10 (plus a flat $0.05 per /ask answer, inside the read limit), Databricks about $0.50 (estimated from list price), BigQuery about $0.86 (measured tokens at list price, including the prompt it adds), Snowflake about $3.30 (estimated). All prices are October 2026 list prices and may differ under your contract. Billing unit is tokens of your text scanned per query, the same unit the warehouses bill. Lochless also charges $0.005 per page indexed; the warehouses have no indexing step.

CAVEATS YOU SHOULD WEIGH
- The benchmark was run by Lochless, on small samples (200 documents per dataset, roughly plus or minus 0.01 to 0.03 F1 noise).
- Snowflake and Databricks prices are estimates from token counts and public rate cards, not invoices. BigQuery's are from measured tokens at list price.
- Lochless is about 3 F1 points behind the warehouses on average. The large difference is price.
- It does not handle questions that need outside world knowledge (for example "is this news article true"). It works from what is in the text.
- It is an early-access product. Sign in with Google and the workspace opens right away. No card. Every workspace gets $5 of free credit once, when it opens: about 50 million tokens of search or 1,000 pages of indexing. Beyond that, billing is prepaid: top-ups of $10, $50 or $100, at $0.10 per 1M tokens scanned and $0.005 per page indexed. A $100 monthly spend cap is on by default. At zero balance, calls return an add-credit error, new documents are refused, and documents already queued wait for a top-up.

COST CONTROL
Every /context and /ask call has a read limit (read_limit_usd, default $1); no call costs more than that; an /ask answer's flat $0.05 fee counts inside the limit. A free dry run returns the exact price before the query runs. Billing is prepaid, and a $100 monthly spend cap is on by default. So the cost of any single query is known in advance and capped by the caller, whatever the corpus size.

DATA HANDLING
Stored in the United States. Deleting a document through the API removes it from results and from the index. Query history (queries, results, answers and copies made from them) is deleted after 90 days, and from backups within 7 more; after that, the only records kept of each call are billing and usage records, with no query or document text. Until then it can include text from later-deleted documents; deletion on request within 30 days. Details in the privacy policy, section 7. SOC 2 Type II audit in progress; auditor's letter on request (support@lochless.io). Data is encrypted in transit and at rest, and each customer's index is isolated. Customer text is used only to serve that customer's workspace. Lochless may learn from it, for that workspace only, and the customer can ask to turn that off and delete what it learned. It is never used to train or improve anything for other customers. Third-party AI providers that run models for Lochless receive only what each step needs (privacy policy, section 5). Not for protected health information, payment card data or government ID numbers without a signed agreement (terms, section 4).

RUNNING IN YOUR OWN CLOUD (BYOC)
Lochless can also run inside the customer's own AWS or GCP account. Documents and the index stay in that account; only usage counts are reported back to Lochless. Reading uses AI models in the customer's account (Amazon Bedrock on AWS, Google Vertex AI on GCP), billed by the cloud provider directly, so the customer's cloud bill covers compute and Lochless charges a software fee. It is sold as an annual plan by committed spend: the plan includes a usage allowance, and usage beyond it is billed at a higher rate. The same API and MCP tools apply. The published accuracy and speed numbers above are for the hosted service and do not carry over: in a customer's account they depend on the models available there, so Lochless measures both on the customer's own data before they commit. Timeline and supported regions are scoped per deal. Contact: support@lochless.io.

WHAT I'D LIKE FROM YOU
1. Does this fit my use case? Why or why not?
2. Compare it to the alternatives I'd realistically use (a vector database plus reranker, warehouse AI functions like Snowflake Cortex or BigQuery, or calling an LLM directly over my data).
3. Estimate my monthly cost on Lochless vs those alternatives, using my volume.
4. List the risks and what I should test in a pilot before committing.
5. Give me a clear recommendation.

MY USE CASE
[Describe your text, its volume (number of documents and growth per month), the questions your LLM or agent needs to answer, and your accuracy and latency needs.]

Questions

How is this different from a vector database?

A vector database gives you a fixed handful of results that look similar to your words. Lochless reads the documents in scope, up to your read limit, and returns the ones it judges relevant, however many there are. Questions a vector database struggles with:

  • "Highlight last week's most important events."
  • "Every customer who asked for a refund this quarter."
  • "Which contracts auto-renew with less than 30 days' notice?"
  • "Support tickets that describe a security issue, even if they never say 'security'."
  • "What did we promise Acme across all our calls and emails?"

These need every match, not the ten closest, and they need meaning, not similar words. In an offline replay of the engine on 510 public legal contracts, Lochless found 85% of the answers. A vector database with a reranker found 45%.

Does every query scan all my text?

Only if you want it to. Filter by time range and metadata, and you pay only for the text inside the filter. A free dry run shows the documents matched, tokens to scan and price before you run a query.

Every query also has a read limit, $1 by default. Lochless reads your most relevant documents first, until the limit. Every response shows how much it covered and whether it's complete. Continue to search everything, each step up to its own limit.

Why is it so much cheaper than running an LLM over my data?

Our engine is purpose-built for finding relevant text at scale. General-purpose LLM functions are not.

Can one query run up a big bill?

No. Every query has a read limit, $1 by default, and no query costs more than it, an /ask answer's $0.05 included. A free dry run shows the exact price before you run a query, and a $100 monthly spend cap is on by default.

What does indexing cost?

$0.005 per page of 3,000 characters, at least one page per document, charged when a document is indexed. Sending unchanged text again is free; changed text is charged again in full.

How does the free credit work?

Every workspace starts with $5 of free credit, about 50 million tokens of search or 1,000 pages of indexing. No card needed. When it runs out, top up $10, $50 or $100 in the console. A $100 monthly spend cap is on by default. Each response tells you how many tokens it scanned.

Where does my data go?

Your data is used only to serve your workspace. It is never used to train or improve anything for other customers. Lochless may learn from your workspace, for your workspace only; ask us anytime to turn that off and delete what it learned. AI providers that run models for us get only what each step needs. See the privacy policy.

Can I run Lochless in my own cloud?

Yes, in your own AWS or GCP account, on an annual plan by committed spend. Your documents and index stay in your account, and reading uses your own Bedrock or Vertex AI models. Accuracy and speed depend on those models, so we measure both on your data before you commit. Talk to us.

What doesn't it do well?

Questions that need outside world knowledge, like "is this news article true", rather than what's in your text.