OpenAI-compatible memory gateway

Give every AI request the right context.

Add structured memory, semantic cache, and provider routing through one OpenAI-compatible endpoint—without rebuilding your application.

Keep application context under one workspace-scoped control plane, even when models or providers change.

A four-step request lifecycle accepts an OpenAI-compatible request, reconstructs workspace context, routes to a provider, and records a request trace.

Platform facts

APIOpenAI-compatible request shape
IsolationWorkspace-scoped memory and cache
Provider accessStored or per-request keys
VisibilityRequest-level logs

Live compatibility

Models your stack can route today

Compare published pricing with verified third-party API offers, then explore the available routes and API model IDs.

Loading available models…

Use cases

Context that stays useful after the first response.

HardCarrx gives applications a durable context layer without coupling product state to one model vendor or one conversation transcript.

Stateful assistants

Carry user preferences, durable facts, and prior decisions into the next useful response instead of restarting every session.

Support copilots

Reuse resolved context and similar intent so agents spend less time reconstructing a customer’s history and repeating model work.

Workflow agents

Store checkpoints and task context so long-running automations can resume reliably when a request, session, or provider changes.

One request lifecycle

Your application sends one request. HardCarrx handles the context path.

Keep the integration familiar while memory, cache, routing, and request visibility remain independently configurable.

  1. 01

    Send your request

    Use the OpenAI-compatible chat completions shape your application already understands.

  2. 02

    Reconstruct context

    Retrieve relevant workspace memory and apply the context controls configured for the request.

  3. 03

    Route and reuse

    Apply provider and cache policy before sending the request to the selected model.

  4. 04

    Inspect the result

    Review provider, model, latency, cache, memory, and request identifiers in your logs.

Quickstart

Make your first request with the tools you already use.

Create a workspace API key, point your application at the HardCarrx endpoint, and keep your existing OpenAI-style request shape.

TerminalPOST /v1/chat/completions
curl -X POST https://api.hardcarrx.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $HARDCARRX_API_KEY" \
  -d '{
    "provider": "openai",
    "model": "gpt-4.1-mini",
    "messages": [
      {"role": "user", "content": "Summarize the next action."}
    ]
  }'

Data controls

Keep context visible, bounded, and portable.

HardCarrx separates application context from model vendors so teams can govern what is stored, where it is scoped, and how each request is handled.

Workspace boundaries

Memory, cache, keys, and request visibility stay scoped to the workspace that owns them.

Context and retention limits

Plan limits make memory-enabled usage, stored context, workspaces, and retention explicit.

Provider control and request evidence

Use stored or per-request provider access, then inspect provider, model, latency, cache, memory, and request identifiers in logs.

Pricing

Start free. Scale context when usage grows.

Compare explicit limits for memory-enabled requests, context items, workspaces, and retention before you integrate.

Free

Best for: POCs and internal validation

$0

  • ✓ 1,000 memory-enabled requests / month
  • ✓ 1 workspace
  • ✓ 10,000 context items
  • ✓ 14-day retention
Start free

Starter

Best for: Early production workloads

$9

/ month

  • ✓ 10,000 memory-enabled requests / month
  • ✓ 2 workspaces
  • ✓ 50,000 context items
  • ✓ 30-day retention
Get started

Pro

Most Popular

Best for: Revenue-critical AI experiences

$29

/ month

  • ✓ 50,000 memory-enabled requests / month
  • ✓ 5 workspaces
  • ✓ 300,000 context items
  • ✓ 180-day retention
Get started

Team

Best for: High-scale production and multi-team ops

$99

/ month

  • ✓ 250,000 memory-enabled requests / month
  • ✓ 10 workspaces
  • ✓ 2,000,000 context items
  • ✓ 1-year retention
Get started
  • Memory-enabled requests: Calls where HardCarrx stores and uses memory to personalize future responses.
  • Context items: Individual saved pieces of memory like preferences, facts, or conversation notes.

Frequently asked questions

Practical answers for engineering and product teams evaluating HardCarrx.

How do I integrate HardCarrx?

Point an OpenAI-compatible chat completions request at the HardCarrx endpoint and authenticate with a workspace API key. Memory, cache, and provider policies can then be configured independently of application code.

Can we run multiple providers at the same time?

Yes. HardCarrx is provider-agnostic and supports traffic strategies based on latency, quality, cost, and reliability targets.

What happens when memory-enabled request quota is exhausted?

Memory-enabled requests are limited by plan, but API passthrough still works after memory quota is exhausted.

What does the memory retention period mean?

Each workspace keeps its most recent memory within one rolling window: 14 days on Free, 30 days on Starter, 180 days on Pro, and one year on Team. Your Billing page shows one boundary date—‘Keeping memory created since [date]’—and that date advances daily. Memory older than the active boundary is excluded from normal retrieval and may later be soft-deleted by retention processing. Soft deletion is not the same as immediate physical erasure.

What happens to memory if I downgrade or cancel?

An upgrade or Retention Extend expansion moves the workspace boundary backward immediately. If a downgrade or Retention Extend removal would shorten the window, the current window remains enforced for seven days; Billing shows the upcoming boundary and exact enforcement date. Cancellation or inactive billing may restrict memory reads and writes. These access changes do not promise immediate physical erasure, and user deletion and context-capacity policies remain separate lifecycle paths.

What happens when I reach the context-item limit?

The configured context-capacity policy determines what happens next: block new memory writes, archive the oldest items, or soft-delete the oldest items. Your current limit and policy are shown in the dashboard so you can act before reaching capacity.

Where is my memory stored?

Memory is stored in HardCarrx-managed production data services and may be processed by model providers you configure, as described in our Privacy Policy.

Start with one endpoint

Give your application context it can reuse.

Create a workspace, make your first OpenAI-compatible request, and add memory, cache, and routing controls when your product needs them.