Skip to content

FastAPI SDK ​

API key authentication, per customer rate limits, quotas and usage billing for your FastAPI app, in one dependency. It needs Python 3.10 or newer. PyPI always has the latest release, and the changelog says what changed in each one.

Install ​

bash
pip install microauth-fastapi

Set your API's SDK secret key (mas_...) in the environment. You see it when you launch the API, and you can reveal it again on the Connect tab. Then protect a route:

python
from fastapi import FastAPI, Security
from microauth_fastapi import Customer, MicroAuth

app = FastAPI()
auth = MicroAuth(app)  # reads MICROAUTH_SECRET_KEY

@app.get("/forecast")
async def forecast(customer: Customer = Security(auth)):
    return {"customer": customer.id}
bash
MICROAUTH_SECRET_KEY=mas_... uvicorn main:app

You can also put MICROAUTH_SECRET_KEY=mas_... in a .env file next to the app. The SDK reads it when the variable isn't set in the environment.

That is the whole integration. Your customers sign up on your portal, create a key and add credit or pick a plan, and every request to /forecast is authenticated, rate limited and billed.

Before you deploy, run the built in check with the same environment. It confirms the secret works and can explain how any key would be treated:

bash
python -m microauth_fastapi check --key map_...

See the check command for what each line means. Every option and error is listed in the SDK reference.

What it does ​

  • API key auth through the X-API-Key header, or the header you set on the Portal tab. The security scheme shows up in your OpenAPI docs, so the Authorize button in Swagger UI works as is.
  • Account, balance and quota checks. Suspended customers and customers waiting for your approval get 403, customers out of prepaid credit get 402 and customers over their monthly quota get 429.
  • Per customer rate limits at each customer's effective requests a second, with 429 and Retry-After.
  • Your MicroAuth plan's allowance. The SDK reserves each request against the monthly allowance from the snapshot and answers 429 once it runs out.
  • Usage reporting with receipts. Every authenticated response is reported for quotas and your allowance. Only billable status codes (200 to 206 by default) are charged, and every charge comes back with a receipt the SDK checks against the price it enforced.

Built for the hot path ​

Known keys are served from memory. The network is only needed for a key created after the latest snapshot, for recovery when the cached data is too old, and for background sync:

  • A snapshot of your customers, keys and limits is cached in memory and refreshed in the background, every 30 seconds by default.
  • A key that isn't in the snapshot yet, because it was created seconds ago, is looked up once. Concurrent lookups for the same key share one call, and unknown keys are remembered as invalid for 30 seconds, so a flood of bad keys never reaches MicroAuth.
  • Usage is written to disk, or to Redis, before the response finishes, and sent in batches when 500 requests are waiting or report_interval (5 seconds) has passed, whichever comes first. Each item keeps the same idempotency key across retries and restarts until MicroAuth acknowledges it.
  • Checking a key costs one SHA-256 and a couple of dictionary lookups.

With fail_open=False, cached data isn't trusted past max_snapshot_age. The default fail_open=True keeps serving known keys beyond that, but never past max_stale_snapshot_age. When the SDK can't reach a trustworthy answer, it returns 503 instead of turning an outage into a false 401.

Multiple workers? Add Redis ​

Rate limits in memory are per process. With 4 uvicorn workers, a customer could reach about 4 times their limit, and the workers can't share their allowance reservations. Add Redis when more than one process serves your API:

bash
pip install 'microauth-fastapi[redis]'
python
auth = MicroAuth(app, redis_url="redis://localhost:6379/0")

The Redis limiter uses one atomic operation and Redis server time, so all workers share one counter per customer. It also shares snapshots, so a cold start reuses a fresh one instead of calling MicroAuth, and it keeps pending usage in a durable queue that any worker can deliver.

If Redis is unreachable, the SDK can't reserve requests safely and answers 503 until Redis is back. Keep Redis close to your API, in the same region.

Optional authentication ​

For endpoints that serve both anonymous and authenticated callers:

python
@app.get("/status")
async def status(customer: Customer | None = Security(auth.optional)):
    return {"authenticated": customer is not None}

The Customer object ​

The dependency resolves to a Customer with the fields a handler might need:

FieldMeaning
idThe customer's ID, a UUID that stays the same across keys and teammates
key_idThe API key that authenticated this request
statusAlways active, because suspended and pending customers are rejected first
billing_modelpayg, subscription or none
rpsEffective requests a second
price_per_request_microEffective price per billable request, in micro-USD
monthly_quotaEffective monthly cap, or None for no cap
credit_balance_microEstimated balance after this request's reservation

Branch on billing_model or the limits rather than on plan names. The full list, including your plan's allowance fields, is in the SDK reference.

Serverless ​

Vercel and AWS Lambda freeze or discard an instance as soon as its work is done, so background timers stop and files on the instance can vanish. The SDK detects both platforms and adapts where it safely can. A good starting point for Vercel with Fluid Compute, with MICROAUTH_SECRET_KEY and MICROAUTH_REDIS_URL set in the project's environment:

python
auth = MicroAuth(app, trailing_flush=True)
  • Use Redis. Without it, usage waits in the instance's own journal, and an instance that gets replaced takes its unreported usage with it. With Redis, usage enters a shared queue before the response is released, and snapshots and limits are shared across instances. Check that the variable is really set in production: a misspelled name fails silently and the SDK falls back to per instance behavior.
  • flush_on_response turns on by itself when VERCEL or AWS_LAMBDA_FUNCTION_NAME is set. After each response the SDK sends any batch that is due while the invocation is still running.
  • trailing_flush=True keeps the last invocation of a burst alive until the batch deadline, at most report_interval seconds, then sends it. Without it, the last few requests before traffic stops wait for the next visitor. Leave it off where the platform buffers the whole response before returning it, because the wait would land on your callers.
  • Keep timeout at 5 seconds. Each call makes up to three attempts, and the first request after a cold start loads a snapshot on the caller's path.
  • Don't count on aclose(). Serverless platforms rarely shut down gracefully, so delivery has to happen while invocations run.

Redis connections under load ​

When you pass only redis_url, the SDK builds a pool of at most 64 connections per process with one second timeouts. On serverless every instance brings its own pool, so the total can pass your Redis plan's connection cap during a burst. Watch your logs for these two lines:

text
microauth: the Redis usage queue sweep failed (...)
microauth: the Redis usage queue is unavailable (...); falling back to the local journal

The first one retries on the next interval. The second means usage is falling back to the instance's disk and could be lost if the instance is recycled. The simplest fix is a smaller pool:

python
auth = MicroAuth(app, redis_max_connections=10)

For more control, pass your own client. A BlockingConnectionPool waits for a free connection instead of failing:

python
import os
from redis.asyncio import BlockingConnectionPool, Redis

pool = BlockingConnectionPool.from_url(
    os.environ["MICROAUTH_REDIS_URL"],
    max_connections=10,
    timeout=5,
    socket_timeout=2.0,
    socket_connect_timeout=2.0,
)
auth = MicroAuth(app, trailing_flush=True, redis_client=Redis(connection_pool=pool))

Multiply max_connections by the instances you run at peak and keep the result under your plan's cap, with room for other clients. A client you pass in stays yours: the SDK never closes it.

Error responses ​

StatusWhen
401Missing or invalid API key
402Prepaid balance used up
403Customer suspended, or waiting for approval
429Rate limit, monthly quota or your plan's allowance reached
503The snapshot, a key lookup or Redis is unavailable

Each one is a subclass of microauth_fastapi.AuthDenied, which is a FastAPI HTTPException. Register a handler to change the response body:

python
from fastapi.responses import JSONResponse
from microauth_fastapi import AuthDenied

@app.exception_handler(AuthDenied)
async def auth_denied(request, exc):
    return JSONResponse(
        status_code=exc.status_code,
        content={"error": exc.detail, "docs": "https://docs.example.com/errors"},
        headers=exc.headers or {},
    )

Receipts and reconciliation ​

MicroAuth answers every usage report with a receipt per item: which customer was charged, how much and when. The SDK compares each charge with what it expected from the prices it enforced when it let the requests in, and logs any difference. Pass hooks to keep your own record:

python
from microauth_fastapi import MicroAuth, UsageReceipt, UsageRejection

async def keep_receipts(receipts: list[UsageReceipt]) -> None:
    for receipt in receipts:
        await ledger.save(receipt.idempotency_key, receipt)

async def flag_rejections(rejections: list[UsageRejection]) -> None:
    for rejection in rejections:
        alerts.send(f"MicroAuth refused {rejection.count} request(s): {rejection.detail}")

auth = MicroAuth(app, on_receipts=keep_receipts, on_rejections=flag_rejections)
  • outcome is accepted when this delivery recorded the item and duplicate when MicroAuth had recorded it already, for example after a retry whose first answer was lost. The customer is charged once either way.
  • charged_micro is what MicroAuth charged and expected_charge_micro what the SDK expected. charge_matches is True when they agree, and None when the SDK can't know, for example for usage another process admitted.
  • Hooks run before an item leaves the queue, so a crash can't skip a receipt. The flip side is that one can arrive twice, so store receipts by idempotency_key.
  • A hook can be async. A plain function runs in a worker thread. Each call gets five seconds, and errors or timeouts are logged and counted without holding up delivery.
  • MicroAuth refuses an item for good only when it can never be recorded, for example usage for a key that isn't part of this API. The SDK does the same for usage older than 45 days. Refused usage is never charged and is kept for inspection, in the journal and in Redis when you use it.

auth.usage_stats() reports the delivery health of the process, ready for your metrics:

python
stats = auth.usage_stats()
metrics.gauge("microauth.queued_requests", stats.queued_requests)
metrics.gauge("microauth.charge_mismatches", stats.charge_mismatches)

Alert when rejected_items or charge_mismatches rise above zero, when consecutive_failures keeps climbing, or when oldest_queued_at keeps getting older. The SDK reference lists every field.

To see what a host still has to deliver, and everything MicroAuth refused, run the usage command where your API runs:

bash
python -m microauth_fastapi usage

Delivery and consistency ​

  • Balance and quota checks use the cached snapshot plus what the process counted since. They keep honest clients within their limits; they aren't a gateway that untrusted code has to pass through.
  • A request's usage is on disk before its response finishes. Without Redis, each process appends to its own journal file in journal_dir (MICROAUTH_JOURNAL_DIR), which defaults to a per user state directory such as ~/.local/state/microauth/journal. In containers, mount a volume there so queued usage outlives the container.
  • Every journal line carries a checksum. A line cut short by a crash belongs to a response that never finished, so it is dropped. Damage anywhere else is logged, and a copy of the file is kept for inspection.
  • When a process dies, the app on the same directory picks up its journal once the file has been idle for two minutes, even if no requests arrive. Running processes keep their files fresh, so a live journal is never taken over. With Redis, the shared queue gives the same guarantee across machines.
  • An item leaves the queue once MicroAuth answers accepted, duplicate or rejected. Items marked retry, and items lost to a network error, stay queued.
  • Each request keeps the usage_policy_id it was admitted under, so delayed reports are charged at the prices that applied when the request was served.
  • Usage hours, policy expiry and snapshot age follow MicroAuth's clock. The SDK measures the offset on every response, so a server whose clock drifts still reports usage in the right hour. The check command shows the offset.
  • One process serves one API, and queued usage is kept per API. Two APIs on one host never mix their usage, and rotating the secret keeps the same journal and Redis queue.
  • Suspensions and revoked keys reach every worker within one sync_interval, 30 seconds by default.
  • The installed middleware loads the first snapshot when the app starts and drains pending usage when it shuts down.

Not using FastAPI? ​

The SDK is a careful client of three HTTP endpoints. The HTTP API guide documents them, so you can build the same thing in Node, Go, Rails or anything else.

MicroAuth is a product of Zyref, LLC.