Lamassu IoT Docs

Troubleshooting

Diagnose access, health, issuance, enrollment and dependencies in a Lamassu deployment.

Start by locating the failing layer. A Gateway response, an unready pod and a rejected issuance can look like the same problem from the console, but they require different evidence.

Diagnostic sequence

Reproduce a single operation

Note the time in UTC, user or principal, resource, endpoint and result. Avoid repeating a destructive action such as revocation or deletion.

Separate client, Gateway and service

Check the HTTP status. A 401 points to authentication; a 403, to authorization; a 404, to route or resource; a 5xx, to the service or a dependency.

Review Kubernetes status

Locate unready pods, restarts, recent events and the state of the release.

Query the affected service

Review its /health endpoint and its logs in the same time window. Then follow the indicated dependency: PostgreSQL, RabbitMQ, the OIDC provider or the cryptographic engine.

Validate the result from outside

For PKI, check the certificate, the chain, OCSP or CRL with the same type of client production uses.

Minimal evidence collection

Replace the placeholders with your real release, namespace and deployment.

helm status <release> -n <namespace>
kubectl get pods -n <namespace> -o wide
kubectl get deployments,statefulsets,jobs -n <namespace>
kubectl get events -n <namespace> --sort-by=.lastTimestamp
kubectl describe pod <pod> -n <namespace>
kubectl logs <pod> -n <namespace> --all-containers --since=30m

Also run the chart's connectivity test:

helm test <release> -n <namespace>

The test checks the health of CA, DMS Manager, Device Manager and VA, and verifies that the UI returns HTML. A passing test confirms basic connectivity, not a full issuance flow.

Probes use /health

The chart's services expose startup, readiness and liveness probes over /health. A pod can be running but receive no traffic if the readiness probe fails.

I cannot sign in

If you receive a 401:

  • check that the UI's OIDC authority is reachable from the browser;
  • verify issuer, audience and expiration of the token;
  • confirm the Gateway can fetch the keys from the JWKS URI;
  • check that the provider's public and internal routes point to the same realm.

If login succeeds but the APIs return 403:

  • inspect the token's real claims;
  • confirm they match an active principal;
  • review the policies granted to that principal;
  • verify that authz reaches PostgreSQL and its JWKS.

On a fresh installation, confirm the migration Job finished. That Job creates schemas, migrates authz, preloads policies and provisions the principals defined in services.authz.bootstrap.

Everything returns 403 or 5xx

External authorization fails closed by default. If authz is unavailable, protected routes are blocked.

  1. Locate the authz pod and Service.
  2. Check /health, logs and its database connection.
  3. Verify the DNS and port configured in externalAuthorization.
  4. Check the check route is /v1/ext_authz/check.
  5. Do not switch to failOpen: true as a permanent workaround: it would allow traffic without authorization during the outage.

A pod is not Ready

Review kubectl describe pod first. Common causes are:

  • wrong credentials or DNS for PostgreSQL or RabbitMQ;
  • unbound persistent volume;
  • inaccessible image or incorrect pull policy;
  • invalid cryptographic engine configuration;
  • CPU or memory limits set too low;
  • an incomplete prior migration.

Generic values request 100 mCPU and 128 MiB and limit to 500 mCPU and 512 MiB. Tune per service when you see OOMKills, throttling or slow startups.

Filesystem-based KMS engines and VA local storage use exclusive-access volumes. Keep a single replica or migrate to shared backends before enabling autoscaling.

Issuance fails

Narrow the error following this order:

  1. The CA exists, is active and has not expired.
  2. The CA has a usable private key and the engine is available.
  3. The CSR has a valid signature.
  4. The profile allows the key type and size.
  5. The requested validity fits within the CA's lifetime.
  6. KU, EKU, subject and extensions comply with the profile.

If the reissuance of a CA fails, remember it is not supported for a revoked or already-expired CA. To change its key, follow CA hierarchy and rotation.

EST enrollment fails

Check:

  • the DMS identifier in the EST path;
  • the configured authentication method;
  • the chain presented by the client;
  • the validation CAs and the level limit;
  • the revocation status of the client certificate;
  • the CSR signature;
  • the CA and profile used for issuance;
  • the re-enrollment window.

During re-enrollment, Lamassu first tries to validate with the enrollment CA and then with the additional CAs. During a migration, keep the old CA in that second set until the overlap ends.

OCSP or CRL do not reflect a change

  • Confirm the revocation was persisted and has reason and timestamp.
  • Review VA configuration and logs.
  • Check the CRL's periodicity and next update.
  • Rule out intermediate caches in proxy and client.
  • Verify the certificate points to the endpoint you are querying.
  • Query by the correct serial number and issuer.

Events are not arriving

Domain and audit events depend on the bus publisher:

  • verify it is enabled in the emitting service;
  • check AMQP connectivity and credentials;
  • review exchange, routing key, queues and consumers;
  • inspect dead-letter queues;
  • alert on orphan durable queues and backlog growth.

A consumer outage does not always prevent the main operation from completing. That is why you must monitor RabbitMQ and independently verify the persistence of audit logs.

Observability

The chart can enable OpenTelemetry instrumentation. Traces are exported over OTLP HTTP and logs can be directed to a compatible endpoint, such as VictoriaLogs. Use a correlation identifier and the same time window to join Gateway, authz, service and dependency.

Before closing an incident, record:

  • the confirmed cause;
  • the time scope and affected resources;
  • the commands or queries used;
  • the applied change and how to revert it;
  • validation from a consumer;
  • the preventive action and its owner.

On this page