ClickHouse
GitGuardian uses ClickHouse, an open-source column-oriented database, to power features that require fast analytics over large volumes of data (for example, machine and endpoint secret occurrence analytics). PostgreSQL remains the primary datastore for the rest of the application; ClickHouse is only involved for the specific features that need it.
Starting with release 2027.1.0, ClickHouse and its object storage backend become mandatory for every self-hosted deployment. See the customer notice for what changes and how to prepare.
Deployment model
ClickHouse runs as a dedicated, in-cluster stateful component, and relies on external object storage (S3-compatible, Azure Blob Storage, or Google Cloud Storage) for durable data storage. It is not a fully externally managed database like PostgreSQL or Redis. See ClickHouse storage for details.
Default topology
ClickHouse is deployed as a single node. This is the simplest and lightest-footprint option, and is sufficient for the vast majority of deployments, though it is not highly available: there is no other replica to fail over to.
Supported version
GitGuardian pins the ClickHouse version shipped with each release to the same version running our SaaS offering, and validates it against our test suite before release. Do not attempt to run a different ClickHouse version than the one shipped with your release.
Storage
ClickHouse splits its storage into three parts: local metadata and filesystem cache volumes, and object storage holding the actual table data, where the bulk of the storage footprint lives and where clickhouse-backup writes to when backing up.
See ClickHouse storage for the local volumes and provider-specific object storage configuration, Sizing and hardening for hardware sizing, Backup and restore for the backup runbook, and Monitoring and alerting to catch a failing backup that pod health won't show you.
Required configuration
A production-ready ClickHouse deployment requires all of the following to be configured:
| What to configure | Details | Where |
|---|---|---|
| Compute sizing and hardening | Node & instance resources, scheduling, node stability | Sizing and hardening |
| ClickHouse credentials | Instance password | Prerequisites |
| Object storage configuration | S3 / Azure Blob / GCS backend, disk authentication | Storage → Configuration |
| In-cluster volumes configuration | Metadata & cache volumes, storage class | Storage → Local volumes |
| Backup configuration | Destination bucket, credentials secret, schedule & retention | Backup and restore |
Prerequisites
Before installing, the ClickHouse admin password and the backup sidecar's local API password should be passed via a Kubernetes secret rather than set inline in values, the same pattern used throughout GitGuardian self-hosted. It can be created directly or synced from your own secret store. See Helm sensitive information management for how to populate it.
If you create it directly:
kubectl create secret generic clickhouse-credentials \
--namespace <namespace> \
--from-literal=admin-password=<a-strong-password> \
--from-literal=api-password=<a-strong-password> \
--from-literal=api-username=<a-username>
clickhouse-credentials, admin-password, api-password, and api-username are wired into the chart's defaults under those exact names: admin-password via clickhouse.auth.existingSecret/existingSecretKey (GitGuardian derives its own CLICKHOUSE_PASSWORD from the same secret), and api-password/api-username consumed directly by the backup sidecar and its ServiceMonitor. Missing or renaming any of the three breaks ClickHouse's own authentication, the backup sidecar, or GitGuardian's connection to it, unless you also override the corresponding values.
With the secret in place, enable ClickHouse itself:
clickhouse:
enabled: true
This alone isn't enough for a working deployment: ClickHouse's object storage backend has no default and must be configured for your provider (see ClickHouse storage), and the backup destination needs the same before backups can run (see Backup and restore). Follow both pages before considering ClickHouse installed.
Complete configuration example
The examples below assemble the pieces from every configuration page of this section into one clickhouse values block.
All providers also require the clickhouse-credentials secret described in Prerequisites above.
- AWS S3
- Azure Blob Storage
- GCS
- MinIO
Assumes the IRSA setup from Storage → Configuration: an IAM OIDC provider on the cluster, and an IAM role trusted for the clickhouse ServiceAccount with a policy covering both buckets.
clickhouse:
enabled: true
# Sizing and hardening (see Sizing → Hardware recommendations)
resources: # Large tier shown; match your own tier, requests = limits
requests:
cpu: 8
memory: 32Gi
limits:
cpu: 8
memory: 32Gi
podAnnotations:
karpenter.sh/do-not-disrupt: 'true' # adjust for your cluster autoscaler
nodeSelector: # assumes a dedicated on-demand node pool (see Sizing → Node stability)
workload: clickhouse
tolerations:
- key: workload
operator: Equal
value: clickhouse
effect: NoSchedule
# Local volumes (see Storage → Local persistent volumes)
persistence:
storageClass: gp3
cache:
storageClass: gp3
size: 100Gi # Large tier shown (see Sizing → Filesystem cache)
# Data bucket, authenticated through IRSA (see Storage → Configuration)
serviceAccount:
annotations:
eks.amazonaws.com/role-arn: 'arn:aws:iam::<account-id>:role/<role-name>'
objectStorage:
provider: s3
s3:
endpoint: 'https://<bucket-name>.s3.<region>.amazonaws.com'
bucket: '<bucket-name>'
region: '<region>'
credentialless: true
# Backup destination, a separate bucket (see Backup and restore → Object storage destination)
backup:
objectStorage:
provider: s3
s3:
bucket: '<backup-bucket-name>'
region: '<region>'
prefix: 'backups'
credentialless: true
Assumes the Microsoft Entra Workload ID setup from Storage → Configuration: OIDC issuer and Workload Identity enabled on the cluster, a managed identity with a federated credential for the clickhouse ServiceAccount, and an RBAC role assignment covering both containers.
clickhouse:
enabled: true
# Sizing and hardening (see Sizing → Hardware recommendations)
resources: # Large tier shown; match your own tier, requests = limits
requests:
cpu: 8
memory: 32Gi
limits:
cpu: 8
memory: 32Gi
podAnnotations:
cluster-autoscaler.kubernetes.io/safe-to-evict: 'false' # adjust for your cluster autoscaler
nodeSelector: # assumes a dedicated on-demand node pool (see Sizing → Node stability)
workload: clickhouse
tolerations:
- key: workload
operator: Equal
value: clickhouse
effect: NoSchedule
# Local volumes (see Storage → Local persistent volumes)
persistence:
storageClass: managed-csi-premium
cache:
storageClass: managed-csi-premium
size: 100Gi # Large tier shown (see Sizing → Filesystem cache)
# Data container, authenticated through Workload ID (see Storage → Configuration)
serviceAccount:
annotations:
azure.workload.identity/client-id: '<identity-client-id>'
podLabels:
azure.workload.identity/use: 'true'
objectStorage:
provider: azblob
azblob:
storageAccountUrl: 'https://<storage-account-name>.blob.core.windows.net'
containerName: '<container-name>'
credentialless: true
# Backup destination, a separate container (see Backup and restore → Object storage destination)
backup:
objectStorage:
provider: azblob
azblob:
storageAccountName: '<storage-account-name>'
containerName: '<backup-container-name>'
prefix: 'backups'
credentialless: true
Assumes the two credential secrets from Storage → Configuration and Backup and restore: an HMAC key pair for ClickHouse's own disk (clickhouse-gcs-credentials), and a native service account JSON key for the backup sidecar (clickhouse-backup-gcs-credentials), which cannot use HMAC.
clickhouse:
enabled: true
# Sizing and hardening (see Sizing → Hardware recommendations)
resources: # Large tier shown; match your own tier, requests = limits
requests:
cpu: 8
memory: 32Gi
limits:
cpu: 8
memory: 32Gi
podAnnotations:
cluster-autoscaler.kubernetes.io/safe-to-evict: 'false' # adjust for your cluster autoscaler
nodeSelector: # assumes a dedicated on-demand node pool (see Sizing → Node stability)
workload: clickhouse
tolerations:
- key: workload
operator: Equal
value: clickhouse
effect: NoSchedule
# Local volumes (see Storage → Local persistent volumes)
persistence:
storageClass: premium-rwo
cache:
storageClass: premium-rwo
size: 100Gi # Large tier shown (see Sizing → Filesystem cache)
# Data bucket, via GCS's S3-compatible endpoint and an HMAC key pair (see Storage → Configuration)
objectStorage:
provider: gcs
gcs:
bucket: '<bucket-name>'
existingSecret: clickhouse-gcs-credentials
# Backup destination, a separate bucket and its own credential type (see Backup and restore)
backup:
objectStorage:
provider: gcs
gcs:
bucket: '<backup-bucket-name>'
prefix: 'backups'
existingSecret: 'clickhouse-backup-gcs-credentials'
Assumes the static credential secrets from Storage → Configuration and Backup and restore (access-key-id/secret-access-key keys; a single MinIO key pair can back both).
clickhouse:
enabled: true
# Sizing and hardening (see Sizing → Hardware recommendations)
resources: # Large tier shown; match your own tier, requests = limits
requests:
cpu: 8
memory: 32Gi
limits:
cpu: 8
memory: 32Gi
podAnnotations:
cluster-autoscaler.kubernetes.io/safe-to-evict: 'false' # adjust for your cluster autoscaler
nodeSelector: # assumes a dedicated on-demand node pool (see Sizing → Node stability)
workload: clickhouse
tolerations:
- key: workload
operator: Equal
value: clickhouse
effect: NoSchedule
# Local volumes (see Storage → Local persistent volumes)
persistence:
storageClass: '<ssd-backed-storage-class>'
cache:
storageClass: '<ssd-backed-storage-class>'
size: 100Gi # Large tier shown (see Sizing → Filesystem cache)
# Data bucket, path-style S3 with static keys (see Storage → Configuration)
objectStorage:
provider: s3
s3:
endpoint: 'http://<minio-host>:<minio-port>'
bucket: '<bucket-name>'
region: 'us-east-1' # MinIO ignores the value but the field is required
forcePathStyle: true
existingSecret: clickhouse-minio-credentials
# Backup destination, a separate bucket (see Backup and restore → Object storage destination)
backup:
objectStorage:
provider: s3
s3:
endpoint: 'http://<minio-host>:<minio-port>'
bucket: '<backup-bucket-name>'
region: 'us-east-1' # required by the S3 client even though MinIO ignores it
prefix: 'backups'
forcePathStyle: true
existingSecret: 'clickhouse-backup-minio-credentials'
config:
extraVars:
S3_DISABLE_SSL: 'true' # omit and use an https:// endpoint above instead if MinIO terminates TLS