Skip to main content

ClickHouse storage

ClickHouse's storage splits into three parts:

  • Object storage: the bucket you provide (S3-compatible, Azure Blob Storage, or GCS), holding the actual table data. This is where the bulk of the storage footprint lives.
  • Filesystem cache: a local persistent volume acting as an LRU cache in front of object storage, for fast reads of recently accessed data (the same pattern used by ClickHouse Cloud).
  • Metadata: a separate local persistent volume holding table structure and part manifests for the data held in object storage.

This page covers the two local volumes first, then the object storage backends and their configuration.

Local persistent volumes​

The metadata and cache volumes are both provisioned locally, independently of your object storage backend, and kept apart from each other: as shown in the diagram above, ClickHouse writes to both, but only the cache ever talks to object storage.

Metadata volume​

The chart already provisions the metadata volume via clickhouse.persistence (enabled, 20Gi, ReadWriteOnce), mounted at ClickHouse's data directory, where it holds table structure and part manifests for the data held in object storage. In your values you only set the storage class (the chart leaves it unset, so your cluster's default StorageClass applies, which may not be SSD-backed) and, if needed, adjust the size:

clickhouse:
persistence:
storageClass: gp3 # set an SSD/NVMe-backed class, see Storage class below
size: 20Gi # optional; defaults to 20Gi, see Metadata volume sizing below

The size is illustrative, see Metadata volume for sizing guidance; unlike the cache, its footprint barely grows with data volume.

Filesystem cache​

Provisioned natively via clickhouse.cache, as a second, separate volume mounted at /cache, kept apart from the metadata volume so a full cache doesn't starve the metadata disk (and vice versa). The chart wires this volume into the object storage disk config for you (see Configuration below): you only set the storage class and size:

clickhouse:
cache:
storageClass: gp3 # adapt to your provider, use an SSD/NVMe-backed class, see Storage class below
size: 50Gi # must stay above cacheMaxSize (see Filesystem cache below); pairs with the chart's auto-derived default

The size above is illustrative, see Filesystem cache for sizing guidance: it only needs to hold your working set, not your full dataset. clickhouse.serverConfig.cacheMaxSize (the soft limit ClickHouse enforces on top of this volume) auto-derives from cache.size minus a safety margin; see Filesystem cache to override it.

Storage class​

Use the same class of storage for both volumes: high IOPS is required for metadata as well as the cache:

ClassExamplesNotes
Network-attached SSD (recommended)AWS EBS gp3 / io2, Azure Premium SSD (Premium_LRS), GCP pd-ssd / pd-balancedSurvives node loss or reschedule: the volume follows the pod to a new node.
Local NVMeAWS instance store (i3 / i4i instance types), an on-prem local-volume StorageClassFastest option, but pinned to the node it was provisioned on, see the caution below.
Not supportedHDD-class (st1 / sc1, pd-standard, Standard HDD), network filesystems (NFS, EFS, Azure Files, Filestore)Latency is incompatible with ClickHouse's read/write patterns.
Local NVMe pins the pod to its node

A local-path StorageClass ties the volume to the specific node it was provisioned on: unlike network-attached storage, it cannot simply follow the pod to a different node. If your cluster autoscaler consolidates or drains that node, the pod cannot be rescheduled without losing the volume. If you use local NVMe, prevent your autoscaler from disrupting the node (see Node stability), and plan to restore from backup rather than expect a simple reschedule if the node is lost anyway.

Object storage backends​

ClickHouse stores its table data on an object storage bucket that you provide, configured under clickhouse.objectStorage. Each provider below has a single supported authentication path: credential-less (credentialless: true) on AWS S3 and Azure (no static keys to create, store, or rotate), and static access-key/secret-key credentials on GCS, IBM Cloud Object Storage and MinIO.

Bucket requirements
  • Two buckets, one per purpose. Give ClickHouse a bucket of its own, and give the backup sidecar a genuinely separate one.
  • Server-side encryption on both. It is transparent to ClickHouse and the sidecar, so nothing changes in the chart. AWS S3, Azure Blob Storage, and GCS encrypt at rest by default; MinIO needs its own setup (KES backed by a KMS). With customer-managed keys (SSE-KMS and equivalents), the identity ClickHouse and the sidecar authenticate with must also be allowed to use the key: on AWS, grant kms:Decrypt and kms:GenerateDataKey on top of the S3 policy shown in the AWS S3 tab.
  • Outbound access. If your cluster restricts egress, ClickHouse and the sidecar need to reach the object storage endpoint, and in a credential-less setup the provider's token exchange endpoint too: AWS STS (sts.amazonaws.com or the regional endpoint) for IRSA, Microsoft Entra ID (login.microsoftonline.com) for Workload Identity.

Bucket lifecycle policies​

ClickHouse's own data bucket and the backup sidecar's bucket need different, in places opposite lifecycle rules and protections. Both are covered here. Read the subsection for each bucket you're configuring.

ClickHouse data bucket​

Don't apply a lifecycle rule to the bucket (or prefix) backing a ClickHouse disk that transitions or expires objects: ClickHouse expects every object it wrote to remain immediately readable indefinitely. A transition to an archive storage tier (S3 Glacier/Deep Archive and equivalents, including S3 Intelligent-Tiering's optional Archive Access tiers) makes the object unreadable until restored, which can make ClickHouse fail to start or throw errors reading existing parts; an expiration rule deletes data ClickHouse still expects to find, risking data loss. Only ClickHouse itself (via TTL / DROP PARTITION) or an informed administrator should ever remove these objects.

The one lifecycle rule that is safe: aborting incomplete multipart uploads after a few days. This only cleans up orphaned upload fragments left behind by interrupted uploads (never live data), and prevents them from accumulating unnoticed in your bucket.

Don't add versioning or Object Lock to this bucket either

ClickHouse deletes and rewrites objects continuously as part of normal operation (merges, TTL, DROP PARTITION). Neither adds a real disaster-recovery benefit here (that's what the backup bucket is for), and both work against normal operation instead: versioning piles up delete markers with no purge plan of their own, and Object Lock blocks ClickHouse's own legitimate deletes outright, not just accidental ones.

Backup bucket​

The same core rule as the data bucket applies here too, for the same reason: don't let a lifecycle rule transition or expire objects that clickhouse-backup still manages. It tracks its own retention (backupsToKeepRemote) and chain dependencies (rebaseBeforeRemoveOldRemote, see Tunable settings) directly, and expects every backup inside that window to stay immediately readable and to be deleted on its own terms, not S3's. An external expiration rule deletes backups it still expects to find; a transition to an archive tier can break a rebase's server-side copy from a still-referenced older chain. The same multipart-abort rule above is safe here too.

Beyond that baseline, this bucket exists specifically to survive a disaster that takes out ClickHouse's own bucket, so hardening it further is worth the extra cost and complexity that isn't justified on the data bucket:

  • Versioning: recommended. Defense in depth against a compromised credential or a bug deleting backups outright, on top of clickhouse-backup's own retention logic.
  • Object Lock (Governance or Compliance mode): worth it if you have a ransomware or compliance requirement to defend against, but set the retention period shorter than your effective backupsToKeepRemote window. Otherwise it blocks clickhouse-backup's own routine pruning too, not just malicious deletes. We haven't verified how it handles a blocked delete, so the safe design goal is that the two windows never actually collide.
  • Cross-region or cross-account replication: recommended for a real disaster-recovery posture. A separate bucket already protects against a shared misconfiguration or credential; replicating it further protects against losing the account or region it lives in.
  • Archive-tier transition: only safe for backups you've explicitly exported outside of clickhouse-backup's own tracked path (a manual copy to a bucket or prefix it doesn't manage). Never transition objects inside the bucket or prefix it actively manages: the same read-availability assumption that rules this out on the data bucket applies here.

Configuration​

On AWS EKS, ClickHouse authenticates to S3 using IAM Roles for Service Accounts (IRSA): there is no access key to create, store, or rotate. This requires:

  1. An IAM OIDC identity provider registered for your EKS cluster (most EKS clusters already have one for other workloads); see AWS's IAM roles for service accounts documentation if you need to set one up.
  2. A dedicated Kubernetes ServiceAccount for ClickHouse, created by the chart, annotated with an IAM role ARN.
  3. An IAM role that trusts your cluster's OIDC provider, scoped to that ServiceAccount.
  4. A permission policy granting only the S3 access ClickHouse and its backup sidecar need.

Chart configuration: create a dedicated ServiceAccount for ClickHouse and annotate it with the role ARN:

clickhouse:
serviceAccount:
annotations:
eks.amazonaws.com/role-arn: 'arn:aws:iam::<account-id>:role/<role-name>'

Trust policy: scoped to the ServiceAccount above (replace <account-id>, <oidc-provider-url>, and <namespace>):

{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam::<account-id>:oidc-provider/<oidc-provider-url>"
},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"<oidc-provider-url>:sub": "system:serviceaccount:<namespace>:clickhouse",
"<oidc-provider-url>:aud": "sts.amazonaws.com"
}
}
}
]
}

Permission policy: scoped to the bucket dedicated to ClickHouse:

{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Action": ["s3:ListBucket", "s3:GetBucketLocation", "s3:ListBucketMultipartUploads"],
"Resource": [
"arn:aws:s3:::<bucket-name>",
"arn:aws:s3:::<backup-bucket-name>"
]
},
{
"Effect": "Allow",
"Action": [
"s3:GetObject",
"s3:PutObject",
"s3:DeleteObject",
"s3:AbortMultipartUpload",
"s3:ListMultipartUploadParts"
],
"Resource": [
"arn:aws:s3:::<bucket-name>/*",
"arn:aws:s3:::<backup-bucket-name>/*"
]
}
]
}
tip

Scope the policy to ClickHouse's own bucket and, since IRSA gives a single role to the whole pod, the separate backup bucket too (see Bucket requirements). <bucket-name> must match objectStorage.s3.bucket below; <backup-bucket-name> is the one you'll set in the backup sidecar's own destination config (see Backup and restore).

With the ServiceAccount in place, set credentialless: true in your values. There are no static credentials and no secret to create:

clickhouse:
objectStorage:
provider: s3
s3:
endpoint: 'https://<bucket-name>.s3.<region>.amazonaws.com'
bucket: '<bucket-name>'
region: '<region>'
credentialless: true

Target bucket: s3.bucket, independent from the bucket the backup sidecar writes to. s3.endpoint must already embed the bucket as a subdomain, matching AWS's own virtual-hosted-style URLs (the default forcePathStyle: false); see the MinIO tab if your S3-compatible store needs a bucket-less, path-style endpoint instead.

Credentials​

AWS S3 and Azure support credential-less authentication (credentialless: true, backed by IRSA and Microsoft Entra Workload ID respectively, see the tabs above). GCS, IBM Cloud Object Storage and MinIO are the exceptions: ClickHouse's own disk always needs a static access-key/secret-key pair, referenced via objectStorage.<provider>.existingSecret, regardless of cluster identity setup. That static credential is passed the same way as other sensitive GitGuardian configuration, see Helm sensitive information management.