Search
CLD Tutorial

Kubecost + OpenCost: Cut Kubernetes Costs in 13 Steps [2026]

Yusuf Demir
Yusuf DemirCloud & Software Reporter
22 min read
Kubecost + OpenCost: Cut Kubernetes Costs in 13 Steps [2026]

Cloud bills used to arrive as a single AWS or Azure invoice that a finance team could squint at once a month. That math breaks down completely inside Kubernetes. A single cluster can run hundreds of pods across a dozen namespaces, shared nodes, autoscaling groups, and increasingly, GPU pools for AI inference. When the bill lands, nobody can say which team, which microservice, or which model actually drove the spend. That blind spot is exactly what Kubecost and its open-source engine OpenCost were built to close, and as of October 2026 they are among the most widely used tools for Kubernetes cost visibility.

This tutorial walks through a complete, from-scratch Kubecost and OpenCost deployment: installing the Helm chart, wiring up Prometheus, connecting real cloud billing data from AWS, Azure, or GCP, building namespace- and label-based cost allocation views, setting budget alerts, and exporting the kind of chargeback report a FinOps team can actually use. By the end you will have a working cost dashboard running against your own cluster, plus the troubleshooting knowledge to keep it running when metrics stop flowing or numbers look wrong.

Google · Preferred Sources

Don't miss new tech stories on Google

Add TrendinTech once in the Google app and our stories appear in your news suggestions.

Add Now

Why Kubernetes Cost Monitoring Became Urgent in 2026

Kubernetes cost monitoring is not a new problem, but it got dramatically harder in the last two years. AI inference workloads now live on the same clusters as ordinary microservices, and GPU time does not show up cleanly in a standard cloud bill broken down by service. According to Analytics Insight’s 2026 enterprise IT trends report, 98% of surveyed organizations now actively manage AI-related cloud spending, compared with just 31% two years earlier. That shift pushed FinOps teams to demand pod-level, label-level, and GPU-level cost attribution instead of account-level totals.

OpenCost answers that need as a Cloud Native Computing Foundation (CNCF) project focused specifically on real-time Kubernetes and cloud cost monitoring. Kubecost builds a commercial product and UI on top of the same open cost model, adding governance features like budgets, alerts, and savings recommendations. Matt Ray, community manager for the CNCF OpenCost project, put it simply: “OpenCost is cloud and Kubernetes cost monitoring.” He added that the project “keeps track of how much everything in your Kubernetes cluster costs, and what’s going on in your cloud bills.”

Webb Brown, CEO of Kubecost, described the origin of the open standard behind both tools: “We’ve come together with a group of contributors to build the first open standard or open spec for Kubernetes or container-based cost allocation and cost monitoring.” That shared spec is why Kubecost and OpenCost numbers line up: OpenCost is the free, vendor-neutral cost engine, and Kubecost is built directly on it with a paid tier for multi-cluster and long-term retention.

How OpenCost Actually Calculates Pod Costs

Before trusting any dashboard, it helps to understand the math underneath it. OpenCost’s cost model starts with the hourly price of each node, either pulled from your cloud billing data or estimated from public on-demand rates for that instance type. It then divides that node price across every pod scheduled on it, proportional to each pod’s CPU and memory requests relative to the node’s total allocatable capacity, not actual usage. This matters because two pods with identical CPU usage but different requested limits will be billed differently: the pod that requested more capacity absorbs a larger share of the node’s cost, even if it never consumes what it asked for.

Unallocated node capacity, the portion of a node’s CPU and memory that no pod has requested, gets bucketed as idle cost. By default, OpenCost reports idle cost separately rather than silently spreading it across every running pod, which is why the “Idle” line item in a cluster-wide report is often one of the largest single categories on an underutilized cluster. GPU costs follow a similar allocatable-capacity model, but weighted by GPU request count rather than CPU millicores, and shared infrastructure costs like load balancers or block storage attached to multiple pods get split using a shared-cost ratio you can configure under cluster settings. Grasping this model early saves a lot of confusion later, when a lightly-loaded pod on an expensive, mostly-empty node shows a surprisingly high cost simply because it is absorbing that node’s idle capacity.

Kubecost vs OpenCost: Picking the Right Starting Point

Before installing anything, decide which tool fits your cluster. Both share the same cost-allocation math, so switching later does not throw away your historical understanding of spend, but the deployment paths differ.

FactorOpenCost (CNCF)Kubecost FreeKubecost Enterprise
LicenseApache 2.0, fully open sourceFree tier, proprietary UICommercial
Clusters supportedUnlimited, self-managed1 cluster for full historyUnlimited, multi-cluster
Metrics retentionDepends on your Prometheus setup15 days rolling by defaultUnlimited, long-term storage
Cloud billing integrationManual/API-basedAWS, Azure, GCP built-inAWS, Azure, GCP, plus reserved/committed-use reconciliation
Governance (budgets, alerts)Not includedBasic alertsBudgets, policies, SSO, audit logs
Best forPlatform teams wanting a free, embeddable cost APISmall teams, single clusterMulti-team, multi-cluster FinOps programs

Kubecost’s latest upstream release, version 3.3.0, shipped in late September 2026 and bundles Cluster Controller v0.16.35. Note that the AWS Marketplace listing for IBM Kubecost currently tracks version 3.2.4, so the packaged marketplace build can lag a minor version or two behind the upstream Helm chart. If reproducibility matters for your rollout, deploy from the upstream Helm repository rather than the marketplace listing, and pin the exact chart version in your values file.

Prerequisites and Versions

This walkthrough assumes an existing Kubernetes cluster. It works identically on managed clusters (EKS, AKS, GKE) and self-managed clusters, with small differences noted where the cloud billing integration diverges. Confirm you have the following before starting:

  • A running Kubernetes cluster on version 1.28 or newer (tested here against 1.31)
  • kubectl configured and authenticated against the target cluster
  • Helm 3.14 or newer installed locally
  • Cluster-admin or namespace-admin rights to create a new namespace
  • At least 2 vCPU and 4 GB of free memory available for the cost-analyzer and Prometheus pods combined
  • Kubecost Helm chart targeting the cost-analyzer app version released with the 3.3.0 line (September 2026), or OpenCost Helm chart 1.20.0 if you are going fully open source
  • Optional but recommended: an AWS IAM role with Cost and Usage Report (CUR) access, an Azure Enterprise Agreement or Cost Management API key, or a GCP BigQuery billing export, depending on your cloud
  • Outbound internet access from the cluster to pull container images, unless you are mirroring images internally

You do not need an existing Prometheus deployment. The standard Kubecost Helm chart bundles its own Prometheus instance by default, which is the simplest path for a first install. If your cluster already runs Prometheus for other monitoring, you can point Kubecost at it instead and skip the bundled one, which the troubleshooting section covers later.

Step 1: Create a Dedicated Namespace

Keep cost-monitoring components isolated from application workloads. This makes RBAC scoping, resource quotas, and later troubleshooting far simpler.

kubectl create namespace kubecost
kubectl label namespace kubecost app.kubernetes.io/part-of=finops

The label is optional but useful later when you want OpenCost or Kubecost itself to exclude its own namespace from cost reports, since monitoring infrastructure shouldn’t inflate the numbers you’re trying to measure.

Step 2: Add the Kubecost Helm Repository

Add the official chart repository and refresh your local index so Helm can see the latest release.

helm repo add kubecost https://kubecost.github.io/cost-analyzer/
helm repo update
helm search repo kubecost/cost-analyzer --versions | head -5

You should see the 3.3.x chart line at the top of the output. If you want the fully open-source OpenCost install instead of the Kubecost UI, add the OpenCost chart repository in parallel:

helm repo add opencost https://opencost.github.io/opencost-helm-chart
helm repo update
helm search repo opencost/opencost --versions | head -5

Pin the exact chart version you tested, rather than letting Helm silently pull the newest tag on your next upgrade. This tutorial pins 3.3.0 for Kubecost and 1.20.0 for OpenCost, matching the releases that shipped in September 2026.

Step 3: Install Kubecost with Helm

Run the install command against the namespace you created. The bundled Prometheus and the cost-analyzer pod will both deploy from this single release.

helm install kubecost kubecost/cost-analyzer \
  --namespace kubecost \
  --version 3.3.0 \
  --set kubecostToken="" \
  --set prometheus.server.persistentVolume.size=32Gi \
  --wait --timeout 10m

The kubecostToken field is left blank here because the free tier does not require a token for a single-cluster deployment; enterprise multi-cluster setups will need a token from the Kubecost portal. The prometheus.server.persistentVolume.size flag matters more than it looks: the default volume size is too small for clusters with more than a few dozen nodes, and running out of disk is one of the most common reasons Kubecost dashboards go blank after a few weeks.

Watch the rollout and confirm every pod reaches a Running state before moving on:

kubectl get pods -n kubecost -w

Expect to see pods named roughly kubecost-cost-analyzer, kubecost-prometheus-server, kubecost-kube-state-metrics, and kubecost-prometheus-node-exporter (one per node). The cost-analyzer pod runs an Nginx front end alongside the cost-model backend, which pulls metrics from Prometheus, applies pricing data, and serves the dashboard.

Step 4: Expose and Access the Dashboard

For a first look, port-forward the cost-analyzer service rather than exposing it publicly. Production deployments should sit behind an authenticated ingress instead.

kubectl port-forward --namespace kubecost deployment/kubecost-cost-analyzer 9090:9090

Open http://localhost:9090 in a browser. The first load can take two to five minutes while Prometheus scrapes enough data points to populate the initial charts. If the page loads but every graph is empty, that delay is normal and not a bug, give it another few minutes before troubleshooting.

Step 5: Connect Real Cloud Billing Data

Out of the box, Kubecost and OpenCost estimate costs using public on-demand list prices for your detected node types. That is a reasonable starting point, but it ignores reserved instances, savings plans, committed-use discounts, and spot pricing, all of which can make your real bill 20 to 60% cheaper than list price. Connecting your actual cloud billing account closes that gap.

AWS: Cost and Usage Report Integration

Create an IAM role with read access to an S3 bucket containing your AWS Cost and Usage Report (CUR), then pass the role ARN and bucket details into the Helm values:

helm upgrade kubecost kubecost/cost-analyzer \
  --namespace kubecost \
  --version 3.3.0 \
  --set kubecostProductConfigs.athenaBucketName="s3://your-cur-bucket" \
  --set kubecostProductConfigs.athenaRegion="us-east-1" \
  --set kubecostProductConfigs.athenaDatabase="athenacurcfn_cur_report" \
  --set kubecostProductConfigs.athenaTable="cur_report" \
  --set kubecostProductConfigs.projectID="123456789012"

Replace the bucket name, region, Athena database, table, and account ID with your own values from the CUR setup in the AWS console. Full details on EKS-specific pricing dimensions are documented on the AWS EKS pricing page.

Azure and GCP Integration

Azure integration uses either an Enterprise Agreement billing export or the Cost Management API, configured through a service principal with Cost Management Reader rights. GCP integration points Kubecost at a BigQuery dataset populated by your GCP billing export. Both follow the same pattern as AWS: create a read-only credential scoped to billing data, then reference it in the Helm values under kubecostProductConfigs. Expect this step to take 15 to 30 minutes longer than AWS, mostly spent waiting for the first billing export to populate the BigQuery table or Azure export container.

Step 6: Build Namespace and Label-Based Cost Allocation Views

This is the step that actually answers “who is spending what.” In the dashboard, open the Allocations view and group by namespace first to get a baseline per-team or per-environment breakdown.

# Example: query the cost-analyzer API directly for a 7-day namespace breakdown
curl -s "http://localhost:9090/model/allocation" \
  --data-urlencode "window=7d" \
  --data-urlencode "aggregate=namespace" \
  --data-urlencode "accumulate=true" | jq '.data[0] | keys'

Once namespace-level numbers look right, switch the aggregation to label and filter on a label your teams already apply consistently, such as team, app, or cost-center. This is where most clusters hit their first real gap: cost allocation is only as good as your labeling discipline. If pods are not consistently labeled, Kubecost rolls uncategorized spend into an “__unallocated__” bucket, which is a strong signal to fix your label conventions before trusting the chargeback numbers.

A good practice: require a team and app label on every Deployment, StatefulSet, and Job via an admission policy (OPA Gatekeeper or Kyverno), so new workloads can’t land on the cluster without cost-attribution metadata.

Step 7: Identify Idle and Over-Provisioned Resources

Kubecost’s Savings view surfaces pods that request far more CPU or memory than they actually use. A common pattern: a pod requesting 2 vCPUs but averaging 0.3 vCPU of real usage over a week is wasting roughly 85% of its reserved capacity. Multiply that across a few hundred pods and the waste adds up fast.

curl -s "http://localhost:9090/model/savings/requestSizingRecommendations" \
  --data-urlencode "namespace=payments" | jq '.data[] | {pod: .podName, cpuRequest: .currentCpuRequest, cpuUsed: .cpuUsageAvg}'

Run this per namespace, starting with your highest-spend namespaces from Step 6. Resist the urge to blanket-apply recommended request sizes across the whole cluster in one pass; validate a handful of changes in a staging namespace first, since over-aggressive downsizing can trigger throttling or OOMKills under traffic spikes.

Step 8: Set Up Horizontal Pod Autoscaling to Act on the Data

Cost visibility without a response mechanism just produces a prettier bill. Pair your Kubecost findings with the Kubernetes Horizontal Pod Autoscaler (HPA) so GPU and CPU-heavy workloads scale down automatically during low-traffic windows instead of sitting idle around the clock.

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: inference-service-hpa
  namespace: ml-inference
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: inference-service
  minReplicas: 1
  maxReplicas: 8
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 65

Apply it with kubectl apply -f inference-service-hpa.yaml, then watch the Savings view in Kubecost over the following week to confirm the idle-GPU and idle-CPU flags for that namespace drop. The full HPA reference, including custom and external metrics support for GPU-based scaling, is documented on the official Kubernetes autoscaling page.

Step 9: Configure Budget Alerts

Inside the Kubecost Govern section, create a budget scoped to a namespace, label, or cluster, with a monthly ceiling and an alert threshold. A typical configuration sends a Slack or email alert at 80% of budget, giving a team roughly a week of runway to react before the hard ceiling at 100%.

curl -s -X POST "http://localhost:9090/model/alerts" \
  -H "Content-Type: application/json" \
  -d '{
    "type": "budget",
    "window": "30d",
    "aggregation": "namespace",
    "filter": "payments",
    "threshold": 5000,
    "triggers": [80, 100],
    "notification": "slack-webhook"
  }'

Replace the Slack webhook configuration with your own integration details under the Settings panel before this alert can actually deliver anywhere. An alert rule with no configured notification channel will silently do nothing, which is a common first-week mistake.

Step 10: Export Chargeback Reports for Finance

FinOps teams rarely want to log into the Kubecost UI themselves. Export a scheduled CSV or push allocation data into your existing BI tool via the API instead.

curl -s "http://localhost:9090/model/allocation" \
  --data-urlencode "window=month" \
  --data-urlencode "aggregate=namespace,label:team" \
  --data-urlencode "accumulate=true" \
  -o monthly-chargeback.json

python3 -c "
import json, csv
data = json.load(open('monthly-chargeback.json'))['data'][0]
with open('monthly-chargeback.csv', 'w', newline='') as f:
    writer = csv.writer(f)
    writer.writerow(['allocation', 'totalCost', 'cpuCost', 'ramCost', 'gpuCost'])
    for key, v in data.items():
        writer.writerow([key, v.get('totalCost', 0), v.get('cpuCost', 0), v.get('ramCost', 0), v.get('gpuCost', 0)])
"

Schedule this as a CronJob inside the cluster so the CSV lands in a shared bucket on the first of every month, ready for finance to import without needing cluster access.

Step 11: Add GPU and Carbon Cost Tracking

GPU time is the fastest-growing line item on most 2026 Kubernetes bills, and standard cloud cost reports frequently fail to break it down below the node-pool level. OpenCost and Kubecost both expose GPU-specific cost allocation by reading NVIDIA DCGM or node-level GPU utilization metrics, letting you attribute a shared GPU node pool back to the specific inference workloads consuming it.

OpenCost has also added carbon-cost tracking. Matt Ray described the addition directly: “This week we announced we are now supporting carbon costs, so you can see what the carbon footprint of your Kubernetes workloads is.” Enable it by setting the carbon-cost feature flag in your Helm values and pointing it at your cloud provider’s published carbon-intensity data for the regions your cluster runs in.

helm upgrade kubecost kubecost/cost-analyzer \
  --namespace kubecost \
  --version 3.3.0 \
  --set kubecostModel.carbonCost.enabled=true \
  --reuse-values

Step 12: Secure the Dashboard Before Going to Production

Cost data is sensitive. It reveals team headcount proxies, infrastructure architecture, and which products are getting investment, so treat the Kubecost or OpenCost dashboard with the same care as any internal admin tool rather than leaving it on an open port-forward. Three changes matter most before wider rollout.

First, put the cost-analyzer service behind an ingress with authentication (OAuth2 Proxy, an identity-aware proxy, or basic auth at minimum) instead of relying on kubectl port-forward, which only ever worked as a local debugging shortcut. Second, apply a NetworkPolicy restricting which namespaces can reach the cost-analyzer and Prometheus pods directly, since the underlying Prometheus instance can expose raw cluster metrics beyond just cost data if queried directly. Third, scope the service account Kubecost runs under to read-only ClusterRole permissions; it needs to list pods, nodes, and namespaces across the cluster, but never needs write access to anything.

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: kubecost-restrict-ingress
  namespace: kubecost
spec:
  podSelector: {}
  policyTypes:
    - Ingress
  ingress:
    - from:
        - namespaceSelector:
            matchLabels:
              app.kubernetes.io/part-of: finops-viewers
      ports:
        - protocol: TCP
          port: 9090

This policy restricts inbound access to the cost-analyzer port to only pods in namespaces explicitly labeled as approved viewers, closing off the default behavior of any pod on the cluster being able to reach the dashboard’s internal service address.

Step 13: Run OpenCost Standalone (Fully Open-Source Path)

If governance features, budgets, and multi-cluster support are not a current need, OpenCost alone gives you the same underlying cost model with zero licensing considerations. Install it directly instead of the Kubecost chart:

helm install opencost opencost/opencost \
  --namespace opencost \
  --create-namespace \
  --version 1.20.0 \
  --wait --timeout 10m

kubectl port-forward --namespace opencost service/opencost 9003:9003

OpenCost exposes the same /allocation API shape used throughout this tutorial, so every curl example above works against it with only the port number changed. This makes it straightforward to prototype with OpenCost and migrate to full Kubecost later if you need multi-cluster rollups, since both tools speak the same cost-allocation spec referenced by Webb Brown earlier in this piece. More detail on the project’s scope is available on the CNCF OpenCost project page and the OpenCost documentation site.

Common Pitfalls

These five mistakes account for most of the support questions in Kubecost and OpenCost community channels.

  • Undersized Prometheus storage. The default persistent volume fills up within a few weeks on clusters with more than 50 nodes, silently truncating historical data. Size it for at least 30 days of retention up front.
  • Inconsistent labeling. Cost allocation is only as accurate as your label coverage. Clusters without a team or app label standard end up with large “__unallocated__” buckets that make chargeback reports useless.
  • Skipping the cloud billing integration. Relying on on-demand list pricing when your organization actually runs reserved instances or savings plans can overstate real costs by 30% or more.
  • Deploying in the same namespace as monitored workloads. This inflates the very numbers you are trying to measure and complicates RBAC. Always isolate into a dedicated namespace.
  • Treating marketplace chart versions as current. Cloud marketplace listings (like the AWS Marketplace IBM Kubecost listing) can lag the upstream Helm chart by a minor version or more. Always check the upstream GitHub releases page before assuming you have the latest fixes.

Expected Output: What a Healthy Dashboard Looks Like

After a successful install and a week of data collection, the Allocations view should show a breakdown resembling the table below for a mid-sized cluster running roughly 40 nodes.

NamespaceCPU Cost (7d)RAM Cost (7d)GPU Cost (7d)Total (7d)
ml-inference$412$188$2,340$2,940
payments$890$410$0$1,300
checkout-api$520$260$0$780
kube-system$140$95$0$235
__unallocated__$60$40$0$100

An “__unallocated__” row under 5% of total spend is a reasonable target. If that row is closer to 20 or 30%, go back to Step 6 and tighten your labeling policy before trusting the rest of the report.

Troubleshooting

  • Dashboard loads but all charts are empty. Prometheus typically needs 5 to 15 minutes to accumulate enough scrape intervals for the first graphs to render. Wait and refresh before assuming something is broken.
  • Cost-analyzer pod is CrashLoopBackOff. Check kubectl logs -n kubecost deployment/kubecost-cost-analyzer for an out-of-memory kill; the default resource limits are sometimes too tight for clusters with thousands of pods. Raise the memory limit in your Helm values and redeploy.
  • Numbers don’t match my actual AWS bill. Confirm the CUR integration from Step 5 actually completed; without it, Kubecost falls back to on-demand list pricing and ignores your reserved instance or savings plan discounts.
  • Prometheus disk fills up and data stops flowing. Check the PVC usage with kubectl exec into the Prometheus pod and run df -h. If it’s full, increase the PVC size (most storage classes support online expansion, but confirm yours does before resizing) and reduce the retention window if needed.
  • GPU costs show as zero despite running GPU workloads. This almost always means the NVIDIA device plugin or DCGM exporter isn’t installed on the GPU node pool, so OpenCost has no GPU utilization metric to read. Install the DCGM exporter DaemonSet on GPU nodes first.
  • Budget alerts never fire. Check that a notification channel (Slack webhook, email SMTP relay) is actually configured under Settings; an alert rule with no destination silently does nothing.
  • Helm upgrade hangs at “Waiting.” This is usually a PodDisruptionBudget blocking a rolling update on a single-replica deployment. Temporarily scale to 2 replicas, complete the upgrade, then scale back if needed.
  • Allocation API returns a 401 or 403. If you’ve placed the cost-analyzer service behind an ingress with authentication, make sure your CronJob or BI integration is using a service account token or API key rather than hitting the port-forward address, which only works locally.
  • “__unallocated__” costs are unexpectedly high. Besides missing labels, this can also happen when nodes run system DaemonSets (like logging or security agents) that aren’t tied to a specific workload. Create a dedicated namespace for cluster-wide infrastructure pods so their cost is attributed clearly rather than lumped into “unallocated.”

Advanced Tips: Scaling Beyond a Single Cluster

Once the single-cluster setup is stable, a few advanced moves get you closer to a production-grade FinOps practice:

  • Federate multiple clusters. Kubecost’s enterprise tier supports a central aggregator that rolls up allocation data from every cluster into one view, which matters once you run separate clusters per environment or region.
  • Reconcile with committed-use discounts. Feed your AWS Savings Plans or GCP committed-use contract data back into the cost model so dashboards reflect your real effective hourly rate, not list price.
  • Automate request-sizing recommendations into CI. Pull the requestSizingRecommendations endpoint from Step 7 into a pre-merge check on Helm chart pull requests, so new deployments start with sane resource requests instead of copy-pasted defaults.
  • Layer in Prometheus’s own retention and federation options. For clusters with heavy metric cardinality, review the Prometheus overview documentation for remote-write and federation patterns that keep long-term storage costs under control while still feeding Kubecost accurate data.
  • Tie budgets to your organization’s FinOps maturity model. The FinOps Foundation’s framework, outlined on their introduction to FinOps page, is a useful reference for structuring how engineering, finance, and product teams share ownership of the numbers this dashboard surfaces.

Complete Working Project: Minimal GitOps Deployment

Putting the full setup together as a single, repeatable Helm values file makes the deployment reproducible across environments. Save this as kubecost-values.yaml:

kubecostProductConfigs:
  clusterName: "production-east"
  currencyCode: "USD"
  athenaBucketName: "s3://your-cur-bucket"
  athenaRegion: "us-east-1"
  athenaDatabase: "athenacurcfn_cur_report"
  athenaTable: "cur_report"
  projectID: "123456789012"

prometheus:
  server:
    persistentVolume:
      size: 32Gi
    retention: "35d"

kubecostModel:
  carbonCost:
    enabled: true

networkCosts:
  enabled: true

kubecostAlerts: []  # configured via UI/API in Step 9

Deploy with a single reproducible command referencing this file:

helm upgrade --install kubecost kubecost/cost-analyzer \
  --namespace kubecost --create-namespace \
  --version 3.3.0 \
  -f kubecost-values.yaml \
  --wait --timeout 10m

Commit this values file to the same Git repository as your cluster’s other infrastructure manifests, and let your existing GitOps controller (Argo CD or Flux) manage upgrades the same way it manages every other workload. That removes the temptation to run ad hoc helm upgrade commands from a laptop, which is how version drift between clusters usually starts. If you haven’t installed Helm itself yet, the official Helm installation guide covers every supported platform.

Kubecost vs OpenCost vs Cloud-Native Billing Dashboards

It’s worth comparing this stack against the billing dashboards AWS, Azure, and GCP already provide natively, since some teams wonder if a dedicated tool is even necessary.

CapabilityNative Cloud Billing ConsoleOpenCost / Kubecost
Pod-level cost attributionNoYes
Label/namespace chargebackNoYes
GPU utilization-based cost splitNo, node-pool level onlyYes, per-pod
Real-time (sub-24hr) visibilityNo, billing data lags 24-48hrsYes, near real-time
Idle/over-provisioning recommendationsLimitedYes, built-in Savings view
Multi-cloud single paneNo, each console is siloedYes

Native billing consoles remain useful for reconciling the final invoice, but they were never designed to answer “which pod in which namespace caused this.” That gap is exactly why Kubernetes-native cost tools have moved from a nice-to-have to a standard part of platform engineering toolchains in 2026.

Frequently Asked Questions

Is Kubecost free to use?
Yes, Kubecost offers a free tier that covers a single cluster with 15 days of rolling metrics retention, which is enough for most of the workflow in this tutorial, including allocation views, savings recommendations, and basic alerts. Multi-cluster federation, unlimited historical retention, enterprise budgets with approval workflows, and SSO all require the paid Enterprise tier, priced per node and negotiated directly with Kubecost’s sales team rather than published as a flat rate.

What’s the actual difference between Kubecost and OpenCost?
OpenCost is the free, CNCF-hosted, open-source cost-allocation engine that Webb Brown’s team open-sourced specifically so the underlying cost math would be vendor-neutral and auditable. Kubecost is built on that same spec and engine, then wraps it in a packaged UI, cloud billing integrations for AWS, Azure, and GCP, governance features like budgets and policies, and a commercial support tier. In practice, choosing OpenCost means you accept more manual integration work in exchange for zero licensing dependency; choosing Kubecost trades a bit of vendor lock-in for a faster path to a polished dashboard and multi-cluster rollups.

Does Kubecost require its own Prometheus instance?
Not necessarily. The default Helm chart bundles Prometheus for convenience, which is the fastest path for a first install, but you can point it at an existing Prometheus deployment instead if your cluster already runs one for other monitoring. Reusing an existing Prometheus avoids running two separate time-series databases on the same cluster, though you will need to confirm your existing instance retains the cluster, pod, and node-level metrics Kubecost’s cost model depends on, specifically kube-state-metrics and node-exporter data.

How accurate are the cost numbers without a cloud billing integration?
Reasonably close for on-demand pricing, but they will overstate real spend for any organization using reserved instances, savings plans, or committed-use discounts, sometimes by 30% or more. Connecting your actual billing data (Step 5) closes that gap by reconciling the estimated node price against what you were actually charged, which is also the only way to see the benefit of existing discount commitments reflected in per-namespace or per-team cost reports.

Can Kubecost or OpenCost track GPU costs for AI workloads?
Yes, both tools can attribute GPU spend down to individual pods when the NVIDIA device plugin and DCGM exporter are running on GPU node pools. Without those exporters installed, GPU costs will show as zero even on GPU-backed workloads, which is one of the most common “why isn’t this working” questions in the community, since the dashboard gives no obvious error when the exporter is simply missing. Once DCGM metrics are flowing, GPU cost attribution follows the same allocatable-capacity model described earlier in this tutorial, just weighted by GPU request count instead of CPU millicores.

How long does a full production deployment take?
The basic Helm install takes about 10 minutes. Connecting cloud billing data typically adds 30 to 60 minutes depending on your cloud provider, with GCP and Azure usually taking longer than AWS because of billing export propagation delays. Tuning labels, budgets, and alerts for a real multi-team organization is usually a multi-week rollout rather than a single sitting, mostly because getting consistent labeling adopted across every team’s Helm charts and manifests takes longer than any of the technical steps.

What happens if my cluster’s labeling is inconsistent?
Costs that can’t be mapped to a known label or namespace fall into an “__unallocated__” bucket. A small unallocated percentage, generally under 5-10%, is normal and often reflects legitimate shared infrastructure costs. Anything above roughly 15-20% signals a labeling policy gap worth fixing, ideally by enforcing required labels through an admission controller like OPA Gatekeeper or Kyverno, before trusting chargeback reports enough to put them in front of a finance team.

Does this setup work on managed Kubernetes services like EKS, AKS, and GKE?
Yes. The Helm install steps are identical across managed and self-managed clusters; only the cloud billing integration step (Step 5) differs based on which provider’s billing export or API you’re connecting. One managed-service nuance worth knowing: EKS, AKS, and GKE all charge a per-cluster control-plane fee on top of node costs, and Kubecost’s allocation model generally attributes that control-plane fee at the cluster level rather than trying to split it across namespaces, since no single workload is responsible for it.

Should a small team with one cluster bother with any of this, or is it overkill?
If a single cluster runs under a dozen services and one team owns all of it, basic namespace-level allocation from OpenCost alone is probably enough, and the governance features in Kubecost Enterprise would be solving a problem that doesn’t exist yet. The moment a second team, a second cluster, or a GPU-backed inference workload enters the picture, label-based chargeback and budget alerts earn their setup cost quickly, usually within the first month of catching one over-provisioned deployment or one forgotten GPU node pool left running overnight.

Related Coverage

Yusuf Demir

Yusuf Demir

Cloud & Software Reporter

Yusuf Demir is the Cloud & Software Reporter at TrendinTech, where he covers cloud infrastructure, enterprise platforms, developer tools and the digital transformation of businesses in the UK and the United States. He previously reported on enterprise technology for The Register in London and covered the cloud and SaaS beat for TechCrunch, following the competition between AWS, Microsoft Azure and Google Cloud, the open source licensing disputes and the rise of Kubernetes. Yusuf holds an MEng in Computing from Imperial College London and speaks regularly at KubeCon and AWS re:Invent, where he moderates conversations with engineers and chief technology officers. He is most interested in the gap between vendor roadmaps and the systems engineers actually run, and in what the cloud bill looks like once the free credits expire.

All stories by Yusuf Demir (327)

Related Articles