Terraform + Kubernetes 1.37: Multi-Cloud in 14 Steps [2026]
![Terraform + Kubernetes 1.37: Multi-Cloud in 14 Steps [2026]](https://trendintech.com/wp-content/uploads/2026/10/terraform-kubernetes-1-37-multi-cloud-tutorial-2026-gen.webp)
Kubernetes 1.37 shipped its first patch, 1.37.1, on September 15, 2026, and the clock is already running out on 1.34, which enters full end-of-life on October 27, 2026. If your clusters are still on 1.34, you have about two weeks left before the upstream project stops shipping security fixes for that line. At the same time, Terraform 1.16.5 landed on September 30, 2026, and the 1.17 beta line is previewing a general-availability Terraform Policy engine plus a new -minimal-refresh planning mode. Put those two timelines together and October 2026 is a genuinely awkward moment to be running Kubernetes by hand across three clouds.
This tutorial walks through provisioning matched Kubernetes 1.37 clusters on AWS, Azure, and GCP with a single Terraform codebase, upgrading an existing cluster safely, and turning on the new 1.37 features that actually change how your workloads behave: beta Memory QoS, PVC last-used tracking, and the storage security additions to emptyDir. Everything below uses Terraform 1.16.5 (the current stable release as of this writing), kubectl 1.37.1, and provider versions that were live on the Terraform Registry in the first week of October 2026. By the end you will have a working multi-cloud module structure you can drop into your own repository.
Don't miss new tech stories on Google
Add TrendinTech once in the Google app and our stories appear in your news suggestions.
Why Terraform and Kubernetes 1.37 Together, Right Now
Kubernetes keeps showing up as the default answer for running containers at scale, and the 2025 Stack Overflow Developer Survey put it at 28.5% adoption among developers working on cloud infrastructure, with Terraform close behind at 17.8%. Those numbers only tell part of the story. The Cloud Native Computing Foundation’s annual survey found Kubernetes already running in production for 66% of potential or actual cloud-native users, and multi-cloud setups in place at 56% of organizations. Running the same workload on two or three clouds is not a rare edge case anymore, it is closer to the median setup at any company past a certain size.
The problem is that provisioning three clouds by hand means three sets of console clicks, three sets of drift, and three ways for a cluster to quietly fall out of sync with the others. Terraform solves the provisioning half of that problem by giving you one HCL codebase that targets the AWS, Azure, and Google providers at once. Kubernetes 1.37 solves the other half by shipping scheduling and storage behavior that used to require hand-rolled workarounds, now built into the control plane as beta features enabled by default. Combining them is less about novelty and more about catching up with where both projects already are in October 2026.
There are usually three reasons a team ends up here instead of staying on a single cloud. The first is a contractual one: enterprise customers in regulated industries increasingly ask vendors to demonstrate they are not fully dependent on a single cloud provider, and a working multi-cloud deployment is the concrete evidence that satisfies procurement reviews. The second is cost arbitrage, since spot and reserved pricing shifts between AWS, Azure, and GCP often enough that workloads with flexible scheduling can meaningfully cut compute spend by running wherever capacity is cheapest that quarter. The third, and the one this tutorial is built around, is simple resilience: a regional outage on one provider should not take down a service that otherwise has no technical reason to live on only one cloud. None of these reasons require an exotic setup, they just require the provisioning layer to treat all three clouds as interchangeable targets instead of three separate codebases maintained by three separate people.
Prerequisites: Exact Versions You Need Before Starting
Version mismatches are the single biggest source of confusing errors in a multi-cloud Terraform setup, so pin everything before you write a line of HCL. Here is the baseline this tutorial assumes:
- Terraform core 1.16.5 (released September 30, 2026) for the main walkthrough, or Terraform 1.17.0-beta2 if you want to try the new
-minimal-refreshflag and provider-requirement variables early - kubectl 1.37.1, matched to the server version you are deploying, since the Kubernetes project only guarantees compatibility within one minor version of skew
- hashicorp/aws provider 6.68.0 (released October 7, 2026) or newer
- hashicorp/azurerm provider pulled after October 8, 2026, to get the newest list resources and data sources
- hashicorp/google provider, any release from the active 6.x line that supports
google_container_clusterwith release channel pinning - AWS CLI v2, Azure CLI 2.6x or newer, and the Google Cloud SDK, each authenticated against a non-production account first
- A terminal with at least 4 GB of free memory for local plan/apply runs, since Terraform state refresh on three clouds at once can get memory-hungry
You do not need a Kubernetes cluster running yet, Terraform will create them. You do need existing AWS, Azure, and GCP accounts with billing enabled, since managed Kubernetes control planes are not free on any of the three. Budget roughly $0.10 per hour per cluster for the control plane fee on AWS and GCP, with Azure AKS waiving that fee on its free tier.
Step 1: Install and Pin Your Terraform Version
Start by locking the Terraform version at the repository level so nobody on the team accidentally applies with a newer or older binary. Create a versions.tf file before anything else.
terraform {
required_version = "~> 1.16.5"
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 6.68.0"
}
azurerm = {
source = "hashicorp/azurerm"
version = "~> 4.0"
}
google = {
source = "hashicorp/google"
version = "~> 6.0"
}
}
}
Run terraform version to confirm your local binary matches. If it does not, use tfenv install 1.16.5 or download the binary directly from the official Terraform releases page. Terraform 1.16.4 added import blocks that work inside modules and a new terraform_data resource with a store block, both of which this tutorial’s module structure relies on later, so do not skip below 1.16.4 even if you are not chasing the absolute latest patch.
Step 2: Configure Provider Credentials for All Three Clouds
Keep credentials out of your HCL files entirely. Use environment variables or a secrets manager, never hardcoded keys in providers.tf. A typical local setup looks like this before you run any Terraform command.
# AWS
export AWS_ACCESS_KEY_ID="..."
export AWS_SECRET_ACCESS_KEY="..."
export AWS_DEFAULT_REGION="us-east-1"
# Azure (service principal)
export ARM_CLIENT_ID="..."
export ARM_CLIENT_SECRET="..."
export ARM_SUBSCRIPTION_ID="..."
export ARM_TENANT_ID="..."
# Google Cloud
export GOOGLE_APPLICATION_CREDENTIALS="$HOME/.gcp/terraform-sa.json"
export GOOGLE_PROJECT="your-gcp-project-id"
Then declare the three provider blocks in providers.tf, each scoped to its own region or location variable so the same module set can be reused for staging and production without duplicating code.
provider "aws" {
region = var.aws_region
}
provider "azurerm" {
features {}
subscription_id = var.azure_subscription_id
}
provider "google" {
project = var.gcp_project
region = var.gcp_region
}
Step 3: Design the Repository Structure Before Writing Resources
A multi-cloud Kubernetes repository gets unmanageable fast if every resource lives in one flat main.tf. Split each cloud into its own module and keep a thin root module that wires variables through. This is the structure used for the rest of this tutorial and for the complete working project at the end.
multicloud-k8s/
├── versions.tf
├── providers.tf
├── variables.tf
├── outputs.tf
├── terraform.tfvars
├── modules/
│ ├── eks/
│ │ ├── main.tf
│ │ ├── variables.tf
│ │ └── outputs.tf
│ ├── aks/
│ │ ├── main.tf
│ │ ├── variables.tf
│ │ └── outputs.tf
│ └── gke/
│ ├── main.tf
│ ├── variables.tf
│ └── outputs.tf
└── kubeconfig/
└── merge.sh
Each cloud module is self-contained, which means you can apply just one of them during testing with terraform apply -target=module.eks without touching the other two. That matters a lot when you are iterating on one cloud’s node pool configuration and do not want to trigger a plan against Azure or GCP every time.
Step 4: Provision the AWS EKS Cluster on Kubernetes 1.37
Inside modules/eks/main.tf, pin the cluster version explicitly to 1.37 rather than leaving it on “latest,” since EKS control plane versions roll out to each region on its own schedule and an unpinned version can produce a different cluster than you tested locally.
resource "aws_eks_cluster" "this" {
name = var.cluster_name
role_arn = aws_iam_role.eks_cluster.arn
version = "1.37"
vpc_config {
subnet_ids = var.subnet_ids
endpoint_private_access = true
endpoint_public_access = true
}
access_config {
authentication_mode = "API_AND_CONFIG_MAP"
}
}
resource "aws_eks_node_group" "default" {
cluster_name = aws_eks_cluster.this.name
node_group_name = "default-pool"
node_role_arn = aws_iam_role.eks_node.arn
subnet_ids = var.subnet_ids
instance_types = ["m6i.large"]
scaling_config {
desired_size = 3
max_size = 6
min_size = 2
}
}
Run terraform plan -target=module.eks first. A clean plan against a fresh account typically proposes somewhere between 12 and 18 resources once IAM roles, the VPC, and node groups are counted, so do not be alarmed by a long resource list on the first run.
Step 5: Provision the Azure AKS Cluster
AKS uses a slightly different pattern, with the Kubernetes version set on the cluster resource directly and node pools attached separately. Azure’s AKS release channel setting matters here, set it to patch so Azure applies security patches automatically without bumping your minor version out from under you.
resource "azurerm_kubernetes_cluster" "this" {
name = var.cluster_name
location = var.location
resource_group_name = azurerm_resource_group.this.name
dns_prefix = var.dns_prefix
kubernetes_version = "1.37"
automatic_upgrade_channel = "patch"
default_node_pool {
name = "default"
node_count = 3
vm_size = "Standard_D2s_v5"
}
identity {
type = "SystemAssigned"
}
}
One thing worth checking before you apply: not every Azure region offers Kubernetes 1.37 on day one of the upstream release, since managed providers typically validate a new minor version against their own control plane image before exposing it. If terraform apply fails with an unsupported version error, check which versions are currently available in your target region with az aks get-versions --location <region> -o table rather than assuming the Terraform provider is wrong.
Step 6: Provision the Google GKE Cluster
GKE’s release channel model is the most opinionated of the three. Instead of pinning an exact patch version, you generally pick a channel (rapid, regular, or stable) and let Google manage the patch cadence, though you can still pin a minimum minor version inside that channel.
resource "google_container_cluster" "this" {
name = var.cluster_name
location = var.gcp_region
release_channel {
channel = "regular"
}
min_master_version = "1.37"
remove_default_node_pool = true
initial_node_count = 1
}
resource "google_container_node_pool" "default" {
name = "default-pool"
cluster = google_container_cluster.this.name
location = var.gcp_region
node_count = 3
node_config {
machine_type = "e2-standard-4"
}
}
Setting remove_default_node_pool = true and defining your own node pool resource is a common pattern because the auto-created default pool cannot be customized as flexibly as one you declare yourself. Skipping this step is one of the pitfalls covered later in this guide.
Step 7: Apply All Three Clusters and Merge Kubeconfig
With all three modules wired into the root main.tf, run a full plan before applying anything in production.
terraform init -upgrade
terraform plan -out=multicloud.tfplan
terraform apply multicloud.tfplan
Expect this apply to take between 12 and 20 minutes, since EKS and GKE control plane creation typically runs 10 to 15 minutes on its own, and AKS tends to finish a few minutes faster. Once all three report success, merge the generated kubeconfig entries so you can switch contexts with one command.
aws eks update-kubeconfig --name $CLUSTER_NAME --region $AWS_REGION
az aks get-credentials --resource-group $RG_NAME --name $CLUSTER_NAME
gcloud container clusters get-credentials $CLUSTER_NAME --region $GCP_REGION
kubectl config get-contexts
kubectl config use-context $EKS_CONTEXT
kubectl get nodes -o wide
A healthy kubectl get nodes output on a freshly provisioned 1.37 cluster looks like this:
NAME STATUS ROLES AGE VERSION
ip-10-0-1-23.ec2.internal Ready <none> 4m v1.37.1
ip-10-0-1-45.ec2.internal Ready <none> 4m v1.37.1
ip-10-0-1-67.ec2.internal Ready <none> 4m v1.37.1
If the VERSION column shows anything other than a 1.37.x patch, the node pool bootstrapped against a stale AMI or image family and needs a forced node pool replacement, not a Terraform re-apply.
Step 8: Turn On and Verify Kubernetes 1.37’s Beta Memory QoS
Memory QoS graduated to beta in Kubernetes 1.37 and ships enabled by default, which means you likely do not need to flip a feature gate at all on a fresh 1.37 cluster. What you do need to do is confirm it is actually active on your nodes, since managed providers occasionally lag the upstream default by a patch release or two while they validate the feature against their own kubelet image.
kubectl get --raw /api/v1/nodes/$NODE_NAME/proxy/configz | grep -i memoryqos
# If you need to confirm feature-gate status directly on the node
ssh $NODE_NAME "cat /var/lib/kubelet/config.yaml | grep -A2 featureGates"
If MemoryQoS is reported as disabled on a node you expected to have it on by default, the fix is almost always a kubelet config override, not a Terraform change. Set it explicitly in your node pool’s kubelet config block and roll the node pool, which is covered in the troubleshooting section further down.
Step 9: Configure Storage Security and PVC Tracking
Two more Kubernetes 1.37 features matter for anyone running stateful workloads. Tracking of when a PersistentVolumeClaim was last used graduated to beta and ships enabled by default, giving you a built-in signal for identifying abandoned volumes instead of writing a custom cron job to check timestamps. The storage security additions let you set permission modes and bind-mount options directly on emptyDir volumes, which closes a long-standing gap where temporary volumes had looser defaults than persistent ones.
apiVersion: v1
kind: Pod
metadata:
name: secure-scratch-pod
spec:
containers:
- name: app
image: your-registry/app:1.0
volumeMounts:
- name: scratch
mountPath: /tmp/scratch
volumes:
- name: scratch
emptyDir:
sizeLimit: 2Gi
Apply it with kubectl apply -f secure-scratch-pod.yaml and check kubectl describe pod secure-scratch-pod for the mount options reported under the volume section. On a 1.37 cluster you should see the permission mode reflected without any additional admission webhook, which is new behavior compared to 1.36 and earlier.
Step 10: Automate Drift Detection Across All Three Clouds
Once three clusters are live, manual console changes on any one of them will silently drift your Terraform state. Schedule a read-only plan on a cron job or CI pipeline, not an apply, so you catch drift before it compounds.
#!/bin/bash
# drift-check.sh — run on a schedule, never auto-apply
terraform plan -detailed-exitcode -out=drift.tfplan
EXIT_CODE=$?
if [ $EXIT_CODE -eq 2 ]; then
echo "DRIFT DETECTED — review drift.tfplan before anyone applies"
exit 1
elif [ $EXIT_CODE -eq 0 ]; then
echo "No drift"
fi
The -detailed-exitcode flag is what makes this usable in CI: exit code 0 means no changes, 1 means an error, and 2 means Terraform found a difference between state and reality. Wire that exit code into a Slack or email alert rather than letting a drift-detection job fail silently in a CI dashboard nobody checks.
Step 11: Upgrade an Existing Cluster to 1.37 Safely
If you are not starting from scratch, upgrading in place is riskier than provisioning new. The safest pattern is a blue-green node pool swap: create a new 1.37 node pool alongside the existing one, cordon and drain the old pool, then delete it once workloads have rescheduled cleanly.
# 1. Add a new node pool resource at 1.37 in Terraform, apply only that addition
terraform apply -target=aws_eks_node_group.pool_v137
# 2. Cordon the old pool so no new pods schedule there
kubectl cordon -l eks.amazonaws.com/nodegroup=default-pool
# 3. Drain gradually, respecting PodDisruptionBudgets
kubectl drain -l eks.amazonaws.com/nodegroup=default-pool \
--ignore-daemonsets --delete-emptydir-data --timeout=300s
# 4. Once empty, remove the old node group from Terraform and apply
terraform apply -target=aws_eks_node_group.default
Upgrade the control plane itself before the node pools, since Kubernetes only supports nodes running up to two minor versions behind the control plane, never ahead of it. If you manage the control plane with kubeadm rather than a managed service, the official kubeadm upgrade documentation has the exact command sequence for upgrading the control plane one minor version at a time, which you cannot skip even to get to 1.37 faster.
Remember that Kubernetes 1.34 enters full end-of-life on October 27, 2026, with 1.34.12 as its final patch released September 15. Clusters still on 1.34 after that date stop receiving CVE fixes from upstream, which also means your managed provider’s own security patching for that line typically winds down on a similar timeline. If you run a CVE remediation pipeline as part of your patch process, the CVE patch pipeline built around Nuclei and the CISA KEV catalog is a reasonable template for layering vulnerability scanning on top of this Kubernetes upgrade cadence.
Step 12: Lock Down IAM and Network Policy Across All Three Clusters
A multi-cloud cluster that is reachable but not locked down is worse than a single-cloud cluster, since you now have three attack surfaces to patch instead of one. Start with least-privilege IAM: the Terraform service account or service principal you used to provision each cluster should not be the same identity your CI pipeline uses to deploy workloads into it. Separating provisioning credentials from deployment credentials means a compromised CI token cannot be used to tear down the cluster itself, only to deploy into namespaces it already has access to.
Inside the cluster, apply a default-deny NetworkPolicy to every namespace before deploying workloads, then open only the specific paths your services actually need. This single policy, applied consistently across all three clouds, closes the most common lateral-movement path in a freshly provisioned cluster where every pod can otherwise talk to every other pod by default.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny-all
namespace: production
spec:
podSelector: {}
policyTypes:
- Ingress
- Egress
---
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-app-to-db
namespace: production
spec:
podSelector:
matchLabels:
app: api-server
policyTypes:
- Egress
egress:
- to:
- podSelector:
matchLabels:
app: postgres
ports:
- protocol: TCP
port: 5432
Apply this to each of the three clusters with the same kubectl context loop used earlier for the smoke test, and verify with kubectl get networkpolicy -n production that all three report the same two policies. A cluster where only one or two clouds picked up the policy is a sign the apply step failed silently on one context, usually because the active kubectl context was not what you expected when you ran the command.
For secrets, avoid storing them directly as Kubernetes Secret objects base64-encoded in etcd without an external secrets backend. Each cloud has its own native option (AWS Secrets Manager, Azure Key Vault, Google Secret Manager) and the External Secrets Operator can sync all three into the same Kubernetes-native Secret format, so your application manifests do not need to know which cloud they are running on to find their credentials.
Kubernetes Release Support Window as of October 2026
| Version | Latest Patch | Patch Release Date | Next Scheduled Patch | Status |
|---|---|---|---|---|
| 1.37 | 1.37.1 | 2026-09-15 | 1.37.2 (target 2026-10-13) | Actively Supported |
| 1.36 | 1.36.5 | 2026-09-15 | 1.36.6 (target 2026-10-13) | Actively Supported |
| 1.35 | 1.35.9 | 2026-09-15 | 1.35.10 (target 2026-10-13) | Actively Supported |
| 1.34 | 1.34.12 | 2026-09-15 | None planned | Maintenance mode, EOL 2026-10-27 |
Source: Kubernetes patch release schedule and the individual Kubernetes 1.37 release page, both current as of October 10, 2026.
Terraform and Provider Versions Used in This Tutorial
| Component | Version | Release Date | Why It Matters Here |
|---|---|---|---|
| Terraform core (stable) | 1.16.5 | 2026-09-30 | Baseline used throughout this tutorial |
| Terraform core (beta) | 1.17.0-beta2 | 2026-09-23 | Adds -minimal-refresh and provider-requirement variables |
| hashicorp/aws | 6.68.0 | 2026-10-07 | EKS resource support used in Step 4 |
| hashicorp/azurerm | Oct 2026 build | 2026-10-08 | New list resources and data sources for AKS and related services |
| kubectl | 1.37.1 | 2026-09-15 | Matched to the cluster server version to avoid skew errors |
Full changelogs for each of these are on the Terraform GitHub releases page and the hashicorp/aws provider documentation on the Terraform Registry.
Complete Working Project: Full Repository Layout
Putting every step together, the root main.tf that ties the three cloud modules into one apply looks like this. This is the file you run terraform apply against once all three module directories from Step 3 are filled in with the resources from Steps 4 through 6.
module "eks" {
source = "./modules/eks"
cluster_name = "${var.project_name}-eks"
subnet_ids = module.aws_network.subnet_ids
}
module "aks" {
source = "./modules/aks"
cluster_name = "${var.project_name}-aks"
location = var.azure_location
dns_prefix = var.project_name
}
module "gke" {
source = "./modules/gke"
cluster_name = "${var.project_name}-gke"
gcp_region = var.gcp_region
}
output "eks_endpoint" {
value = module.eks.cluster_endpoint
}
output "aks_endpoint" {
value = module.aks.cluster_endpoint
}
output "gke_endpoint" {
value = module.gke.cluster_endpoint
}
Checked into a repository with the module structure from Step 3, this gives you a single terraform apply that stands up matched Kubernetes 1.37 clusters across all three clouds, with each cloud’s networking, IAM, and node pool logic isolated in its own module so a change to the GKE node pool size never touches the AWS or Azure plan.
Advanced Tips: Terraform Policy, Minimal-Refresh, and Cost Control
Terraform 1.17’s release candidate line makes Terraform Policy generally available, removing the need for the -allow-experimental-features flag that earlier adopters had to pass. If you have been holding off on policy-as-code because it felt experimental, the GA milestone is the signal to actually adopt it, since the flag requirement going away usually means HashiCorp considers the feature’s API stable going forward.
The new -minimal-refresh planning option in the 1.17 beta line is worth testing on large multi-cloud state files specifically because refreshing three clouds’ worth of resources on every plan is often the slowest part of the workflow. Minimal-refresh skips re-reading resources Terraform has high confidence are unchanged, which can meaningfully cut plan time on a repository with 100-plus resources spread across three providers.
On the cost side, remember that EKS and GKE both charge a per-cluster control plane fee while AKS does not on its free tier, so running three test clusters around the clock for a proof of concept costs more on AWS and Google than on Azure purely from the control plane line item, before you even count compute. Tear down proof-of-concept clusters with terraform destroy -target=module.eks (scoped per module) rather than leaving all three running between work sessions.
Finally, keep an eye on the Kubernetes project blog, which published detailed posts on the shift to cgroup v2 and on scaling workloads with node swap on October 6 and October 5, 2026 respectively. Both change how memory-constrained nodes behave under pressure, and both interact directly with the Memory QoS behavior you turned on in Step 8.
Common Pitfalls When Automating Multi-Cloud Kubernetes with Terraform
- Leaving the GKE default node pool in place. Forgetting
remove_default_node_pool = trueleaves an unmanaged, uncustomizable node pool running alongside the one you actually configured, quietly doubling your node count and your bill. - Applying all three clouds in one plan during initial testing. A typo in the Azure module will block your entire apply, including the AWS and GCP resources that were ready to go. Use
-targetper module while iterating, and only run the combined apply once each module is individually verified. - Pinning the Kubernetes version on only the control plane, not the node pool. EKS and AKS node groups can silently default to a different image version than the control plane, producing a cluster where
kubectl get nodesreports a mismatched VERSION column. - Skipping the control-plane-before-nodes upgrade order. Kubernetes nodes cannot run a newer minor version than the control plane. Upgrading node pools first breaks the API compatibility contract and produces scheduling failures that look unrelated to the actual cause.
- Not accounting for regional version availability. A brand-new Kubernetes minor version does not land in every cloud region simultaneously. Always check regional version availability before writing a version number into Terraform, rather than assuming the provider API will simply reject an unavailable version cleanly (some fail with unrelated-looking errors instead).
- Running drift-detection plans with auto-apply enabled. A scheduled job that both plans and applies without human review can reverse legitimate emergency changes an on-call engineer made directly in the console during an incident.
- Forgetting that Terraform state for three clouds in one backend is a single point of failure. Use a remote backend with locking (S3 plus DynamoDB, Azure Storage with blob leasing, or Terraform Cloud) so two engineers running applies at the same time against three clouds cannot corrupt each other’s state.
Troubleshooting Terraform and Kubernetes 1.37 Errors
| Symptom | Likely Cause | Fix |
|---|---|---|
| “version not supported” on AKS apply | Kubernetes 1.37 not yet available in target Azure region | Run az aks get-versions --location <region> and pick an available patch, or change region |
| kubectl reports version skew warning | Local kubectl is more than one minor version ahead or behind the cluster | Install kubectl 1.37.1 to match, or use kubectl version --client to confirm before connecting |
| Terraform apply hangs on EKS node group creation | Subnets lack available IP addresses or NAT gateway misconfigured | Check subnet CIDR sizing and confirm outbound internet access for node bootstrap |
| GKE cluster created but default node pool unexpectedly persists | remove_default_node_pool omitted or set to false | Add the flag, re-apply; may require manual deletion of the orphaned pool first |
| MemoryQoS reports disabled despite 1.37 upgrade | Managed provider’s kubelet image lags the upstream default | Explicitly set the feature gate in the node pool’s kubelet config and roll the pool |
| terraform plan shows unexpected diff on every run | A cloud console change created configuration drift | Run terraform plan -detailed-exitcode, review the diff, then terraform apply to reconcile or terraform import if the resource was created outside Terraform |
| Azure provider authentication failure mid-apply | Service principal secret expired | Rotate the ARM_CLIENT_SECRET and re-export before retrying |
| Node pool drain never completes | PodDisruptionBudget blocking eviction of the last replica | Temporarily scale the deployment up by one replica or adjust the PDB’s minAvailable before draining |
| Terraform state lock stuck after a crashed apply | Previous run was interrupted (Ctrl+C, CI timeout) before releasing the lock | Confirm no other apply is actually running, then use terraform force-unlock <LOCK_ID> |
Validating the Whole Setup End to End
Before calling the migration done, run the same smoke test against all three clusters to confirm behavior is actually consistent, not just that three clusters exist.
for ctx in $EKS_CONTEXT $AKS_CONTEXT $GKE_CONTEXT; do
echo "=== $ctx ==="
kubectl --context=$ctx get nodes -o custom-columns=NAME:.metadata.name,VERSION:.status.nodeInfo.kubeletVersion
kubectl --context=$ctx run smoke-test --image=busybox --rm -it --restart=Never -- echo "ok"
done
Every context should report a 1.37.x kubelet version and a successful “ok” from the smoke test pod. If one cloud reports “ok” slower than the other two by more than a few seconds, that is usually a signal the node pool is still warming up rather than a real failure, so re-run once before treating it as a bug.
Where This Fits Against Serverless Alternatives
Not every workload needs a Kubernetes cluster on three clouds. If what you are actually running is a handful of event-driven functions rather than long-lived services, it is worth comparing the operational overhead of this setup against a serverless approach. The breakdown of AWS Lambda, Azure Functions, and Google Cloud Run pricing is a useful reference point for deciding whether multi-cloud Kubernetes is actually solving a problem you have, or just a pattern you inherited from a larger team’s architecture.
For teams that do need the control Kubernetes provides, the module structure in this tutorial scales reasonably well up to a few dozen clusters before you will want to introduce a proper platform layer like Crossplane or a GitOps controller on top of it. That is a natural next step once the Terraform-only approach here starts to feel like it is fighting you on every new cluster.
Frequently Asked Questions
Do I need Kubernetes 1.37 specifically, or will 1.36 work for this tutorial?
The Terraform patterns in this tutorial work against 1.36 as well, since the provisioning logic does not change between minor versions. The Memory QoS, PVC tracking, and storage security steps are specific to 1.37, however, since those features graduated to beta in that release. If you are on 1.36, those three steps will not apply until you upgrade.
Is Terraform 1.17 safe to use in production yet?
As of October 10, 2026, 1.17 is still in its beta and release-candidate line, not a stable release. Use 1.16.5 for production work and reserve 1.17 betas for testing the new Terraform Policy GA and minimal-refresh features in a non-production environment first.
Can I use this same module structure for more than three clouds?
Yes. The pattern of one module per cloud with a thin root module wiring them together extends to any number of providers. The main constraint is state file size and plan time, which is exactly what the troubleshooting and advanced-tips sections above address with scoped applies and the minimal-refresh option.
What happens if I do not upgrade off Kubernetes 1.34 before October 27, 2026?
Your cluster keeps running, nothing shuts off automatically. What stops is upstream security patching for that minor version. Any CVE discovered in the Kubernetes codebase after end-of-life will not get a 1.34 backport, leaving you to either upgrade reactively under pressure or accept the exposure.
Why does my Azure cluster show a different available Kubernetes version than AWS or GCP?
Each managed Kubernetes provider validates a new upstream minor version against its own control plane image before making it selectable, and that validation does not happen on the same schedule across AWS, Azure, and Google. Check each provider’s own version-availability command (covered in Steps 4 through 6) rather than assuming all three expose 1.37 at the same time.
Does enabling Memory QoS change existing pod behavior without any action on my part?
Since it is a beta feature enabled by default in 1.37, yes, it can affect memory-constrained pods immediately after upgrade without any YAML changes from you. Review your workloads’ memory requests and limits before upgrading production clusters, since Memory QoS changes how the kernel enforces those boundaries under pressure.
What is the minimum team size where multi-cloud Kubernetes with Terraform actually makes sense?
There is no hard number, but the operational overhead of managing three cloud providers’ worth of IAM, networking, and node pool quirks tends to only pay off once you have a dedicated platform or infrastructure function, rather than a single generalist engineer handling it alongside feature work. For smaller teams, the serverless comparison linked earlier in this guide is often the more realistic starting point.
Should I use Terraform modules from the public registry instead of writing my own?
Public modules like terraform-aws-modules/eks can save time on boilerplate, but they also add an external dependency with its own release cadence to track alongside the Kubernetes and Terraform versions covered in this tutorial. Writing the thin, self-contained modules shown here gives you more direct control over exactly which Kubernetes version and feature flags are set, which matters more in a multi-cloud setup where consistency across three providers is the whole point.
Related Coverage
Yusuf Demir
Yusuf Demir is the Cloud & Software Reporter at TrendinTech, where he covers cloud infrastructure, enterprise platforms, developer tools and the digital transformation of businesses in the UK and the United States. He previously reported on enterprise technology for The Register in London and covered the cloud and SaaS beat for TechCrunch, following the competition between AWS, Microsoft Azure and Google Cloud, the open source licensing disputes and the rise of Kubernetes. Yusuf holds an MEng in Computing from Imperial College London and speaks regularly at KubeCon and AWS re:Invent, where he moderates conversations with engineers and chief technology officers. He is most interested in the gap between vendor roadmaps and the systems engineers actually run, and in what the cloud bill looks like once the free credits expire.
All stories by Yusuf Demir (296)