rezachalak.site
02 / Services

Services

Six areas where I do the most useful work. Most engagements combine a few of them — the goal is always the same: fewer manual steps between a commit and production, and fewer surprises after it.

01

Kubernetes Platform Engineering

Clusters that survive contact with production — on bare metal with Talos, or managed on EKS and AKS.

  • Bare-metal and air-gapped clusters: Talos, RKE, K3s, Kubespray
  • Networking and ingress: Cilium, MetalLB, cert-manager, mTLS
  • Storage and stateful workloads: Longhorn, NetApp Trident, CNPG, Kafka/Strimzi
  • Migrations off Docker Swarm, systemd, and VM-based deployments
02

GitOps & CI/CD

One source of truth for every environment, so releases stop depending on who is at the keyboard.

  • ArgoCD and FluxCD delivery, including ad-hoc preview environments
  • GitLab CI and GitHub Actions pipelines with Helm and Terraform
  • Private registries and artifact supply chains: Harbor, Nexus
  • Secrets handling with External Secrets Operator and Keycloak-backed auth
03

Cloud & Cost Optimization

Right-sized infrastructure on AWS, Azure, and Hetzner — backed by load-test data, not guesswork.

  • Node pool sizing, lifecycle policies, and workload placement
  • Terraform-managed landing zones and reproducible environments
  • Lift-and-reshape migrations from on-premises to EKS and AKS
  • Track record: 30% off an AWS bill without cutting capacity
04

Observability & Reliability

Know what broke before the customer tells you, and know how to get it back.

  • Prometheus and Grafana stacks, OpenTelemetry instrumentation
  • Log pipelines with Fluentd and the ELK stack
  • Disaster recovery plans with defined RTO/RPO targets and tested restores
  • Backup and restore automation with Velero and custom tooling
05

Air-gapped & Regulated Environments

Infrastructure for sites with no internet, strict audits, and no room for improvisation.

  • Offline registries, mirrored dependencies, and reproducible installers
  • 300+ air-gapped servers managed with Ansible playbooks and roles
  • pfSense firewalls, network segmentation, and internal PKI with SmallStep
  • ISO 27001 implementation support alongside security teams
06

AI Agents in Production

Getting agents past the demo and onto the same release train as everything else.

  • Agent workloads packaged, deployed, and released through GitOps
  • RAG and orchestration with LangChain, LangGraph, and ADK
  • Self-hosted inference and data paths that stay inside your network
  • Observability and cost controls for model-backed services

Ways to work together

how we start
01

Build

A platform from scratch — cluster, networking, storage, delivery pipeline, and the runbooks to operate it.

02

Migrate

Move what already runs: Swarm to Kubernetes, VMs to EKS/AKS, or a cloud bill to something defensible.

03

Review

A focused audit of an existing setup — reliability gaps, cost leaks, supply-chain risk — with a prioritized plan.

Not sure which one you need?

Describe the setup and what hurts about it. If it isn't something I should take on, I'll say so and point you somewhere better.