Loading blogs...
Posts tagged Kubernetes.
Pair Kubeflow and Argo CD for MLOps on Kubernetes: train and register models in Kubeflow, promote InferenceServices with GitOps, roll back with git revert.

What are Large Language Models and how do they work? A clear, non-technical explainer for managers and engineers — tokens, embeddings, transformers, training and inference — plus production Kubernetes YAML to deploy your own LLM with Ollama and vLLM.
A support-friendly, SRE-ready overview of Couchbase Server rolling upgrades under the Couchbase Autonomous Operator (CAO), using spec.paused to gate swap-rebalance one node at a time—so you can stabilize, validate health, and keep rollback options open.
Complete tour of NebulaDNS features: API-first control plane, Prometheus metrics, propagation gates, peer fingerprinting, DNSSEC roadmap, Helm/k3s operator story, and how to integrate with AWS Route 53 and CoreDNS.
NebulaDNS is a safe-Rust authoritative DNS server with an API-first control plane, Prometheus metrics, propagation gates, and Kubernetes-native operations—built for visibility when AI services and agents depend on DNS more than ever.
Professional overview of NebulaCR: Rust OCI registry with nebula-registry and nebula-auth, zero-trust OIDC for humans and CI, SCIM lifecycle, Prometheus and OpenTelemetry, multi-region replication, pull-through cache, and Kubernetes Helm deployment—synthesised from the upstream docs/ and README.
How NebulaCB, the open-source Couchbase mission control from nebulacb.org, helps enterprises upgrade fearlessly, validate XDCR and data integrity, run Kubernetes-native operations, and use local AI (Ollama) for RCA—so teams ship changes without losing data or sleep.
Complete guide to running Couchbase Server in high availability production environments with XDCR multi-region replication, Couchbase Autonomous Operator on Kubernetes, across AWS, Azure, GCP, and bare metal k3s with Rancher.
Battle-tested database backup strategies for MySQL and PostgreSQL in production. Covers physical and logical backups, PITR, WAL archiving, automated scheduling, cloud storage, and disaster recovery across AWS, Azure, GCP, and k3s.
Complete production guide for Longhorn distributed storage on k3s/Rancher clusters. Covers dynamic PVC expansion, multi-region replication, backup to S3/GCS/Azure, disaster recovery, and performance tuning for database workloads.
Complete production guide for MongoDB high availability using Replica Sets, Sharding, MongoDB Community Kubernetes Operator, and Percona Operator across AWS, Azure, GCP, and bare metal k3s with Rancher.
Complete guide to running MySQL in high availability production environments using InnoDB Cluster, Group Replication, MySQL Operator, and Percona XtraDB Cluster across AWS, Azure, GCP, and bare metal k3s with Rancher.
Production guide for deploying highly available PostgreSQL using Crunchy Data PGO (Postgres Operator) across AWS, Azure, GCP, and bare metal k3s/Rancher with comprehensive backup, monitoring, and multi-region strategies.