Kubeflow (train / evaluate / register) ਨੂੰ Argo CD (GitOps deliver / rollback) ਨਾਲ ਜੋੜਨ ਦੀ ਅਮਲੀ ਗਾਈਡ. ਪੂਰੀ architecture, manifests ਅਤੇ anti-patterns ਲਈ ਪੜ੍ਹੋ ਲੰਮਾ article.
InferenceService, pipeline definition ਅਤੇ platform component Git ਤੋਂ reconcile ਹੁੰਦਾ ਹੈ. Kubeflow ਵਿੱਚ train ਤੇ evaluate ਕਰੋ; Git manifest ਅੱਪਡੇਟ ਕਰਕੇ promote ਕਰੋ; roll back ਕਰੋ git revert. Models ਨੂੰ object storage ਵਿੱਚ ਰੱਖੋ — Git URIs ਅਤੇ config ਸਟੋਰ ਕਰਦਾ ਹੈ, weights ਨਹੀਂ.
ਦੋਵੇਂ tools ਕਿਉਂ?
Kubeflow ਜਵਾਬ ਦਿੰਦਾ ਹੈ: GPUs ਉੱਤੇ reproducible training ਅਤੇ evaluation ਕਿਵੇਂ ਚਲਾਈਏ? Argo CD ਜਵਾਬ ਦਿੰਦਾ ਹੈ: resulting model ਨੂੰ snowflake clusters ਤੋਂ ਬਿਨਾਂ production ਵਿੱਚ ਕਿਵੇਂ ਭੇਜੀਏ? ਸਿਰਫ਼ Kubeflow ਵਰਤਣਾ ਅਕਸਰ manual ਉੱਤੇ ਖਤਮ ਹੁੰਦਾ ਹੈ kubectl apply serving ਲਈ. ਸਿਰਫ਼ Argo CD ਵਰਤਣ ਨਾਲ ਤੁਹਾਡੇ ਕੋਲ first-class ML pipeline ਅਤੇ experiment system ਨਹੀਂ ਰਹਿੰਦਾ. ਇਕੱਠੇ ਉਹ ਇੱਕ ਬੰਦ MLOps loop ਬਣਾਉਂਦੇ ਹਨ.
ਮਿਹਨਤ ਦੀ ਵੰਡ
| ਚਿੰਤਾ | ਟੂਲ | Git ਵਿੱਚ ਕੀ ਰਹਿੰਦਾ ਹੈ |
|---|---|---|
| Pipeline DAGs, TrainJobs, experiments | Kubeflow Pipelines / Trainer | Pipeline YAML / component specs |
| Model artifacts ਅਤੇ lineage | Model Registry + object storage | URI + metadata — weights file ਨਹੀਂ |
| Serving, canaries, rollbacks | KServe + Argo CD | InferenceService overlays |
| Platform install ਅਤੇ upgrades | Argo CD Applications | Kubeflow ਖੁਦ ਲਈ Helm/Kustomize |
End-to-end loop (ਚੰਗਾ ਕਿਵੇਂ ਲੱਗਦਾ ਹੈ)
- Data scientist pipeline ਬਦਲਾਅ merge ਕਰਦਾ ਹੈ → Argo CD pipeline CRDs sync ਕਰਦਾ ਹੈ.
- Kubeflow GPU nodes ਉੱਤੇ training ਚਲਾਉਂਦਾ ਹੈ (ਵਿਕਲਪਿਕ ਤੌਰ ਤੇ Kueue ਨਾਲ queued).
- Evaluation gates ਪਾਸ → model URI registered; ਇੱਕ PR prod ਅੱਪਡੇਟ ਕਰਦਾ ਹੈ
InferenceServicestorageUri. - Argo CD OutOfSync ਪਛਾਣਦਾ ਹੈ ਅਤੇ rollout ਕਰਦਾ ਹੈ (ਪਹਿਲਾਂ canary traffic).
- Drift ਜਾਂ ਖਰਾਬ metrics →
git revertਪਿਛਲਾ URI restore ਕਰਦਾ ਹੈ; Argo CD reconcile ਕਰਦਾ ਹੈ.
ਪੰਜ practices ਜੋ ਦਰਦ ਬਚਾਉਂਦੀਆਂ ਹਨ
- ਵੱਖਰੇ repos ਜਾਂ folders: platform (Kubeflow install) ਬਨਾਮ apps (pipelines + InferenceServices).
- Model binaries ਕਦੇ commit ਨਾ ਕਰੋ — ਹਵਾਲਾ
s3:///gs:/// PVC URIs. - Sync waves: KFP components ਤੋਂ ਪਹਿਲਾਂ MySQL/MinIO (ਜਾਂ managed equivalents).
- ਹਰ step ਉੱਤੇ resource limits — ਬਿਨਾਂ limits ਦੇ training jobs cluster ਨੂੰ ਭੁੱਖਾ ਛੱਡ ਦੇਣਗੀਆਂ.
- Promotion = PR, notebook ਵਿੱਚ ਬਟਨ ਨਹੀਂ. Audit trail ਮੁਫ਼ਤ ਮਿਲਦਾ ਹੈ.
ਜਦੋਂ ਇਹ stack overkill ਹੋਵੇ
SME chatbot ਲਈ Ollama ਚਲਾਉਂਦਾ ਇੱਕਲਾ Mac Studio Kubeflow + Argo CD ਨਹੀਂ ਚਾਹੁੰਦਾ. ਇਸ pattern ਨੂੰ ਉਦੋਂ ਅਪਣਾਓ ਜਦੋਂ ਕਈ models ਹੋਣ, GPU contention ਹੋਵੇ, compliance ਲੋੜਾਂ ਹੋਣ, ਜਾਂ ਇੱਕ ਤੋਂ ਵੱਧ environments (dev/staging/prod) ਜੋ ਇੱਕੋ ਜਿਹੇ ਰਹਿਣੇ ਚਾਹੀਦੇ ਹਨ.
ਲੰਮਾ ਸੰਸਕਰਣ ਪੜ੍ਹੋ
The ਲੰਮਾ article 2026 Kubernetes MLOps map (KFP v2, Trainer, KServe, Kueue), Argo CD Application patterns, promotion gates, canary serving, GPU FinOps ਅਤੇ rollout checklist ਕਵਰ ਕਰਦਾ ਹੈ. ਪ੍ਰਕਾਸ਼ਿਤ Workstation.