Opleiding: AI Workloads on Kubernetes (English) (Virtual)
Lesmethode :
Virtueel
Algemeen :
In this hands-on, two-day Kubernetes course, you learn how to prepare and operate a Kubernetes platform for GPU-accelerated AI workloads. AI and machine learning workloads need scalable compute, GPU acceleration and flexible resource management. You see how GPUs and other accelerators are made available to Kubernetes, how AI applications request and share these resources, and how Kubernetes schedules workloads based on their requirements.
The course focuses on the infrastructure and platform engineering side of AI. You work with GPU-enabled Kubernetes nodes and make them available through the NVIDIA GPU Operator, device plugins and GPU runtime. You then manage those resources with requests and limits, GPU sharing, Dynamic Resource Allocation, ResourceClaims, workload priorities and quotas. On top of that you scale workloads with the Horizontal Pod Autoscaler and KEDA, serve Large Language Models with vLLM, run AI and ML pipelines with Kubeflow, and monitor everything with Prometheus and Grafana.
Theory and practice alternate throughout the course. Every topic is followed by hands-on exercises in your own Kubernetes environment, where you deploy, schedule, scale, monitor and troubleshoot AI workloads. You do not train machine learning models yourself; instead, you learn how to provide a Kubernetes environment in which data scientists, ML engineers and AI applications can run their workloads reliably.
Doel :
After this AI Workloads on Kubernetes course, you can prepare and operate a Kubernetes platform for GPU-accelerated AI workloads, and deploy, schedule, scale and monitor these workloads reliably. You also know how technologies such as the NVIDIA GPU Operator, Dynamic Resource Allocation, KEDA, vLLM and Kubeflow fit into your Kubernetes environment.
Doelgroep :
Engineers who run or support AI and machine learning workloads on Kubernetes, such as DevOps Engineers, Platform Engineers, Kubernetes Administrators, Cloud Engineers and Infrastructure Engineers. The course focuses on the platform and operational side of AI workloads and is not a machine learning or data science course.
Voorkennis :
The following prior knowledge is required:
- Kubernetes Fundamentals
- Experience with AI or machine learning is a plus; experience training machine learning models is not required.
Onderwerpen :
- Introduction to AI workloads on Kubernetes
+nbsp;+nbsp;+nbsp;- AI platform architecture
+nbsp;+nbsp;+nbsp;- CNCF AI platform architecture and ecosystem
+nbsp;+nbsp;+nbsp;- CPU versus GPU workloads
- GPU-enabled Kubernetes nodes
+nbsp;+nbsp;+nbsp;- From hardware to Kubernetes
+nbsp;+nbsp;+nbsp;- NVIDIA GPU Operator
+nbsp;+nbsp;+nbsp;- NVIDIA device plugins and GPU runtime
- Deploying applications with GPU requirements
+nbsp;+nbsp;+nbsp;- Deploying GPU workloads
+nbsp;+nbsp;+nbsp;- Kubernetes resource requests and limits
- GPU resource management
+nbsp;+nbsp;+nbsp;- GPU sharing
+nbsp;+nbsp;+nbsp;- Dynamic Resource Allocation (DRA)
+nbsp;+nbsp;+nbsp;- ResourceClaims
+nbsp;+nbsp;+nbsp;- Workload priorities and quotas
- Scheduling workloads
+nbsp;+nbsp;+nbsp;- Scheduling AI workloads
+nbsp;+nbsp;+nbsp;- Batch workloads
- Autoscaling workloads
+nbsp;+nbsp;+nbsp;- Autoscaling AI workloads
+nbsp;+nbsp;+nbsp;- Horizontal Pod Autoscaler
+nbsp;+nbsp;+nbsp;- Event-driven scaling with KEDA
+nbsp;+nbsp;+nbsp;- AI-aware scheduling and KAI Scheduler concepts
- AI use cases
+nbsp;+nbsp;+nbsp;- Large Language Model inference
+nbsp;+nbsp;+nbsp;- Deploying LLMs on Kubernetes
+nbsp;+nbsp;+nbsp;- Model serving with vLLM
+nbsp;+nbsp;+nbsp;- Concurrent inference workloads
+nbsp;+nbsp;+nbsp;- Distributed inference concepts
+nbsp;+nbsp;+nbsp;- Introduction to Kubeflow
+nbsp;+nbsp;+nbsp;- AI and ML pipelines on Kubernetes
- Monitoring
+nbsp;+nbsp;+nbsp;- Monitoring GPU workloads
+nbsp;+nbsp;+nbsp;- Prometheus metrics
+nbsp;+nbsp;+nbsp;- Grafana dashboards
+nbsp;+nbsp;+nbsp;- Troubleshooting AI workloads