2026年,AI不上云等于没上
模型训练在实验室,推理服务在云端——这是AI工程化的铁律。
一个残酷现实:你的7B模型在笔记本上跑得飞快,上了生产环境却因为GPU调度不当、容器配置错误、扩缩容失灵而全线崩溃。AI部署不是写个Dockerfile那么简单。
云原生AI部署全景图
┌─────────────────────────────────────────────────────────────────────┐
│ 云原生AI部署全景架构 │
├─────────────┬─────────────┬──────────────┬─────────────────────────┤
│ 模型服务层 │ 向量服务层 │ 微服务层 │ 基础设施层 │
│ │ │ │ │
│ vLLM/Triton │ Milvus集群 │ Spring Boot │ K8s + GPU Operator │
│ TGI/Ollama │ Elastic- │ AI Gateway │ NVIDIA MIG + 时间分片 │
│ 模型路由 │ search │ Ingress │ Prometheus + Grafana │
│ HPA扩缩容 │ Embedding │ Service Mesh │ ArgoCD + Kustomize │
├─────────────┴─────────────┴──────────────┴─────────────────────────┤
│ Docker 容器化 + K8s 编排 │
│ 多阶段构建 · GPU感知调度 · 弹性伸缩 · GitOps交付 │
└─────────────────────────────────────────────────────────────────────┘
AI部署三大挑战
挑战一:GPU稀缺——算力是硬通货
┌──────────────────────────────────────────────────┐
│ GPU资源供需矛盾 │
│ │
│ 需求:100个推理服务 × 1 GPU = 100 GPU │
│ 现实:集群只有 8 × A100 = 8 GPU │
│ 缺口:92 GPU ❌ │
│ │
│ ┌──────┐ MIG分片 ┌──────────────┐ │
│ │ A100 │ ────────→ │ 7 × MIG实例 │ │
│ │ 80GB │ 时间分片 │ + 时间复用 │ │
│ └──────┘ ────────→ │ = 14+ 逻辑GPU │ │
│ └──────────────┘ │
└──────────────────────────────────────────────────┘
| GPU型号 |
显存 |
MIG分片数 |
适合模型规模 |
单卡成本(月) |
| A100 80GB |
80GB |
7 × 10GB |
7B-13B |
¥15,000 |
| A100 40GB |
40GB |
4 × 10GB |
7B |
¥10,000 |
| H100 80GB |
80GB |
7 × 10GB |
7B-70B |
¥25,000 |
| L40S 48GB |
48GB |
不支持MIG |
7B-13B |
¥6,000 |
| T4 16GB |
16GB |
不支持MIG |
1B-3B |
¥2,000 |
挑战二:模型体积爆炸——存储与传输的噩梦
| 模型 |
参数量 |
FP16体积 |
INT8体积 |
INT4体积 |
| Qwen2.5-7B |
7B |
14GB |
7GB |
3.5GB |
| Llama3.1-13B |
13B |
26GB |
13GB |
6.5GB |
| Qwen2.5-72B |
72B |
144GB |
72GB |
36GB |
| DeepSeek-V3-671B |
671B |
1.3TB |
671GB |
335GB |
模型拉取时间:72B模型FP16从Docker Registry拉取需要30分钟+,INT4量化后仅需8分钟。
挑战三:延迟与吞吐矛盾——鱼与熊掌不可兼得
┌─────────────────────────────────────────────────────┐
│ 延迟 vs 吞吐 权衡曲线 │
│ │
│ 吞吐 │ ╱ │
│ (req/│ ╱ │
│ sec) │ ╱ ← 批处理增大,吞吐提升 │
│ │ ╱ │
│ │ ╱ │
│ │ ╱ │
│ │───╱──────────────────────── │
│ │ ╱ ← 但延迟也增大! │
│ └─────────────────────── 延迟(ms) │
│ │
│ 最优策略:动态批处理 + Continuous Batching │
└─────────────────────────────────────────────────────┘
| 策略 |
延迟 |
吞吐 |
适用场景 |
| 单请求串行 |
最低 |
最低 |
实时对话 |
| 静态批处理 |
中等 |
中等 |
离线推理 |
| Continuous Batching |
低 |
高 |
在线服务 |
| Speculative Decoding |
更低 |
高 |
实时+高吞吐 |
AI模型服务化:四大框架对比
vLLM vs Triton vs TGI vs Ollama
| 维度 |
vLLM |
Triton Inference Server |
TGI (Text Generation Inference) |
Ollama |
| 开发方 |
UC Berkeley |
NVIDIA |
HuggingFace |
自社区 |
| 核心优势 |
PagedAttention |
多框架支持 |
Flash Attention |
极简部署 |
| Continuous Batching |
✅ 原生 |
✅ 动态批处理 |
✅ 原生 |
❌ |
| Tensor Parallel |
✅ |
✅ |
✅ |
❌ |
| 流式输出 |
✅ SSE |
✅ |
✅ SSE |
✅ |
| OpenAI兼容API |
✅ |
需适配 |
✅ |
✅ |
| 量化支持 |
AWQ/GPTQ/GGUF |
INT8/FP8 |
AWQ/GPTQ/bitsandbytes |
GGUF |
| GPU利用率 |
90%+ |
80%+ |
85%+ |
60%+ |
| K8s友好度 |
⭐⭐⭐⭐⭐ |
⭐⭐⭐⭐ |
⭐⭐⭐⭐ |
⭐⭐⭐ |
| 生产就绪 |
✅ |
✅ |
✅ |
⚠️ 开发/测试 |
| 多模型服务 |
单实例单模型 |
单实例多模型 |
单实例单模型 |
单实例多模型 |
| 模型热加载 |
❌ |
✅ |
❌ |
✅ |
架构对比图
┌─────────────────────────────────────────────────────────────────┐
│ vLLM 架构 │
│ ┌─────────┐ ┌──────────────┐ ┌─────────────────┐ │
│ │ Request │───→│ Scheduler │───→│ PagedAttention │ │
│ │ Queue │ │ (Continuous │ │ (KV Block Mgmt) │ │
│ │ │ │ Batching) │ │ │ │
│ └─────────┘ └──────────────┘ └─────────────────┘ │
│ ↑ ↑ ↑ │
│ OpenAI API Dynamic Batch GPU HBM │
│ Compatible Size Control KV Cache │
└─────────────────────────────────────────────────────────────────┘
┌─────────────────────────────────────────────────────────────────┐
│ Triton 架构 │
│ ┌─────────┐ ┌──────────────┐ ┌─────────────────┐ │
│ │ Model │───→│ Dynamic │───→│ Backend │ │
│ │ Repo │ │ Batcher │ │ (PyTorch/TF/ │ │
│ │ │ │ │ │ ONRT/TensorRT) │ │
│ └─────────┘ └──────────────┘ └─────────────────┘ │
│ ↑ ↑ ↑ │
│ 多模型配置 优先级队列 多框架推理引擎 │
└─────────────────────────────────────────────────────────────────┘
vLLM快速启动
from vllm import LLM, SamplingParams
llm = LLM(
model="Qwen/Qwen2.5-7B-Instruct",
tensor_parallel_size=2,
gpu_memory_utilization=0.9,
max_model_len=8192,
quantization="awq",
)
params = SamplingParams(
temperature=0.7,
top_p=0.9,
max_tokens=2048,
)
outputs = llm.generate(["解释云原生AI部署的核心挑战"], params)
for output in outputs:
print(output.outputs[0].text)
TGI启动示例
docker run --gpus all -p 8080:80 \
-v $PWD/data:/data \
ghcr.io/huggingface/text-generation-inference:2.3.0 \
--model-id Qwen/Qwen2.5-7B-Instruct \
--quantize awq \
--max-input-length 4096 \
--max-total-tokens 8192 \
--cuda-memory-fraction 0.9
选型建议
┌──────────────────────────────────────────────────────┐
│ 模型服务框架选型决策树 │
│ │
│ 需要GPU极致利用率? │
│ ├─ 是 → vLLM (PagedAttention, 90%+ 利用率) │
│ └─ 否 ↓ │
│ 需要多框架/多模型共存? │
│ ├─ 是 → Triton (单实例多模型) │
│ └─ 否 ↓ │
│ 需要HuggingFace生态无缝集成? │
│ ├─ 是 → TGI (Flash Attention + HF Hub) │
│ └─ 否 ↓ │
│ 只是本地开发/测试? │
│ └─ 是 → Ollama (一条命令启动) │
└──────────────────────────────────────────────────────┘
Docker化AI应用:多阶段构建 + Compose
多阶段Dockerfile
# ===== Stage 1: 模型下载 =====
FROM python:3.11-slim AS model-downloader
RUN pip install --no-cache-dir huggingface-hub
ARG MODEL_ID=Qwen/Qwen2.5-7B-Instruct
ARG QUANTIZATION=awq
RUN huggingface-cli download \
${MODEL_ID} \
--local-dir /models/${MODEL_ID} \
--exclude "*.safetensors" \
&& echo "Model config downloaded"
# ===== Stage 2: 推理运行时 =====
FROM nvidia/cuda:12.4.1-runtime-ubuntu22.04 AS runtime
ENV DEBIAN_FRONTEND=noninteractive
ENV PYTHONUNBUFFERED=1
RUN apt-get update && apt-get install -y --no-install-recommends \
python3.11 python3-pip python3.11-venv \
&& rm -rf /var/lib/apt/lists/*
RUN python3.11 -m venv /opt/vllm
ENV PATH="/opt/vllm/bin:${PATH}"
RUN pip install --no-cache-dir \
vllm==0.8.0 \
transformers>=4.45.0 \
accelerate
COPY --from=model-downloader /models /models
ARG MODEL_ID=Qwen/Qwen2.5-7B-Instruct
ENV MODEL_ID=${MODEL_ID}
EXPOSE 8000
HEALTHCHECK --interval=30s --timeout=10s --retries=3 \
CMD curl -f http://localhost:8000/health || exit 1
ENTRYPOINT ["python3.11", "-m", "vllm.entrypoints.openai.api_server"]
CMD ["--model", "/models/Qwen/Qwen2.5-7B-Instruct", \
"--host", "0.0.0.0", \
"--port", "8000", \
"--tensor-parallel-size", "2", \
"--gpu-memory-utilization", "0.9", \
"--max-model-len", "8192"]
Docker Compose本地开发环境
version: "3.8"
services:
vllm-server:
build:
context: .
dockerfile: Dockerfile
args:
MODEL_ID: Qwen/Qwen2.5-7B-Instruct
QUANTIZATION: awq
ports:
- "8000:8000"
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: 2
capabilities: [gpu]
environment:
- VLLM_WORKER_MULTIPROC_METHOD=spawn
volumes:
- model-cache:/root/.cache/huggingface
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 5
restart: unless-stopped
milvus-standalone:
image: milvusdb/milvus:v2.4.17
ports:
- "19530:19530"
- "9091:9091"
environment:
- ETCD_USE_EMBED=true
- COMMON_STORAGETYPE=local
volumes:
- milvus-data:/var/lib/milvus
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:9091/healthz"]
interval: 15s
timeout: 5s
retries: 3
elasticsearch:
image: elasticsearch:8.15.0
ports:
- "9200:9200"
environment:
- discovery.type=single-node
- xpack.security.enabled=false
- ES_JAVA_OPTS=-Xms1g -Xmx1g
volumes:
- es-data:/usr/share/elasticsearch/data
healthcheck:
test: ["CMD-SHELL", "curl -f http://localhost:9200/_cluster/health"]
interval: 15s
timeout: 5s
retries: 3
springboot-ai:
build:
context: ../springboot-ai-service
dockerfile: Dockerfile
ports:
- "8080:8080"
environment:
- SPRING_AI_OPENAI_BASE-URL=http://vllm-server:8000/v1
- SPRING_AI_OPENAI_API-KEY=not-needed
- MILVUS_HOST=milvus-standalone
- MILVUS_PORT=19530
- ELASTICSEARCH_URIS=http://elasticsearch:9200
depends_on:
vllm-server:
condition: service_healthy
milvus-standalone:
condition: service_healthy
elasticsearch:
condition: service_healthy
volumes:
model-cache:
milvus-data:
es-data:
Docker镜像优化技巧
| 优化手段 |
效果 |
示例 |
| 多阶段构建 |
镜像体积减少60%+ |
编译工具不进运行时 |
| 模型量化 |
体积减少50%-75% |
FP16→INT4 |
| .dockerignore |
构建上下文减少 |
排除.git, pycache |
| 层缓存复用 |
构建时间减少80% |
依赖层不变则复用 |
| 基础镜像选择 |
体积差异10倍 |
slim vs full |
| 模型外挂Volume |
镜像不含模型数据 |
运行时挂载 |
K8s GPU调度:NVIDIA GPU Operator + MIG + 时间分片
NVIDIA GPU Operator安装
apiVersion: v1
kind: Namespace
metadata:
name: gpu-operator
---
apiVersion: helm.cattle.io/v1
kind: HelmChart
metadata:
name: gpu-operator
namespace: gpu-operator
spec:
repo: https://helm.ngc.nvidia.com/nvidia
chart: gpu-operator
targetNamespace: gpu-operator
valuesContent: |
driver:
enabled: true
version: "550.54.15"
toolkit:
enabled: true
version: "v1.15.0"
devicePlugin:
enabled: true
config:
name: mig-config
shared:
timeSlicing:
resources:
- name: nvidia.com/gpu
replicas: 4
migManager:
enabled: true
config:
name: mig-partition-config
MIG分区配置
apiVersion: v1
kind: ConfigMap
metadata:
name: mig-partition-config
namespace: gpu-operator
data:
config.yaml: |
version: v1
mig-configs:
all-balanced:
- devices: all
mig-enabled: true
mig-devices:
"1g.10gb": 7
all-1g.5gb:
- devices: all
mig-enabled: true
mig-devices:
"1g.5gb": 14
mixed:
- devices: [0]
mig-enabled: true
mig-devices:
"2g.20gb": 2
"1g.10gb": 3
- devices: [1]
mig-enabled: false
GPU时间分片配置
apiVersion: v1
kind: ConfigMap
metadata:
name: time-slicing-config
namespace: gpu-operator
data:
a100-80gb.yaml: |
version: v1
sharing:
timeSlicing:
resources:
- name: nvidia.com/gpu
replicas: 4
devices: all
---
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: nvidia-device-plugin-daemonset
namespace: gpu-operator
spec:
selector:
matchLabels:
name: nvidia-device-plugin-ds
template:
spec:
containers:
- name: nvidia-device-plugin
image: nvcr.io/nvidia/k8s-device-plugin:v0.16.2
args:
- --config-file=/etc/nvidia-plugin/config.yaml
volumeMounts:
- name: config
mountPath: /etc/nvidia-plugin
volumes:
- name: config
configMap:
name: time-slicing-config
Pod使用MIG实例
apiVersion: v1
kind: Pod
metadata:
name: vllm-mig-pod
spec:
containers:
- name: vllm
image: vllm/vllm-openai:v0.8.0
resources:
limits:
nvidia.com/mig-1g.10gb: 1
command:
- python3
- -m
- vllm.entrypoints.openai.api_server
- --model
- Qwen/Qwen2.5-7B-Instruct-AWQ
- --max-model-len
- "4096"
- --gpu-memory-utilization
- "0.9"
GPU调度策略对比
| 策略 |
原理 |
优点 |
缺点 |
适用场景 |
| 独占GPU |
1 Pod = 1 GPU |
零干扰,性能稳定 |
资源浪费 |
大模型推理 |
| MIG分片 |
硬件级隔离 |
真正隔离,可配不同规格 |
仅A100/H100支持 |
多小模型服务 |
| 时间分片 |
软件级复用 |
所有GPU都支持 |
上下文切换开销 |
低优先级任务 |
| MPS |
共享GPU上下文 |
低开销共享 |
无隔离,一方崩溃全部影响 |
同团队多任务 |
模型推理服务部署:vLLM + HPA自动扩缩容
vLLM Deployment
apiVersion: apps/v1
kind: Deployment
metadata:
name: vllm-qwen7b
namespace: ai-inference
labels:
app: vllm-qwen7b
spec:
replicas: 2
selector:
matchLabels:
app: vllm-qwen7b
template:
metadata:
labels:
app: vllm-qwen7b
annotations:
prometheus.io/scrape: "true"
prometheus.io/port: "8000"
prometheus.io/path: "/metrics"
spec:
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchLabels:
app: vllm-qwen7b
topologyKey: kubernetes.io/hostname
containers:
- name: vllm
image: vllm/vllm-openai:v0.8.0
ports:
- containerPort: 8000
resources:
limits:
nvidia.com/gpu: 2
requests:
nvidia.com/gpu: 2
cpu: "4"
memory: 16Gi
env:
- name: MODEL_NAME
value: "Qwen/Qwen2.5-7B-Instruct-AWQ"
command:
- python3
- -m
- vllm.entrypoints.openai.api_server
args:
- --model
- Qwen/Qwen2.5-7B-Instruct-AWQ
- --host
- "0.0.0.0"
- --port
- "8000"
- --tensor-parallel-size
- "2"
- --gpu-memory-utilization
- "0.9"
- --max-model-len
- "8192"
- --enable-prefix-caching
livenessProbe:
httpGet:
path: /health
port: 8000
initialDelaySeconds: 120
periodSeconds: 30
readinessProbe:
httpGet:
path: /health
port: 8000
initialDelaySeconds: 60
periodSeconds: 10
volumeMounts:
- name: model-cache
mountPath: /root/.cache/huggingface
- name: shm
mountPath: /dev/shm
volumes:
- name: model-cache
persistentVolumeClaim:
claimName: model-cache-pvc
- name: shm
emptyDir:
medium: Memory
sizeLimit: 8Gi
nodeSelector:
gpu-type: a100-80gb
tolerations:
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
HPA自动扩缩容
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: vllm-qwen7b-hpa
namespace: ai-inference
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: vllm-qwen7b
minReplicas: 2
maxReplicas: 8
metrics:
- type: Pods
pods:
metric:
name: vllm_num_requests_running
target:
type: AverageValue
averageValue: "10"
- type: Pods
pods:
metric:
name: vllm_gpu_cache_usage_perc
target:
type: AverageValue
averageValue: "70"
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
behavior:
scaleUp:
stabilizationWindowSeconds: 60
policies:
- type: Pods
value: 2
periodSeconds: 120
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Pods
value: 1
periodSeconds: 120
Service + PodMonitor
apiVersion: v1
kind: Service
metadata:
name: vllm-qwen7b-svc
namespace: ai-inference
spec:
selector:
app: vllm-qwen7b
ports:
- name: http
port: 8000
targetPort: 8000
type: ClusterIP
---
apiVersion: monitoring.coreos.com/v1
kind: PodMonitor
metadata:
name: vllm-qwen7b-monitor
namespace: ai-inference
spec:
selector:
matchLabels:
app: vllm-qwen7b
podMetricsEndpoints:
- port: http
path: /metrics
interval: 15s
vLLM关键监控指标
| 指标 |
含义 |
告警阈值 |
vllm:num_requests_running |
正在处理的请求数 |
> 80% max_num_seqs |
vllm:num_requests_waiting |
等待队列中的请求数 |
> 10 持续5分钟 |
vllm:gpu_cache_usage_perc |
KV Cache使用率 |
> 90% |
vllm:avg_generation_throughput |
平均生成吞吐量 |
< 100 tok/s |
vllm:e2e_request_latency_seconds |
端到端请求延迟 |
P99 > 10s |
vllm:num_preemptions |
抢占次数 |
> 0 |
RAG向量服务K8s部署:Milvus集群 + Elasticsearch
整体架构
┌──────────────────────────────────────────────────────────────────┐
│ RAG向量服务架构 │
│ │
│ ┌──────────┐ ┌──────────────┐ ┌──────────────────┐ │
│ │ 用户查询 │───→│ Spring Boot │───→│ vLLM Embedding │ │
│ │ │ │ AI Gateway │ │ (向量生成) │ │
│ └──────────┘ └──────┬───────┘ └──────────────────┘ │
│ │ │
│ ┌──────────┼──────────┐ │
│ ▼ ▼ │
│ ┌──────────────┐ ┌──────────────┐ │
│ │ Milvus │ │Elasticsearch │ │
│ │ (向量检索) │ │ (全文检索) │ │
│ │ HNSW/IVF │ │ BM25+稠密 │ │
│ └──────────────┘ └──────────────┘ │
│ │ │ │
│ └────────┬───────────┘ │
│ ▼ │
│ ┌──────────────┐ ┌──────────────────┐ │
│ │ 混合检索结果 │───→│ vLLM LLM生成 │ │
│ │ Rerank重排 │ │ (答案生成) │ │
│ └──────────────┘ └──────────────────┘ │
└──────────────────────────────────────────────────────────────────┘
Milvus集群部署
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: milvus
namespace: ai-rag
spec:
serviceName: milvus-headless
replicas: 3
selector:
matchLabels:
app: milvus
template:
metadata:
labels:
app: milvus
spec:
containers:
- name: milvus
image: milvusdb/milvus:v2.4.17
ports:
- containerPort: 19530
- containerPort: 9091
resources:
requests:
cpu: "4"
memory: 8Gi
limits:
cpu: "8"
memory: 16Gi
env:
- name: ETCD_ENDPOINTS
value: "etcd-0.etcd:2379,etcd-1.etcd:2379,etcd-2.etcd:2379"
- name: MINIO_ADDRESS
value: "minio:9000"
- name: COMMON_STORAGETYPE
value: "remote"
volumeMounts:
- name: milvus-data
mountPath: /var/lib/milvus
livenessProbe:
httpGet:
path: /healthz
port: 9091
initialDelaySeconds: 30
periodSeconds: 15
readinessProbe:
httpGet:
path: /healthz
port: 9091
initialDelaySeconds: 15
periodSeconds: 10
volumeClaimTemplates:
- metadata:
name: milvus-data
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 100Gi
---
apiVersion: v1
kind: Service
metadata:
name: milvus-svc
namespace: ai-rag
spec:
selector:
app: milvus
ports:
- name: grpc
port: 19530
targetPort: 19530
- name: metrics
port: 9091
targetPort: 9091
type: ClusterIP
Elasticsearch部署
apiVersion: apps/v1
kind: StatefulSet
metadata:
name: elasticsearch
namespace: ai-rag
spec:
serviceName: elasticsearch-headless
replicas: 3
selector:
matchLabels:
app: elasticsearch
template:
metadata:
labels:
app: elasticsearch
spec:
initContainers:
- name: sysctl
image: busybox
command: ["sysctl", "-w", "vm.max_map_count=262144"]
securityContext:
privileged: true
containers:
- name: elasticsearch
image: elasticsearch:8.15.0
ports:
- containerPort: 9200
resources:
requests:
cpu: "2"
memory: 4Gi
limits:
cpu: "4"
memory: 8Gi
env:
- name: cluster.name
value: "ai-rag-cluster"
- name: discovery.seed_hosts
value: "elasticsearch-0.elasticsearch-headless,elasticsearch-1.elasticsearch-headless,elasticsearch-2.elasticsearch-headless"
- name: cluster.initial_master_nodes
value: "elasticsearch-0,elasticsearch-1,elasticsearch-2"
- name: xpack.security.enabled
value: "false"
- name: ES_JAVA_OPTS
value: "-Xms4g -Xmx4g"
volumeMounts:
- name: es-data
mountPath: /usr/share/elasticsearch/data
volumeClaimTemplates:
- metadata:
name: es-data
spec:
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 200Gi
Milvus Collection创建
from pymilvus import MilvusClient, DataType
client = MilvusClient(uri="http://milvus-svc.ai-rag.svc:19530")
schema = client.create_schema(auto_id=True, enable_dynamic_field=True)
schema.add_field(field_name="id", datatype=DataType.INT64, is_primary=True)
schema.add_field(field_name="text", datatype=DataType.VARCHAR, max_length=65535)
schema.add_field(field_name="embedding", datatype=DataType.FLOAT_VECTOR, dim=1024)
schema.add_field(field_name="source", datatype=DataType.VARCHAR, max_length=512)
schema.add_field(field_name="timestamp", datatype=DataType.INT64)
index_params = client.prepare_index_params()
index_params.add_index(
field_name="embedding",
index_type="HNSW",
metric_type="COSINE",
params={"M": 32, "efConstruction": 256}
)
index_params.add_index(
field_name="text",
index_type="INVERTED"
)
client.create_collection(
collection_name="knowledge_base",
schema=schema,
index_params=index_params,
)
向量检索 vs 全文检索 vs 混合检索
| 检索方式 |
原理 |
优势 |
劣势 |
适用场景 |
| 向量检索 |
语义相似度 |
理解语义 |
精确匹配弱 |
语义问答 |
| 全文检索 |
BM25关键词 |
精确匹配 |
无语义理解 |
关键词搜索 |
| 混合检索 |
向量+全文融合 |
兼顾语义与精确 |
计算开销大 |
生产RAG |
| Rerank |
交叉编码器重排 |
精度最高 |
速度慢 |
高精度场景 |
Spring Boot AI微服务上云
Spring Boot AI应用Dockerfile
FROM eclipse-temurin:21-jdk AS build
WORKDIR /app
COPY gradle/ gradle/
COPY gradlew build.gradle settings.gradle ./
RUN ./gradlew dependencies --no-daemon
COPY src/ src/
RUN ./gradlew bootJar --no-daemon -x test
FROM eclipse-temurin:21-jre-alpine
WORKDIR /app
COPY --from=build /app/build/libs/*.jar app.jar
ENV JAVA_OPTS="-XX:+UseZGC -XX:MaxRAMPercentage=75.0"
ENV SPRING_PROFILES_ACTIVE=k8s
EXPOSE 8080
HEALTHCHECK --interval=10s --timeout=3s --retries=3 \
CMD wget -qO- http://localhost:8080/actuator/health || exit 1
ENTRYPOINT ["sh", "-c", "java ${JAVA_OPTS} -jar app.jar"]
K8s Deployment + Service + Ingress
apiVersion: apps/v1
kind: Deployment
metadata:
name: springboot-ai-service
namespace: ai-app
spec:
replicas: 3
selector:
matchLabels:
app: springboot-ai-service
template:
metadata:
labels:
app: springboot-ai-service
annotations:
prometheus.io/scrape: "true"
prometheus.io/port: "8080"
prometheus.io/path: "/actuator/prometheus"
spec:
containers:
- name: app
image: registry.example.com/springboot-ai-service:1.0.0
ports:
- containerPort: 8080
resources:
requests:
cpu: "1"
memory: 2Gi
limits:
cpu: "2"
memory: 4Gi
env:
- name: SPRING_AI_OPENAI_BASE-URL
value: "http://vllm-qwen7b-svc.ai-inference.svc:8000/v1"
- name: SPRING_AI_OPENAI_API-KEY
valueFrom:
secretKeyRef:
name: ai-api-keys
key: openai-key
- name: MILVUS_HOST
value: "milvus-svc.ai-rag.svc"
- name: MILVUS_PORT
value: "19530"
- name: ELASTICSEARCH_URIS
value: "http://elasticsearch-svc.ai-rag.svc:9200"
livenessProbe:
httpGet:
path: /actuator/health/liveness
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
readinessProbe:
httpGet:
path: /actuator/health/readiness
port: 8080
initialDelaySeconds: 15
periodSeconds: 5
volumeMounts:
- name: config
mountPath: /app/config
volumes:
- name: config
configMap:
name: springboot-ai-config
---
apiVersion: v1
kind: Service
metadata:
name: springboot-ai-svc
namespace: ai-app
spec:
selector:
app: springboot-ai-service
ports:
- name: http
port: 80
targetPort: 8080
type: ClusterIP
---
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: springboot-ai-ingress
namespace: ai-app
annotations:
nginx.ingress.kubernetes.io/proxy-read-timeout: "300"
nginx.ingress.kubernetes.io/proxy-send-timeout: "300"
nginx.ingress.kubernetes.io/proxy-body-size: "10m"
nginx.ingress.kubernetes.io/configuration-snippet: |
more_set_headers "X-AI-Service: springboot-ai";
spec:
ingressClassName: nginx
tls:
- hosts:
- ai-api.example.com
secretName: ai-api-tls
rules:
- host: ai-api.example.com
http:
paths:
- path: /api/ai
pathType: Prefix
backend:
service:
name: springboot-ai-svc
port:
number: 80
Spring Boot AI核心配置
spring:
ai:
openai:
base-url: http://vllm-qwen7b-svc.ai-inference.svc:8000/v1
api-key: ${OPENAI_API_KEY}
chat:
options:
model: Qwen/Qwen2.5-7B-Instruct-AWQ
temperature: 0.7
max-tokens: 2048
vectorstore:
milvus:
client:
host: ${MILVUS_HOST:milvus-svc.ai-rag.svc}
port: ${MILVUS_PORT:19530}
database-name: default
collection-name: knowledge_base
embedding-dimension: 1024
metric-type: COSINE
index-type: HNSW
elasticsearch:
uris: ${ELASTICSEARCH_URIS:http://elasticsearch-svc.ai-rag.svc:9200}
management:
endpoints:
web:
exposure:
include: health,info,prometheus,metrics
metrics:
tags:
application: springboot-ai-service
tracing:
sampling:
probability: 1.0
otlp:
metrics:
export:
url: http://otel-collector.observability.svc:4318/v1/metrics
tracing:
endpoint: http://otel-collector.observability.svc:4318/v1/traces
可观测性:Prometheus + Grafana + OpenTelemetry
监控架构
┌──────────────────────────────────────────────────────────────────┐
│ 全栈可观测性架构 │
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ vLLM │ │ Milvus │ │ES │ │Spring │ │
│ │ /metrics │ │ /metrics │ │/_prom │ │/actuator │ │
│ └────┬─────┘ └────┬─────┘ └────┬─────┘ └────┬─────┘ │
│ │ │ │ │ │
│ └──────────────┴──────────────┴──────────────┘ │
│ │ │
│ ┌────────▼────────┐ │
│ │ OTel Collector │ │
│ │ (接收+处理+导出)│ │
│ └────────┬────────┘ │
│ ┌─────────────┼─────────────┐ │
│ ▼ ▼ ▼ │
│ ┌──────────────┐ ┌──────────┐ ┌──────────┐ │
│ │ Prometheus │ │ Jaeger │ │ Loki │ │
│ │ (指标存储) │ │ (链路追踪)│ │ (日志) │ │
│ └──────┬───────┘ └──────────┘ └──────────┘ │
│ ▼ │
│ ┌──────────────┐ │
│ │ Grafana │ │
│ │ (统一面板) │ │
│ └──────────────┘ │
└──────────────────────────────────────────────────────────────────┘
OpenTelemetry Collector部署
apiVersion: apps/v1
kind: Deployment
metadata:
name: otel-collector
namespace: observability
spec:
replicas: 2
selector:
matchLabels:
app: otel-collector
template:
metadata:
labels:
app: otel-collector
spec:
containers:
- name: otel-collector
image: otel/opentelemetry-collector-contrib:0.110.0
ports:
- containerPort: 4318
- containerPort: 4317
resources:
requests:
cpu: "500m"
memory: 512Mi
limits:
cpu: "1"
memory: 1Gi
volumeMounts:
- name: config
mountPath: /etc/otelcol-contrib
volumes:
- name: config
configMap:
name: otel-collector-config
---
apiVersion: v1
kind: ConfigMap
metadata:
name: otel-collector-config
namespace: observability
data:
config.yaml: |
receivers:
otlp:
protocols:
http:
endpoint: 0.0.0.0:4318
grpc:
endpoint: 0.0.0.0:4317
prometheus:
config:
scrape_configs:
- job_name: "vllm"
scrape_interval: 15s
static_configs:
- targets: ["vllm-qwen7b-svc.ai-inference.svc:8000"]
- job_name: "milvus"
scrape_interval: 15s
static_configs:
- targets: ["milvus-svc.ai-rag.svc:9091"]
processors:
batch:
timeout: 10s
send_batch_size: 1024
memory_limiter:
check_interval: 5s
limit_mib: 400
exporters:
prometheusremotewrite:
endpoint: http://prometheus.observability.svc:9090/api/v1/write
otlphttp:
endpoint: http://jaeger-collector.observability.svc:4318
loki:
endpoint: http://loki.observability.svc:3100/loki/api/v1/push
service:
pipelines:
metrics:
receivers: [otlp, prometheus]
processors: [memory_limiter, batch]
exporters: [prometheusremotewrite]
traces:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [otlphttp]
logs:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [loki]
Prometheus告警规则
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: ai-inference-alerts
namespace: observability
spec:
groups:
- name: vllm-alerts
rules:
- alert: VLLMHighQueueDepth
expr: vllm_num_requests_waiting > 20
for: 5m
labels:
severity: warning
annotations:
summary: "vLLM请求队列深度过高"
description: "实例 {{ $labels.instance }} 等待队列 {{ $value }} 个请求"
- alert: VLLMHighLatency
expr: histogram_quantile(0.99, rate(vllm_e2e_request_latency_seconds_bucket[5m])) > 10
for: 3m
labels:
severity: critical
annotations:
summary: "vLLM P99延迟过高"
description: "实例 {{ $labels.instance }} P99延迟 {{ $value }}s"
- alert: VLLMGPUCacheFull
expr: vllm_gpu_cache_usage_perc > 0.9
for: 5m
labels:
severity: warning
annotations:
summary: "vLLM GPU缓存使用率过高"
description: "实例 {{ $labels.instance }} 缓存使用率 {{ $value }}%"
- alert: VLLMPreemptions
expr: increase(vllm_num_preemptions[5m]) > 0
for: 1m
labels:
severity: critical
annotations:
summary: "vLLM发生请求抢占"
description: "实例 {{ $labels.instance }} 5分钟内抢占 {{ $value }} 次"
- name: milvus-alerts
rules:
- alert: MilvusHighQueryLatency
expr: histogram_quantile(0.99, rate(milvus_proxy_sq_latency_bucket[5m])) > 2
for: 5m
labels:
severity: warning
annotations:
summary: "Milvus查询延迟过高"
- alert: MilvusInsertRateLow
expr: rate(milvus_proxy_insert_vectors_count[5m]) < 10
for: 10m
labels:
severity: info
annotations:
summary: "Milvus插入速率异常低"
Grafana Dashboard关键面板
| 面板 |
指标 |
可视化类型 |
| 推理QPS |
rate(vllm_num_requests_total[1m]) |
折线图 |
| P50/P90/P99延迟 |
histogram_quantile |
折线图 |
| GPU缓存使用率 |
vllm_gpu_cache_usage_perc |
仪表盘 |
| 请求队列深度 |
vllm_num_requests_waiting |
柱状图 |
| 向量检索延迟 |
milvus_proxy_sq_latency |
折线图 |
| Spring Boot JVM |
jvm_memory_used_bytes |
面积图 |
| HTTP错误率 |
rate(http_server_requests_seconds_count{status=~"5.."}[5m]) |
折线图 |
| 链路追踪 |
Jaeger集成 |
表格+链接 |
成本优化:量化 + MIG复用 + 动态扩缩容 + Spot实例
模型量化对比
| 量化方法 |
精度损失 |
体积压缩比 |
推理加速 |
兼容性 |
| FP16 (基线) |
0% |
1x |
1x |
全部 |
| INT8 (PTQ) |
<1% |
2x |
1.5x |
A100/H100 |
| INT4 (GPTQ) |
1-3% |
4x |
1.8x |
需校准数据 |
| INT4 (AWQ) |
1-2% |
4x |
2x |
需校准数据 |
| INT4 (GGUF) |
2-5% |
4x |
1.5x |
CPU/GPU通用 |
| FP8 |
<0.5% |
2x |
2x+ |
H100原生 |
成本优化策略全景
┌──────────────────────────────────────────────────────────────────┐
│ 成本优化四板斧 │
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────┐ │
│ │ 1.模型量化 │ │ 2.MIG复用 │ │ 3.动态扩缩容 │ │4.Spot│ │
│ │ │ │ │ │ │ │实例 │ │
│ │ FP16→INT4 │ │ 1GPU→7MIG │ │ HPA+VPA │ │竞价 │ │
│ │ 体积↓75% │ │ 利用率↑7x │ │ 闲时缩容 │ │成本 │ │
│ │ 推理↑2x │ │ 成本↓85% │ │ 成本↓40% │ │↓70% │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ └──────┘ │
│ │
│ 综合效果:成本降至原来的 5%-15% │
└──────────────────────────────────────────────────────────────────┘
成本计算实例
| 方案 |
GPU配置 |
月成本 |
说明 |
| 基线:FP16独占 |
4 × A100 80GB |
¥60,000 |
2副本 × 2GPU |
| INT4量化 |
2 × A100 80GB |
¥30,000 |
2副本 × 1GPU |
| INT4 + MIG |
1 × A100 80GB |
¥15,000 |
1副本 × 1MIG实例 |
| INT4 + MIG + HPA |
1 × A100 80GB(闲时0.5) |
¥10,000 |
闲时缩到1副本 |
| INT4 + MIG + HPA + Spot |
1 × A100 Spot |
¥3,000 |
Spot折扣70% |
Spot实例使用策略
apiVersion: apps/v1
kind: Deployment
metadata:
name: vllm-qwen7b-spot
namespace: ai-inference
spec:
replicas: 2
selector:
matchLabels:
app: vllm-qwen7b-spot
template:
metadata:
labels:
app: vllm-qwen7b-spot
spec:
affinity:
nodeAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
preference:
matchExpressions:
- key: cloud.google.com/gke-preemptible
operator: In
values: ["true"]
tolerations:
- key: cloud.google.com/gke-preemptible
operator: Equal
value: "true"
effect: NoSchedule
- key: nvidia.com/gpu
operator: Exists
effect: NoSchedule
containers:
- name: vllm
image: vllm/vllm-openai:v0.8.0
resources:
limits:
nvidia.com/gpu: 1
requests:
nvidia.com/gpu: 1
env:
- name: GRACEFUL_SHUTDOWN
value: "true"
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "sleep 30"]
量化模型加载配置
from vllm import LLM
llm_awq = LLM(
model="Qwen/Qwen2.5-7B-Instruct-AWQ",
quantization="awq",
gpu_memory_utilization=0.85,
max_model_len=4096,
enforce_eager=True,
)
llm_gptq = LLM(
model="Qwen/Qwen2.5-7B-Instruct-GPTQ-Int4",
quantization="gptq",
gpu_memory_utilization=0.85,
max_model_len=4096,
)
GitOps:ArgoCD + Kustomize自动化交付
GitOps工作流
┌──────────────────────────────────────────────────────────────────┐
│ GitOps 交付流水线 │
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ 开发提交 │───→│ CI构建 │───→│ 镜像推送 │───→│ Git仓库 │ │
│ │ 代码变更 │ │ 测试通过 │ │ Registry │ │ 清单更新 │ │
│ └──────────┘ └──────────┘ └──────────┘ └────┬─────┘ │
│ │ │
│ ┌────────▼───────┐ │
│ │ ArgoCD │ │
│ │ 检测到变更 │ │
│ │ 自动Sync │ │
│ └────────┬───────┘ │
│ │ │
│ ┌──────────────────┼──────┐ │
│ ▼ ▼ ▼ │
│ ┌──────────┐ ┌──────────┐ ┌──┐ │
│ │ Dev │ │ Staging │ │PR││
│ │ 环境 │ │ 环境 │ │OD│ │
│ └──────────┘ └──────────┘ └──┘ │
└──────────────────────────────────────────────────────────────────┘
Kustomize目录结构
k8s/
├── base/
│ ├── kustomization.yaml
│ ├── deployment.yaml
│ ├── service.yaml
│ ├── hpa.yaml
│ └── ingress.yaml
├── overlays/
│ ├── dev/
│ │ ├── kustomization.yaml
│ │ └── patch-replicas.yaml
│ ├── staging/
│ │ ├── kustomization.yaml
│ │ └── patch-resources.yaml
│ └── production/
│ ├── kustomization.yaml
│ ├── patch-replicas.yaml
│ ├── patch-resources.yaml
│ └── patch-hpa.yaml
Base Kustomization
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
resources:
- deployment.yaml
- service.yaml
- hpa.yaml
- ingress.yaml
commonLabels:
app.kubernetes.io/part-of: ai-inference-platform
app.kubernetes.io/managed-by: kustomize
images:
- name: vllm-server
newName: registry.example.com/vllm-openai
newTag: v0.8.0
configMapGenerator:
- name: vllm-config
literals:
- MODEL_NAME=Qwen/Qwen2.5-7B-Instruct-AWQ
- MAX_MODEL_LEN=8192
- GPU_MEMORY_UTILIZATION=0.9
secretGenerator:
- name: ai-api-keys
literals:
- openai-key=placeholder
Production Overlay
apiVersion: kustomize.config.k8s.io/v1beta1
kind: Kustomization
namespace: ai-inference-production
resources:
- ../../base
patches:
- target:
kind: Deployment
name: vllm-qwen7b
patch: |
apiVersion: apps/v1
kind: Deployment
metadata:
name: vllm-qwen7b
spec:
replicas: 4
template:
spec:
containers:
- name: vllm
resources:
limits:
nvidia.com/gpu: 2
requests:
nvidia.com/gpu: 2
cpu: "4"
memory: 16Gi
- target:
kind: HorizontalPodAutoscaler
name: vllm-qwen7b-hpa
patch: |
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: vllm-qwen7b-hpa
spec:
minReplicas: 4
maxReplicas: 16
ArgoCD Application
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: ai-inference-platform
namespace: argocd
annotations:
notifications.argoproj.io/subscribe.on-deployed.slack: ai-platform-deploys
spec:
project: ai-platform
source:
repoURL: https://git.example.com/ai-platform/k8s-manifests.git
targetRevision: main
path: k8s/overlays/production
destination:
server: https://kubernetes.default.svc
namespace: ai-inference-production
syncPolicy:
automated:
prune: true
selfHeal: true
allowEmpty: false
syncOptions:
- CreateNamespace=true
- ServerSideApply=true
retry:
limit: 3
backoff:
duration: 30s
factor: 2
maxDuration: 5m
ArgoCD App of Apps模式
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: ai-platform-apps
namespace: argocd
spec:
project: default
source:
repoURL: https://git.example.com/ai-platform/argocd-apps.git
targetRevision: main
path: apps
destination:
server: https://kubernetes.default.svc
namespace: argocd
---
apiVersion: argoproj.io/v1alpha1
kind: AppProject
metadata:
name: ai-platform
namespace: argocd
spec:
description: AI推理平台项目
sourceRepos:
- https://git.example.com/ai-platform/*
destinations:
- namespace: ai-inference-*
server: https://kubernetes.default.svc
- namespace: ai-rag-*
server: https://kubernetes.default.svc
- namespace: ai-app-*
server: https://kubernetes.default.svc
clusterResourceWhitelist:
- group: ""
kind: Namespace
namespaceResourceBlacklist:
- group: ""
kind: ResourceQuota
总结:云原生AI部署检查清单
| 阶段 |
检查项 |
状态 |
| 模型服务 |
vLLM/Triton选型确定 |
☐ |
| 模型服务 |
量化策略(INT4/AWQ)确定 |
☐ |
| Docker化 |
多阶段Dockerfile编写 |
☐ |
| Docker化 |
Docker Compose本地验证 |
☐ |
| K8s GPU |
GPU Operator安装 |
☐ |
| K8s GPU |
MIG/时间分片配置 |
☐ |
| 推理部署 |
vLLM Deployment部署 |
☐ |
| 推理部署 |
HPA自动扩缩容配置 |
☐ |
| RAG服务 |
Milvus集群部署 |
☐ |
| RAG服务 |
Elasticsearch部署 |
☐ |
| 微服务 |
Spring Boot AI上云 |
☐ |
| 微服务 |
Ingress路由配置 |
☐ |
| 可观测性 |
Prometheus+Grafana |
☐ |
| 可观测性 |
OpenTelemetry链路追踪 |
☐ |
| 可观测性 |
告警规则配置 |
☐ |
| 成本优化 |
模型量化验证 |
☐ |
| 成本优化 |
Spot实例策略 |
☐ |
| GitOps |
ArgoCD + Kustomize |
☐ |
| GitOps |
多环境Overlay配置 |
☐ |
云原生AI部署不是终点,而是起点。 从Docker容器化到K8s编排,从GPU调度到可观测性,每一步都是工程化的积累。记住:没有监控的部署就是盲飞,没有自动化的运维就是手工活。 用GitOps把一切自动化,让AI推理服务真正成为云原生的一等公民。