Blog

1. 创建Deployment工作流程

sequenceDiagram
    participant User as 用户 (kubectl)
    participant API as APIServer
    participant Etcd as etcd
    participant DC as Deployment Controller
    participant RSC as ReplicaSet Controller
    participant Sched as Scheduler
    participant Kubelet as 目标 Node Kubelet

    User->>API: 1. kubectl apply -f deployment.yaml
    API->>API: 2. 认证、鉴权、准入控制 (Admission)
    API->>Etcd: 3. 写入 Deployment 对象
    Etcd-->>API: 4. 返回写入成功
    Etcd-->>API: 5. Watch 事件:Deployment 创建
    API-->>DC: 6. 推送 Deployment 事件 (通过 Informer)
    
    rect rgb(230, 240, 255)
    Note over DC: 阶段一:Deployment 控制循环
    DC->>DC: 7. 对比期望(1个RS)与实际(0个RS)
    DC->>API: 8. 请求创建 ReplicaSet
    API->>Etcd: 9. 写入 ReplicaSet
    Etcd-->>API: Watch 事件:RS 创建
    API-->>RSC: 10. 推送 RS 事件 (通过 Informer)
    end

    rect rgb(255, 245, 230)
    Note over RSC: 阶段二:ReplicaSet 控制循环
    RSC->>RSC: 11. 对比期望(3个Pod)与实际(0个Pod)
    RSC->>API: 12. 请求创建 3个 Pod (此时 nodeName 为空)
    API->>Etcd: 13. 写入 3个 Pod
    Etcd-->>API: Watch 事件:Pod 创建 (未调度状态)
    API-->>Sched: 14. 推送未调度 Pod 事件 (通过 Informer)
    end

    rect rgb(230, 255, 230)
    Note over Sched: 阶段三:Scheduler 调度流程
    Sched->>Sched: 15. 过滤 (Predicates): 排除资源不足/有污点的 Node
    Sched->>Sched: 16. 打分 (Priorities): 对可用 Node 评分 (如负载均衡)
    Sched->>Sched: 17. 绑定 (Binding): 选出最优 Node (假设为 Node-A)
    Sched->>API: 18. 调用 Binding API (更新 Pod.spec.nodeName=Node-A)
    API->>Etcd: 19. 更新 Pod 绑定状态
    Etcd-->>API: Watch 事件:Pod 已分配 Node
    API-->>Kubelet: 20. 推送 Pod 分配事件 (Node-A 的 Kubelet 监听到)
    end

    rect rgb(255, 230, 240)
    Note over Kubelet: 阶段四:Kubelet 启动与状态闭环
    Kubelet->>Kubelet: 21. 调用 CRI 拉取镜像并启动容器
    Kubelet->>Kubelet: 22. 调用 CNI 配置网络, CSI 挂载存储
    Kubelet->>API: 23. 持续汇报 Pod 状态 (Running/Ready)
    API->>Etcd: 24. 更新 Pod 最终运行状态
    Etcd-->>API: Watch 事件:Pod 状态变更
    API-->>RSC: 25. 推送状态更新 (RS 确认 3/3 Ready)
    API-->>DC: 26. 推送状态更新 (Deployment 确认 Available)
    end

2. 创建Service工作流程

sequenceDiagram
    participant User as 用户 (kubectl)
    participant API as APIServer
    participant Etcd as etcd
    participant ESC as EndpointSlice Controller
    participant KProxy as 节点 kube-proxy
    participant Pod as 后端 Pod

    rect rgb(230, 240, 255)
    Note over User, Etcd: 阶段一:API 提交与持久化
    User->>API: 1. kubectl apply -f service.yaml
    API->>API: 2. 认证、鉴权、准入控制
    API->>API: 3. 分配 ClusterIP (从预分配池中获取)
    API->>Etcd: 4. 写入 Service 对象 (包含 ClusterIP 和 Selector)
    Etcd-->>API: 5. Watch 事件:Service 创建
    end

    rect rgb(255, 245, 230)
    Note over ESC, Etcd: 阶段二:EndpointSlice Controller (计算后端)
    API-->>ESC: 6. 推送 Service 事件 (通过 Informer)
    ESC->>ESC: 7. 解析 Service 的 selector
    ESC->>ESC: 8. 遍历集群,查找带有匹配 Label 且状态为 Ready 的 Pod
    ESC->>API: 9. 创建/更新 EndpointSlice 对象 (包含后端 Pod 的 IP:Port)
    API->>Etcd: 10. 写入 EndpointSlice 对象
    Etcd-->>API: 11. Watch 事件:EndpointSlice 变更
    end

    rect rgb(230, 255, 230)
    Note over KProxy, Pod: 阶段三:kube-proxy (配置网络规则)
    API-->>KProxy: 12. 推送 Service & EndpointSlice 事件 (通过 Informer)
    KProxy->>KProxy: 13. 计算本节点需要的 iptables/IPVS 转发规则
    KProxy->>KProxy: 14. 调用系统接口,更新内核网络规则 (如 iptables-restore)
    Note right of KProxy: 规则生效:发往 ClusterIP 的流量将被 DNAT 到 Pod IP
    end

3. 创建Ingress工作流程

sequenceDiagram
    participant User as 用户 (kubectl)
    participant API as APIServer
    participant Etcd as etcd
    participant IC as Ingress Controller (Nginx)
    participant Nginx as 底层 Nginx 进程
    participant Client as 外部客户端

    rect rgb(230, 240, 255)
    Note over User, Etcd: 阶段一:API 提交与持久化
    User->>API: 1. kubectl apply -f ingress.yaml
    API->>API: 2. 认证、鉴权、准入控制
    API->>Etcd: 3. 写入 Ingress 对象 (包含 Host, Path, Backend Service)
    Etcd-->>API: 4. Watch 事件:Ingress 创建/更新
    end

    rect rgb(255, 245, 230)
    Note over IC, Nginx: 阶段二:Ingress Controller 解析与配置下发
    API-->>IC: 5. 推送 Ingress 事件 (通过 Informer)
    IC->>IC: 6. 解析 Ingress 规则 (提取域名、路径、目标 Service 名称)
    
    Note over IC: 高级特性:绕过 kube-proxy
    IC->>API: 7. 监听目标 Service 关联的 EndpointSlice
    API-->>IC: 8. 返回后端 Pod 的真实 IP 列表
    
    IC->>IC: 9. 结合 Ingress 规则和 Pod IP,生成 Nginx upstream 配置
    IC->>Nginx: 10. 写入/更新 nginx.conf 配置文件
    IC->>Nginx: 11. 发送 reload 信号 (如 nginx -s reload)
    Nginx->>Nginx: 12. 热加载新配置,新路由规则生效
    end

    rect rgb(230, 255, 230)
    Note over IC, Nginx: 阶段三:状态回写 (可选)
    IC->>API: 13. 更新 Ingress 对象的 status 字段 (如写入外部 LB 的 IP)
    end

    rect rgb(255, 230, 240)
    Note over Client, Nginx: 阶段四:外部流量验证
    Client->>Nginx: 14. HTTP 请求 (GET http://a.com/api)
    Nginx->>Nginx: 15. 匹配 server_name 和 location 规则
    Nginx->>Nginx: 16. 从 upstream 中选择一个真实的 Pod IP
    Nginx->>Nginx: 17. 建立与后端 Pod 的连接,转发请求
    end

4. 底层组件讲解

  • Watch 监听机制
    它是 K8s 提供的一种底层通信机制/协议特性:

    • Watch模式
      • Watch 本质上是一种“基于长连接的事件推送机制”。
      • 客户端向 APIServer 发送一个特殊的 HTTP 请求(带上 watch=true 参数)。
      • APIServer 不关闭这个 HTTP 连接,而是保持它(长连接)。
      • 当底层数据(etcd)发生变化时,APIServer 主动通过这个长连接,把变化的事件(Added, Modified, Deleted)推送给客户端。
  • Informer 机制(K8s 高性能的秘密)
    Controller 并不直接去轮询 APIServer,而是通过 Informer 机制工作。是 K8s 官方提供的 Go 语言客户端库 client-go 中的一个核心代码模块/框架,它的核心定位是:Controller(控制器)的“数据中枢”与“缓冲层”。它的终极使命是让 Controller 能够以极低的延迟、极高的效率、零 APIServer 压力的方式,获取并监听集群状态 Informer 包含三个核心部分:

    • Reflector
      • 通过 Watch 机制监听 APIServer,获取数据变更事件。
    • Local Cache (Indexer)
      • 将 APIServer 的数据在本地内存中做一份镜像缓存。
      • Controller 读取数据时直接读本地缓存,极大减轻了 APIServer 的压力。
    • Work Queue (工作队列)
      • Reflector 收到事件后,不直接处理,
      • 而是将资源的 Key 放入工作队列,
      • 然后触发回调函数进行处理。
  • 控制循环工作流(Reconciliation)

    • 从本地缓存中获取资源的当前状态。
    • 对比“期望状态”和“实际状态”。
    • 如果不一致,计算出差异,并调用 APIServer 去创建、更新或删除资源,以消除差异。
    • :K8s 采用的是 Level-based(基于状态)而非 Edge-based(基于事件)的控制循环。
      即 Controller 关心的是“当前状态是什么”,而不是“发生了什么事件”,这保证了系统的最终一致性和强大的自愈能力。

  • Scheduler:集群的“决策中心” (Pod 调度器)
    Scheduler 负责监听集群中新创建的、尚未分配节点的 Pod,并为它们挑选最合适的 Node。

    • 核心工作流(两阶段调度)
      Scheduler 同样基于 Informer 机制,监听 spec.nodeName 为空的 Pod。一旦发现有未调度的 Pod,进入调度流程:
      • Predicates(过滤/预选)
        • 硬限制。遍历所有 Node,排除掉不满足 Pod 运行条件的节点。例如:
          • 资源(CPU/内存)不足
          • 端口冲突
          • 节点有污点(Taint)而 Pod 没有容忍(Toleration)
          • 节点选择器(NodeSelector)不匹配等
      • Priorities(打分/优选)
        • 软限制。对经过过滤后剩下的可用 Node 进行打分。例如:
          • 节点负载情况(负载均衡)
          • 数据局部性(Pod 尽量和它需要的数据在同一个节点)
          • 镜像是否已缓存等
      • Binding(绑定)
        • 选出得分最高的 Node,Scheduler 调用 APIServer 的 Binding API,
        • 将 Pod 和 Node 的绑定关系写入 etcd(即更新 Pod 的 spec.nodeName 字段)。

Comments & discussion

The first comment in each thread opens a topic. Signed-in readers can keep the conversation going under that topic.

No comments yet. Sign in to start a topic.

Start a new topic

Sign in to start a topic or join the discussion.