Easy Book Translation翻译你自己的书 ↗
双语书摘

副本不等于独立故障

Betsy Beyer, Niall Richard Murphy, David K. Rensin, Kent Kawahara and Stephen Thorne (editors) · 站点可靠性工程工作手册

第2章 实施 SLO · Steven Thurgood、David Ferguson

问题与方法

本章把服务级别目标(SLO)作为可靠性决策工具:从用户关心的结果选择指标,确定目标和统计窗口,再与产品、开发及运维负责人共同约定错误预算政策。游戏服务的示例演示测量方法与取舍,后半章讨论持续校准、用户旅程和依赖。多个副本并不自动意味着独立故障,目标也必须随用户实际体验调整。

英文原文

It can be tempting to try to math your way out of these problems. If you have a service that offers 99.9% availability in a single zone, and you need 99.95% availability, simply deploying the service in two zones should solve that requirement. The probability that both services will experience an outage at the same time is so low that two zones should provide 99.9999% availability. However, this reasoning assumes that both services are wholly independent, which is almost never the case. The two instances of your app will have common dependencies, common failure domains, shared fate, and global control planes—all of which can cause an outage in both systems, no matter how carefully it is designed and managed. Unless each of these dependencies and failure patterns is carefully enumerated and accounted for, any such calculations will be deceptive.

中文译文

人们可能很想靠计算来摆脱这些问题。如果某个服务在单个可用区提供 99.9% 的可用性,而你需要 99.95%,那么只要把服务部署到两个可用区,似乎就能满足要求。两个服务同时发生故障的概率很低,因此两个可用区应该能提供 99.9999% 的可用性。然而,这种推理假定两个服务完全独立,而实际情况几乎从来不是这样。应用的两个实例会有共同的依赖、共同的故障域、共同的命运以及全局控制平面。无论设计和管理多么谨慎,这些因素都可能让两个系统同时中断。除非逐一列出并考虑每项依赖和故障模式,否则这样的计算都会产生误导。

Betsy Beyer, Niall Richard Murphy, David K. Rensin, Kent Kawahara and Stephen Thorne (editors) · 站点可靠性工程工作手册 · 第2章 实施 SLO · Steven Thurgood、David Ferguson · ¶ 340

本页中文译文由 AI 对照英文原文新编,非官方译本。章节概括、情境与赏析为编辑文字,不属于作者引文。

翻译你自己的书

翻译你自己的书 ↗