ポストモーテムとは

ポストモーテムを書く目的は、そのインシデントがドキュメント化されること、影響を及ぼした根本原因が理解されること。 そして、将来の再発の可能性を削減するための効果的な予防策が確実に導入されることです。

ポストモーテムには何を書くか?

末尾にテンプレートを探しましたので、ご確認ください。

ポストモーテムはどう使われるべきか

ポストモーテムは以下のような場面で使われます。

  • 今月のポストモーテム : ニュースレターとして、うまく書かれた興味深いポストモーテムを組織全体へ共有
  • Google+ ポストモーテムグループ : 社内/社外含めてポストモーテムを共有し、コメント機能を用いて議論をします
  • ポストモーテム読書会 : 定期的なポストもーてむの読書会を開催するグループです。対象となるポストモーテムは数ヶ月前、あるいは数年前のものであることもあります。
  • ポストモーテムロールプレイング : 新人SREへ訓練の一環として、インシデントが起きた状態をエンジニアたちに演じてもらって過去のポストモーテムを再現します。

レビューサれないポストモーテムは存在しないも同じです。 SREのミーティングの中で必ず確認するようにしましょう。

ポストモーテムのテンプレート

以下はポストモーテムのテンプレートです。

https://gist.github.com/mlafeldt/6e02ea0caeebef1205b47f31c2647966

# シェイクスピア 定型詩障害 ポストモーテム (インシデント #465)

## 日付

2015-10-21

## 著者

*   jennifer 
*   martym   
*   agoogler
    

## ステータス

完了(アクションアイテムは対応中)

## 概要

新種のソネット=定型詩(十四行詩)の発見によりシェイクスピアへの関心が極めて高まった時間帯に、シェイクスピアサーチ が66分間にわたってダウンした。

## 影響

推定12.1億件のクエリが失われた。収益への影響はなし。

## 根本原因

例外的な高負荷と、検索語句がシェイクスピアのコーパス(全作品データ)に存在しない場合に発生するリソースリークの組み合わせによる連鎖的障害(カスケード障害)。新しく発見されたソネットには、これまでのシェイクスピア作品に一度も登場したことがない単語が含まれており、ユーザーがまさにその単語を検索していた。通常の状況下では、リソースリークによるタスクの失敗率は十分に低く、気づかれないレベルであった。

## トリガー

トラフィックの急増によって引き起こされた潜在的バグ。

## 解決策

連鎖的障害を緩和するため、トラフィックを犠牲クラスターに誘導し、容量を10倍に増強した。更新されたインデックスをデプロイしたことで、潜在的バグとの相互作用が解消された。新しいソネットに対する世間の関心の高まりが収まるまで、追加容量を維持している。リソースリークを特定し、修正をデプロイした。

## 検知

Borgmon が高レベルの HTTP 500 エラーを検知し、オンコール担当者にページャー(オンコール呼び出し)を送信した。


## アクションアイテム


| Action Item | Type | Owner | Bug |
| ----------- | ---- | ----- | --- |
| Update playbook with instructions for responding to cascading failure | mitigate | jennifer | n/a **DONE** |
| Use flux capacitor to balance load between clusters | prevent | martym | Bug 5554823 **TODO** |
| Schedule cascading failure test during next DiRT | process | docbrown | n/a **TODO** |
| Investigate running index MR/fusion continuously | prevent | jennifer | Bug 5554824 **TODO** |
| Plug file descriptor leak in search ranking subsystem | prevent | agoogler | Bug 5554825 **DONE** |
| Add load shedding capabilities to シェイクスピアサーチ | prevent | agoogler | Bug 5554826 **TODO** |
| Build regression tests to ensure servers respond sanely to queries of death | prevent | clarac | Bug 5554827 **TODO** |
| Deploy updated search ranking subsystem to prod | prevent | jennifer | n/a **DONE** |
| Freeze production until 2015-11-20 due to error budget exhaustion, or seek exception due to grotesque, unbelievable, bizarre, and unprecedented circumstances | other | docbrown | n/a **TODO** |


## 得られた教訓

### うまくいったこと

*   監視システムが、HTTP 500 エラーの発生率が非常に高いこと(ほぼ100%に達していた)を迅速にアラートした
*   更新されたシェイクスピアコーパスを全クラスターに迅速に配信できた

### うまくいかなかったこと

*   連鎖的障害への対応手順に不慣れであった
*   失敗に終わった例外的なトラフィック急増により、利用可能性(アベイラビリティ)のエラー予算を大幅に(数桁上回る規模で)超過した
    

### 運がよかったこと

*   シェイクスピア愛好家のメーリングリストが新しいソネットのテキストを保有していた
*   サーバーログにクラッシュの原因がファイル記述子の枯渇であることを示すスタックトレースが残っていた
*   人気の検索単語を含む新しいインデックスをプッシュしたことで「死のクエリ」が解消された
    

## タイムライン


2015-10-21 (*all times UTC*)

| Time  | Description |
| ----- | ----------- |
| 14:51 | News reports that a new Shakespearean sonnet has been discovered in a Delorean's glove compartment |
| 14:53 | Traffic to シェイクスピアサーチ increases by 88x after post to */r/shakespeare* points to シェイクスピアサーチ engine as place to find new sonnet (except we don't have the sonnet yet) |
| 14:54 | **OUTAGE BEGINS** -- Search backends start melting down under load |
| 14:55 | docbrown receives pager storm, `ManyHttp500s` from all clusters |
| 14:57 | All traffic to シェイクスピアサーチ is failing: see <http://monitor/shakespeare?end_time=20151021T145700> |
| 14:58 | docbrown starts investigating, finds backend crash rate very high |
| 15:01 | **INCIDENT BEGINS** docbrown declares incident #465 due to cascading failure, coordination on #shakespeare, names jennifer incident commander |
| 15:02 | someone coincidentally sends email to shakespeare-discuss@ re sonnet discovery, which happens to be at top of martym's inbox |
| 15:03 | jennifer notifies shakespeare-incidents@ list of the incident |
| 15:04 | martym tracks down text of new sonnet and looks for documentation on corpus update |
| 15:06 | docbrown finds that crash symptoms identical across all tasks in all clusters, investigating cause based on application logs |
| 15:07 | martym finds documentation, starts prep work for corpus update |
| 15:10 | martym adds sonnet to Shakespeare's known works, starts indexing job |
| 15:12 | docbrown contacts clarac & agoogler (from Shakespeare dev team) to help with examining codebase for possible causes |
| 15:18 | clarac finds smoking gun in logs pointing to file descriptor exhaustion, confirms against code that leak exists if term not in corpus is searched for |
| 15:20 | martym's index MapReduce job completes |
| 15:21 | jennifer and docbrown decide to increase instance count enough to drop load on instances that they're able to do appreciable work before dying and being restarted |
| 15:23 | docbrown load balances all traffic to USA-2 cluster, permitting instance count increase in other clusters without servers failing immediately |
| 15:25 | martym starts replicating new index to all clusters |
| 15:28 | docbrown starts 2x instance count increase |
| 15:32 | jennifer changes load balancing to increase traffic to nonsacrificial clusters |
| 15:33 | tasks in nonsacrificial clusters start failing, same symptoms as before |
| 15:34 | found order-of-magnitude error in whiteboard calculations for instance count increase |
| 15:36 | jennifer reverts load balancing to resacrifice USA-2 cluster in preparation for additional global 5x instance count increase (to a total of 10x initial capacity) |
| 15:36 | **OUTAGE MITIGATED**, updated index replicated to all clusters |
| 15:39 | docbrown starts second wave of instance count increase to 10x initial capacity |
| 15:41 | jennifer reinstates load balancing across all clusters for 1% of traffic |
| 15:43 | nonsacrificial clusters' HTTP 500 rates at nominal rates, task failures intermittent at low levels |
| 15:45 | jennifer balances 10% of traffic across nonsacrificial clusters |
| 15:47 | nonsacrificial clusters' HTTP 500 rates remain within SLO, no task failures observed |
| 15:50 | 30% of traffic balanced across nonsacrificial clusters |
| 15:55 | 50% of traffic balanced across nonsacrificial clusters |
| 16:00 | **OUTAGE ENDS**, all traffic balanced across all clusters |
| 16:30 | **INCIDENT ENDS**, reached exit criterion of 30 minutes' nominal performance |

page:https://minegishirei.hatenablog.com/entry/2026/10/01/160940

Dear My Frends.: 個人開発宣伝ラボ - 個人開発者が間違った施策で時間を溶かさないための、心理学と実務知見のナレッジ共有コミュニティ