Start here. This is the direct spoken answer to practice first.
Why this question matters
Bad alerts train teams to ignore production. Good alerts wake people for user-impacting symptoms or imminent risk.
I would alert on symptoms users care about: high error rate, high latency, failed critical workflows, queue age, or dependency failure that affects the product. I would avoid paging people for every CPU blip or single log error. Alerts need severity, ownership, runbook, and a clear action. If nobody knows what to do, the alert is not ready.