In classical machine learning we assume training data is IID: independent and identically distributed, as if all examples were drawn from one big shared distribution. In federated learning this assumption breaks completely. Your typing habits, a hospital's patient population, a bank's regional customers — every client's data reflects its own local reality.
We call client data non-IID (heterogeneous) when distributions differ across clients. This can take several forms: label distribution skew (one hospital sees mostly pneumonia cases, another mostly fractures), feature distribution skew (different phone cameras produce different image statistics), and concept drift (the same input means different things for different users).
Heterogeneity has a concrete, measurable consequence for FedAvg: client drift. When each client trains locally on its own skewed data, local models drift toward local optima. Averaging drifted models does not recover the model you would get from centralized training — the global model converges to a worse solution, and the problem worsens with more local epochs.
Much of intermediate and advanced FL research is a response to this single phenomenon: methods that correct client drift, personalize models instead of forcing one global model, or cluster clients with similar data. Keep the phrase 'client drift' in mind — you will see it again and again.