Skip to content
FedLab
Sign in

Communication and Systems Constraints

The engineering reality: bandwidth, stragglers, and dropouts.

Federated learning is as much a systems problem as a machine-learning problem. In cross-device settings, the server may coordinate millions of clients over slow, metered, unreliable connections. Upload bandwidth is typically the scarcest resource, and clients can drop out at any moment — battery dies, Wi-Fi drops, the user closes the app.

Several practical techniques address the communication bottleneck. Client selection limits each round to a small fraction of clients. Compression reduces the size of updates through quantization (fewer bits per parameter) or sparsification (sending only the largest changes). Local epochs, as in FedAvg, trade computation on the client for fewer rounds of communication.

Dropouts and stragglers force design decisions too. If the server waits for every selected client, the slowest client dictates the round time; if it proceeds with whoever has responded, the aggregate is biased toward clients with fast devices and good connections — which often correlates with wealthier users and newer hardware. Fairness in FL begins at the systems layer.

This systems perspective is what separates FL papers from FL deployments. As you progress through FedLab, keep asking of every method: what does it cost in communication, and what happens when clients disappear?