Home Work Personalized privacy-aware federated weather forecasting
Project
Personalized privacy-aware federated weather forecasting
I built a federated learning system that forecasts next-day weather across city-level clients, trained on non-IID station data in Flower's simulation framework, and benchmarked it against centralized, local-only and FedAvg baselines with custom per-client and global metric aggregation.
The problem
Weather stations in different cities do not produce interchangeable data. Each one carries its own climate, its own sensor quirks and its own seasonal profile, which makes the data non-IID in exactly the way that breaks naive distributed training. Pooling every station into one central dataset gives a model that is accurate on average and mediocre where it matters; training a separate model per city throws away everything the cities have in common.
The approach
I treated each city as a federated client and used Flower's simulation framework to run next-day forecasting rounds across them. Clients train locally on their own station records and share model updates rather than raw observations, which is what makes the setup privacy-aware by construction. On top of that, I added personalization so each city keeps a model adapted to its own conditions instead of accepting the global average wholesale.
What I benchmarked
The point of the project was the comparison, not a single number. I ran four training regimes over the same data and the same forecasting target:
- Centralized: all station data pooled, the conventional upper reference.
- Local-only: each city trains in isolation, no sharing at all.
- FedAvg: standard federated averaging across clients.
- Personalized federated: federated training with per-client adaptation.
Measuring it honestly
A global average error hides the failure mode that matters in federated learning: the client everyone else's data is pulling away from. I wrote custom metric aggregation so forecasting error was tracked both per client and globally across rounds, which is what makes the four regimes comparable and shows where personalization actually earns its place.
Scope
This is a simulation study run on historical station data, not a deployed service. The value is in the comparison and the measurement discipline: the same pipeline, the same split, four training strategies, and error reported where it can be inspected per client.