Replicating—and extending—the paper
We began by reproducing Predicting Food Crises Using News Streams, the 2023 Science Advances study by Ananth Balashankar, Lakshminarayanan Subramanian, and Samuel P. Fraiberger. We rebuilt its news, traditional, and combined-feature Random Forest experiments, then adapted the methodology beyond its English-language global dataset to Pakistan's local news ecosystem.
- Reproduced the paper's Random Forest forecasting pipeline and temporal evaluation setup.
- Implemented the data path in Pandas and Polars, then moved training to cuDF and cuML.
- Cut Random Forest training from 22 minutes 22 seconds to 14.7 seconds—about a 90× speedup.
Comparing multiple model families
The project did not rely on one forecasting model. We compared tree-based, linear, ordinal, and language-model approaches to understand both predictive performance and the trade-offs between speed, interpretability, and localized reasoning.
- Random Forest models in scikit-learn and cuML replicated the original paper's primary forecasting approach.
- OLS and Lasso panel regressions provided interpretable fixed-effects baselines, while LogisticAT modeled the ordered IPC scale directly.
- GPT-4o and Claude 3.7 Sonnet powered single-model and two-agent summarization-and-prediction workflows for Pakistan.
From news to usable signals
To localize the original methodology, we collected 162,653 Urdu articles from Dawn, Jang, and UrduPoint, translated and classified them, and combined their food-security signals with district-level weather data.
- Collected 27,167 Dawn, 108,693 Jang, and 26,793 UrduPoint articles.
- Structured 5,569 relevant articles by geography and food-security risk factors.
- Added monthly rainfall, temperature, soil-moisture, and other weather signals.
From replication to 94% accuracy
The Pakistan extension progressively improved IPC phase prediction by enriching local news with environmental evidence and separating article summarization from final forecasting.
- News-only GPT-4o prediction reached 69% accuracy.
- Adding weather raised accuracy to 91%; a two-agent workflow reached 93%.
- The final Claude 3.7 Sonnet approach reached 94% accuracy.