Behind the Feed Data Challenge
How can we use data to make digital information environments more transparent, diverse, and fair?
Data Starting Points
Teams may use publicly available datasets, APIs or other appropriately sourced data. The following resources are recommended starting points:
- Microsoft MIND News Recommendation Dataset: News articles and user behaviour data designed for recommendation systems research.
- GDELT: Global news and media data that can be used to explore differences in topic, geographic and source representation.
- MovieLens: User ratings and movie metadata that can be used to build and evaluate recommendation systems and investigate personalization and content diversity.
- Google Trends: Aggregated search-interest data over time and across locations. Useful for exploring consumer interests, trends, and differences between search demand and product or content visibility.
- Other public datasets: Teams may identify their own dataset related to recommendation, information exposure, media representation or algorithmic fairness, provided that the data source and methodology are clearly documented.
Teams are encouraged to consider whether their dataset contains sufficient information to support the conclusions they want to draw. Students should not infer demographic bias when demographic information is not present in the data.
Every project should identify:
- Where the data came from
- What the data represents and does not represent
- What proxy they are using for "information exposure"
- What limitations or biases exist in the dataset
- Why their chosen fairness metric or comparison is appropriate
