NYC supermarket secrets
What does the language of 1,331 Google Maps reviews reveal about how New Yorkers relate to where they shop? This project uses NLP to surface linguistic asymmetries across 267 stores, 14 chains, and 3 boroughs (Manhattan, Brooklyn, and Queens) — asking whether vocabulary, sentiment, and complaint patterns vary by chain, neighbourhood, and socioeconomic context.
The pipeline:
- Google Places API for collection across 14 NYC neighbourhoods
- VADER for sentence-level sentiment scoring
- TF-IDF at store and chain level to surface distinctive vocabulary
- Keyword-based categorisation for complaint and praise types (quality, service, price, checkout, availability)
- Vocabulary complexity metrics (diversity, word length, sentence length) as socioeconomic proxies
- Heuristic fake review detector using multi-factor scoring (generic language patterns, extreme sentiment, review length, absence of product specifics)
Key findings:
- Whole Foods reviewers use the most distinctive and complex vocabulary
- Astoria and Harlem have the highest complaint-to-praise ratios
- Long Island City concentrates complaints almost entirely in quality and service
- Food Bazaar and Trader Joe's score highest on positive sentiment
- Vocabulary complexity correlates with neighbourhood income level — language is a legible socioeconomic signal
- 2.6% of reviews flagged as potentially fake — below industry average