A fictional movie-review startup, "Ripe Pumpkins," wants a "Pumpkinmeter" score — a personalized recommendation engine like the ones streaming platforms use.
I built and tested that recommender using Apache Spark's MLlib ALS (Alternating Least Squares) algorithm on the full MovieLens dataset: roughly 33 million ratings across 86,000 movies from 330,000+ users. Two real people (myself and a family member) each rated 10 movies, and the model generated personalized top-15 recommendations for each of us under two different filtering thresholds, to show the board how the Pumpkinmeter would actually behave in production.
Ratings and movie data were loaded as RDDs and parsed with map/filter transformations. I tuned the model's rank parameter on the small MovieLens dataset first, picking whichever rank gave the lowest RMSE on a validation split, then trained the final model on the complete dataset and evaluated it on a held-out test split.
To generate personal recommendations, each new user's 10 ratings were unioned into the full dataset, the model was retrained, and predictAll() scored every movie that user hadn't already rated. Candidates were then filtered by a minimum ratings-count threshold (25 vs. 100) and ranked by predicted rating to produce the final top 15.
The first run failed immediately — a Py4JJavaError claiming the input path didn't exist. The tutorial hardcodes a path that assumes one specific zip-extraction layout, but depending on how the archive unpacks, the real files can end up nested one folder deeper than expected.
Instead of hand-fixing the path once and hoping it held, I wrote a small find_dataset_folder() helper that checks both possible locations — and falls back to walking the folder tree — for both the small and complete datasets. If the files genuinely aren't there, it raises a clear error naming exactly which paths it checked, instead of a cryptic Java stack trace.
def find_dataset_folder(base_path, dataset_name):
"""
Locate the ratings/movies CSVs regardless of how the zip
extracted. Checks the expected path first, then the common
'nested folder' variant, then falls back to a tree walk.
Raises a clear error listing every path checked if nothing
is found -- no cryptic Py4JJavaError.
"""
candidates = [
os.path.join(base_path, dataset_name),
os.path.join(base_path, dataset_name, dataset_name),
]
for path in candidates:
if os.path.exists(os.path.join(path, "ratings.csv")):
return path
for root, dirs, files in os.walk(base_path):
if "ratings.csv" in files:
return root
raise FileNotFoundError(
f"Could not find ratings.csv. Checked: {candidates}"
)
A second, smaller issue: MovieLens titles with an internal comma get quoted in the CSV (e.g. "Shawshank Redemption, The (1994)"), but naive line.split(",") parsing ignores those quotes, truncating several titles in the raw output. It's a display-only issue — the predicted rating and rating count are unaffected — so I used the corrected full titles throughout this report while leaving the raw output flagged as-is in the notebook.
Top recommendations for User 1, compared across both filtering thresholds:
| Scenario 1 (≥25 ratings) | Predicted | Scenario 2 (≥100 ratings) | Predicted |
|---|---|---|---|
| Wallace & Gromit: Best of Aardman | 4.71 | The Shawshank Redemption | 4.70 |
| The Shawshank Redemption | 4.70 | Star Wars: Episode IV | 4.58 |
| Patton | 4.62 | Schindler's List | 4.58 |
| The Philadelphia Story | 4.59 | The Princess Bride | 4.55 |
| Star Wars: Episode IV | 4.58 | The Godfather | 4.53 |
Full top-15 lists for both users and both scenarios are in the notebook and report.