← Back
▸ GRADUATE COURSEWORK · BIG DATA / MACHINE LEARNING

Building a Movie Recommender with Apache Spark

OMSBA 5240, Big Data Analytics · Seattle University

THE SETUP

A fictional movie-review startup, "Ripe Pumpkins," wants a "Pumpkinmeter" score — a personalized recommendation engine like the ones streaming platforms use.

I built and tested that recommender using Apache Spark's MLlib ALS (Alternating Least Squares) algorithm on the full MovieLens dataset: roughly 33 million ratings across 86,000 movies from 330,000+ users. Two real people (myself and a family member) each rated 10 movies, and the model generated personalized top-15 recommendations for each of us under two different filtering thresholds, to show the board how the Pumpkinmeter would actually behave in production.

THE TECHNICAL APPROACH

Ratings and movie data were loaded as RDDs and parsed with map/filter transformations. I tuned the model's rank parameter on the small MovieLens dataset first, picking whichever rank gave the lowest RMSE on a validation split, then trained the final model on the complete dataset and evaluated it on a held-out test split.

To generate personal recommendations, each new user's 10 ratings were unioned into the full dataset, the model was retrained, and predictAll() scored every movie that user hadn't already rated. Candidates were then filtered by a minimum ratings-count threshold (25 vs. 100) and ranked by predicted rating to produce the final top 15.

DEBUGGING: WHEN THE TUTORIAL'S ASSUMPTIONS DON'T HOLD

The first run failed immediately — a Py4JJavaError claiming the input path didn't exist. The tutorial hardcodes a path that assumes one specific zip-extraction layout, but depending on how the archive unpacks, the real files can end up nested one folder deeper than expected.

Instead of hand-fixing the path once and hoping it held, I wrote a small find_dataset_folder() helper that checks both possible locations — and falls back to walking the folder tree — for both the small and complete datasets. If the files genuinely aren't there, it raises a clear error naming exactly which paths it checked, instead of a cryptic Java stack trace.

def find_dataset_folder(base_path, dataset_name):
    """
    Locate the ratings/movies CSVs regardless of how the zip
    extracted. Checks the expected path first, then the common
    'nested folder' variant, then falls back to a tree walk.
    Raises a clear error listing every path checked if nothing
    is found -- no cryptic Py4JJavaError.
    """
    candidates = [
        os.path.join(base_path, dataset_name),
        os.path.join(base_path, dataset_name, dataset_name),
    ]
    for path in candidates:
        if os.path.exists(os.path.join(path, "ratings.csv")):
            return path

    for root, dirs, files in os.walk(base_path):
        if "ratings.csv" in files:
            return root

    raise FileNotFoundError(
        f"Could not find ratings.csv. Checked: {candidates}"
    )

A second, smaller issue: MovieLens titles with an internal comma get quoted in the CSV (e.g. "Shawshank Redemption, The (1994)"), but naive line.split(",") parsing ignores those quotes, truncating several titles in the raw output. It's a display-only issue — the predicted rating and rating count are unaffected — so I used the corrected full titles throughout this report while leaving the raw output flagged as-is in the notebook.

RESULTS

Top recommendations for User 1, compared across both filtering thresholds:

Scenario 1 (≥25 ratings)PredictedScenario 2 (≥100 ratings)Predicted
Wallace & Gromit: Best of Aardman4.71The Shawshank Redemption4.70
The Shawshank Redemption4.70Star Wars: Episode IV4.58
Patton4.62Schindler's List4.58
The Philadelphia Story4.59The Princess Bride4.55
Star Wars: Episode IV4.58The Godfather4.53

Full top-15 lists for both users and both scenarios are in the notebook and report.

INSIGHTS FOR THE BOARD

Discovery vs. reliability is a product lever, not a fixed setting. The lower threshold (25 ratings) surfaces niche, cult titles a user might never find on their own — but with far less social proof behind them. The higher threshold (100) trades that discovery value for titles hundreds of other users have already validated. Ripe Pumpkins could offer both as modes: a "Discover" setting for exploration, a "Safe Bet" setting for users who bounce off unfamiliar recommendations.
The model matches rating behavior, not genre. User 1 rated only animated/family films highly, yet the model's top recommendations were almost entirely adult dramas and classics. ALS collaborative filtering finds people with similar rating patterns across the whole catalog, not similar genres. That's by design, but it needs in-product framing ("customers with similar taste also loved...") so it doesn't read as a mismatch to end users.
Raw predicted ratings aren't comparable across users. User 2 rates everything lower on average, so even their best recommendations only reached ~3.1–3.4, versus User 1's 4.5–4.7. A single platform-wide "recommend if predicted > X" cutoff would systematically under-serve naturally harsher raters. Scores should be normalized against each user's own rating history, not compared on a shared raw scale.