Predicting Heavy Gamers with Machine Learning and an LLM Data Analyst
A machine learning project that predicts heavy gamers from personal and behavioural data, comparing logistic regression and a decision tree, plus a LangChain bot that answers questions about the data in plain English.
On this page 6 sections
Overview
This project builds a model that predicts whether someone is a heavy gamer, meaning they play online games for 3 or more hours a day, based on personal and behavioural data.
Why it matters
Understanding how gamers behave helps with better game design, smarter marketing and tools for healthier digital habits. Spotting high engagement profiles can help design features that keep games fun while encouraging good time management.
The dataset
The data comes from 118 online gamers in Pakistan and covers demographic and gaming details, such as:
- Age, gender, education, income and occupation
- Daily play time, how often they play, and game difficulty
- Motivations like stress relief, achievement or social connection
After cleaning and preprocessing, 12 useful features were kept for training. That meant handling missing values, encoding categories, converting time units and standardising text answers.
Feature engineering
- Target variable. A binary
heavy_gamercolumn marks anyone who plays 3 or more hours a day as 1. - Categorical encoding. Text fields like gender, education and occupation were turned into numbers.
- Motivation columns. Multi answer responses (for example "Stress Relief, Achievement") were split into separate yes or no columns.
- Clean up. Duplicate and unneeded columns were removed.
Models and results
Model 1: Logistic regression
A classic binary classifier came first. With an 80/20 train and test split and feature scaling, it reached:
- Accuracy: 91.67%
- AUC score: 0.84
- Evaluated with a ROC curve and a confusion matrix
View the model on Google Colab
Model 2: Decision tree classifier
This model classified the test data perfectly:
- Accuracy: 100%
- AUC score: 1.00
- The ROC curve hit the ideal top left corner and the confusion matrix showed zero errors
View the model on Google Colab
Comparison
Both models did well. The decision tree's perfect score is likely overfitting because the dataset is small. Logistic regression scored a little lower but should generalise better to new data.
Bonus: an LLM powered data analyst
To make the data easier to explore, I built an AI data analyst bot with LangChain and OpenAI GPT-3.5. Anyone can ask questions about the dataset in plain English, no code needed.
You can ask things like:
- "Which gender has more heavy gamers?"
- "Show average hours by education level"
- "Plot heavy gamers by occupation"
- "Bar chart of average playtime by age"
It answers with written insights and live charts made with matplotlib and seaborn. It shows how LLMs can make data analysis far more approachable for people who do not write code.

Want something like this built?
I build web apps, AI tools and custom software. Tell me your idea.

