Amazon Purchase Predictor Regression Model

A random forest regression workflow for predicting customer spending from purchase and demographic features.
Author

Khang Thai

Published

July 28, 2024

R Random Forest Tidyverse Caret ggplot2

Introduction

The Amazon Purchase Predictor Regression Model analyzes purchasing behaviors based on demographic and transaction data. The primary goal was to predict log_total, the total spending amount in log scale, using explanatory variables from the dataset. By implementing regression techniques, this study identifies key factors influencing customer spending habits and provides insight into consumer behavior.

Process

  1. Data Preprocessing:
    • Cleaned and handled missing values.
    • Normalized and transformed variables where necessary.
    • Selected relevant features based on correlation and importance.
  2. Model Selection and Training:
    • Experimented with linear regression, ridge regression, lasso regression, random forest, and KNN regression.
    • Tuned hyperparameters using cross-validation for optimal performance.
  3. Feature Engineering:
    • Identified key explanatory variables such as count_hh4 and count_G, which showed strong correlation with log_total.
    • Considered interactions between income level and purchasing frequency to refine the model.

Outcome

The final model selected was a Random Forest regression model, which achieved:

  • Root Mean Squared Error (RMSE): 0.1128934
  • Standard Error (SE): 0.003002182
  • R-squared Score: Higher than 0.9587, indicating strong predictive power.
  • Household size, order count, and income level were significant predictors of Amazon purchases.
  • KNN regression was tested but did not outperform tree-based models in terms of accuracy.