Predictive people analytics · public sample data

Which Employees Are Most Likely to Leave, and Why

An attrition risk model built and tested in Python, with the drivers behind it and the guardrails needed before a model like this is used on real employees.

Data: IBM HR Analytics sample (fictional, 1,470 employees)Models: logistic regression · random forestTools: Python · scikit-learn · pandas

Executive summary

  • 16.1%attrition in the dataset
  • 0.79ROC AUC on unseen employees
  • 64%of leavers flagged by the model
  • 3×attrition with regular overtime

A logistic regression correctly ranks a leaver above a stayer 79% of the time on employees it has not seen. At the default threshold it flags 38 of the 59 leavers in the test set, alongside 64 false alarms. Overtime, frequent travel, short tenure and entry-level roles such as sales representative and laboratory technician carry the highest risk.

Recommendation: use the model to target retention work at groups and conditions (overtime, travel, early tenure), not to label individuals, and validate it on the organisation's own data before any use.

01

Business question

Replacing an employee costs recruitment, onboarding and lost productivity. HR teams want to know which parts of the workforce are most at risk so they can act before people resign. This project tests how well standard HR data predicts attrition and which factors matter most.

The dataset is a fictional sample created by IBM data scientists and published for learning. It is not from a real organisation, and the results show the method rather than facts about any employer.

02

Data and preparation

1,470 employees and 35 columns covering role, department, level, pay, tenure, overtime, travel, satisfaction and performance ratings. 237 employees (16.1%) left.

  • Removed as uninformativeEmployeeCount, Over18 and StandardHours have one value for everyone. EmployeeNumber is an ID.
  • Excluded on purposeAge, gender and marital status are protected characteristics. They were left out so the model cannot base risk on them.
  • EncodingNumbers were standardised so coefficients can be compared. Categories such as job role were one-hot encoded, with one level held as the baseline.
  • Train and test split75% for training, 25% (368 employees, 59 leavers) held back for testing, keeping the same leaver share in both.
03

The model explained

Pipeline

prep = ColumnTransformer([
    ("num", StandardScaler(), numeric),
    ("cat", OneHotEncoder(handle_unknown="ignore", drop="first"), categorical),
])
model = Pipeline([("prep", prep),
                  ("clf", LogisticRegression(max_iter=2000, class_weight="balanced"))])

Preparation and model sit in one pipeline, so scaling is learned from training data only and applied unchanged to the test set. class_weight="balanced" gives leavers more weight, because only one in six employees left and an unweighted model would learn to predict "stays" for almost everyone.

Validation

cv = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)
cross_val_score(model, X_train, y_train, cv=cv, scoring="roc_auc")

Five-fold cross-validation on the training data checks that performance is stable, before a single final check on the held-back test set.

A random forest was trained the same way for comparison. It captures interactions between factors but is harder to explain.

04

How well it works

ModelCV ROC AUCTest ROC AUCPrecisionRecallF1
Logistic regression0.834 ± 0.0420.7940.3730.6440.472
Random forest0.799 ± 0.0450.7580.4560.5250.488
  • ROC AUCThe chance the model ranks a random leaver above a random stayer. 0.5 is guessing, 1.0 is perfect.
  • RecallShare of actual leavers the model flags. 64% means 38 of 59.
  • PrecisionShare of flagged employees who actually left. 37% means most flags are false alarms.
ROC curves on the held-out test set. The logistic regression curve, AUC 0.79, sits slightly above the random forest curve, AUC 0.76. Both are well above the diagonal random-guess line.
ROC curves on the 368 test employees.
Confusion matrix for the logistic regression at a 0.5 threshold: 245 correctly predicted to stay, 64 predicted to leave who stayed, 21 predicted to stay who left, 38 correctly predicted to leave.
Logistic regression at a 0.5 threshold: 38 leavers caught, 21 missed, 64 false alarms.

The logistic regression was chosen. It ranks employees better on unseen data and each coefficient can be explained to HR and managers. The threshold can be moved: a higher threshold means fewer false alarms but more missed leavers.

05

What drives attrition

Two bar charts. Left: attrition is 10.4 percent for employees without overtime and 30.5 percent with overtime. Right: attrition by job role, highest for sales representatives at 39.8 percent and lowest for research directors at 2.5 percent.
Attrition rate by overtime and by job role.

Employees working regular overtime left at 30.5%, three times the rate of those who did not.

GroupAttritionCompared withAttrition
Overtime30.5%No overtime10.4%
Travels frequently24.9%No travel8.0%
0 to 1 year at the company34.9%11+ years8.1%
Job level 126.3%Job level 44.7%
Sales representative39.8%Research director2.5%
Bar chart of the ten largest logistic regression coefficients. Laboratory technician, overtime, frequent travel and sales representative increase risk most. Education field other, research director and total working years reduce risk.
The ten largest standardised coefficients. Right of zero raises risk, left lowers it.
  • Overtime and frequent travel raise risk even after role, pay and tenure are taken into account. Both are working conditions the organisation can change.
  • Job role coefficients are measured against healthcare representatives, the baseline role.
  • Leavers had a median monthly income of 3,202 against 5,204 for stayers (units as supplied in the dataset). Pay overlaps with job level and experience, so the model spreads the effect across these, and single coefficients such as job level should not be read alone.
06

Limitations and responsible use

  • Fictional dataResults describe a sample dataset. A real organisation would need to rebuild and test the model on its own records.
  • Association, not causeThe model finds patterns linked to leaving. It does not prove that reducing overtime will keep a particular person.
  • False alarmsMost flagged employees stay. Treating a flag as a fact could damage trust or lead to unfair decisions.
  • Fairness checksProtected characteristics were excluded, but other fields can act as proxies. Results should be checked for different outcomes by group before use.
  • PrivacyEmployees should be told how their data is used, and a data protection impact assessment should be completed under UK GDPR.
  • DriftPatterns change with the labour market and the organisation. The model needs retesting at least every six months.
07

Recommendations

  1. Review overtime patterns first. It is the strongest controllable driver. Report overtime by team monthly and set a threshold for review.
  2. Strengthen the first year. A third of employees with a year or less of service left. Structured onboarding and check-ins at 30, 90 and 180 days target the highest-risk period.
  3. Look at entry-level roles in sales and the laboratory, including pay against the market and progression routes.
  4. Use risk scores at group level in workforce planning and retention budgets, not as individual labels in performance or promotion decisions.
  5. Pilot on real data with governance: HR, legal and data protection sign-off, fairness testing and a regular review of accuracy.
08

How to reproduce it

  1. The IBM HR Analytics Employee Attrition & Performance dataset is in attrition-risk-model/data/.
  2. Run python python/attrition_model.py. It trains both models with a fixed random seed and writes outputs/metrics.json and the four charts.

View the Python, data and outputs on GitHub

Prepared by Olajumoke O. Medunoye