← Back to portfolio

Employee Satisfaction Analysis

Python · Pandas · Scikit-learn · NLTK / TextBlob

Brief

A 142-row HR dataset — demographics, tenure, income, a 0–100 satisfaction score, and an open-ended text review per employee. Three angles on the same question — what's actually driving satisfaction: unsupervised clustering to find natural employee segments, a decision tree to rank which features predict dissatisfaction, and sentiment analysis on the free-text reviews themselves.

Employees
142
9 fields each
Age
37 yrs
range 19–58
Monthly income
$8.35K
range $1.56K–$20.24K
Satisfaction
59.5
range 20–100

Cleaning

Two missing values in the free-text review column were filled with "No feedback provided" rather than dropped, to keep every row for the numeric analysis. No duplicate rows. Columns were renamed to consistent snake_case, and the row_id identifier was dropped before modelling since it carries no signal.

Exploring the data

Business travel splits roughly 52% non-travel, 37% frequent travellers, and the rest rarely-travel. Education skews toward A-Level and Higher Certificate holders, with Bachelors, Diploma, Masters and Doctorate each in the minority. The workforce is 62% male, 38% female.

Histograms of age, monthly income, total working years and satisfaction score
Fig 1. — Age and tenure skew younger/junior; income is bimodal (a distinct lower- and higher-earning group); satisfaction is fairly evenly spread rather than clustered near one end.
Bar charts of business travel, education and gender counts
Fig 2. — Category counts for business travel, education level and gender.

What correlates with satisfaction

The standout relationship in the data isn't with satisfaction at all: age and total working years are strongly correlated (0.93) — unsurprising, since tenure accumulates with age. Income correlates moderately with both age (0.57) and tenure (0.55). Satisfaction score, though, barely moves with any of them — age (0.03), tenure (0.11), income (0.22) — which suggests seniority and pay aren't what's driving how satisfied people say they are.

Correlation heatmap of age, monthly income, total working years and satisfaction score
Fig 3. — Correlation heatmap: strong age↔tenure relationship, weak links to satisfaction.
Scatter plot of monthly income against satisfaction score
Fig 4. — Income vs. satisfaction isn't a smooth trend — it reads more like distinct groups, which is what motivated formal clustering rather than eyeballing the scatter.

Preparing for clustering

Numeric features (age, income, tenure, satisfaction) were Min-Max scaled to a 0–1 range, and the three categorical fields (business travel, education, gender) were label-encoded, since K-Means needs purely numeric input.

K-Means: how many segments?

Running the elbow method across k = 1 to 10, inertia drops sharply through k = 4 and flattens out after that. A silhouette analysis over the same range agrees — k = 4 scores highest at 0.74, well clear of every other option — so both methods point to the same number of segments.

Elbow method plot showing inertia against number of clusters
Fig 5. — Elbow method: the bend at k=4 marks where adding more clusters stops meaningfully reducing inertia.
Silhouette score plot peaking at 4 clusters
Fig 6. — Silhouette analysis agrees: 4 clusters scores highest (0.74).

Four employee segments

Fitting K-Means with k = 4 on scaled age, income, tenure and satisfaction splits the workforce into four fairly distinct groups when plotted on income against satisfaction — and the chart makes one thing obvious straight away: income alone doesn't predict satisfaction.

Scatter plot of the four K-Means clusters by monthly income and satisfaction score
Fig 7. — Four segments by income and satisfaction. Two of the higher-income groups aren't the most satisfied ones.

Predicting satisfaction with a decision tree

Beyond describing segments, a decision tree classifier was trained to predict whether an employee sits above or below the satisfaction midpoint, using demographic and role features. On a held-out 20% test split it reached 86.2% accuracy — 90% precision and recall on "Satisfied", 78% on "Dissatisfied" (the softer number likely reflects the smaller class in a 142-row dataset).

Bar chart of decision tree feature importances, led by frequent business travel
Fig 8. — What the tree actually splits on: frequent travel dominates, then tenure, then income.

Frequent business travel is the single strongest predictor of dissatisfaction — by a wide margin over everything else the model had access to, including pay. Total working years and monthly income follow well behind it.

What employees are actually saying

The free-text reviews were cleaned, tokenised and scored for sentiment with TextBlob. Read across all 142 reviews, feedback skews positive but not overwhelmingly so:

Pie chart of sentiment distribution across employee reviews: 55% positive, 27.5% neutral, 17.6% negative
Fig 9. — 55% positive, 27.5% neutral, 17.6% negative across all reviews.
Word cloud of the most frequent words in employee feedback, dominated by work, pay, travel, management and good
Fig 10. — Most frequent words after stopword removal — "pay," "travel" and "management" sit right alongside "good" and "happy."

Cutting sentiment by business travel frequency lines up with the decision tree's top predictor almost exactly: frequent travellers post the highest share of negative reviews by a clear margin, while non-travellers and rare travellers skew far more positive.

Stacked bar chart of sentiment share by business travel frequency, with Travel Frequently showing the most negative sentiment
Fig 11. — Sentiment by travel frequency — the "Travel Frequently" group carries the most negative and least positive feedback of the three.

The two clearest extremes: "ServiceFirst is a great start for my career, learning a lot" (most positive) against "Office politics are the worst" (most negative) — career growth and workplace culture are pulling in opposite directions.

Where this points

Three independent methods — clustering, a decision tree, and sentiment analysis of free text — converge on the same story: