Every organization already assesses skills. The manager who decides who gets the hard project is assessing. The team lead who signs off a new hire on the line is assessing. The partner who staffs a client engagement from memory is assessing. The problem is that none of it is written down in a form anyone else can use, compare, or defend six months later.
That informal layer works until it has to scale. Three people on a team, fine. Three hundred across four sites, and the questions start arriving: who is actually qualified on this procedure, how do we know, and when did we last check. The answer to all three is a skills assessment that follows a method, lands in a record, and gets repeated.
This post is about the choosing part. Which method fits which skill, where the line sits between assessing employees and screening candidates, and what to demand from any tool that claims to do it. If you want the argument for why assessment data has to be trustworthy before you build anything on it, that is a separate post. This one is the practical menu.
What a Skills Assessment Is
A skills assessment is a structured judgment of how well a specific person can perform a specific skill, recorded against a defined proficiency scale at a known point in time. Four elements make it an assessment rather than an opinion. The skill is named precisely, not "communication" but "runs a customer escalation call to resolution." The scale is defined in advance with observable behaviors at each level. The assessor and method are recorded, so a self-rating and a manager rating are never confused. And the date is captured, because a skill assessed three years ago tells you something different from one assessed last quarter. Together those four elements turn a scattered set of impressions into a comparable data point, which is what allows an organization to see gaps across a team, plan succession, or prove competence to an auditor. Without them, you have a survey.
The Six Assessment Methods
There are six ways to assess a skill in a workforce setting. Most organizations should use three or four of them, matched to the skill, not one method applied to everything.
| Method | How it works | Validity | Cost to run | Use it when |
|---|---|---|---|---|
| Self-assessment | Employee rates their own proficiency against the scale | Low alone; useful as a starting point and for engagement | Very low | Building an initial inventory; low-risk skills; paired with another method |
| Manager assessment | Direct manager rates proficiency based on observed work | Moderate; depends on manager exposure and calibration | Low | Skills the manager sees weekly; performance-linked skills |
| Peer or 360 | Colleagues and internal customers rate observed behavior | Moderate to high for behavioral and collaborative skills | Moderate | Leadership, communication, cross-team skills the manager cannot observe |
| Test-based | Knowledge or practical test with scored outcome | High for knowledge; moderate for applied skill | Moderate to high | Regulated knowledge, technical certifications, tool proficiency |
| Evidence-based | Proficiency confirmed by artifacts: certifications, completed work, supervised sign-off | High when evidence is specific and dated | Moderate | Safety-critical procedures, licensed roles, compliance |
| Inferred | Proficiency estimated from work signals: systems used, projects completed, job history | Low to moderate; a hypothesis, not a verdict | Low once set up | Bootstrapping a skills inventory at scale before human verification |
The table hides the most important point, so here it is directly. No single method is right for every skill. A forklift certification should be evidence-based. A senior engineer's system design ability should be assessed by peers who have reviewed their designs. A new analyst's spreadsheet proficiency can start as a self-rating and be validated by a test. The method is a property of the skill, decided when the skill is defined.
On inference specifically
Inferring skills from résumés, project history, and system activity has become common, and it is genuinely useful for one thing: getting a first draft of the inventory without asking 5,000 people to fill out forms. Treat that draft as a set of hypotheses. Then verify the ones that matter with a human method. Organizations that stop at the inferred layer end up with an impressive-looking skills map that nobody trusts when a real decision depends on it. Bootstrap with inference, verify with assessment.
Skills Assessment vs. Pre-Employment Testing
Two things share the phrase "skills assessment" and they are not the same product.
Pre-employment testing screens candidates. Its job is to rank strangers on a narrow set of skills in a single sitting, with the emphasis on cheat-resistance and speed. The output is a hiring decision. Once the candidate is hired, the test result is rarely looked at again.
Workforce skills assessment measures people you already employ, repeatedly, against the full set of skills their role requires. Its job is to produce a living record that feeds gap analysis, development plans, staffing, and compliance. The output is a data layer the organization runs on.
The confusion matters because the tools built for the first job are poor at the second. A screening platform optimizes for proctoring and question banks. A workforce assessment system optimizes for scales, multi-rater input, evidence, currency, and what happens after the score is recorded. If a vendor demo spends most of its time on question libraries, you are looking at a screening tool.
Designing an Assessment That Produces Comparable Data
The value of an assessment is that it can be compared: this person to that person, this quarter to last quarter, this site to that site. Four design decisions determine whether comparison is possible.
Pick one scale and define every level
A 0 to 4 or 1 to 5 scale is standard. What makes it work is the behavioral anchor at each level. "Level 3: performs the task independently under normal conditions; escalates exceptions" is comparable across assessors. "Level 3: proficient" is not. Anchors should describe what you would see the person do, not how good they are. The full treatment of proficiency scales is in the proficiency levels post.
Calibrate the assessors
Two managers using the same anchors will still drift. A short calibration exercise, where five managers independently rate the same three anonymized cases and then discuss the differences, closes most of the gap. Repeat it annually and whenever a new manager starts assessing. This is the step most programs skip, and it is why their data varies more by assessor than by employee.
Set the cadence per skill
Compliance and safety skills need reassessment on a fixed schedule, often annually, sometimes quarterly. Fast-moving technical skills need reassessment when the underlying tool changes. Stable behavioral skills can go longer. Skills decay at different rates, so a single organization-wide cadence is either too frequent for some skills or dangerously infrequent for others. Set the interval when you define the skill.
Record the method with the score
A record that says "Level 3" is incomplete. A record that says "Level 3, manager-assessed by J. Ortiz on 2026-08-14, evidence attached" is usable. When the method travels with the score, you can weight self-ratings differently from evidence-backed ratings in every downstream report.
The Assessment Record: A Template
Whatever tool you use, every assessment should capture the same fields. If you are starting in a spreadsheet, these are your columns.
| Field | Why it matters |
|---|---|
| Person | The individual assessed |
| Skill | The precise skill, from the shared library, not free text |
| Required level for role | So the gap is visible in the same row |
| Assessed level | The score on the defined scale |
| Method | Self, manager, peer, test, evidence, inferred |
| Assessor | Who made the judgment |
| Date assessed | Drives currency and reassessment |
| Reassess by | Calculated from the skill's decay interval |
| Evidence | Link or attachment: certificate, test result, sign-off, work sample |
| Notes | Context an auditor or the next manager would need |
A spreadsheet with these ten columns is a legitimate skills assessment system for a team of twenty. It stops being one somewhere around a hundred people, or the first time an auditor asks to see the history of a single row.
What to Look for in Skills Assessment Software
The category is crowded. Point solutions do one method well. Learning platforms bolt on a self-rating. HCM suites add a skills field. The test below separates tools that assess from tools that collect ratings.
Scale flexibility
You should be able to define your own levels and anchors per skill family. A tool that ships one fixed five-point scale for every skill in the organization will be wrong for most of them.
Multiple assessment methods on the same skill
A single skill should be able to carry a self-rating, a manager rating, and an evidence record at once, each stored separately. Tools that overwrite one with another destroy the comparison you need.
Evidence attachments
Certificates, test results, supervisor sign-offs, and work samples should attach to the specific assessment record, not to the person's profile in general. This is what makes a record defensible.
Currency and expiry
Every assessment should carry a reassess-by date driven by the skill's own interval. The system should surface what is expiring, not wait for someone to remember. Trained and current are different states, and the software should know the difference.
Gap output
The point of assessing is the gap. The tool should compute required minus assessed at the person, team, role, and site level without an export to a spreadsheet. If gap analysis is a separate module or a separate product, the assessment data is stranded.
Audit trail
Who changed what, when, and from what value. Regulated organizations need this to pass audits. Everyone else needs it the first time a promotion decision is challenged.
SkillsDB was built around these six because they are what assessment data has to have before anything else, staffing, succession, learning plans, can rely on it. The proficiency assessment and surveys capabilities are two ways in; the shared record underneath is the same.
Three Industries, Three Method Mixes
Manufacturing line qualification. A packaging line has forty operators and twelve stations. Each station has three or four skills with a required level. The method mix is evidence-based for the initial qualification (supervised sign-off, documented), manager-assessed at ninety days, and test-based on a fixed annual cycle for the safety-critical steps. The output is a matrix the shift supervisor can read at 6 a.m. to see who can cover station seven.
Data center operations. Technicians work across sites on procedures where a mistake takes down a customer. Critical procedures are evidence-based with a named qualifying engineer and a twelve-month currency window. Certifications from vendors are evidence records with expiry dates pulled from the certificate. Self-assessment is used only for the initial inventory of adjacent skills, and every self-rated critical skill is flagged until verified.
Professional services bench. A consultancy staffing client work needs a defensible view of who can lead a workstream. Peer assessment from prior project leads carries the most weight. Manager assessment covers delivery behaviors. Inferred skills from project history seed the inventory for new joiners and are verified after their first engagement. The resource manager staffs from verified levels, not from the résumé.
The Mistakes That Ruin Assessment Programs
Assessing everything at once. A full-workforce assessment of every skill produces a mountain of self-ratings nobody validates. Start with the skills tied to a decision you have to make this quarter.
Using one method because it is easy. Self-assessment for everything produces data that is systematically optimistic and unusable for compliance.
Skipping calibration. The data ends up describing the assessors more than the employees.
No reassessment cadence. The inventory is accurate on the day it is built and decays from there.
Storing scores without methods or evidence. Every downstream report has to treat a guess and a certificate as equal.
Treating the tool as the program. Software stores and computes. The method choices, anchors, and calibration are decisions people make, and they have to be made before the first assessment is entered.
Assessment Is How an Inventory Earns Trust
A skills inventory is a list of claims. Every row says this person can do this thing at this level. Assessment is the process that turns a claim into something a manager will staff against, an auditor will accept, and a leadership team will plan around. The method has to fit the skill, the scale has to be anchored, the assessors have to be calibrated, and the record has to carry its own provenance.
Get those four right and the tool question mostly answers itself. Get them wrong and no tool will save the data.