An operator can score 18 out of 20 on a quiz about their work procedures and still lose their rhythm as soon as they pick up the phone, because knowledge is not the same as practical skill.
This is the discrepancy highlighted by quality standards—from ISO 9001 to NADCAP, including IFS and IATF 16949—when they ask: « How do you know he's really good at it?" .
Assessing skills in a real-world work setting involves providing this evidence by observing actual performance using a standardized evaluation rubric. The system consists of three components that work together: a rubric with observable criteria, a checklist used on the job, and a trained evaluator to minimize bias.
What a practical assessment measures, and what a theoretical assessment overlooks
The theoretical assessment, conducted via multiple-choice questions or an online questionnaire, verifies that the operator is familiar with the procedure, the QHSE guidelines, and how the machine operates. It measures declarative knowledge. It's necessary, but it doesn't say anything about what happens once you've put on your helmet and gloves.
The practical evaluation observes the procedure, performed at the correct pace, either at an actual workstation or at a simulated training station. It validates the operational know-how : line startup, adherence to cycle times, in-process quality control, handling of unexpected events. On a food packaging line, the difference between the two can be significant. A temporary operator may know that the pasteurization temperature is 72°C for 15 seconds, but you have to see them at the control screen to know if they can react when the sensor malfunctions.
Skipping the practical assessment has two direct consequences: the competency matrix shows a level of proficiency that does not exist, and during a NADCAP or IFS audit, the auditor requests proof of real-world application that no one can provide. Clause 7.2 of the ISO 9001 standard states this explicitly: the company must retain documented evidence demonstrating competence, not just the training completed.
What exactly should you look for at a workstation?
A reliable practical assessment is based on three dimensions that should be addressed in this order, starting with the’execution of the technical movement. Is the operator following the machining sequence in the correct order and with the correct parameters? On an aerospace machining station, this means: correctly setting the origins, running the expected CNC program, and performing an intermediate dimensional inspection before releasing the part. This is what the Training Within Industry method refers to as the job instructions : the work standard broken down into key steps and key points.
Next comes the compliance with QHSE rules and PPE requirements : proper use of protective equipment, logging before work begins, reporting of deviations, and adherence to traffic rules in the area. It is in this area that the distinction between job competency and certification comes into play: the former means «he knows how to do it,» while the latter adds «and he is authorized to do it, within a formalized regulatory framework.» An electrical authorization, a CACES forklift operator’s license, and an NF EN 9606 welding certification always include a mandatory practical evaluation.
That leaves the’Autonomy and Collective Behavior, an aspect that is more difficult to assess objectively. Does the operator report a deviation immediately? Do they know when to escalate the issue to the team leader? Do they tidy up their workstation before the shift change? These indicators carry as much weight as the action itself, especially when moving from one level of versatility to another.
Creating a Practical Evaluation Rubric in 5 Steps
A framework that can be used in the field can be developed in a few weeks—not six months of quality assurance work—by following five successive steps.
Mapping the Key Competencies for the Position
An effective assessment grid fits on one to two pages and covers a variable number of competencies depending on the technical nature of the position: approximately 3 to 5 for less technical roles—which are often found in the agri-food industry—up to a good dozen for the most technical and heavily audited roles, such as those in the aerospace industry. Beyond that, the evaluator’s focus becomes diluted, and the results become impossible to interpret.
The sources used to identify these skills are the job description, the work instruction, the TWI work standard (when available), and observation of one or two operators recognized as experts. It is these experts who highlight the steps that are not included in the official procedure but are critical in practice: how to position the part before clamping it, or glancing at the pressure gauge as production ramps up.
Write observable criteria, not intentions
A poorly written evaluation grid can be recognized by its vague wording: «has a good command of the machine,» «works independently,» «demonstrates a good safety attitude.» None of these criteria are observable. Two evaluators assessing the same operator will assign different scores. The rule is to write each criterion with a an action verb and a measurable result : «start up machine X according to the procedure in less than 5 minutes,» «perform the self-checks listed in steps 3, 7, and 12 of the procedure,» «trigger the emergency shutdown in the event of a pressure deviation greater than 0.5 bar .» The criterion must be checkable or uncheckable, without debate.
Select the proficiency level (the 4 ILUO levels)
The most widely used industry standard is the ILUO scale, which originated from the Toyota system and the Lean philosophy. Four levels are sufficient to describe an operator’s progression without complicating its use on the shop floor:
- I, Introduction: The operator is in training and works only under the direct supervision of a mentor or team leader.
- L, Learning: The operator can perform standard operations independently but still needs assistance with special cases.
- U, Understanding: The operator is proficient in all tasks associated with the position, handles common issues, and can assist a beginner.
- O, Ownership: The operator develops the standard, trains other operators, and proposes improvements to workstations.
Why not a 10-point scale? Because beyond four, raters can no longer reliably distinguish between the levels. The skills matrix The resulting document remains easy to understand at a glance for both the team leader and the quality auditor.
At Mercateam, each ILUO level is numbered (I = 1, L = 2, U = 3, O = 4) rather than left as a single word. This approach makes the data analytically versatile: it allows for calculating averages by team or by line, tracking progress over time, and cross-referencing this data with other performance metrics. Performing analytics on words («Initiation,» «Learning,» etc.) rather than on numbers does not allow for this type of calculation, which makes the data difficult to use beyond individual interpretation.
Calculator
Calculate an operator's ILUO versatility average
Assign a level (1 to 4) to each skill being evaluated to calculate an analytical average that can be used for team or line management.
Competency 1
Competency 2
Competency 3
Average level of versatility
3.0 / 4
An average that can be calculated only because each level is numbered rather than named.
Test the grid on a pilot unit
Before rolling it out across the entire workshop, the evaluation grid must undergo a pilot test: a single workstation, two to three operators being evaluated, over a period of 4 to 6 weeks. The goal is not to evaluate the operators; it is to evaluate the grid. Any vague criteria are flagged immediately, as are any overlapping levels. This is also the time to time the actual duration of the evaluation: if it exceeds 90 minutes, the checklist is too cumbersome and no one will use it after it’s rolled out.
Train evaluators to ensure consistency in judgment
No rating scale can completely eliminate the evaluator’s subjectivity. Common biases remain: the halo effect (a likable operator appears more competent), the strictness or leniency effect, and the anchoring effect based on the most recent evaluation. The solution lies in a simple practice: cross-calibration. Two evaluators rate the same operator simultaneously, then compare their scores. The discrepancies reveal which criteria are interpreted differently by each person. Three or four calibration sessions are enough to align a team of evaluators on a given rating scale.
The on-the-job evaluation checklist
In addition to the rating scale (which assigns a score), the observation checklist guides the evaluator step by step throughout the session. It is organized into six consecutive sections that follow the actual sequence of a job handover.
- Job Setup: The operator reviews the work order, verifies the availability of materials and tools, and checks the machine settings before starting production.
- Safety and PPE: Protective equipment was worn as required, lockout procedures were followed, and the condition of safety devices was visually inspected.
- Performance of the task: adherence to the operating procedure and key points, maintaining the target pace, and the quality of the sequence of operations.
- Intermediate quality control: self-inspections conducted at scheduled stages, traceability of readings on the tracking form or tablet, acceptable standard deviation.
- Risk Management: Responding to a caused or observed defect (machine shutdown, reporting, escalation to the appropriate contact).
- Teamwork and Wrap-Up: Passing on instructions to the next shift, tidying up and cleaning the workstation, and reporting on supplies that need to be reordered.
Each item is rated according to one of three statuses: compliant, non-compliant, or not applicable, with a comment field to provide context. The total duration of a typical assessment ranges from 30 and 90 minutes depending on the complexity of the position. Beyond that, the checklist covers too many tasks and should be split into two separate positions.
Who evaluates, and based on what criteria?
The field evaluator is usually the team leader, the job coach, or a mentor AFEST identified within the training program. These three roles are not the same, and confusion is common. The mentor trains and supports the operator over the long term. The evaluator assesses performance at a specific point in time using a defined rubric. The same person may serve in both roles, but not during the same session: a tutor who evaluates their own student during the first official assessment places themselves in a situation of obvious bias.
The simplest rule for preventing bias is to have the first official evaluation conducted by an evaluator who did not train the operator. Subsequent evaluations can be conducted by the direct supervisor, since they focus on progression to higher levels (such as moving from Level L to Level U) rather than on initial skill acquisition. At sites with high staff turnover, such as in seasonal food processing, a matrix of authorized evaluators makes it possible to track who can validate what.
When it comes to training evaluators, two things are key: mastery of the evaluation grid (correctly interpreting each criterion, knowing how to score the status) and approach (asking open-ended questions, explaining the score to the operator, and accepting differing opinions). A half-day of initial training plus a quarterly calibration session are sufficient to maintain a robust system.
Validation, Signing, and Updating of the Competency Matrix
A completed form triggers a three-step approval workflow: immediate feedback to the operator, signatures from stakeholders, and an update to the competency matrix. If any of these three steps is skipped, the form loses its evidential value.
Feedback is provided immediately, either at the workstation or in the team leader’s office, in less than 15 minutes. The evaluator explains each area for improvement, the operator can provide feedback, and any necessary action plan is established immediately, without postponing the discussion. This immediate feedback is essential for the operator to view the evaluation as a tool for growth rather than as a judgment being imposed on them.
The signature involves three parties: the employee being evaluated (who reviews the evaluation), the evaluator (who certifies the rating), and the supervisor (who validates the impact on the employee’s versatility). With paper forms, the time between the evaluation and the final signature often exceeds three weeks, and unsigned evaluation forms pile up in team leaders’ filing cabinets. Electronic signatures on tablets reduce this time to just a few minutes and produce evidence that can be used directly in audits.
The final step: updating the matrix. The transition from Level L to Level U for a given position results in additional versatility for the team, and thus a new option in the staffing schedule. If the matrix is maintained manually, this benefit remains unnoticed for weeks. If it is automatically updated by the evaluation tool, the versatility matrix Starting the following Monday, you can already take advantage of this.
What the auditor looks for when verifying evidence of competence
Quality standards have tightened their requirements regarding proof of competence in recent years. ISO 9001 §7.2, EN 9100 for the aerospace industry, NADCAP for special processes, IATF 16949 for the automotive industry, BRC and IFS for the food industry, FDA 21 CFR 211 for the pharmaceutical industry: all now require proof that competence is evaluated in practice, not just on paper.
The auditor is no longer satisfied with just a training certificate. He or she randomly selects an operator on the line, verifies the operator’s reference position in the competency matrix, and then requests the most recent practical evaluation form: date, evaluator’s name, score by criterion, and signatures. A form that is missing, more than two years old, or present but not signed by the operator is immediately flagged as a non-conformity.
The accepted frequency varies depending on the standard and the risk associated with the position. For positions critical to safety or quality, a Practical evaluation every 12 months is the standard. For standard workstations, every 18 to 24 months is sufficient, provided that the checklist is triggered whenever there is a change in the process or a quality deviation is reported. The authorization management remains more restrictive, with renewal requirements based on regulatory timeframes (3 years for the CACES R489, for example).
Simulator
How often should this position be reassessed?
Recommended frequency
Every 12 months
Additional triggers, regardless of the position
Process Change at the Workstation
Quality Deviation Report Related to the Position
Getting Back on Track After a Long Absence
When to Switch from an Excel Spreadsheet to a Dedicated Tool
An Excel spreadsheet works very well for a single site with one lead evaluator and 20 to 30 operators. Tipping points emerge as the organization grows:
- Several sites or teams each maintain their own version of the file, and the criteria differ without anyone noticing until the group audit.
- Signed paper checklists pile up in binders, making it impossible to find them when an auditor asks, «Show me the last five evaluations for the packaging station.».
- Authorization expiration dates slip through the cracks because no one checks the Excel file every week to verify the expiration dates.
- Updating the skills matrix is delayed by several weeks after each assessment, which means the schedule is consistently suboptimal.
A dedicated tool performs three functions that Excel does not cover. It guides the evaluator on a tablet directly on the shop floor, item by item. It triggers an immediate electronic signature, which is stored with a timestamp. It updates the matrix and schedule in real time, generating alerts for expiring certifications. Across the 300 industrial sites equipped by Mercateam, the time required to prepare for a quality audit has been reduced from several days to just a few hours, and the time spent on administrative tasks related to training has been reduced by a factor of 4 on average.
While the competency matrix is currently the tool used to visualize who can do what, practical assessment is what validates that matrix. It’s what allows us to answer the auditor’s questions without hesitation, ensure the reliability of Monday’s schedule, and transform on-the-job training into a short, traceable cycle. To see how this works on a tablet in a real workshop, Request a demo of Mercateam and we'll use a typical job posting from your site as an example.




