Home / Statistical Tools / MSA / Attribute MSA: Crosstabulations Method / Attribute MSA How-To
Attribute MSA How-To¶
This walkthrough runs a complete attribute study: twelve parts, two appraisers, two trials each, judged pass or fail against a known standard.
It is built around one result worth seeing. The two appraisers score exactly the same on agreement with the standard, and they are wrong in opposite directions. Only two tables on the report can tell them apart, and knowing which two is most of what this tool is for.
The data is on this page rather than in a file to download. Press Copy for Excel, then paste it into a blank worksheet.
The data¶
Twelve moulded housings inspected for a cosmetic flaw against a boundary sample. Each is either acceptable, P, or not, F. Two inspectors, Ivan and Priya, each judged every part twice, in random order, without seeing their own earlier answer.
The Standard column is the right answer, settled beforehand by an engineer with a microscope: of the twelve parts, eight are genuinely good and four are genuinely bad.
| Part | Appraiser | Trial | Rating | Standard |
|---|---|---|---|---|
| 1 | Ivan | 1 | P | P |
| 1 | Ivan | 2 | P | P |
| 1 | Priya | 1 | P | P |
| 1 | Priya | 2 | P | P |
| 2 | Ivan | 1 | P | P |
| 2 | Ivan | 2 | P | P |
| 2 | Priya | 1 | P | P |
| 2 | Priya | 2 | P | P |
| 3 | Ivan | 1 | F | P |
| 3 | Ivan | 2 | F | P |
| 3 | Priya | 1 | P | P |
| 3 | Priya | 2 | P | P |
| 4 | Ivan | 1 | F | F |
| 4 | Ivan | 2 | F | F |
| 4 | Priya | 1 | P | F |
| 4 | Priya | 2 | P | F |
| 5 | Ivan | 1 | P | P |
| 5 | Ivan | 2 | P | P |
| 5 | Priya | 1 | P | P |
| 5 | Priya | 2 | P | P |
| 6 | Ivan | 1 | F | P |
| 6 | Ivan | 2 | F | P |
| 6 | Priya | 1 | P | P |
| 6 | Priya | 2 | P | P |
| 7 | Ivan | 1 | F | F |
| 7 | Ivan | 2 | F | F |
| 7 | Priya | 1 | F | F |
| 7 | Priya | 2 | F | F |
| 8 | Ivan | 1 | P | P |
| 8 | Ivan | 2 | P | P |
| 8 | Priya | 1 | P | P |
| 8 | Priya | 2 | P | P |
| 9 | Ivan | 1 | P | P |
| 9 | Ivan | 2 | P | P |
| 9 | Priya | 1 | P | P |
| 9 | Priya | 2 | P | P |
| 10 | Ivan | 1 | F | F |
| 10 | Ivan | 2 | F | F |
| 10 | Priya | 1 | P | F |
| 10 | Priya | 2 | F | F |
| 11 | Ivan | 1 | P | P |
| 11 | Ivan | 2 | P | P |
| 11 | Priya | 1 | P | P |
| 11 | Priya | 2 | P | P |
| 12 | Ivan | 1 | F | F |
| 12 | Ivan | 2 | F | F |
| 12 | Priya | 1 | F | F |
| 12 | Priya | 2 | F | F |
Forty-eight rows: twelve parts, times two appraisers, times two trials. One row is one judgement, and the standard is repeated on every row of a part. That repetition is required. A part whose rows carry two different standards stops the study.
Before running anything, find the disagreements by eye. Ivan calls parts 3 and 6 bad, twice each, when both are good. Priya calls part 4 good, twice, when it is bad, and cannot decide about part 10. That is the whole of the disagreement in this data set, and every number below is a different way of counting it.
Steps¶
-
Put the data in Excel
Press Copy for Excel above the table. In Excel, open a blank worksheet, click cell A1, and press Ctrl+V. You should have headers in row 1 and data in rows 2 through 49.
-
Select the data and start the analysis
Select A1:E49, including the header row. From the Excel ribbon: QXL Stat Tools > MSA / Gage R&R > Attribute MSA - Crosstabulations Method.
The report is written straight away, from defaults, and the dialog opens on top of it.
-
Assign the columns on the Data tab
- Tick Rating under Rating Columns.
- Tick Part under Part.
- Tick Appraiser under Appraiser.
- Tick Standard under Standard (Optional).
Leave the Trial column unassigned. The study works out the trials from how many times each appraiser rated each part; it does not read a trial number.
-
Check the Options tab
The data holds exactly two levels, so Binomial (two levels) is available and already chosen. Ordinal (three or more ordered levels) is greyed out, with the line Ordinal needs three or more levels; this data holds two. underneath.
The Levels block lists F and P, in that order, and its buttons are greyed because the data is not ordinal.
Conforming level: already reads P. The dialog recognised the words: p is on its list of conforming words and f is on its list of nonconforming ones. The guess is right, so leave it.
Leave Alpha: at
0.05, which shows Confidence level: 95% beside it. -
Finish
Fill in the Gage Info tab if you want the details kept with the sheet, then choose Finish.
Reading the report¶
The worksheet is named Attribute MSA. Columns A and B hold the input echo, the Gage Information block and the Study Summary. The charts and the tables run from column D.
Study Summary should read: Parts 12, Appraisers 2, Trials per Appraiser 2, Data Type Binomial, Levels F, P, Known Standard Yes, Conforming Level P, Alpha 0.05000, Confidence Level 95%.
Underneath it, one note: The conforming level was taken by the default rule (P). Change it in the dialog if F is the conforming level. That is the study telling you it guessed, and inviting you to check the guess. It guessed right.
Assessment Agreement (Within Appraiser)¶
Does each inspector agree with themselves?
| Appraiser | # Inspected | # Matched | Percent | 95% CI Lower | 95% CI Upper |
|---|---|---|---|---|---|
| Ivan | 12 | 12 | 100.00% | 77.91% | 100.00% |
| Priya | 12 | 11 | 91.67% | 61.52% | 99.79% |
Ivan gave the same answer both times on all twelve parts. Priya changed her mind about part 10.
Notice that Ivan's perfect 12 out of 12 does not produce an interval of zero width: the lower bound is 77.91%. Twelve consistent judgements are not proof of a consistent inspector, and the interval says so.
Assessment Agreement (Each Appraiser vs Standard)¶
Does each inspector agree with the right answer?
| Appraiser | # Inspected | # Matched | Percent | 95% CI Lower | 95% CI Upper |
|---|---|---|---|---|---|
| Ivan | 12 | 10 | 83.33% | 51.59% | 97.91% |
| Priya | 12 | 10 | 83.33% | 51.59% | 97.91% |
This is the number to be suspicious of. The two inspectors score identically, down to the interval. Stop here and you would conclude they are equally good and equally in need of the same retraining. They are not.
Notice also what happened to Ivan. He repeats perfectly and agrees with the standard only 83.33% of the time, which is the signature of an inspector who is consistently wrong. Repeatability is not accuracy, and the two tables above are the demonstration.
Assessment Agreement (Between Appraisers) and (All Appraisers vs Standard)¶
| Section | # Inspected | # Matched | Percent | 95% CI Lower | 95% CI Upper |
|---|---|---|---|---|---|
| Between Appraisers | 12 | 8 | 66.67% | 34.89% | 90.08% |
| All Appraisers vs Standard | 12 | 8 | 66.67% | 34.89% | 90.08% |
Both drop to eight parts out of twelve, because these figures need every rating on a part to agree. Parts 3, 4, 6 and 10 each lose the part for the whole study, even though only one inspector was wrong on each of them.
The two sections match here, at 8 parts, only because every part the two inspectors agreed on they also got right. That is a coincidence of this data set, not a rule.
Effectiveness¶
Where the tables above score a whole part, this one scores each judgement on its own.
| Appraiser | # Correct | # Decisions | Percent | 95% CI Lower | 95% CI Upper |
|---|---|---|---|---|---|
| Ivan | 20 | 24 | 83.33% | 62.62% | 95.26% |
| Priya | 21 | 24 | 87.50% | 67.64% | 97.34% |
| Overall | 41 | 48 | 85.42% | 72.24% | 93.93% |
Priya edges ahead here where she was level on the part based figure, because her mistake on part 10 cost her one judgement rather than a whole part. The two ways of counting answer different questions, which is why both are on the report.
Conformance: the table that tells them apart¶
| Appraiser | # False Alarms | # Misses | P(False Alarm) | P(Miss) |
|---|---|---|---|---|
| Ivan | 4 | 0 | 25.00% | 0.00% |
| Priya | 0 | 3 | 0.00% | 37.50% |
There is the difference. Two inspectors with the same agreement percentage, and opposite failure modes:
- Ivan never passes a bad part. Every one of his errors is a good part scrapped. His 25.00% is four false alarms out of the sixteen judgements made on genuinely good parts.
- Priya never scraps a good part. Every one of her errors is a bad part shipped. Her 37.50% is three misses out of the eight judgements made on genuinely bad parts.
These two errors cost entirely different things. Ivan's cost is scrap; Priya's cost is a customer complaint. A single agreement percentage cannot distinguish them, and this table exists so you do not have to try.
The denominators differ on purpose: a false alarm rate is out of the assessments of conforming parts, a miss rate is out of the assessments of nonconforming parts.
Assessment Disagreement (Each Appraiser vs Standard)¶
The same story, counted by part rather than by judgement, and available because the data is two-level.
| Appraiser | # P / F | Percent | # F / P | Percent | # Mixed | Percent |
|---|---|---|---|---|---|---|
| Ivan | 0 | 0.00% | 2 | 25.00% | 0 | 0.00% |
| Priya | 1 | 25.00% | 0 | 0.00% | 1 | 8.33% |
The three notes under the table say what each column counted. # F / P is the parts an appraiser called F on every trial whose standard is P: Ivan's parts 3 and 6. # P / F is the reverse: Priya's part 4. # Mixed is the parts an appraiser rated inconsistently, which is Priya's part 10, and a mixed part is counted in neither of the first two columns, because the appraiser did not consistently call it anything.
The kappa tables¶
Fleiss' Kappa Statistics appears in every section, one row per level plus Overall, with Kappa, SE Kappa, Z and P(vs > 0).
Cohen's Kappa Statistics (Within Appraiser) is printed, because each appraiser gave exactly two trials, which is the shape Cohen's kappa needs.
Cohen's Kappa Statistics (Between Appraisers) is not, and instead carries the sentence You must have two appraisers and single trial per appraiser to compute kappa. This study has two appraisers but two trials each, so the comparison is not between two things.
The two charts¶
Assessment Agreement Within Appraiser and Assessment Agreement Appraiser vs Standard, at columns D and J. Each shows one filled circle per appraiser at the agreement percentage, with crosses at the two confidence bounds and a red vertical line joining them.
Look at the first chart: Ivan's circle is at the top of the axis and his red line still stretches down past 77%. Look at the second: the two circles sit at the same height, and the two red lines are the same length. The chart cannot tell the inspectors apart either, which is why the Conformance table is the one to read next.
Try changing one thing¶
Clear the Standard column on the Data tab. Every table that compares against the right answer disappears: both vs Standard sections, Effectiveness, Agreement Counts, Misclassifications, Conformance and Assessment Disagreement, plus the second chart. What is left says the two inspectors agreed with each other on eight parts out of twelve and nothing at all about whether they were right. Getting a standard is what makes an attribute study worth running.
Switch the conforming level to F. The Conformance table turns inside out: Ivan's four errors become misses and Priya's three become false alarms. Nothing about the data changed. The conforming level is a statement about which category is the good one, and getting it backwards inverts the only table that distinguishes the two failure modes.
Choose Ordinal. The study is refused with Ordinal data needs at least three distinct values in the ratings and the standard together. This study has 2. Choose Binomial or Nominal for two-level data. Pass and fail do not rank.
See Also¶
- Attribute MSA: Crosstabulations Method, what the tool does and what it reports
- Options, every control on the dialog
- Math Details, every formula behind the tables above