Skip to content

Home / Statistical Tools / MSA / Extended / Options

Options

Every control on the Extended dialog. An Extended study uses all four tabs: Data, Model, Options and Gage Info.

The Data, Options and Gage Info tabs are the same ones the Crossed and Nested studies use, with two additions: the Additional factor columns list on the Data tab, and the Extended options group on the Options tab.

Data tab

Two radio buttons at the top of the tab choose where the data come from:

  • Excel reads the range you selected on the worksheet.
  • GroupBy reads the same range but splits it into groups first, and runs a separate complete analysis for each group.

A problem with the current selections is reported in red near the top of the tab, above the role lists. The message stays until the selections satisfy it, so the tab is worth re-reading from the top when Finish does not do what you expect.

Data Types

Data Types lists every column in the selected range with a drop-down setting its type. The type decides which role lists the column is offered in below, so a column that will not appear where you expect is usually typed wrongly here.

The available types are Continuous, Nominal, Count, Integer and DateTime.

Measurement Columns

Measurement Columns is the list of columns holding the measured values. Tick one column per characteristic you want analysed. This list accepts more than one column, unlike the role lists below it.

Each measurement column is a complete analysis of its own

Ticking three measurement columns does not produce one report covering three characteristics. It produces three separate worksheets, each a full analysis of one column: its own Gage R&R table, its own analysis of variance, its own charts.

In GroupBy mode the two multiply. Three measurement columns across four groups produce twelve worksheets.

Part

Part is the column identifying which part each measurement was taken on. It is required, and exactly one column can be chosen.

Operator (Optional)

Operator (Optional) is the column identifying who or what took each measurement. Its label says optional and it means it: a study with no operator column runs, and reports repeatability and part to part variation without a reproducibility component to separate out.

Exactly one column can be chosen. Parts of the report depend on this column being present, and each is simply absent without it: the reproducibility rows, the Measurement by Operator chart, the box plot by operator, and the Part by Operator Interaction chart.

The average and range charts are not among them. Without an operator column they are still drawn, grouped by part alone and carrying a single series each instead of one series per operator.

Reference (Optional)

Reference (Optional) is the column holding the known true value of each part. Exactly one column can be chosen, and it must be numeric.

Supplying it turns on the bias and linearity part of the analysis, which is the only part that can say whether the measurement system reads correctly rather than merely consistently. Without it, the bias and linearity tables and the Linearity and Bias charts are absent.

Specification Limits

Specification Limits is a small grid with one row for each measurement column you ticked, and columns headed Column, LSL and USL. Each characteristic gets its own limits, because each becomes its own worksheet.

Either limit can be left empty. What the limits control:

Supplied Effect
both the tolerance width is the upper limit minus the lower
one the tolerance width is twice the distance from the study mean to that limit
neither every percent of tolerance column, the precision to tolerance ratio, the whole misclassification block, the gage performance curve and the sweep charts are all absent

A row with a problem in it is highlighted, and the reason is available as a tooltip on the row. A lower limit that is not strictly below its upper limit is refused.

Additional factor columns

Additional factor columns is where an Extended study names the factors beyond part and operator: a fixture, a gauge, a laboratory, a day, or anything else you want the analysis to account for. It accepts more than one column.

Every column ticked here appears as a row on the Model tab, where its type and its nesting are set. This list only says which columns are factors; the Model tab says what the analysis does with them.

This list is present only for an Extended study.

GroupBy

Choosing GroupBy reveals the grouping panel: Available Items on the left and GroupBy Order on the right. Between the two lists, > moves the selected column across and >> moves all of them, with < and << moving them back. To the right of the second list, ^ and v change the position of a column within GroupBy Order, which is the order the grouping is applied in. A DateTime column additionally gets a drop-down choosing how its values are grouped.

Each distinct combination of the grouping columns becomes its own complete analysis on its own worksheet, and each worksheet names its group.

This is the same grouping panel used throughout Quantum XL, and it behaves the same way here.

Selecting a column shows a line of information about it at the foot of the tab. It is there to confirm you picked the column you meant and it changes nothing about the analysis.

Model tab

The Model tab is where an Extended study says what its factors are and how they relate. The Data tab named the columns; this tab decides what the analysis does with them.

It appears only for an Extended study. A Crossed or Nested study has a model fixed by its own definition, so there is nothing to declare.

Factors

The Factors group lists every factor in the study, one row each, and each row has a Type: drop-down with two choices:

  • Random means the levels in the study are a sample of many possible levels, and the question is how much variation the whole population of levels produces. Operators are usually random: these three operators stand in for operators generally.
  • Fixed means these levels are the only ones of interest, and the question is how they differ from each other. Two specific machines you own, and no others, are fixed.

The choice changes what is reported for the factor. A random factor gets a variance component. A fixed factor has no variance component, because a fixed effect has no distribution to have a variance; a substitute quantity is reported in its place and the row carries no confidence interval.

When the drop-down is unavailable, the reason is shown as text beside it rather than left to be guessed.

At least one factor must remain random. A model in which every factor is declared fixed leaves REML with nothing to estimate and is refused.

The rest of each factor row declares nesting. Nested in: names the factor this one sits inside, and further and in: drop-downs appear as they are needed.

A factor is nested inside another when its levels mean something only within a level of that other factor. Parts nested within operators is the common case: operator A's part 1 and operator B's part 1 are different physical parts, so the part labels only identify a part within an operator.

The test is always the same question: does this level label mean the same thing under every level of the other factor? If yes, the two are crossed. If no, this one is nested inside it.

A factor can be nested inside as many as five parents. The count includes every ancestor the analysis reaches by following the nesting chain, not only the parents named directly on this row, so a part nested in an operator that is itself nested in a laboratory has used two of the five.

To undo a nesting, choose (None) in the drop-down. It is a real entry in the list, not a blank at the top, so it can be selected again after a parent has been chosen.

Model structure

The Model structure group draws the model the current declarations produce, as an indented tree. Each level of indentation is one level of nesting, and a nested entry is marked with a corner guide:

Operator
└ Part

It is read-only. It is there so the structure can be checked before the analysis is run, which matters because a nesting declared the wrong way round produces a report rather than an error.

Two warnings can appear beneath the tree. Both describe a consequence of the current declarations, and neither prevents the analysis from running.

The first says the model is too deeply nested for chart data to be produced:

With this nesting the engine computes the full analysis but returns no chart data; the report keeps the box plots.

Every number on the report is still computed and reported. What is lost is the charts drawn from the measurements themselves. The box plots are drawn by a different route and survive.

The second says the model is beyond what REML can fit:

A term in this model involves more than three factors once nesting is counted, so the REML method cannot fit it; if REML is selected or chosen automatically, the analysis will use expected mean squares instead.

Note that nesting counts toward the three. A term that reads as two factors on this tab can involve four once its nesting chain is followed, which is why this warning can appear on a model that looks small.

Interactions

The Interactions group has one checkbox per interaction the declared factors allow. An interaction asks whether the effect of one factor depends on the level of another: whether the operators disagree more about some parts than others, for instance.

Only crossed factors can interact. Where an interaction is unavailable the checkbox is locked and the reason is shown as text beside it.

Adding an interaction spends degrees of freedom, so it needs replication to support it. The Remove insignificant interactions option on the Options tab is the automatic counterpart to this group: this decides which interactions are offered to the model, that one drops the ones the data give no evidence for.

Part-to-part variation

The Part-to-part variation group decides which terms count as variation in the parts rather than variation in the measurement system. Its own introduction states the default:

By default, the variation from additional factors and interactions is attributed to reproducibility. Select a term below to attribute its variation to part-to-part variation instead.

This matters because every headline number on the report is a comparison between the two groups. Total Gage R&R is repeatability plus everything counted as reproducibility; moving a term across moves its variance from the numerator of that comparison to the denominator, and changes percent contribution, percent study variation, the number of distinct categories and the misclassification block with it.

The question to ask of each term is whether its variation is something the measurement system is doing, or something the parts are doing. A fixture that biases readings is the measurement system. A characteristic that genuinely differs between the batches the parts came from is the parts.

The operator main effect cannot be moved into part to part variation, and a selection that includes it is refused.

Options tab

Estimation Method

The Estimation Method group chooses how the variance components are computed. All four choices report the same set of quantities where they can compute them; they differ in how they get there and in what they cannot produce.

Automatic (balanced: EMS; unbalanced: REML)

Automatic (balanced: EMS; unbalanced: REML) is the default. It uses the expected mean squares method when the data are balanced and restricted maximum likelihood when they are not.

Balanced means every cell of every model term holds the same number of measurements and none is empty. A single missing measurement is enough to make a study unbalanced, which is the usual reason this setting picks REML.

Restricted Maximum Likelihood (REML)

Restricted Maximum Likelihood (REML) estimates every variance component at once from the data as a whole, under the constraint that a variance cannot be negative.

Two things follow that the other methods do not offer:

  • It reports confidence intervals on unbalanced data as well as balanced.
  • It never reports a negative estimate floored to zero, because it never computes a negative one. A reported zero means the constrained estimate reached the zero boundary.

It also has one limit: it cannot fit a model with a term involving more than three factors once nesting is counted. Where that happens the analysis uses expected mean squares instead, and the Model tab warns before you run it.

Expected Mean Squares (EMS) - Equivalent to ANOVA when balanced

Expected Mean Squares (EMS) - Equivalent to ANOVA when balanced builds the analysis of variance table and solves for the components from the mean squares.

As the label says, on balanced data this gives the same answers as the classical analysis of variance, which is what makes it the method to choose when the numbers need to match a hand calculation or a published example.

Two consequences worth knowing before choosing it:

  • A component's estimate can come out negative. Zero is reported in its place, with a footnote saying so and suggesting REML.
  • On data that is not balanced this method reports no confidence intervals at all. Balanced data under this method keeps its full set of intervals.

XbarR

Not available for an Extended study. The radio button is disabled and the reason is shown beneath it. The method requires a plain crossed study with exactly one part factor and one operator factor, both random, no additional factors and no nesting, so an Extended study never qualifies. Choose one of the three methods above instead.

The report states the method that was actually used

The Estimation method: row near the top of the report names the method the analysis ran with, not the one selected in this group. The two differ when Automatic made the choice, and when a model REML cannot fit caused the analysis to fall back to expected mean squares.

Analysis Options

Misclassification sweep charts

Misclassification sweep charts adds the charts that show how the misclassification probabilities would change if the process mean moved. The mean is swept across a range measured in part to part standard deviations and each probability is recomputed at every step.

The charts need at least one specification limit, so they are absent without one whatever this box says. A Crossed or Nested study gets the two joint-probability sweeps; an Extended study gets all five.

Part Variation is Not Representative of the Population

Part Variation is Not Representative of the Population is for a study whose parts were not chosen to span the range the process actually produces, for example a set of parts gathered from one shift or deliberately picked to be similar. When the parts do not represent the process, every statistic that compares the gage against part to part variation describes the sample of parts rather than the process.

The dialog's own tooltip states exactly what the option does:

Display-only: hides part-to-part variation and every statistic derived from it (total variation, % contribution, % study variation, ndc, misclassification). Does not change what the engine computes.

This option hides, it does not recompute

Every number that remains is the same number it would have been with the box clear. Nothing is re-estimated and no component changes. What is withheld is the part to part variation and the statistics that divide by it, because those are the ones a non-representative set of parts makes misleading.

The percent of tolerance columns are not withheld, because they compare against the tolerance rather than against the parts. That is also why the misclassification block and its sweeps come back when a historical standard deviation was supplied and used: the process variation then comes from the historical value instead of from the parts in the study.

Extended options

This group is shown only for an Extended study.

Historical standard deviation:

Historical standard deviation: supplies a known process standard deviation from outside this study. The dialog's own hint states its purpose and its one requirement:

Used as a replacement for total standard deviation in Percent Misclassified Calculations and Graphs. Must be larger than total gage (measurement) standard deviation.

It is a process standard deviation, so it already includes measurement error. The analysis removes the gage variance from it to recover the part to part variation used by the misclassification calculation, which is why it must be the larger of the two. Supplied but too small, and the misclassification block is refused with a note saying so rather than being computed from a substituted value.

Supplying it also adds a percent of historical process column to the results table. That column divides by the historical value directly, with no subtraction. It does not change the number of distinct categories, which always uses the estimated standard deviations.

Leave it empty when there is no such value; empty is the normal case.

Study variation multiplier (k):

Study variation multiplier (k): is the number of standard deviations a study variation spans. It defaults to 6, which covers plus and minus three standard deviations.

It scales the study variation column and the percent of tolerance columns. It does not affect percent contribution or percent study variation, where it cancels. A value of zero or below selects the default rather than being used.

Confidence level:

Confidence level: sets the level of every confidence bound on the report. It defaults to 0.95.

It must be at least 0.5 and below 1.0. A value outside that range is refused rather than clamped. The lower bound is deliberate: below one half the interval formulas reach a form with no reference to check it against, and a level of exactly 1.0 has no finite bound at all.

An Extended study reports no confidence intervals

The confidence level is still accepted and validated, and it sets the level of the confidence interval on bias. It does not produce variance component intervals, because an Extended study reports none.

Significance level:

Significance level: is the threshold the bias t tests and the linearity acceptance test are judged against. It defaults to 0.05. A value of zero or below selects the default.

It is a different setting from the interaction removal Alpha: below, which is compared against a different set of p values for a different purpose.

Interaction removal runs by default on every study

Quantum XL drops interaction terms the data give no evidence for, pooling each removed term's degrees of freedom and sum of squares into repeatability. This happens by default in a Crossed and a Nested study as well as an Extended one, at a p-value threshold of 0.25.

Only an Extended study can change it. The Remove insignificant interactions checkbox and its Alpha: field sit in the Extended options group, so a Crossed or Nested study always runs with removal on at 0.25.

You can see whether anything was removed: when a term is dropped the report carries two analysis of variance tables rather than one, and the removed term is marked in the first.

Removal is turned off automatically when the data have no replicates, and the report notes that.

Remove insignificant interactions and Alpha:

Remove insignificant interactions is ticked by default. It drops interaction terms the data give no evidence for, pooling each removed term's degrees of freedom and sum of squares into repeatability. Alpha: is the p-value threshold a term must fail to be removed, and it defaults to 0.25. The box must be ticked for the Alpha: field to be editable.

Clearing the box is the only way to fit the model exactly as declared. A Crossed or Nested study has no such control and always runs with removal on.

Note that the default of 0.25 is much looser than the 0.05 the Significance level: field above defaults to. The two are separate settings compared against different p values, and neither changes the other.

When something is removed the report carries two analysis of variance tables, one for all terms and one for the terms actually used, so the removal is visible rather than silent.

A term that an empty cell makes unestimable is dropped whatever this setting says. That is a different event, it is reported separately, and it is noted on the report.

Gage Info tab

The Gage Info tab records seven free-text fields describing the gage and the study: Gage name:, Gage no.:, Gage type:, Part name:, Part no.:, Date: and Performed by:.

They are copied onto the report and enter no calculation. Every one can be left empty. They are there so a printed report identifies the gage and the study it came from without needing a separate record.

See Also