Services
Audit & Certification
ISO Gap AnalysisInternal AuditAudit PreparationAfter the AuditClose Nonconformities
Standards
ISO 9001 Quality ManagementISO 9001:2026 TransitionISO 14001 Environmental ManagementISO 45001 Occupational Health & SafetyISO 27001 Information SecurityISO 42001 AI ManagementISO 13485 Medical Devices
Industries & Support
Industry SolutionsMedical DevicesMechanical Engineering & ProductionIT, SaaS & AIManagement System MaintenanceExternal QMRReview of AI-Generated QMSQM Training
Industries
All IndustriesMechanical Engineering & ProductionAutomotive SuppliersLaser Optics, Photonics & SemiconductorsIndustrial Service ProvidersMedical DevicesIT, SaaS & AI
FundingFAQKnowledgeFree toolsAboutContactSend emailCall now
DE/EN
Free Consultation

Sternberg Consulting · DOE guide

Good experiments start with a clear plan.

Design of experiments explained: from your first question to a confirmed improvement. With a worked example and detailed FAQ.

Everything is free. No registration. Also for workplace projects.

DOE stands for Design of Experiments. You vary several inputs in a planned pattern to investigate their effects and interactions. This guide walks through a full two-level factorial experiment.

1. Define a specific question

Start with a decision: Which settings improve bond strength without exceeding the allowable processing time? “Understand the process” is too broad on its own.

Describe the process, material, baseline and intended scope. Decide in advance what improvement would matter in practice. This helps prevent selecting only striking results after seeing the data.

Your next step: Record the question, response, unit and relevant improvement in your project.

2. Check the measurement

The response is the outcome measured for each run. Use a consistent method, unit and measurement time. Your measurement system must reliably distinguish changes of interest.

Check resolution, repeatability, operator influence and calibration where appropriate. Destructive testing needs comparable specimens. A better design cannot compensate for an unreliable measurement system.

Your next step: Write a short measurement procedure and check measurement variation before the main experiment.

3. Choose factors and levels

Factors are inputs you deliberately change, such as temperature, pressure or material type. Identify plausible causes with the people who work with the process.

A two-level design uses two distinct settings per factor. For numbers these are usually a low and high setting; categories could be two materials. Every combination must be feasible and safe. Individually acceptable settings do not guarantee an acceptable combination.

Your next step: Check every combination for feasibility. Include units in factor names.

Try it in the free app

4. Select an appropriate design

A full two-level factorial tests every combination: 4 for two factors, 8 for three, 16 for four and 32 for five. Each complete replicate multiplies that number. This free app supports two to five factors and one to five independent runs per combination.

Fractional factorial designs reduce runs but can alias effects. Screening can help with many factors. Curvature and an interior optimum require additional levels or suitable response surface designs. Mixture experiments need dedicated models because the proportions sum to a fixed total.

Your next step: Use this app for full two-level factorial designs. Choose a suitable method first for other questions.

Try it in the free app

5. Plan replication and resources

An independent replicate means preparing the experimental unit and conditions again and repeating the run. Measuring the same part three times describes measurement variation; it does not replace three independent process runs.

Two runs per combination allow within-combination variation to be estimated. Adequacy depends on variation, the smallest effect of interest and the required evidence. The app does not calculate power or guarantee detection with a particular number of runs.

Your next step: Budget material, setup time and independent replicates. Use pilot data to refine sample size planning.

6. Consider order and nuisance variables

Random order helps avoid systematically mixing time trends with factor effects. Still record date, material batch, operator and unusual events. Process warm-up can affect results even with randomization.

Different experimental days may call for blocking. Hard-to-change settings, such as a large furnace temperature, may require a split-plot design. This app assumes fully randomizable runs and does not model blocks or split-plot structures.

Your next step: Generate the randomized order before the first run and retain it throughout execution.

7. Run and document the experiment

Follow the planned sequence. Keep constant conditions, waiting times and measurement methods consistent. Record actual settings and deviations rather than just copying target settings.

A missing value is not zero. Document decisions to repeat failed runs and do not remove inconvenient results without a technical reason. The app requires a complete plan for analysis so missing values cannot distort balanced effect calculations.

Your next step: Enter one response per run. Use project notes for deviations and observations.

8. Understand effects and interactions

A main effect here is the mean at the high level minus the mean at the low level, averaged across other factors. An effect of +8 MPa therefore means an average increase of 8 MPa over the tested range.

An interaction means that the effect of one factor depends on another factor’s setting. Different slopes in the interaction plot reveal this dependency. The app shows main effects and two-factor interactions, not higher-order interactions. Bars represent effect size, not a significance test.

Your next step: Read interactions alongside main effects. Consider both magnitude and practical relevance.

Try it in the free app

9. Make a supported decision

The best observed combination has the most favorable mean among the tested settings. It is not a proven global optimum or automatically a technically or economically suitable choice.

With independent replicates the app shows pooled within-combination standard deviation and the standard error of an effect. These require independent errors and comparable variance. Inspect observations and unusual events as well. The app does not provide p-values, confidence intervals, residual diagnostics or formal model validation.

Your next step: Weigh benefits against cost, process limits and uncertainty. Document your decision.

10. Confirm the improvement

Check the selected settings with new independent runs. Where possible compare against the baseline under comparable conditions. Define success criteria before collecting confirmation data.

After confirmation, introduce the setting into operations in a controlled way. Update work instructions and monitor whether the benefit persists across batches, operators and time. One favorable result is insufficient.

Your next step: Record the confirmation plan, owners and results, including the scope and remaining evidence gaps.

Example: improve bond strength

Synthetic teaching example: temperature 160/180 °C, pressure 3/5 bar and time 20/30 s. The aim is greater strength in MPa. These are not recommended process settings. Eight combinations are independently executed twice, giving 16 runs.

Example data in standard order
Temperature °CPressure barTime sResponses MPa
16032045.5 / 46.5
18032047.5 / 48.5
16052043.5 / 44.5
18052057.5 / 58.5
16033047.5 / 48.5
18033049.5 / 50.5
16053045.5 / 46.5
18053059.5 / 60.5

Temperature calculation: mean at 180 °C = 54 MPa; at 160 °C = 46 MPa. The main effect is 54 − 46 = +8 MPa. Pressure has an effect of +4 MPa and time +2 MPa. The temperature-pressure interaction is +6 MPa using ±1 coding.

At low temperature, higher pressure changes the mean by −2 MPa; at high temperature, by +10 MPa. This illustrates why pressure should not be interpreted alone. The best observed combination is 180 °C, 5 bar, 30 s with a mean of 60 MPa.

The grand mean is 50 MPa. The pooled standard deviation is about 0.7071 MPa with 8 degrees of freedom; the standard error of an effect is about 0.3536 MPa. An operational decision would require new confirmation runs and a review of process limits.

Open the app and choose “Load example”

Method and sources

The app calculates balanced contrasts in full two-level factorial designs. Main and pair interaction effects are 2/N × Σ(sᵢ × yᵢ), where sᵢ is the factor code or product of codes. Error variance is estimated only from replicates within identical combinations; an effect’s standard error is 2s/√N.

NIST/SEMATECH: Process Improvement · Modeling DOE data · Confirmatory runs

Frequently asked questions about DOE

Planning, execution, analysis and using the free app.

What is DOE?

DOE means Design of Experiments. Several inputs are varied according to a planned structure to investigate their effects and how they interact.

When is DOE useful?

When several controllable inputs may affect an outcome and you need to learn which settings help. You need a clear question, measurable outcomes and controlled experimental conditions.

Why not change one factor at a time?

One-factor-at-a-time trials do not systematically reveal whether the effect of one factor depends on another. A factorial design deliberately tests those combinations.

What prior knowledge do I need?

You should understand your process and measurement method. The guide explains the basics. Complex designs, formal evidence and reliable optimization require additional statistical expertise.

How many runs do I need?

The app uses 2 to the power of k combinations for k factors. Multiply by the runs per combination. Three factors with two runs each give 16 runs. Whether this is sufficient for a specific detection goal needs separate sample size planning.

What does “runs per combination” mean?

It includes the first run. A value of 2 means each combination is executed twice independently, not two additional runs after the first.

Can I use categorical factors?

Yes. Two materials or methods are possible. Enter their names as level − and level +. The signs are codes, not quality ratings.

How far apart should the levels be?

Far enough to reveal relevant changes while staying within feasible and meaningful limits. Narrow ranges may hide effects; very wide ranges may introduce different mechanisms or unacceptable conditions.

Can I add a third level or center points?

Not in this app. Center points can indicate curvature for quantitative factors. They need an appropriate extended design and analysis.

How does replication differ from repeated measurement?

An independent replicate uses a new experimental unit and re-established conditions. Repeated measurements of the same unit mainly examine measurement variation. Do not treat them as independent process runs.

Why randomize the order?

To reduce systematic alignment of time-related disturbances with factor levels. The app shuffles complete runs before execution. Once results are entered, order is locked.

What if a factor is hard to change?

Choose an appropriate structure such as split-plot. Do not simply rearrange these runs for convenience and then interpret them as a fully randomized experiment.

What happens with missing or invalid values?

Analysis waits until every run has a valid numeric response. Empty fields remain missing; zero is a real measurement. Decimal points, decimal commas and scientific notation are accepted, but thousands separators are not.

Can I analyze multiple responses?

Each project has one response. For multiple outcomes you can use separate project files with the same design. The app does not automatically optimize trade-offs.

What do positive and negative effects mean?

Positive means the response increases on average from level − to level +. Negative means it decreases. Whether this is desirable depends on your objective.

What is an interaction?

The effect of one factor changes depending on another. More pressure might have little benefit at low temperature but a large benefit at high temperature. Pressure must then be interpreted in context.

Are large bars statistically significant?

Not necessarily. Bars are ranked by absolute effect size without a significance test. Even a large observed effect needs to be assessed against variation, experimental conditions and technical plausibility.

Why are there no p-values?

Simplified automatic model selection can create misleading confidence. The app provides descriptive analysis, adding variation and standard error with independent replicates. It does not replace full statistical inference.

Is the best combination the optimum?

It is only the most favorable measured combination, based on its mean. Behavior between levels or outside the tested region may differ. Confirm the choice with new runs.

What if there is no clear effect?

Check measurement quality, level spacing, variation, selected factors and interactions. An inconclusive result does not prove a factor is unimportant under all conditions.

Can I use Excel data?

Yes. Download the plan as CSV. Paste responses from one spreadsheet column in the displayed run order. Use a JSON project file to transfer the entire editable project.

Is everything really free?

Yes. Planning, analysis, examples, guide, FAQ and exports are free. No registration, subscription or locked paid features. Personal consulting is a separate optional service.

Where are experiment data processed?

Calculations and project import run in your browser. The DOE app does not send entered experiment data to a server. The privacy policy covers general website services.

Will my project remain after I close the page?

The app does not automatically save permanently. Download a JSON project file before closing or reloading. Use “Load project” to continue later. CSV and print views document results; JSON preserves the editable project.

Can I use results commercially?

Yes, including workplace projects. Check method suitability, data quality and technical requirements. The app does not approve a process or provide certification evidence.

Which methods are not supported?

Fractional factorial, blocked, split-plot, mixture and response surface designs. No automated power planning, p-values, confidence intervals or model validation. The guide describes the free feature set.

Support for your experiment

Planning a process experiment or making sense of results? Discuss with Jonathan Sternberg what support would suit your project.

Discuss your experiment