ERP pipeline · PSYCH 390 step 9 / 11

Artifact Detection

Which trials should be thrown out?

Figure 1

Artifact Detection on 48 Single Trials at Fp1

flagged by the detector true ERP average of surviving trials average with no rejection

The detector

Where it looks

The trials and their artifacts are generated by the page, which is what makes the grading possible; real artifacts are more varied than the four kinds simulated here.

What You Are Looking At

Each small tile in Panel A is one trial, the epoch from −200 to 800 ms at Fp1, the electrode just above the left eye. Every trial contains the same true ERP plus ordinary background EEG, and about a quarter of them also contain something that should not be averaged: a blink, which at Fp1 is a positive bump of roughly 50 to 150 µV lasting 200 to 400 ms; a saccade, which shifts the voltage in a step; or a burst of muscle activity. Two trials have a large constant offset. It looks alarming but is not an artifact; baseline correction removes it completely. The page generated these trials, so it knows which ones are contaminated and can grade the detector the way a signal-detection task is graded. A coral tile is one the detector flagged; the small mark in the corner says whether that was a hit, a miss, or a false alarm.

In Panel B, the teal waveform is the average of the trials that survived, drawn over the true ERP in gold and the average with no rejection at all in gray. In Panel C, the dot plot tracks hits against false alarms as the threshold is moved, and the little head in Panel D shows why a blink is easy to find at the front. Its field is positive above the eye, negative below it, and fades toward the back of the head.

Try This

  1. Start with the moving-window peak-to-peak detector at 80 µV. Almost every blink is caught, one smaller blink slips through, the two offset trials are left alone, and the surviving average sits close to the true ERP.
  2. Drag the threshold up to 250 µV. Blinks slip through as misses and the teal average bends upward late in the epoch, where the blinks tend to fall.
  3. Drag it down to 30 µV. Now ordinary EEG trips the detector, a third of the trials are false alarms, and the average gets noisier without getting more accurate.
  4. Switch to the simple voltage threshold and turn off Baseline-correct first. The two offset trials are rejected for no good reason. Turn baseline correction back on and they survive.
  5. Choose the step function with a threshold of 20 to 30 µV. The three saccades that peak-to-peak missed at 80 µV are caught, because a saccade shifts the mean voltage in a step even though its size is modest. In practice you run more than one detector, each tuned to one kind of artifact.

Why It Matters for the Pipeline

There are three reasons to throw a trial away. Artifacts are large, so a few of them add more noise to an average than dozens of clean trials remove. Blinks and eye movements also change what reached the retina, so the brain's response on those trials is not the response you designed. And if one condition provokes more blinks than another, the difference between the averages will contain blink voltage rather than brain voltage. Rejection is a signal-detection problem. Pick a measure that is large when the artifact is present, set a threshold, and accept the trade-off between misses and false alarms (Luck, 2014, Chapter 6).

The simple voltage threshold tests every sample against a limit, so it needs baseline correction first and is thrown off by drift. The moving-window peak-to-peak measure looks at the largest swing inside a sliding window of about 200 ms, so a blink is caught wherever the baseline happens to be. The step function compares the mean before and after each point and picks up the sudden shift of a saccade. In ERPLAB the calls are EEG = pop_artextval(EEG, 'Channel', 1:33, 'Threshold', [-100 100], 'Twindow', [-200 800]); and EEG = pop_artmwppth(EEG, 'Channel', 1:33, 'Threshold', 100, 'Windowsize', 200, 'Windowstep', 50);. Afterwards pop_summary_AR_eeg_detection(EEG, 'Report') prints the percentage rejected per bin (Lopez-Calderon & Luck, 2014). Tune the parameters for each participant by scrolling through single trials, never by watching which setting makes the conditions differ. Rejecting 10 to 25% of trials is normal; above 50% points to a recording problem, and below 5% often means the threshold is too lenient. Luck (2014) recommends deciding in advance the percentage at which a participant is excluded, typically 25%. Independent component analysis is the alternative that subtracts the blink instead of dropping the trial, useful when participants blink on most trials.

References

Lopez-Calderon, J., & Luck, S. J. (2014). ERPLAB: An open-source toolbox for the analysis of event-related potentials. Frontiers in Human Neuroscience, 8, Article 213. https://doi.org/10.3389/fnhum.2014.00213

Luck, S. J. (2014). An introduction to the event-related potential technique (2nd ed.). MIT Press.