2026-10-03 — geospatial · pipelines · photogrametry

Automatic frame culling before photogrammetry

Automatic frame culling before photogrammetryAutomatic frame culling before photogrammetry

Three-Arm Experimental Design

Objective

To measure how culling low-quality frames before photogrammetry affects the quality of the resulting orthophoto and the accuracy of georeferencing.

Why this design is needed

The naive way to evaluate the method would be: cull the frames, run photogrammetry, and compare the result with a run without culling. This setup contains a serious confounding variable.

When culling is applied, two things change at once:

  1. The number of processed frames decreases
  2. The composition of the processed frames changes

In a two-arm comparison (no culling / culling), any observed difference is the sum of these two effects and cannot be separated. If the result improves, there is no way to tell whether this comes from the scoring method or simply from the reduced data volume.

A third arm makes this distinction possible.

Experimental arms

A — Baseline

All frames from the flight are processed without any culling. This arm represents the current state and is the reference point against which the other two arms are compared.

B — Method

The proposed quality scorer assigns a score to each frame. Frames scoring below the threshold are removed, and the rest are processed. This is the actual condition under evaluation.

C — Control

The same number of frames removed in arm B are removed, chosen at random. The composition is random; the count is the same as in B.

The purpose of this arm is to isolate the effect of the reduced frame count. The difference between B and C stems only from which frames were culled, because how many frames were culled is the same in both.

Controlled variables

For the three arms to be comparable, the following must stay the same:

  • The same flight and the same pool of raw frames
  • The same photogrammetry engine and the same version
  • The same processing parameters (resolution, matching settings, output format)
  • The same set of ground control points
  • The same hardware (since processing time is compared)

The only independent variable is the culling method: none / scored / random.

Measured variables

Metric What it measures Why it matters
GCP residual error (RMSE) Georeferencing accuracy Based on ground truth, the strongest metric
Tenengrad sharpness Orthophoto sharpness Computed per tile, produces a heat map
Canny edge density Detail preservation Catches oversmoothing
Processing time Computational cost Measure of the practical gain

Sharpness and edge density are no-reference metrics; they are used for relative differences between arms, not for absolute quality.

Interpreting the results

There are four possible result patterns, and all four are meaningful:

Observation Interpretation
B > A and B > C The method works. The improvement comes from which frames were culled.
B > A but B ≈ C The improvement comes only from the reduced frame count. Scoring adds nothing extra.
B ≈ A and B > C The method lowers cost while preserving quality. Still a practical gain.
B < A and C < A Culling is harmful. The loss of observations outweighs the sharpness gain.

The last row is a negative result, but it is not an invalid result. Rejecting the hypothesis is also a finding; what matters is that the measurement is set up correctly.

Threats to validity and countermeasures

Spatial clustering of blurry frames

Low-quality frames are not randomly distributed across the dataset. They occur consecutively, due to causes such as wind, turning maneuvers or speed changes. Since consecutive frames cover spatially adjacent areas, culling becomes concentrated in certain regions.

As a result, nearly all frames that see a ground control point may be removed at once, and that point may fall below the minimum number of observations required by the photogrammetry engine. Looking at the average number of observations hides this; a per-point check is needed.

Countermeasure — constrained culling: A frame is not culled if removing it would drop the observation count of any ground control point below the set threshold (N). Among the frames that see a point, the highest-scoring ones are kept.

Results based on a single flight

If the study is carried out on a single flight, the findings remain dependent on that flight's conditions (lighting, wind, terrain texture). If possible, it should be repeated with a second flight under different conditions; otherwise, this should be stated as a limitation.

Arbitrariness of threshold selection

If the quality score threshold is set arbitrarily, the result can be manipulated. The threshold should be determined on a hand-labeled validation subset, and this process should be reported.

Limitations

  • The sharpness and edge metrics used are no-reference; they do not provide an absolute measure of quality.
  • If the random culling arm is run with a single random selection, it carries a chance effect; if possible, it should be repeated several times and averaged.
  • The results may be specific to the photogrammetry engine used. Repeating them with a second engine shows that the finding is engine-independent.

Journal