How this comparison works
76 models on 25 scenes, 13 of them shipped with Oku3D. How each figure was arrived at, and where the comparison stops claiming anything.
Table of Contents
Why there is no accuracy score
An accuracy figure is a distance between a prediction and a truth. These scenes are generated images - nothing in them was measured with a depth sensor, so there is no truth to measure against. A figure computed against one model's output, or an average of several, would measure agreement with that choice and read as accuracy. A wrong number gets quoted; an absent one gets looked into.
Here instead is what can be established: what each model predicts, what it costs to run, what its license permits, and the same scene under two models at the same place. The one editorial judgment is which models Oku3D ships, and that is already visible.
The scenes
Each of the 25 was composed around one difficulty rather than for how it looks: a range spanning centimeters to kilometers, a surface with no texture, two cues that contradict each other, a drawing with no photographic cues at all. The line under each scene names its difficulty; the marked places say where it shows.
Generated rather than photographed, which cuts both ways: a test set built to order, at the cost of any ground truth. Every note is written from the delivered picture, because a generator does not always deliver what was ordered.
Speed and memory
Measured on an AMD Radeon RX 7900 XTX, on 39 settings - every shipped model at every inference resolution it offers, because a model that follows video at its smallest resolution can be far too slow at its largest.
Speed is the fastest of 40 to 100 timed runs per setting: a best case, not a typical one. Memory is a middle reading, not a peak. Other hardware moves both; how the settings compare is what carries over. Video marks what estimates depth inside one 24 fps frame on that machine - a slower setting still plays, it just refreshes depth less often than the picture.
Unshipped models carry no timings: they were run to produce their maps, but not on the same machine or through the same path, and a number measured differently would be worse than a blank.
What the pictures are
Each map is one model's output on one scene at one inference resolution, stored as neutral grey - dark for far, bright for near. The level the model produced is the pixel value, and each map is normalized over its own image, so the darkest point in a frame is that frame's furthest.
Color goes on between the file and the screen. That is what lets the computed views read the model's own level: recovering a level from a color needs the exact table it was colored with, and a lossy encoder has already moved it. The key beside the view picks the ramp - Plasma by default, plus plain grey (the file itself), grey inverted for the convention where a picture brightens with distance, Viridis, Magma and Batlow, all perceptually uniform and safe for color vision deficiency, and Turbo, which trades that evenness for contrast and will suggest an edge the data does not have. A ramp is a reading aid; the measurement under it does not move.
Levels are cut at the 0.5 and 99.5 percentiles of each frame, so anything beyond arrives flattened onto the first or last level and a map cannot show how much was flattened. The clipping marks paint those two levels in colors no ramp produces: where a range was cut, not how far past it went. Off by default, because a picture with them on has something painted over it.
Models also disagree on what they predict - relative disparity, metric depth in meters, affine-invariant depth. Per-image normalization is what makes them comparable to look at, and also why no absolute distance appears anywhere: a value would be a statement about the picture, not about the scene.
The edge, contour and 3D views
Three views work something out from the map rather than showing it: where a model put a depth discontinuity, lines of equal depth, and the source raised into a surface. All three read the map's brightness and nothing else.
That limit has a consequence worth stating. They can say where something is and what shape it has - whether a depth edge sits on the picture's own, whether a flat wall came out flat - and never how much. Distances, histograms, a value under the pointer: those need the numbers behind the pictures, and the pictures are what is published. Approximating them from brightness would look like a measurement, so they are not offered.
The 3D preview does what Oku3D does with a depth map, artifacts included: stretched pixels where a foreground edge pulls away from what is behind it. That is not a flaw in the preview - it is what a depth-image renderer does, and a still is the cheapest way to judge how a model will behave in motion.
It is also the only place here where anything is changed before being read. Oku3D cleans up depth edges before building a 3D view and ships with that on, so Oku3D edge cleanup starts out checked. It is not a blur - an edge keeps its place and an object its size, so a model that puts an edge in the wrong place still shows that - but it does shrink the halo around a foreground object, and with it a difference between models. Switch it off for the map exactly as the model wrote it: sharper, and torn wherever the surface has to absorb a whole depth step at once.
Corrections
If something here is wrong, or a model belongs on this page and is not, write to jens@oku3d.com. Neither the model list nor the scene set is closed.