MiDaS DPT-Hybrid

MiDaS DPT-Hybrid is one of the depth models Oku3D ships. Below is what it does to each of the 25 test scenes, and what it costs to run.

Table of Contents

What it is

Ships with Oku3DRelative disparityApache 2.0384×384
Released
2021
Predicts
Relative disparity
Parameters
122.4 M
Download
112 MB
License
Apache 2.0

Older architecture with more detail than the other MiDaS builds.

Speed and memory

ResolutionSpeedMemoryVideo
384×38425.2 ms0.47 GBKeeps up

Measured on an AMD Radeon RX 7900 XTX. Speed is the fastest of 40 to 100 timed runs per setting, so a best case, not a typical one; memory is a middle reading, not a peak. Other hardware shifts the numbers; how the settings compare carries over.

On all 25 scenes

Every scene was composed around one difficulty; the line under each map names it. At 384×384, this model's finest setting. Click one to open it beside the source image.

MiDaS DPT-Hybrid depth map of Alpine valleyAlpine valleyCentimeters to kilometers in one frame. Everything past the valley floor fits into the last sliver of a model's output, which is where distant ridges get pressed into one flat wall.MiDaS DPT-Hybrid depth map of Anime, cherry blossomsAnime, cherry blossomsTrained on photographs, asked about a drawing. Flat color fills and hard outlines remove the texture gradients and shading distance is read from, leaving only convention.MiDaS DPT-Hybrid depth map of Anime, ramen barAnime, ramen barThe bowl is drawn at a scale no lens would produce, so the two cues a model relies on contradict each other: by size it is enormous, by perspective it is close.MiDaS DPT-Hybrid depth map of Blizzard whiteoutBlizzard whiteoutLarge areas carry no texture and no contrast - nothing local to estimate from. Featureless white could be snow a meter away or sky a kilometer off, and the fence is the only structure left to lose.MiDaS DPT-Hybrid depth map of CG animation, bedroomCG animation, bedroomFur is thousands of strands with no solid surface behind them, and it was rendered rather than filmed - so the shading and sensor noise distance is read from are not in the pixels.MiDaS DPT-Hybrid depth map of Claymation kitchenClaymation kitchenGenuine photography at table scale: every cue is real and every one is scaled wrong, so reasoning about how big things usually are points the wrong way. The window is a painted flat.MiDaS DPT-Hybrid depth map of Dancer in confettiDancer in confettiHundreds of small objects spread through the entire depth range, each needing its own distance - and the nearest are so far out of focus they have no edge to attach one to.MiDaS DPT-Hybrid depth map of 1950s sitcom1950s sitcomNo color and almost no shadow modeling: two of the cues a model normally leans on are absent by construction. The black surround is an edge with nothing behind it.MiDaS DPT-Hybrid depth map of Fisherman portraitFisherman portraitIndividual strands of hair are thinner than the grid depth is estimated on. They either come out separate from the background or get rounded into the silhouette.MiDaS DPT-Hybrid depth map of Horror forestHorror forestMost of the frame is too dark to carry information, and the only bright things are close to the lens. That makes brightness look like a distance cue, which it is not: the fog behind is bright too.MiDaS DPT-Hybrid depth map of Kitchen dialogueKitchen dialogueFocus falls off towards the camera as well as away from it. A model that reads blur as distance has to put the foreground head somewhere behind the window.MiDaS DPT-Hybrid depth map of Macro, beeMacro, beeA working distance of a few centimeters, a range most models never saw: the training data is rooms and streets. The dew drops refract what is behind them.MiDaS DPT-Hybrid depth map of Metro corridorMetro corridorEvery column is identical, so how large one appears is the only thing saying how far away it is. Nothing else varies: even lighting, uniform tiling, the whole depth range stacked into the middle.MiDaS DPT-Hybrid depth map of Motorcycle chaseMotorcycle chaseThe background is blurred by movement, not distance, and a model that learned blur means far cannot tell them apart. The sharp thing and the streaked thing are almost equally close.MiDaS DPT-Hybrid depth map of News studioNews studioTwo flat surfaces busy showing depth that is not there: a screen with a globe on it, a floor with the whole set in it. Both invite a model to carve geometry out of a plane.MiDaS DPT-Hybrid depth map of Overhead cookingOverhead cookingShot straight down, so there is no perspective to read, and the whole scene is a hand's width deep. Either those small differences come out, or the table becomes one flat plane.MiDaS DPT-Hybrid depth map of Parking garage lineupParking garage lineupSix people standing within a few meters of each other. Telling near from far is easy here; getting their order right is not, and one swapped pair is visible immediately.MiDaS DPT-Hybrid depth map of Pixel art castlePixel art castleDepth exists only as discrete layers with nothing in between, and there is not a single photographic cue in the frame. Whatever comes out was inferred from convention alone.MiDaS DPT-Hybrid depth map of Prison mesh gatesPrison mesh gatesA mesh is mostly hole, and each hole opens onto a corridor meters behind it - so the answer flips between two distances dozens of times across a single row. A second gate repeats it at a third the pitch.MiDaS DPT-Hybrid depth map of Pure blackPure blackOne value across the whole frame: no texture, no edge, no horizon, nothing to read distance from. Paired with pure white, which carries the same information - none - and differs only in level.MiDaS DPT-Hybrid depth map of Pure whitePure whiteThe blizzard with the fence taken out, carried to where no pixel differs from its neighbour. There is no correct answer here, so whatever a map shows was not read from the picture.MiDaS DPT-Hybrid depth map of Rainy night streetRainy night streetThe wet road carries a mirror image of everything above it. A reflection has the structure of a real scene and none of its distance, and here there is more reflection than road.MiDaS DPT-Hybrid depth map of Watercolor illustrationWatercolor illustrationPainted edges bleed rather than end, so object boundaries are a matter of degree instead of a line. And the kite hangs in empty sky with nothing beside it to be measured against.MiDaS DPT-Hybrid depth map of Wedding, soap bubblesWedding, soap bubblesWhere a bubble crosses the couple, two surfaces share one pixel. A depth map holds one number per pixel, so transparency has no correct answer - only a choice, and models make different ones.MiDaS DPT-Hybrid depth map of Western duelWestern duelA wide empty ground plane whose only cue is how quickly the gravel texture gets finer. That gradient is subtle, and a small misreading moves the second man by tens of meters.

The rest of MiDaS

Seeing it against another model

A map on its own is hard to judge; the same scene under two models is not. Hold the space bar to cut between them - only what differs appears to move. No accuracy score anywhere on this site: these are generated images with no ground-truth depth to score against.