Skip to content

Review

Review shows many annotations at once as a grid of cropped image chips, so you can audit a whole class (or everything carrying an attribute) across one or more datasets and fix wrong types in place, without stepping through each sequence in the annotation viewer. It is available on both DIVE Web and DIVE Desktop as the Review tab in the navigation bar.

Opening the Review page

Web

  • Click the Review tab in the top navigation.
  • Or select one or more datasets on the Data home page and click Review.
  • Direct URL: /review?datasetIds=<id1>,<id2>.

Desktop

  • Click the Review tab in the navigation bar.
  • Or on the Library page, select projects and click Review.

Choosing datasets

The page has three views, toggled with the Results / Statistics / Datasets buttons at the top left; only one is shown at a time to keep the grid uncluttered.

The Datasets view has the same dataset picker as the Training and Pipelines pages: search the library, add datasets one at a time or with Select all, and drop them again with Remove all (on the web, Browse opens the folder picker instead). The selected datasets are listed underneath with their type, how many tracks they hold and their load state; reload a dataset’s annotations or remove it from there. Datasets picked here are only queued: their annotations are read when you switch to Results, so picking many costs nothing until you look at them.

Notes:

  • Multicamera datasets are added one camera at a time; adding the parent expands it into its cameras automatically.
  • Tiled large-image datasets load, but their media cannot be cropped into chips yet; their entries show a placeholder.
  • On web, the dataset’s default annotation set is used.

The grid

The Datasets panel opens on a first visit with nothing loaded; coming back with loaded datasets opens Results. The Results grid shows the annotations that match the current query. Until a dataset has been added and loaded it only says so. Each entry shows:

  • the annotation’s box, outlined, cropped out of the image or video frame with some extra context around it;
  • the confidence of the shown type (top left);
  • an editable type field underneath, with every type seen across the loaded datasets offered as suggestions;
  • the dataset name (when more than one is loaded), track id, and frame or frame count;
  • any polygon outline and head/tail points the detection carries, drawn over the chip.

Hover an entry for its actions (they grow under the mouse): mark correct (sets the shown type’s confidence to 1 and drops other candidate types), delete (a red X; the annotation is removed on the next save), edit geometry, and open in viewer. Double clicking the image also opens the annotation viewer on that dataset, seeks to the frame the entry is showing, and selects the track. While a review session is open, the viewer’s top bar shows a Review tab so you can jump back to the same grid and page.

Tracks cycle through their sampled frames at the dataset’s real-time rate (sparser samples wait proportionally longer, so a loop lasts about as long as the track does). The arrows in the filmstrip badge step through them by hand, which pauses the cycling on that frame until the play button resumes it; starting an edit pauses it too.

The type field and caption grow a little as the grid shows fewer entries, so a 3 by 3 grid is comfortably readable while a dense grid stays compact.

Stereo and multi-camera datasets

Each camera of a multi-camera (or stereo) dataset is loaded as its own sequence, but a track that appears in several cameras is one entry, showing a chip per camera side by side with the camera named on it. The chips show the same frames on every side, and zooming or panning one side moves the others with it. Where one camera has no detection on a frame the track has elsewhere, that side is still cropped at a position interpolated from its own neighbouring boxes and shows no box; its add box action creates a detection there, at the interpolated position, ready to be adjusted. The type field applies to the track in every camera, and opening the viewer opens the whole rig.

Editing boxes, polygons and points in place

Right click an entry (or use its edit action) to adjust the frame it is showing without opening the viewer. The cycling pauses on that frame, the box gains the same handles the annotator uses (in the type’s colour, red while dragged), and any polygon vertices and head/tail points can be dragged too. The mouse wheel zooms into the chip about the cursor and dragging empty space or middle-dragging anywhere pans it, as in the annotator; the zoom stays until you wheel back out. Right click again, or press Enter or Apply, to keep the change; Esc or Cancel drops it. The chip keeps its crop after an edit, with the box drawn over it at its new position. Edits are held with the type edits until you Save; when auto-save is enabled in the settings, review edits are saved after the same delay the annotator uses.

Tracks spanning several frames first show their first box, then, once the extra frames have loaded, cycle through up to eight boxes evenly sampled along the track. The object stays centred in the entry as it cycles. A filmstrip badge shows which sampled frame is on screen.

Querying

Two query modes are available in the toolbar:

  • Type: pick a type (or Any type) and a minimum confidence. The list offers only the types found on the selected datasets at that confidence or above, so every choice has results. Every track whose matching type meets the threshold is listed once.
  • Attribute: pick an attribute key, optionally a value, and whether to look at track attributes, detection (per-frame) attributes, or both. Matches on detection attributes show the frames that carry the attribute.

Changes to the query take effect as soon as they settle. The grid otherwise keeps its entries, so editing a type never reshuffles the page you are working on.

The gear next to Save opens the same settings as the annotator, including the auto-save switch and delay.

Entries are sorted by confidence, highest first, by default; the sort field also offers lowest first, dataset and track id, or frame.

Grid size and zoom

The mouse wheel over any entry zooms into it about the cursor, editing or not, and dragging the zoomed image (with the left or the middle button) pans it; wheel back out to return to the full chip.

The default grid is 5 columns by 4 rows. Set the columns and rows directly, or use the zoom buttons: zoom in shows fewer, larger entries and zoom out shows more, smaller ones, keeping the grid’s shape. The Context slider controls how much image is shown around each box, as a fraction of the box size (30% by default). These settings are remembered per browser.

Use the arrow keys, Page Up / Page Down, Home and End to page through the results. Paging quickly only loads the page you stop on; pages passed over are skipped.

Videos are read through a hidden player, or, for videos imported for native playback without transcoding, through the desktop’s frame extractor. Chips are rendered at the resolution of the cell they fill. A box a few pixels across still has only a few pixels of image behind it, but it is resampled once at full cell size rather than being stretched by the browser, so small objects come out as sharp as the source allows.

Editing types

Type into an entry’s type field and press Enter (or click away), or open its dropdown with the arrow button (which also closes it) and pick one of the known types, to reassign the annotation’s type; the new type becomes the top confidence pair with confidence 1, the same as changing the type in the viewer’s track list. Edited entries get an amber border and a pencil badge until saved.

Page actions applies to every entry on the current page: set them all to one type, or mark them all correct.

Nothing is written until you press Save; the Save button stays disabled until there is something to write, and a badge on it shows how many annotations are queued. Discard reloads the affected datasets. Leaving the page (or opening the viewer) with unsaved changes asks whether to Save and Leave, Discard and Leave, or Stay. After a clean leave, coming back to Review resumes the same datasets and page, opening Results when any are loaded. Starting Review from the library with a different selection begins a fresh session. Closing the browser tab or the application with unsaved changes asks for confirmation.

When you return from the viewer, Review reloads saved annotations before showing the chips, so a later type edit preserves geometry changed in the viewer. The query and page are retained. If a dataset cannot be refreshed, its stale annotations are not editable; retry loading it on the Datasets panel. If Discard and Leave cannot reload the original annotations, Review stays open with your changes pending and shows an error.

For multicamera queries, a match in any camera selects the logical track. Type changes, acceptance and deletion apply to all its displayed camera tracks, including cameras whose confidence or attributes did not match the query.

Web resource use and shared deployments

Review crops images and decodes playable videos in each user’s browser. It uses the existing authenticated media endpoints; opening a grid does not launch a pipeline, worker job, or server-side frame extraction. Videos must already be playable in the browser, as in the annotation viewer.

Each review session limits API operations to three at a time, including loading and saving across datasets. A web camera config also reads the parent configuration so hierarchy-aware type edits use the same taxonomy as the viewer. The web track reader skips groups and annotation-set listings, which the grid does not use. Chip rendering has a separate four-operation limit and prioritizes the visible page and the next page’s primary chips. Queued work for skipped pages is dropped. Image loads and video metadata/seeks time out after 15 seconds, and disposing a session cancels pending media work and prevents queued API calls from starting.

Review retains at most two decoded frames and 16 MiB of decoded pixels per dataset. Rendered chips are pruned when paging to retain 256 recent items, with the visible and prefetched pages protected. Selected annotations remain in browser memory, so the number and size of selected datasets still affect memory usage. These are per-browser limits, not a server-wide quota.

A retained review session belongs to the signed-in account and is cleared on logout or account change. Saves use the existing dataset write permissions and update only changed tracks. Review does not add collaborative locking or conflict detection: coordinate assignments when multiple people edit the same tracks, since a later save can replace another person’s edits.

When opening Review from a sequence with an empty dataset list, DIVE selects that sequence and opens Results after it loads. This also applies after clearing a previous selection. A nonempty selection is retained.

Stereo and multicamera sequences appear as one entry in the dataset list. Reloading or removing that entry affects all cameras. Annotations with the same track ID across cameras share one grid cell, showing each camera’s view.

Search results

On DIVE Desktop, the Video Search panel launches searches on the Query page, which shows ranked similarity results across indexed datasets in the same chip grid, with the same editing. Accept or reject entries with their corner buttons and Refine to re-rank; Hide reviewed (right of the context slider) takes the accepted and rejected entries out of the grid so only the unmarked ones remain.

Every result is also an annotation in waiting. Typing a type under a chip, or editing its box (the edit action or a right click), gives the result a track of its own in its dataset, built from the result’s frames and boxes; the dataset’s existing annotations are never edited in place. Such entries show their track id and can be deleted.

Save on the results toolbar writes the results that count as annotations: every accepted result (with its type, or unknown when none was given), every result given a type, rejected results only when they were given a type, and unmarked results that were typed or box-edited. When any of them overlaps an annotation the dataset already has on the same frame, a dialog asks how to proceed: Keep originals saves the other results and leaves the overlapping ones out, Replace overlapping deletes just the overlapped annotations before saving, Replace all annotations deletes every annotation in those sequences and saves only the results, and Discard results saves nothing. Discard drops the unsaved results, and leaving the page with unsaved changes asks first. Opening a saved entry selects its track in the viewer. Grid shape, zoom and context settings are shared with Review.

Stereo review on the web

Select the stereo sequence in the library or in Review → Datasets → Browse. Selecting an individual camera folder or entering review from a camera link also loads the whole sequence, so both sides appear together. The review picker accepts whole stereo/multicamera sequences; scoring still requires a single camera. Cameras appear together as one entry per track, in the sequence’s configured camera order, with synchronized frame cycling and zoom/pan.

Type assignment, acceptance, and deletion apply across the entry’s cameras. Geometry edits and adding a missing box affect only the chosen camera. Saves target the individual camera folders; if one save fails, its edits remain pending for retry. Opening an entry in the viewer opens the whole rig at the selected frame and track.

Statistics

Choose Review → Statistics to summarize every selected sequence. Queued sequences load when this tab opens; partial results are labelled while loading or if a sequence fails. Retry failed loads on Datasets.

Categories counts each track once per category whose confidence meets that sequence’s saved type-specific threshold (or its default threshold). Without saved filters, the viewer default of 0.1 applies. Unclassified tracks have their own row. A track with multiple qualifying labels can appear in multiple category counts. Attributes separates track and detection occurrences by name and value, including false and zero. Only attributes on qualifying tracks are counted. Stereo cameras contribute independently. Search, sorting, and pagination keep large category and attribute lists manageable.

Statistics use the current loaded annotations, including unsaved edits, and do not inherit the Results tab’s type, attribute, or confidence query. Changing the selected sequences, editing, deleting, or reloading annotations updates totals.

Timeline rows show qualifying track spans as colored, per-type steps in 200 bins, with the total number of tracks present as a grey outline. Each row has its own count axis, labelled in tracks and reaching that sequence’s peak, so a quiet sequence is as readable as a busy one; compare rows by their axis labels, not by the height of their plots. All cameras of a sequence share one plot; overlapping spans for the same track ID are counted once in each type series. Each row uses elapsed seconds, converting each camera by its own FPS (frames when any camera lacks FPS). Image sequences include their full frame range; video rows use the annotated extent because review metadata does not provide the full video duration. No media is decoded to construct these plots. Capture timestamps on images or parsed from sequence names order rows newest first; undated sequences follow alphabetically. Import dates are not treated as capture dates. Scroll inside the timeline list to browse more rows. Click a sequence name to open its viewer, or click a point on its plot to open the viewer at that point in the sequence.