Skip to content

Quick-Start Guide

VIAME (Video and Image Analytics for Multiple Environments) is a do-it-yourself AI system for analyzing imagery and video, primarily targeting marine species analytics but also useful as a general computer vision toolkit.

There are 5 types of documentation: this quick-start guide, tutorial videos, user forums, example readmes, and the full manual. Installers (pre-built binaries), docker images, and source code are hosted on GitHub. Pre-built binaries are for users, while the source code and build instructions are for developers.

VIAME Flavors

VIAME comes in a few different interfaces with slightly different capabilities. Not listed in this document thoroughly are programming APIs for developers.

DIVE – Web and Desktop Annotator

Dive annotator Dive dataset list

Originally created as the VIAME-Web interface (with a public server hosted at https://viame.kitware.com), a desktop version of this web annotator and model trainer is also available in both Windows .msi installers and .zip release formats.

This tool is currently the most general purpose annotator, and supports polygons, lines, points, or boxes, and can train models over multiple videos or image sequences using standard models. See the DIVE interface pages for how to use it, and user interfaces for the ways of launching it.

Project Files

Project files are a collection of scripts targeting either groups of images or videos. They are documented on the project folders page. Project files are also used to launch some of the annotation GUIs in the desktop version of the software, or to train models across multiple sequences headless (without a GUI) to prevent the GUI from using any system resources while training, e.g. VRAM, reserving more for the training process.

Example Folders

In the “examples” folder of a VIAME install are a series of standalone .bat (Windows) or .sh (Linux) launchers broken down based on functionality covering all aspects of the system. See scripts and example folders for what each folder contains.

Command Line Interface

Everything VIAME does can be run from a terminal through the viame command, which suits batch processing, scripted workflows and machines without a display. See the command line interface page for its commands.

Deprecated Interfaces

Three older desktop tools remain in the installers for a few specialized cases, but are no longer developed now that DIVE covers what they do:

  • VIEW: the original C++ annotator for boxes and polygons, still quick on very high resolution imagery, see VIEW interface
  • SEARCH: standalone image and video search, refined by feedback on the results
  • SEAL: box annotation across 2 to 4 camera views side by side

Capabilities Breakdown

Y Supported P Partial N Not supported

FeatureExamplesProject FilesCommand LineDIVE DesktopDIVE Web
Platform
Included in the desktop installersYYYYN
Runs in a web browser from a remote serverNNNNY
Runs without a display, for batch scriptsYYYNP¹
Docker images providedYYYNY
Annotation
Boxes, polygons, keypoints and linesP²P²NYY
Point-click segmentationNNNYP³
Multi-camera and stereo annotationNNNYY
Tiled large images (GeoTIFF)NNNYY
Review grid across datasetsNNNYY
Processing
Detection and tracking pipelinesYYYYY
Stereo measurement pipelinesYNYYY
Interactive stereo measurementNNNYP³
Image and video search with refinementYYYYP⁴
Text queryYNYYN
Image enhancement outputYNYYY
Registration and mosaicingYYYP⁵P⁵
Scoring and evaluationYNYYY
Annotation format conversionYNYYY
Training
Detector training over multiple sequencesYYYYY
Frame classifier trainingYNYYY
Tracker trainingYNYYY
Add-on model pack downloadsNNYYP⁶

¹ Through the REST API
² Launches DIVE
³ Smaller models, run in the browser
⁴ Image queries only, models cannot be saved
⁵ Registration only, no mosaic output
⁶ Server administrators only

GPU vs CPU Installations

VIAME is designed to run on 8 Gb+ VRAM NVIDIA Graphics cards (1 or more), but:

  • Many algorithms can run with less and a generic 4 Gb patch is available on the install page
  • Also depends on if talking about just inference (pre-trained model running, uses less) or training
  • This is just for algorithms and processing pipelines; annotation GUIs can be run on CPU
  • The GPU and CPU installers are both listed in installing VIAME from binaries

Additionally:

  • Some algorithms are meant to run on CPU (motion tracker, baseline pixel classification)
  • Some algorithms are meant to run on GPU, but can run on CPU (deep frame classification)
  • Some are designed for GPU, and can run on but take forever on CPU (most deep CNN detectors, many deep learning training routines)

How do I know if I have a GPU?

Device manager gpu

On Windows, look in Device Manager. Sometimes computers have more than one card (one embedded on the motherboard, then a 2nd in a plugin slot). Next, search for the card to know its specifications. On Linux, many terminal commands can tell you which GPU you have (e.g. nvidia-smi, lspci | grep -i nvidia).

Types of Annotation and Detection Models

There are four main types of annotations and detection models:

Annotation box level

Box-Level: A bounding box around the object of interest.

Annotation frame level

Frame-Level: The entire frame is classified (e.g. the whole image has a label).

Annotation pixel level

Pixel-Level: Pixel masks or polygons tracing the exact outline of objects.

Annotation keypoints

Keypoints: Specific points of interest on objects (e.g. head, tail).

Each type has a page of its own, see object detection, frame level classification and scene segmentation, alongside detector training for training models of each.

Detections vs Tracks

Detections and tracks are synonymous across examples and user interfaces. A track is a (temporal) sequence of single-frame detections, but a detection can also be viewed as a track with just a single state. See object tracking for the tracking pipelines.

Detection example

Detection

Track example

Track

Annotation Formats

For details on annotation file formats, see the Detection File Formats section.

VIAME-CSV is the primary input/output format supported by default, with a single line for either each detection, or each detection state in a track. It has 9 required fields comma separated, with optional additional columns for keypoints, attributes, polygons, and masks.

COCO JSON adaptation is also supported by some GUIs, with added track support.

Annotation Best Practices

Annotation best practices Annotation gui example

When creating bounding box annotations:

  • The goal is for the center of bounding box to remain over the center of the tracked object without clipping too many extremity pixels
  • Attempt to avoid dramatic box size changes that aren’t associated with an object’s movement or overly large boxes
  • Need to consider efficiency (time) vs quality tradeoffs when deciding to do boxes vs pixel masks, box quality, keypoints + boxes, etc.

The DIVE annotation quickstart covers drawing each of these in the interface.

Model Generation Workflows

There are several routes from raw imagery to a working model, from annotating everything by hand to searching by text or example. See model generation workflows for the steps of each and how they compare.