VIAME as a Python Package¶
pip install viame installs two things: the viame command line tool, and
the viame python package that the tool itself is built on. Anything the
pipelines do can be driven directly from python, because the pipelines are
assembled out of the same algorithm interfaces this package exposes.
pip install viame
Linux and Windows, CPython 3.10 through 3.14. On Linux, GPU support comes
from whichever CUDA torch pulls in and no separate CUDA installation is
needed; on Windows the torch pins below are part of the install command.
Every example here was run against the published wheel with nothing
downloaded beyond the package itself.
Windows¶
Two commands, in this order:
pip install "torch<2.11" torchvision --index-url https://download.pytorch.org/whl/cu128
pip install viame
The wheels are built for CUDA 12, for the Pascal and Volta cards CUDA 13 dropped.
Why torch goes first, and capped. PyPI’s Windows torch is the CPU build
– torch marks its nvidia-* requirements platform_system == "Linux" – so
pip install viame on its own gives you a CPU torch. Installing a CUDA one
first avoids that, because pip leaves a requirement alone once it is
satisfied. The cap is needed because VIAME asks for torch<2.11 while
pytorch.org’s unpinned command installs 2.11: one release out of range, so
pip replaces it with 2.10 +cpu, reports success, and leaves
cuda.is_available() False with nothing to say why. torchvision needs no
pin of its own.
A CPU-only install – plain pip install viame – is supported. VIAME’s own
CUDA kernels come from the nvidia-*-cu12 wheels it declares, not from
torch, so they work either way; nothing in torch will use the GPU.
What the GPU has to be. Torch’s CUDA 12 builds start at sm_70, so on
anything older than Volta – a GTX 1080 Ti is sm_61 – torch reports
cuda.is_available() as True and then fails the first kernel launch with
no kernel image is available for execution on the device. VIAME’s own
kernels are built from sm_60 and do run on those cards.
Loading files¶
viame.open recognizes the same file formats as viame inspect and loads
its reader plugins automatically:
import viame
image = viame.open("image.png")
array = image.image().asarray() # RGB, preserving the reader's bit depth
for index, image in enumerate(viame.open("image_list.txt")):
process(image)
with viame.open("example.mp4", 5) as frames: # sample at 5 Hz
for image in frames:
process(image)
print(frames.timestamp.get_frame(), frames.timestamp.get_time_seconds())
# Equivalent keyword form:
with viame.open("example.mp4", frame_rate=5) as frames:
first_image = next(frames)
tracks = viame.open("annotations.csv") # VIAME CSV
tracks = viame.open("annotations.json") # recognizes DIVE or COCO JSON
for track in tracks.tracks():
print(track.id, len(track))
| Input | Return value |
|---|---|
| Still image | Native ImageContainer |
Image list (.txt) |
Lazy ImageSequence of image containers |
| Directory | Lazy image sequence, sorted by filename; subdirectories are skipped |
| Video | Lazy VideoSequence of image containers |
| VIAME CSV, DIVE JSON, COCO JSON | Native ObjectTrackSet, loaded through that format’s reader |
.pipe, supported model file or ZIP bundle |
Pipeline handle, prepared using the same model wrappers as viame run |
An image list contains one path per line; blank lines and comments beginning
with # are ignored. Relative paths resolve against the list’s directory
first, then against the current working directory for older lists. A directory
loads its immediate image files. Frames are loaded as iteration advances,
so opening a sequence does not load all its images into memory.
Video sampling keeps the first frame and then the first available frame at
or after each requested time. It uses presentation timestamps, including for
variable-rate video, and falls back to the source frame rate if timestamps
are unavailable. Requesting more than the source rate returns available
frames without duplicating them. The rate must be positive and finite and
is accepted only for video inputs. Native frame numbers and timestamps are
available as frames.timestamp; image lists have one-based frame numbers.
Iterators close at end of input or on a read error. Use with or call
close() when stopping early. Each iterator is single-pass; call viame.open
again to start over. Unsupported formats and invalid options raise errors;
opening a file never runs a pipeline or executes a model. Annotation files
are loaded as a complete track set, preserving the respective reader’s
track IDs, frame numbering and detection data. COCO annotations without a
track ID become individual single-state tracks.
Pipeline and model files are prepared at open time, then executed explicitly:
with viame.open("detector.pipe") as detector:
result = detector.run("example.mp4", output_dir="results", frame_rate=5)
# A ZIP may contain a pipeline and its weights, or any model package that
# viame run recognizes (ONNX, netharn, or weights with companion files).
with viame.open("model.zip") as detector:
detector.run("image_list.txt", output_dir="results")
print(detector.path) # prepared .pipe; valid until the handle is closed
# Multiple pipelines require an exact archive member name; no stdin prompt.
with viame.open("models.zip", pipeline="configs/detector.pipe") as detector:
detector.run("images/", output_dir="results")
# A self-contained pipe can supply its own input and output configuration.
with viame.open("complete.pipe") as pipeline:
pipeline.run()
Bare .pt, .pth, .ckpt, .weights and .onnx files use the same
identification and templates as viame run. Opening prepares configuration;
model weights load when execution begins. The installed algorithms must
support the selected model. Set up the VIAME environment as for the CLI.
Pipeline.run calls viame run synchronously and returns a
subprocess.CompletedProcess. It accepts file or directory inputs and uses
the command’s defaults, including its default sampling rate. Additional CLI
arguments go in args=["--no-reset-prompt", ...]; subprocess options such as
capture_output=True and timeout=60 are also accepted. Nonzero exit codes
raise subprocess.CalledProcessError unless check=False is supplied.
frame_rate belongs on .run() for pipelines; the second argument to
viame.open remains reserved for opening videos. Extracted files and rendered
templates are removed on close() or context exit. A handle can run multiple
inputs before closing. For in-memory processing, use embedded=True as described below, or the
lower-level native EmbeddedPipeline interface.
Algorithm .create() and native EmbeddedPipeline() automatically initialize
plugins when first used. viame.open also initializes the readers it needs.
The plugin manager remembers completed registration, so later calls do not
repeat it. The examples below need no explicit initialization call.
Explicit loading remains available for eager initialization or for inspecting
the registry (for example, registered_names()) before creating anything:
from viame.modules import modules
modules.load_known_modules()
Running a pipeline¶
Use the file loader for automatic adaptation, or the lower-level interfaces below.
Open a normal pipeline for in-memory processing¶
Use embedded=True to replace file readers and standard output writers with
native memory adapters. The conversion uses the shared C++ API;
includes, configuration substitutions and relative model paths are resolved
by the native pipeline parser. Models and ZIP
bundles use the same preparation as file-based viame.open.
with viame.open("detector.pipe", embedded=True) as detector:
print(detector.input_names) # e.g. ('input',)
print(detector.output_ports) # original writer process.port names
detector.send(viame.open("image.png")) # also accepts a numpy image array
result = detector.receive(timeout=30)
detections = result["detector_writer.detected_object_set"]
with viame.open("stereo.pipe", embedded=True) as stereo:
stereo.send({"input1": left_image, "input2": right_image})
result = stereo.receive(timeout=30)
One send supplies a synchronized set of camera frames. It accepts native
image containers without converting them to arrays, or 2D/3D NumPy arrays.
Camera names come from the original reader processes, including three-camera
and larger graphs. input_ports lists the connected reader ports. Standard
image, timestamp, filename and frame-rate ports are filled automatically:
pipeline.send(image, timestamp=source_timestamp, frame_rate=30,
values={"input.file_name": "frame00042.png"})
Without a timestamp, frame numbers start at one and time starts at zero at
the supplied frame_rate (default 1 Hz). Extra connected ports, such as
metadata or custom inputs, must be supplied through values using the
original process.port name.
Video readers and existing input adapters are recognized by process type.
Standard detection, track, image, video, homography and track-descriptor
writers, plus existing output adapters, are replaced. Custom source and sink
processes can be selected explicitly with
inputs=["camera_a", "camera_b"] and outputs=["custom_writer"].
Selected inputs must be sources and selected outputs must be sinks. Other
processes retain their configuration and behavior, including any other file
I/O they perform. Unsupported graphs fail during preparation or native setup.
receive() returns a dictionary keyed by the original writer input ports,
for example detector_writer.detected_object_set. Result values use the
native adapter types. Sampling and batching remain active: a send need
not produce a result, and output branches must have compatible output rates
because the output adapter synchronizes their ports. receive(timeout=...)
raises TimeoutError if no result arrives; the handle remains usable.
Send and receive calls use bounded native queues, so interleave them instead
of sending an entire video before reading results. Calls on a handle should
be made from one calling thread.
Opening an embedded pipeline configures its algorithms, loads its models and
starts it waiting for input. Use with or call close() to send end-of-input,
finish queued work, discard unread results and release the native pipeline
and temporary files. Embedded handles use send/receive; their file-based
run method is disabled.
In memory, through an embedded pipeline¶
EmbeddedPipeline runs a .pipe in process and lets you push data in and
pull results out, one frame at a time, with nothing written to disk. The
pipeline needs an input_adapter and an output_adapter where it would
otherwise have a reader and a writer:
# detect.pipe
process in
:: input_adapter
process detector
:: image_object_detector
:detector:type hough_circle
process out
:: output_adapter
connect from in.image to detector.image
connect from detector.detected_object_set to out.detected_object_set
connect from in.image to out.image
Then drive it:
import cv2, numpy as np
from viame.adapters import EmbeddedPipeline, AdapterDataSet
from viame.types import Image, ImageContainer
pipeline = EmbeddedPipeline()
pipeline.build_pipeline("detect.pipe", ".")
pipeline.start()
for frame in frames: # any numpy array
data = AdapterDataSet.create()
data["image"] = ImageContainer(Image(frame))
pipeline.send(data)
output = pipeline.receive()
detections = output["detected_object_set"]
print(len(detections))
pipeline.send_end_of_input()
pipeline.wait()
input_port_names() and output_port_names() report what the adapters
expose, which is how you find out what a given pipeline wants to be fed.
build_pipeline’s second argument anchors relativepath config entries to
the pipe file rather than the working directory.
Through the command line tool¶
For a pipeline that reads and writes files anyway – a whole video in, a csv out – there is nothing to gain from holding it in memory, and the tool already handles input types, output directories and batching:
import subprocess, pathlib
result = subprocess.run(
["viame", "run", "filter_enhance", "clip.mp4", "-o", "output"],
capture_output=True, text=True)
frames = sorted(pathlib.Path("output").rglob("frame*.png"))
print(result.returncode, len(frames))
The shipped pipelines live in <sys.prefix>/configs/pipelines, so a
pipeline can be named directly as above or given as a path. viame run
takes a video, an image list, or a folder.
Running a detector on your own imagery¶
ImageObjectDetector takes an ImageContainer and returns a detected
object set. Any numpy array will do, so the imagery need not come from a
file:
import cv2, numpy as np
from viame.algo import ImageObjectDetector
from viame.types import Image, ImageContainer
detector = ImageObjectDetector.create('hough_circle')
frame = np.zeros((240, 320, 3), dtype=np.uint8)
cv2.circle(frame, (110, 120), 40, (255, 255, 255), 2)
detections = detector.detect(ImageContainer(Image(frame)))
for d in detections:
box = d.bounding_box
print(box.min_x(), box.min_y(), box.max_x(), box.max_y(), d.confidence)
hough_circle is used here because it needs no model, so the example runs
on a bare install. ImageObjectDetector.registered_names() lists what this
build has; netharn, rf_detr, ultralytics, mmdet, onnx, darknet
and the rest need a model, which comes from an add-on pack:
viame add-ons
A detector that needs a model is configured the same way any algorithm is,
through its config block – see get_configuration and set_configuration
below.
Running a tracker¶
TrackObjects consumes per frame detections and returns the track set so
far. It is stateful: call it once per frame, in order.
from viame.algo import ImageObjectDetector, TrackObjects
from viame.types import Image, ImageContainer, Timestamp
detector = ImageObjectDetector.create('hough_circle')
tracker = TrackObjects.create('ocsort')
tracker.set_configuration(tracker.get_configuration())
for number, frame in enumerate(frames):
image = ImageContainer(Image(frame))
timestamp = Timestamp()
timestamp.set_frame(number)
timestamp.set_time_seconds(number / 30.0)
detections = detector.detect(image)
tracks = tracker.track(timestamp, image, detections)
print(number, len(detections), len(tracks.tracks()))
The set_configuration(get_configuration()) line is not optional. It is
what applies an implementation’s defaults; without it a tracker raises an
AttributeError on its first frame for a parameter it never initialised.
ocsort, bytetrack and botsort all work against detections alone.
srnn, siammask, deepsort, motr and sam3_tracker need models.
Reading video¶
VideoInput is the same reader the pipelines use, so it handles whatever
they do – video files through PyAV, or a list of images:
from viame.algo import VideoInput
reader = VideoInput.create('vidl_ffmpeg')
reader.open('clip.mp4')
while reader.next_frame():
image = reader.frame_image()
print(image.width(), image.height(), image.depth())
reader.close()
vidl_ffmpeg, ffmpeg and pyav are the same PyAV backed reader under
three names, which is where the FFmpeg comes from – this package bundles
none of its own. image_list reads a text file of image paths instead, and
VideoInput.registered_names() lists them all.
To go the other way, viame.algo.ImageIO reads and writes single images
and VideoOutput writes video.
Finding your way around¶
Every algorithm interface follows the same shape:
Interface.registered_names() # what this build can create
algo = Interface.create('name') # make one
config = algo.get_configuration() # its parameters, with defaults
algo.set_configuration(config) # apply them, after any edits
viame.algo holds the interfaces – detectors, trackers, readers, writers,
filters, stereo and measurement among them. viame.types holds what they
exchange: Image, ImageContainer, DetectedObject, Timestamp,
BoundingBoxD and so on.