Skip to content

video

sleap_io.model.video

Data model for videos.

The Video class is a SLEAP data structure that stores information regarding a video and its components used in SLEAP.

Classes:

Name Description
HDF5Video

Video backend for reading videos stored in HDF5 files.

ImageVideo

Video backend for reading videos stored as image files.

MediaVideo

Video backend for reading videos stored as common media files.

Video

Video class used by sleap to represent videos and data associated with them.

VideoBackend

Base class for video backends.

VideoWriter

Simple video writer using imageio and FFMPEG.

Functions:

Name Description
is_file_accessible

Check if a file is accessible.

Attributes:

Name Type Description
__cached__

str(object='') -> str

__doc__

str(object='') -> str

__file__

str(object='') -> str

__name__

str(object='') -> str

__package__

str(object='') -> str

__cached__ = '/home/runner/work/sleap-io/sleap-io/sleap_io/model/__pycache__/video.cpython-313.pyc' module-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

__doc__ = 'Data model for videos.\n\nThe `Video` class is a SLEAP data structure that stores information regarding\na video and its components used in SLEAP.\n' module-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

__file__ = '/home/runner/work/sleap-io/sleap-io/sleap_io/model/video.py' module-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

__name__ = 'sleap_io.model.video' module-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

__package__ = 'sleap_io.model' module-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

HDF5Video

Bases: sleap_io.io.video_reading.VideoBackend

Video backend for reading videos stored in HDF5 files.

This backend supports reading videos stored in HDF5 files, both in rank-4 datasets as well as in datasets with lists of binary-encoded images.

Embedded image datasets are used in SLEAP when exporting package files (.pkg.slp) with videos embedded in them. This is useful for bundling training or inference data without having to worry about the videos (or frame images) being moved or deleted. It is expected that these types of datasets will be in a Group with a int8 variable length dataset called "video". This dataset must also contain an attribute called "format" with a string describing the image format (e.g., "png" or "jpg") which will be used to decode it appropriately.

If a frame_numbers dataset is present in the group, it will be used to map from source video frames to the frames in the dataset. This is useful to preserve frame indexing when exporting a subset of frames in the video. It will also be used to populate frame_map and source_inds attributes.

Attributes:

Name Type Description
filename

Path to HDF5 file (.h5, .hdf5 or .slp).

grayscale

Whether to force grayscale. If None, autodetect on first frame load.

keep_open

Whether to keep the video reader open between calls to read frames. If False, will close the reader after each call. If True (the default), it will keep the reader open and cache it for subsequent calls which may enhance the performance of reading multiple frames.

dataset

Name of dataset to read from. If None, will try to find a rank-4 dataset by iterating through datasets in the file. If specifying an embedded dataset, this can be the group containing a "video" dataset or the dataset itself (e.g., "video0" or "video0/video").

input_format

Format of the data in the dataset. One of "channels_last" (the default) in (frames, height, width, channels) order or "channels_first" in (frames, channels, width, height) order. Embedded datasets should use the "channels_last" format.

frame_map

Mapping from frame indices to indices in the dataset. This is used to translate between the frame indices of the images within their source video and the indices of the images in the dataset. This is only used when reading embedded image datasets.

source_filename

Path to the source video file. This is metadata and only used when reading embedded image datasets.

source_inds

Indices of the frames in the source video file. This is metadata and only used when reading embedded image datasets.

image_format

Format of the images in the embedded dataset. This is metadata and only used when reading embedded image datasets.

channel_order

Channel order of embedded images, either "RGB" or "BGR". This is used to ensure consistent color channel ordering when decoding embedded images. If the encoding and decoding plugins have different channel orders, the channels will be automatically flipped during decoding.

plugin

Plugin to use for decoding embedded images. One of "opencv" or "FFMPEG". If None, uses the global default or auto-detects based on available packages. Note that "pyav" is automatically mapped to "FFMPEG" since PyAV doesn't support image decoding.

Methods:

Name Description
__attrs_post_init__

Auto-detect dataset and frame map heuristically.

__eq__

Method generated by attrs for class HDF5Video.

__init__

Method generated by attrs for class HDF5Video.

__replace__

Method generated by attrs for class HDF5Video.

__repr__

Method generated by attrs for class HDF5Video.

__setattr__

Method generated by attrs for class HDF5Video.

decode_embedded

Decode an embedded image string into a numpy array.

get_frame_raw_bytes

Get raw encoded bytes for a frame without decoding.

has_frame

Check if a frame index is contained in the video.

read_test_frame

Read a single frame from the video to test for grayscale.

Source code in sleap_io/io/video_reading.py
@attrs.define
class HDF5Video(VideoBackend):
    """Video backend for reading videos stored in HDF5 files.

    This backend supports reading videos stored in HDF5 files, both in rank-4 datasets
    as well as in datasets with lists of binary-encoded images.

    Embedded image datasets are used in SLEAP when exporting package files (`.pkg.slp`)
    with videos embedded in them. This is useful for bundling training or inference data
    without having to worry about the videos (or frame images) being moved or deleted.
    It is expected that these types of datasets will be in a `Group` with a `int8`
    variable length dataset called `"video"`. This dataset must also contain an
    attribute called "format" with a string describing the image format (e.g., "png" or
    "jpg") which will be used to decode it appropriately.

    If a `frame_numbers` dataset is present in the group, it will be used to map from
    source video frames to the frames in the dataset. This is useful to preserve frame
    indexing when exporting a subset of frames in the video. It will also be used to
    populate `frame_map` and `source_inds` attributes.

    Attributes:
        filename: Path to HDF5 file (.h5, .hdf5 or .slp).
        grayscale: Whether to force grayscale. If None, autodetect on first frame load.
        keep_open: Whether to keep the video reader open between calls to read frames.
            If False, will close the reader after each call. If True (the default), it
            will keep the reader open and cache it for subsequent calls which may
            enhance the performance of reading multiple frames.
        dataset: Name of dataset to read from. If `None`, will try to find a rank-4
            dataset by iterating through datasets in the file. If specifying an embedded
            dataset, this can be the group containing a "video" dataset or the dataset
            itself (e.g., "video0" or "video0/video").
        input_format: Format of the data in the dataset. One of "channels_last" (the
            default) in `(frames, height, width, channels)` order or "channels_first" in
            `(frames, channels, width, height)` order. Embedded datasets should use the
            "channels_last" format.
        frame_map: Mapping from frame indices to indices in the dataset. This is used to
            translate between the frame indices of the images within their source video
            and the indices of the images in the dataset. This is only used when reading
            embedded image datasets.
        source_filename: Path to the source video file. This is metadata and only used
            when reading embedded image datasets.
        source_inds: Indices of the frames in the source video file. This is metadata
            and only used when reading embedded image datasets.
        image_format: Format of the images in the embedded dataset. This is metadata and
            only used when reading embedded image datasets.
        channel_order: Channel order of embedded images, either "RGB" or "BGR". This is
            used to ensure consistent color channel ordering when decoding embedded
            images. If the encoding and decoding plugins have different channel orders,
            the channels will be automatically flipped during decoding.
        plugin: Plugin to use for decoding embedded images. One of "opencv" or
            "FFMPEG". If None, uses the global default or auto-detects based on
            available packages. Note that "pyav" is automatically mapped to "FFMPEG"
            since PyAV doesn't support image decoding.
    """

    dataset: str | None = None
    input_format: str = attrs.field(
        default="channels_last",
        validator=attrs.validators.in_(["channels_last", "channels_first"]),
    )
    frame_map: dict[int, int] = attrs.field(init=False, default=attrs.Factory(dict))
    source_filename: str | None = None
    source_inds: np.ndarray | None = None
    image_format: str = "hdf5"
    channel_order: str = "RGB"
    plugin: str | None = None

    EXTS = ("h5", "hdf5", "slp")

    def __attrs_post_init__(self):
        """Auto-detect dataset and frame map heuristically."""
        # Check if the file accessible before applying heuristics.
        try:
            f = h5py.File(self.filename, "r")
        except OSError:
            return

        if self.dataset is None:
            # Iterate through datasets to find a rank 4 array.
            def find_movies(name, obj):
                if isinstance(obj, h5py.Dataset) and obj.ndim == 4:
                    self.dataset = name
                    return True

            f.visititems(find_movies)

        if self.dataset is None:
            # Iterate through datasets to find an embedded video dataset.
            def find_embedded(name, obj):
                if isinstance(obj, h5py.Dataset) and name.endswith("/video"):
                    self.dataset = name
                    return True

            f.visititems(find_embedded)

        if self.dataset is None:
            # Couldn't find video datasets.
            return

        if isinstance(f[self.dataset], h5py.Group):
            # If this is a group, assume it's an embedded video dataset.
            if "video" in f[self.dataset]:
                self.dataset = f"{self.dataset}/video"

        if self.dataset.split("/")[-1] == "video":
            # This may be an embedded video dataset. Check for frame map.
            ds = f[self.dataset]

            if "format" in ds.attrs:
                self.image_format = ds.attrs["format"]

            # Read channel_order, with backwards compatibility
            if "channel_order" in ds.attrs:
                self.channel_order = ds.attrs["channel_order"]
            else:
                # Backwards compatibility: Check format_id for older files
                # Prior to format 1.4, embedded images were primarily encoded with
                # OpenCV which uses BGR, so default to BGR for older formats
                if "metadata" in f and "format_id" in f["metadata"].attrs:
                    format_id = f["metadata"].attrs["format_id"]
                    if format_id < 1.4:
                        self.channel_order = "BGR"  # Legacy default
                # If no format_id found, assume BGR (safest legacy default)
                # since most embedded images before this change used OpenCV

            if "frame_numbers" in ds.parent:
                frame_numbers = ds.parent["frame_numbers"][:].astype(int)
                self.frame_map = {frame: idx for idx, frame in enumerate(frame_numbers)}
                self.source_inds = frame_numbers

            if "source_video" in ds.parent:
                self.source_filename = json.loads(
                    ds.parent["source_video"].attrs["json"]
                )["backend"]["filename"]

            # Read FPS from attributes if present
            if "fps" in ds.attrs:
                self._fps = float(ds.attrs["fps"])
            elif "fps" in ds.parent.attrs:
                self._fps = float(ds.parent.attrs["fps"])

        f.close()

        # Set default plugin if not specified (use image plugin, not video plugin)
        if self.plugin is None:
            # Check image plugin default first (for embedded images)
            if _default_image_plugin is not None:
                self.plugin = _default_image_plugin
            # Otherwise auto-detect (for embedded image decoding)
            elif "cv2" in sys.modules:
                self.plugin = "opencv"
            else:
                self.plugin = "imageio"  # imageio fallback

    @property
    def num_frames(self) -> int:
        """Number of frames in the video."""
        with h5py.File(self.filename, "r") as f:
            return f[self.dataset].shape[0]

    @property
    def img_shape(self) -> tuple[int, int, int]:
        """Shape of a single frame in the video as `(height, width, channels)`."""
        with h5py.File(self.filename, "r") as f:
            ds = f[self.dataset]

            img_shape = None
            if "height" in ds.attrs:
                # Try to get shape from the attributes.
                img_shape = (
                    ds.attrs["height"],
                    ds.attrs["width"],
                    ds.attrs["channels"],
                )

                if img_shape[0] == 0 or img_shape[1] == 0:
                    # Invalidate the shape if the attributes are zero.
                    img_shape = None

            if img_shape is None and self.image_format == "hdf5" and ds.ndim == 4:
                # Use the dataset shape if just stored as a rank-4 array.
                img_shape = ds.shape[1:]

                if self.input_format == "channels_first":
                    img_shape = img_shape[::-1]

        if img_shape is None:
            # Fall back to reading a test frame.
            return super().img_shape

        return int(img_shape[0]), int(img_shape[1]), int(img_shape[2])

    def read_test_frame(self) -> np.ndarray:
        """Read a single frame from the video to test for grayscale."""
        if self.frame_map:
            frame_idx = list(self.frame_map.keys())[0]
        else:
            frame_idx = 0
        return self._read_frame(frame_idx)

    @property
    def has_embedded_images(self) -> bool:
        """Return True if the dataset contains embedded images."""
        return self.image_format is not None and self.image_format != "hdf5"

    @property
    def embedded_frame_inds(self) -> list[int]:
        """Return the frame indices of the embedded images."""
        return list(self.frame_map.keys())

    def decode_embedded(self, img_string: np.ndarray) -> np.ndarray:
        """Decode an embedded image string into a numpy array.

        Args:
            img_string: Binary string of the image as a `int8` numpy vector with the
                bytes as values corresponding to the format-encoded image.

        Returns:
            The decoded image as a numpy array of shape `(height, width, channels)`. If
            a rank-2 image is decoded, it will be expanded such that channels will be 1.

            This method does not apply grayscale conversion as per the `grayscale`
            attribute. Use the `get_frame` or `get_frames` methods of the `VideoBackend`
            to apply grayscale conversion rather than calling this function directly.
        """
        # Decode based on plugin
        if self.plugin == "opencv":
            img = cv2.imdecode(img_string, cv2.IMREAD_UNCHANGED)
            decoder_order = "BGR"  # OpenCV decodes to BGR
        else:
            # Use imageio for FFMPEG or any other plugin
            img = iio.imread(BytesIO(img_string), extension=f".{self.image_format}")
            decoder_order = "RGB"  # imageio decodes to RGB

        if img.ndim == 2:
            img = np.expand_dims(img, axis=-1)

        # Convert channel order if needed
        # If the stored order doesn't match the decoder order, flip channels
        if img.shape[-1] == 3 and self.channel_order != decoder_order:
            img = img[..., ::-1]  # Flip RGB <-> BGR

        return img

    def has_frame(self, frame_idx: int) -> bool:
        """Check if a frame index is contained in the video.

        Args:
            frame_idx: Index of frame to check.

        Returns:
            `True` if the index is contained in the video, otherwise `False`.
        """
        if self.frame_map:
            return frame_idx in self.frame_map
        else:
            return frame_idx < len(self)

    def get_frame_raw_bytes(self, frame_idx: int) -> np.ndarray | None:
        """Get raw encoded bytes for a frame without decoding.

        This method reads the raw compressed image data (PNG/JPEG bytes) directly
        from the HDF5 dataset without decoding it. This is useful for fast copying
        of embedded images when the target format matches the source format.

        Args:
            frame_idx: Index of the frame to read.

        Returns:
            Raw encoded bytes as int8 numpy array, or None if:
            - The backend doesn't have embedded images (including "hdf5" format which
              stores raw numpy arrays, not encoded images)
            - The frame index is not available

        Notes:
            For variable-length datasets, returns the raw bytes directly.
            For fixed-length datasets, returns bytes with trailing zeros stripped.
        """
        if not self.has_embedded_images:
            return None

        if not self.has_frame(frame_idx):
            return None

        # Get the internal index (handle frame_map)
        internal_idx = (
            self.frame_map.get(frame_idx, frame_idx) if self.frame_map else frame_idx
        )

        # Read directly from dataset
        if self.keep_open:
            if self._open_reader is None:
                self._open_reader = h5py.File(self.filename, "r")
            f = self._open_reader
        else:
            f = h5py.File(self.filename, "r")

        ds = f[self.dataset]
        raw_bytes = ds[internal_idx]

        # Handle fixed-length padding (strip trailing zeros)
        is_vlen = h5py.check_vlen_dtype(ds.dtype) is not None
        if not is_vlen:
            # Find last non-zero byte
            non_zero_mask = raw_bytes != 0
            if non_zero_mask.any():
                last_non_zero = np.where(non_zero_mask)[0][-1]
                raw_bytes = raw_bytes[: last_non_zero + 1]

        if not self.keep_open:
            f.close()

        return raw_bytes

    def _read_frame(self, frame_idx: int) -> np.ndarray:
        """Read a single frame from the video.

        Args:
            frame_idx: Index of frame to read.

        Returns:
            The frame as a numpy array of shape `(height, width, channels)`.

        Notes:
            This does not apply grayscale conversion. It is recommended to use the
            `get_frame` method of the `VideoBackend` class instead.
        """
        if self.keep_open:
            if self._open_reader is None:
                self._open_reader = h5py.File(self.filename, "r")
            f = self._open_reader
        else:
            f = h5py.File(self.filename, "r")

        ds = f[self.dataset]

        if self.frame_map:
            frame_idx = self.frame_map[frame_idx]

        img = ds[frame_idx]

        if self.has_embedded_images:
            img = self.decode_embedded(img)

        if self.input_format == "channels_first":
            img = np.transpose(img, (2, 1, 0))

        if not self.keep_open:
            f.close()
        return img

    def _read_frames(self, frame_inds: list) -> np.ndarray:
        """Read a list of frames from the video.

        Args:
            frame_inds: List of indices of frames to read.

        Returns:
            The frame as a numpy array of shape `(frames, height, width, channels)`.

        Notes:
            This does not apply grayscale conversion. It is recommended to use the
            `get_frames` method of the `VideoBackend` class instead.
        """
        if self.keep_open:
            if self._open_reader is None:
                self._open_reader = h5py.File(self.filename, "r")
            f = self._open_reader
        else:
            f = h5py.File(self.filename, "r")

        if self.frame_map:
            frame_inds = [self.frame_map[idx] for idx in frame_inds]

        ds = f[self.dataset]
        imgs = ds[frame_inds]

        if "format" in ds.attrs:
            imgs = np.stack(
                [self.decode_embedded(img) for img in imgs],
                axis=0,
            )

        if self.input_format == "channels_first":
            imgs = np.transpose(imgs, (0, 3, 2, 1))

        if not self.keep_open:
            f.close()

        return imgs

EXTS = ('h5', 'hdf5', 'slp') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__annotations__ = {'dataset': 'str | None', 'input_format': 'str', 'frame_map': 'dict[int, int]', 'source_filename': 'str | None', 'source_inds': 'np.ndarray | None', 'image_format': 'str', 'channel_order': 'str', 'plugin': 'str | None'} class-attribute

dict() -> new empty dictionary dict(mapping) -> new dictionary initialized from a mapping object's (key, value) pairs dict(iterable) -> new dictionary initialized as if via: d = {} for k, v in iterable: d[k] = v dict(**kwargs) -> new dictionary initialized with the name=value pairs in the keyword argument list. For example: dict(one=1, two=2)

__attrs_own_setattr__ = True class-attribute

Returns True when the argument is true, False otherwise. The builtins True and False are the only two instances of the class bool. The class bool is a subclass of the class int, and cannot be subclassed.

__attrs_props__ = ClassProps(is_exception=False, is_slotted=True, has_weakref_slot=True, is_frozen=False, kw_only=<KeywordOnly.NO: 'no'>, collected_fields_by_mro=True, added_init=True, added_repr=True, added_eq=True, added_ordering=False, hashability=<Hashability.UNHASHABLE: 'unhashable'>, added_match_args=True, added_str=False, added_pickling=True, on_setattr_hook=<function pipe.<locals>.wrapped_pipe at 0x7f41e68fca40>, field_transformer=None) class-attribute

Effective class properties as derived from parameters to attr.s() or define() decorators.

This is the same data structure that attrs uses internally to decide how to construct the final class.

Warning:

This feature is currently **experimental** and is not covered by our
strict backwards-compatibility guarantees.

Attributes:

Name Type Description
is_exception bool

Whether the class is treated as an exception class.

is_slotted bool

Whether the class is slotted <slotted classes>.

has_weakref_slot bool

Whether the class has a slot for weak references.

is_frozen bool

Whether the class is frozen.

kw_only KeywordOnly

Whether / how the class enforces keyword-only arguments on the __init__ method.

collected_fields_by_mro bool

Whether the class fields were collected by method resolution order. That is, correctly but unlike dataclasses.

added_init bool

Whether the class has an attrs-generated __init__ method.

added_repr bool

Whether the class has an attrs-generated __repr__ method.

added_eq bool

Whether the class has attrs-generated equality methods.

added_ordering bool

Whether the class has attrs-generated ordering methods.

hashability Hashability

How hashable <hashing> the class is.

added_match_args bool

Whether the class supports positional match <match> over its fields.

added_str bool

Whether the class has an attrs-generated __str__ method.

added_pickling bool

Whether the class has attrs-generated __getstate__ and __setstate__ methods for pickle.

on_setattr_hook Callable[[Any, Attribute[Any], Any], Any] | None

The class's __setattr__ hook.

field_transformer Callable[[Attribute[Any]], Attribute[Any]] | None

The class's field transformers <transform-fields>.

.. versionadded:: 25.4.0

__doc__ = 'Video backend for reading videos stored in HDF5 files.\n\nThis backend supports reading videos stored in HDF5 files, both in rank-4 datasets\nas well as in datasets with lists of binary-encoded images.\n\nEmbedded image datasets are used in SLEAP when exporting package files (`.pkg.slp`)\nwith videos embedded in them. This is useful for bundling training or inference data\nwithout having to worry about the videos (or frame images) being moved or deleted.\nIt is expected that these types of datasets will be in a `Group` with a `int8`\nvariable length dataset called `"video"`. This dataset must also contain an\nattribute called "format" with a string describing the image format (e.g., "png" or\n"jpg") which will be used to decode it appropriately.\n\nIf a `frame_numbers` dataset is present in the group, it will be used to map from\nsource video frames to the frames in the dataset. This is useful to preserve frame\nindexing when exporting a subset of frames in the video. It will also be used to\npopulate `frame_map` and `source_inds` attributes.\n\nAttributes:\n filename: Path to HDF5 file (.h5, .hdf5 or .slp).\n grayscale: Whether to force grayscale. If None, autodetect on first frame load.\n keep_open: Whether to keep the video reader open between calls to read frames.\n If False, will close the reader after each call. If True (the default), it\n will keep the reader open and cache it for subsequent calls which may\n enhance the performance of reading multiple frames.\n dataset: Name of dataset to read from. If `None`, will try to find a rank-4\n dataset by iterating through datasets in the file. If specifying an embedded\n dataset, this can be the group containing a "video" dataset or the dataset\n itself (e.g., "video0" or "video0/video").\n input_format: Format of the data in the dataset. One of "channels_last" (the\n default) in `(frames, height, width, channels)` order or "channels_first" in\n `(frames, channels, width, height)` order. Embedded datasets should use the\n "channels_last" format.\n frame_map: Mapping from frame indices to indices in the dataset. This is used to\n translate between the frame indices of the images within their source video\n and the indices of the images in the dataset. This is only used when reading\n embedded image datasets.\n source_filename: Path to the source video file. This is metadata and only used\n when reading embedded image datasets.\n source_inds: Indices of the frames in the source video file. This is metadata\n and only used when reading embedded image datasets.\n image_format: Format of the images in the embedded dataset. This is metadata and\n only used when reading embedded image datasets.\n channel_order: Channel order of embedded images, either "RGB" or "BGR". This is\n used to ensure consistent color channel ordering when decoding embedded\n images. If the encoding and decoding plugins have different channel orders,\n the channels will be automatically flipped during decoding.\n plugin: Plugin to use for decoding embedded images. One of "opencv" or\n "FFMPEG". If None, uses the global default or auto-detects based on\n available packages. Note that "pyav" is automatically mapped to "FFMPEG"\n since PyAV doesn\'t support image decoding.\n' class-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

__firstlineno__ = 883 class-attribute

int([x]) -> integer int(x, base=10) -> integer

Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.int(). For floating-point numbers, this truncates towards zero.

If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by '+' or '-' and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal.

int('0b100', base=0) 4

__match_args__ = ('filename', 'grayscale', 'keep_open', '_cached_shape', '_open_reader', '_fps', 'dataset', 'input_format', 'source_filename', 'source_inds', 'image_format', 'channel_order', 'plugin') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__module__ = 'sleap_io.io.video_reading' class-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

__slots__ = ('dataset', 'input_format', 'frame_map', 'source_filename', 'source_inds', 'image_format', 'channel_order', 'plugin') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__static_attributes__ = ('_fps', '_open_reader', 'channel_order', 'dataset', 'frame_map', 'image_format', 'plugin', 'source_filename', 'source_inds') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

embedded_frame_inds property

Return the frame indices of the embedded images.

has_embedded_images property

Return True if the dataset contains embedded images.

img_shape property

Shape of a single frame in the video as (height, width, channels).

num_frames property

Number of frames in the video.

__attrs_post_init__()

Auto-detect dataset and frame map heuristically.

Source code in sleap_io/io/video_reading.py
def __attrs_post_init__(self):
    """Auto-detect dataset and frame map heuristically."""
    # Check if the file accessible before applying heuristics.
    try:
        f = h5py.File(self.filename, "r")
    except OSError:
        return

    if self.dataset is None:
        # Iterate through datasets to find a rank 4 array.
        def find_movies(name, obj):
            if isinstance(obj, h5py.Dataset) and obj.ndim == 4:
                self.dataset = name
                return True

        f.visititems(find_movies)

    if self.dataset is None:
        # Iterate through datasets to find an embedded video dataset.
        def find_embedded(name, obj):
            if isinstance(obj, h5py.Dataset) and name.endswith("/video"):
                self.dataset = name
                return True

        f.visititems(find_embedded)

    if self.dataset is None:
        # Couldn't find video datasets.
        return

    if isinstance(f[self.dataset], h5py.Group):
        # If this is a group, assume it's an embedded video dataset.
        if "video" in f[self.dataset]:
            self.dataset = f"{self.dataset}/video"

    if self.dataset.split("/")[-1] == "video":
        # This may be an embedded video dataset. Check for frame map.
        ds = f[self.dataset]

        if "format" in ds.attrs:
            self.image_format = ds.attrs["format"]

        # Read channel_order, with backwards compatibility
        if "channel_order" in ds.attrs:
            self.channel_order = ds.attrs["channel_order"]
        else:
            # Backwards compatibility: Check format_id for older files
            # Prior to format 1.4, embedded images were primarily encoded with
            # OpenCV which uses BGR, so default to BGR for older formats
            if "metadata" in f and "format_id" in f["metadata"].attrs:
                format_id = f["metadata"].attrs["format_id"]
                if format_id < 1.4:
                    self.channel_order = "BGR"  # Legacy default
            # If no format_id found, assume BGR (safest legacy default)
            # since most embedded images before this change used OpenCV

        if "frame_numbers" in ds.parent:
            frame_numbers = ds.parent["frame_numbers"][:].astype(int)
            self.frame_map = {frame: idx for idx, frame in enumerate(frame_numbers)}
            self.source_inds = frame_numbers

        if "source_video" in ds.parent:
            self.source_filename = json.loads(
                ds.parent["source_video"].attrs["json"]
            )["backend"]["filename"]

        # Read FPS from attributes if present
        if "fps" in ds.attrs:
            self._fps = float(ds.attrs["fps"])
        elif "fps" in ds.parent.attrs:
            self._fps = float(ds.parent.attrs["fps"])

    f.close()

    # Set default plugin if not specified (use image plugin, not video plugin)
    if self.plugin is None:
        # Check image plugin default first (for embedded images)
        if _default_image_plugin is not None:
            self.plugin = _default_image_plugin
        # Otherwise auto-detect (for embedded image decoding)
        elif "cv2" in sys.modules:
            self.plugin = "opencv"
        else:
            self.plugin = "imageio"  # imageio fallback

__eq__(other)

Method generated by attrs for class HDF5Video.

Source code in sleap_io/io/video_reading.py
    import cv2
except ImportError:
    pass

try:
    import imageio_ffmpeg  # noqa: F401
except ImportError:
    pass

try:
    import av  # noqa: F401
except ImportError:
    pass


# Track available backends (populated on module import)
_AVAILABLE_VIDEO_BACKENDS = {
    "opencv": "cv2" in sys.modules,
    "FFMPEG": "imageio_ffmpeg" in sys.modules,

__init__(filename, grayscale=None, keep_open=True, cached_shape=None, open_reader=None, fps=None, dataset=None, input_format='channels_last', source_filename=None, source_inds=None, image_format='hdf5', channel_order='RGB', plugin=None)

Method generated by attrs for class HDF5Video.

Source code in sleap_io/io/video_reading.py
    "pyav": "av" in sys.modules,
}

_AVAILABLE_IMAGE_BACKENDS = {
    "opencv": "cv2" in sys.modules,
    "imageio": True,  # Always available (core dependency)
}


# Global default video plugin
_default_video_plugin: str | None = None


def normalize_plugin_name(plugin: str) -> str:
    """Normalize plugin names to standard format.

    Args:
        plugin: Plugin name or alias (case-insensitive).

__replace__(**changes)

Method generated by attrs for class HDF5Video.

Source code in sleap_io/io/video_reading.py
Returns:

__repr__()

Method generated by attrs for class HDF5Video.

Source code in sleap_io/io/video_reading.py
"""Backends for reading videos."""

from __future__ import annotations

import sys
from io import BytesIO
from pathlib import Path

import attrs
import h5py
import imageio.v3 as iio
import numpy as np
import simplejson as json

try:

__setattr__(name, val)

Method generated by attrs for class HDF5Video.

Source code in sleap_io/io/video_reading.py
    f = self._open_reader
else:
    f = h5py.File(self.filename, "r")

ds = f[self.dataset]
raw_bytes = ds[internal_idx]

# Handle fixed-length padding (strip trailing zeros)
is_vlen = h5py.check_vlen_dtype(ds.dtype) is not None

decode_embedded(img_string)

Decode an embedded image string into a numpy array.

Parameters:

Name Type Description Default
img_string ndarray

Binary string of the image as a int8 numpy vector with the bytes as values corresponding to the format-encoded image.

required

Returns:

Type Description
ndarray

The decoded image as a numpy array of shape (height, width, channels). If a rank-2 image is decoded, it will be expanded such that channels will be 1.

This method does not apply grayscale conversion as per the grayscale attribute. Use the get_frame or get_frames methods of the VideoBackend to apply grayscale conversion rather than calling this function directly.

Source code in sleap_io/io/video_reading.py
def decode_embedded(self, img_string: np.ndarray) -> np.ndarray:
    """Decode an embedded image string into a numpy array.

    Args:
        img_string: Binary string of the image as a `int8` numpy vector with the
            bytes as values corresponding to the format-encoded image.

    Returns:
        The decoded image as a numpy array of shape `(height, width, channels)`. If
        a rank-2 image is decoded, it will be expanded such that channels will be 1.

        This method does not apply grayscale conversion as per the `grayscale`
        attribute. Use the `get_frame` or `get_frames` methods of the `VideoBackend`
        to apply grayscale conversion rather than calling this function directly.
    """
    # Decode based on plugin
    if self.plugin == "opencv":
        img = cv2.imdecode(img_string, cv2.IMREAD_UNCHANGED)
        decoder_order = "BGR"  # OpenCV decodes to BGR
    else:
        # Use imageio for FFMPEG or any other plugin
        img = iio.imread(BytesIO(img_string), extension=f".{self.image_format}")
        decoder_order = "RGB"  # imageio decodes to RGB

    if img.ndim == 2:
        img = np.expand_dims(img, axis=-1)

    # Convert channel order if needed
    # If the stored order doesn't match the decoder order, flip channels
    if img.shape[-1] == 3 and self.channel_order != decoder_order:
        img = img[..., ::-1]  # Flip RGB <-> BGR

    return img

get_frame_raw_bytes(frame_idx)

Get raw encoded bytes for a frame without decoding.

This method reads the raw compressed image data (PNG/JPEG bytes) directly from the HDF5 dataset without decoding it. This is useful for fast copying of embedded images when the target format matches the source format.

Parameters:

Name Type Description Default
frame_idx int

Index of the frame to read.

required

Returns:

Type Description
ndarray | None

Raw encoded bytes as int8 numpy array, or None if: - The backend doesn't have embedded images (including "hdf5" format which stores raw numpy arrays, not encoded images) - The frame index is not available

Notes

For variable-length datasets, returns the raw bytes directly. For fixed-length datasets, returns bytes with trailing zeros stripped.

Source code in sleap_io/io/video_reading.py
def get_frame_raw_bytes(self, frame_idx: int) -> np.ndarray | None:
    """Get raw encoded bytes for a frame without decoding.

    This method reads the raw compressed image data (PNG/JPEG bytes) directly
    from the HDF5 dataset without decoding it. This is useful for fast copying
    of embedded images when the target format matches the source format.

    Args:
        frame_idx: Index of the frame to read.

    Returns:
        Raw encoded bytes as int8 numpy array, or None if:
        - The backend doesn't have embedded images (including "hdf5" format which
          stores raw numpy arrays, not encoded images)
        - The frame index is not available

    Notes:
        For variable-length datasets, returns the raw bytes directly.
        For fixed-length datasets, returns bytes with trailing zeros stripped.
    """
    if not self.has_embedded_images:
        return None

    if not self.has_frame(frame_idx):
        return None

    # Get the internal index (handle frame_map)
    internal_idx = (
        self.frame_map.get(frame_idx, frame_idx) if self.frame_map else frame_idx
    )

    # Read directly from dataset
    if self.keep_open:
        if self._open_reader is None:
            self._open_reader = h5py.File(self.filename, "r")
        f = self._open_reader
    else:
        f = h5py.File(self.filename, "r")

    ds = f[self.dataset]
    raw_bytes = ds[internal_idx]

    # Handle fixed-length padding (strip trailing zeros)
    is_vlen = h5py.check_vlen_dtype(ds.dtype) is not None
    if not is_vlen:
        # Find last non-zero byte
        non_zero_mask = raw_bytes != 0
        if non_zero_mask.any():
            last_non_zero = np.where(non_zero_mask)[0][-1]
            raw_bytes = raw_bytes[: last_non_zero + 1]

    if not self.keep_open:
        f.close()

    return raw_bytes

has_frame(frame_idx)

Check if a frame index is contained in the video.

Parameters:

Name Type Description Default
frame_idx int

Index of frame to check.

required

Returns:

Type Description
bool

True if the index is contained in the video, otherwise False.

Source code in sleap_io/io/video_reading.py
def has_frame(self, frame_idx: int) -> bool:
    """Check if a frame index is contained in the video.

    Args:
        frame_idx: Index of frame to check.

    Returns:
        `True` if the index is contained in the video, otherwise `False`.
    """
    if self.frame_map:
        return frame_idx in self.frame_map
    else:
        return frame_idx < len(self)

read_test_frame()

Read a single frame from the video to test for grayscale.

Source code in sleap_io/io/video_reading.py
def read_test_frame(self) -> np.ndarray:
    """Read a single frame from the video to test for grayscale."""
    if self.frame_map:
        frame_idx = list(self.frame_map.keys())[0]
    else:
        frame_idx = 0
    return self._read_frame(frame_idx)

ImageVideo

Bases: sleap_io.io.video_reading.VideoBackend

Video backend for reading videos stored as image files.

This backend supports reading videos stored as a list of images.

Attributes:

Name Type Description
filename

Path to image files.

grayscale

Whether to force grayscale. If None, autodetect on first frame load.

plugin

Image plugin to use for reading. One of "opencv" or "imageio". If None, uses global default from get_default_image_plugin(), or auto-detects.

Methods:

Name Description
__eq__

Method generated by attrs for class ImageVideo.

__init__

Method generated by attrs for class ImageVideo.

__replace__

Method generated by attrs for class ImageVideo.

__repr__

Method generated by attrs for class ImageVideo.

__setattr__

Method generated by attrs for class ImageVideo.

find_images

Find images in a folder and return a list of filenames.

Source code in sleap_io/io/video_reading.py
@attrs.define
class ImageVideo(VideoBackend):
    """Video backend for reading videos stored as image files.

    This backend supports reading videos stored as a list of images.

    Attributes:
        filename: Path to image files.
        grayscale: Whether to force grayscale. If None, autodetect on first frame load.
        plugin: Image plugin to use for reading. One of "opencv" or "imageio".
            If None, uses global default from get_default_image_plugin(), or
            auto-detects.
    """

    EXTS = ("png", "jpg", "jpeg", "tif", "tiff", "bmp")

    plugin: str = attrs.field()

    @plugin.validator
    def _validate_plugin(self, attribute, value):
        """Validate and normalize plugin name."""
        normalized = normalize_image_plugin_name(value)
        object.__setattr__(self, attribute.name, normalized)

    @plugin.default
    def _default_plugin(self) -> str:
        """Get default plugin, checking global default first."""
        # Check global default first
        if _default_image_plugin is not None:
            # Warn if preferred plugin not available
            if not _AVAILABLE_IMAGE_BACKENDS.get(_default_image_plugin, False):
                import warnings

                available = get_available_image_backends()
                install_cmd = get_installation_instructions(
                    _default_image_plugin, "image"
                )
                warnings.warn(
                    f"Preferred image plugin '{_default_image_plugin}' is not "
                    f"available. Available plugins: {available}\n"
                    f"Install with: {install_cmd}"
                )
                # Fall through to auto-detection
            else:
                return _default_image_plugin

        # Otherwise auto-detect
        if "cv2" in sys.modules:
            return "opencv"
        else:
            return "imageio"

    @staticmethod
    def find_images(folder: str) -> list[str]:
        """Find images in a folder and return a list of filenames."""
        folder = Path(folder)
        return sorted(
            [f.as_posix() for f in folder.glob("*") if f.suffix[1:] in ImageVideo.EXTS]
        )

    @property
    def num_frames(self) -> int:
        """Number of frames in the video."""
        return len(self.filename)

    def _read_frame(self, frame_idx: int) -> np.ndarray:
        """Read a single frame from the video.

        Args:
            frame_idx: Index of frame to read.

        Returns:
            The frame as a numpy array of shape `(height, width, channels)` in RGB
            order.

        Notes:
            This does not apply grayscale conversion. It is recommended to use the
            `get_frame` method of the `VideoBackend` class instead.

            Images are always returned in RGB order regardless of plugin:
            - imageio: Returns RGB natively
            - opencv: Returns BGR, automatically flipped to RGB
        """
        if self.plugin == "opencv":
            # OpenCV reads as BGR, flip to RGB
            img = cv2.imread(self.filename[frame_idx], cv2.IMREAD_UNCHANGED)
            if img is None:
                raise ValueError(f"Failed to read image: {self.filename[frame_idx]}")
            if img.ndim == 3 and img.shape[-1] == 3:
                img = img[..., ::-1]  # BGR -> RGB
        else:  # imageio
            # imageio reads as RGB natively
            img = iio.imread(self.filename[frame_idx])

        if img.ndim == 2:
            img = np.expand_dims(img, axis=-1)

        return img

EXTS = ('png', 'jpg', 'jpeg', 'tif', 'tiff', 'bmp') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__annotations__ = {'plugin': 'str'} class-attribute

dict() -> new empty dictionary dict(mapping) -> new dictionary initialized from a mapping object's (key, value) pairs dict(iterable) -> new dictionary initialized as if via: d = {} for k, v in iterable: d[k] = v dict(**kwargs) -> new dictionary initialized with the name=value pairs in the keyword argument list. For example: dict(one=1, two=2)

__attrs_own_setattr__ = True class-attribute

Returns True when the argument is true, False otherwise. The builtins True and False are the only two instances of the class bool. The class bool is a subclass of the class int, and cannot be subclassed.

__attrs_props__ = ClassProps(is_exception=False, is_slotted=True, has_weakref_slot=True, is_frozen=False, kw_only=<KeywordOnly.NO: 'no'>, collected_fields_by_mro=True, added_init=True, added_repr=True, added_eq=True, added_ordering=False, hashability=<Hashability.UNHASHABLE: 'unhashable'>, added_match_args=True, added_str=False, added_pickling=True, on_setattr_hook=<function pipe.<locals>.wrapped_pipe at 0x7f41e68fca40>, field_transformer=None) class-attribute

Effective class properties as derived from parameters to attr.s() or define() decorators.

This is the same data structure that attrs uses internally to decide how to construct the final class.

Warning:

This feature is currently **experimental** and is not covered by our
strict backwards-compatibility guarantees.

Attributes:

Name Type Description
is_exception bool

Whether the class is treated as an exception class.

is_slotted bool

Whether the class is slotted <slotted classes>.

has_weakref_slot bool

Whether the class has a slot for weak references.

is_frozen bool

Whether the class is frozen.

kw_only KeywordOnly

Whether / how the class enforces keyword-only arguments on the __init__ method.

collected_fields_by_mro bool

Whether the class fields were collected by method resolution order. That is, correctly but unlike dataclasses.

added_init bool

Whether the class has an attrs-generated __init__ method.

added_repr bool

Whether the class has an attrs-generated __repr__ method.

added_eq bool

Whether the class has attrs-generated equality methods.

added_ordering bool

Whether the class has attrs-generated ordering methods.

hashability Hashability

How hashable <hashing> the class is.

added_match_args bool

Whether the class supports positional match <match> over its fields.

added_str bool

Whether the class has an attrs-generated __str__ method.

added_pickling bool

Whether the class has attrs-generated __getstate__ and __setstate__ methods for pickle.

on_setattr_hook Callable[[Any, Attribute[Any], Any], Any] | None

The class's __setattr__ hook.

field_transformer Callable[[Attribute[Any]], Attribute[Any]] | None

The class's field transformers <transform-fields>.

.. versionadded:: 25.4.0

__doc__ = 'Video backend for reading videos stored as image files.\n\nThis backend supports reading videos stored as a list of images.\n\nAttributes:\n filename: Path to image files.\n grayscale: Whether to force grayscale. If None, autodetect on first frame load.\n plugin: Image plugin to use for reading. One of "opencv" or "imageio".\n If None, uses global default from get_default_image_plugin(), or\n auto-detects.\n' class-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

__firstlineno__ = 1275 class-attribute

int([x]) -> integer int(x, base=10) -> integer

Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.int(). For floating-point numbers, this truncates towards zero.

If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by '+' or '-' and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal.

int('0b100', base=0) 4

__match_args__ = ('filename', 'grayscale', 'keep_open', '_cached_shape', '_open_reader', '_fps', 'plugin') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__module__ = 'sleap_io.io.video_reading' class-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

__slots__ = ('plugin',) class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__static_attributes__ = () class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

num_frames property

Number of frames in the video.

__eq__(other)

Method generated by attrs for class ImageVideo.

Source code in sleap_io/io/video_reading.py
    import cv2
except ImportError:
    pass

try:
    import imageio_ffmpeg  # noqa: F401
except ImportError:
    pass

try:
    import av  # noqa: F401
except ImportError:

__init__(filename, grayscale=None, keep_open=True, cached_shape=None, open_reader=None, fps=None, plugin=NOTHING)

Method generated by attrs for class ImageVideo.

Source code in sleap_io/io/video_reading.py
    pass


# Track available backends (populated on module import)
_AVAILABLE_VIDEO_BACKENDS = {
    "opencv": "cv2" in sys.modules,
    "FFMPEG": "imageio_ffmpeg" in sys.modules,
    "pyav": "av" in sys.modules,
}

_AVAILABLE_IMAGE_BACKENDS = {
    "opencv": "cv2" in sys.modules,
    "imageio": True,  # Always available (core dependency)
}

__replace__(**changes)

Method generated by attrs for class ImageVideo.

Source code in sleap_io/io/video_reading.py
Returns:

__repr__()

Method generated by attrs for class ImageVideo.

Source code in sleap_io/io/video_reading.py
"""Backends for reading videos."""

from __future__ import annotations

import sys
from io import BytesIO
from pathlib import Path

import attrs
import h5py
import imageio.v3 as iio
import numpy as np
import simplejson as json

try:

__setattr__(name, val)

Method generated by attrs for class ImageVideo.

Source code in sleap_io/io/video_reading.py
    f = self._open_reader
else:
    f = h5py.File(self.filename, "r")

ds = f[self.dataset]
raw_bytes = ds[internal_idx]

# Handle fixed-length padding (strip trailing zeros)
is_vlen = h5py.check_vlen_dtype(ds.dtype) is not None

find_images(folder) staticmethod

Find images in a folder and return a list of filenames.

Source code in sleap_io/io/video_reading.py
@staticmethod
def find_images(folder: str) -> list[str]:
    """Find images in a folder and return a list of filenames."""
    folder = Path(folder)
    return sorted(
        [f.as_posix() for f in folder.glob("*") if f.suffix[1:] in ImageVideo.EXTS]
    )

MediaVideo

Bases: sleap_io.io.video_reading.VideoBackend

Video backend for reading videos stored as common media files.

This backend supports reading through FFMPEG (the default), pyav, or OpenCV. Here are their trade-offs:

- "opencv": Fastest video reader, but only supports a limited number of codecs
    and may not be able to read some videos. It requires `opencv-python` to be
    installed. It is the fastest because it uses the OpenCV C++ library to read
    videos, but is limited by the version of FFMPEG that was linked into it at
    build time as well as the OpenCV version used.
- "FFMPEG": Slowest, but most reliable. This is the default backend. It requires
    `imageio-ffmpeg` and a `ffmpeg` executable on the system path (which can be
    installed via conda). The `imageio` plugin for FFMPEG reads frames into raw
    bytes which are communicated to Python through STDOUT on a subprocess pipe,
    which can be slow. However, it is the most reliable and feature-complete. If
    you install the conda-forge version of ffmpeg, it will be compiled with
    support for many codecs, including GPU-accelerated codecs like NVDEC for
    H264 and others.
- "pyav": Supports most codecs that FFMPEG does, but not as complete or reliable
    of an implementation in `imageio` as FFMPEG for some video types. It is
    faster than FFMPEG because it uses the `av` package to read frames directly
    into numpy arrays in memory without the need for a subprocess pipe. These
    are Python bindings for the C library libav, which is the same library that
    FFMPEG uses under the hood.

Attributes:

Name Type Description
filename

Path to video file.

grayscale

Whether to force grayscale. If None, autodetect on first frame load.

keep_open

Whether to keep the video reader open between calls to read frames. If False, will close the reader after each call. If True (the default), it will keep the reader open and cache it for subsequent calls which may enhance the performance of reading multiple frames.

plugin

Video plugin to use. One of "opencv", "FFMPEG", or "pyav". If None, will use the first available plugin in the order listed above.

Methods:

Name Description
__eq__

Method generated by attrs for class MediaVideo.

__init__

Method generated by attrs for class MediaVideo.

__replace__

Method generated by attrs for class MediaVideo.

__repr__

Method generated by attrs for class MediaVideo.

__setattr__

Method generated by attrs for class MediaVideo.

Source code in sleap_io/io/video_reading.py
@attrs.define
class MediaVideo(VideoBackend):
    """Video backend for reading videos stored as common media files.

    This backend supports reading through FFMPEG (the default), pyav, or OpenCV. Here
    are their trade-offs:

        - "opencv": Fastest video reader, but only supports a limited number of codecs
            and may not be able to read some videos. It requires `opencv-python` to be
            installed. It is the fastest because it uses the OpenCV C++ library to read
            videos, but is limited by the version of FFMPEG that was linked into it at
            build time as well as the OpenCV version used.
        - "FFMPEG": Slowest, but most reliable. This is the default backend. It requires
            `imageio-ffmpeg` and a `ffmpeg` executable on the system path (which can be
            installed via conda). The `imageio` plugin for FFMPEG reads frames into raw
            bytes which are communicated to Python through STDOUT on a subprocess pipe,
            which can be slow. However, it is the most reliable and feature-complete. If
            you install the conda-forge version of ffmpeg, it will be compiled with
            support for many codecs, including GPU-accelerated codecs like NVDEC for
            H264 and others.
        - "pyav": Supports most codecs that FFMPEG does, but not as complete or reliable
            of an implementation in `imageio` as FFMPEG for some video types. It is
            faster than FFMPEG because it uses the `av` package to read frames directly
            into numpy arrays in memory without the need for a subprocess pipe. These
            are Python bindings for the C library libav, which is the same library that
            FFMPEG uses under the hood.

    Attributes:
        filename: Path to video file.
        grayscale: Whether to force grayscale. If None, autodetect on first frame load.
        keep_open: Whether to keep the video reader open between calls to read frames.
            If False, will close the reader after each call. If True (the default), it
            will keep the reader open and cache it for subsequent calls which may
            enhance the performance of reading multiple frames.
        plugin: Video plugin to use. One of "opencv", "FFMPEG", or "pyav". If `None`,
            will use the first available plugin in the order listed above.
    """

    plugin: str = attrs.field()

    @plugin.validator
    def _validate_plugin(self, attribute, value):
        # Normalize the plugin name
        normalized = normalize_plugin_name(value)
        # Update the actual value to the normalized version
        object.__setattr__(self, attribute.name, normalized)

    EXTS = ("mp4", "avi", "mov", "mj2", "mkv")

    @plugin.default
    def _default_plugin(self) -> str:
        # Check global default first
        if _default_video_plugin is not None:
            # Warn if preferred plugin not available
            if not _AVAILABLE_VIDEO_BACKENDS.get(_default_video_plugin, False):
                import warnings

                available = get_available_video_backends()
                install_cmd = get_installation_instructions(_default_video_plugin)
                warnings.warn(
                    f"Preferred video plugin '{_default_video_plugin}' is not "
                    f"available. Available plugins: {available}\n"
                    f"Install with: {install_cmd}"
                )
                # Fall through to auto-detection
            else:
                return _default_video_plugin

        # Auto-detect based on what's available
        if "cv2" in sys.modules:
            return "opencv"
        elif "imageio_ffmpeg" in sys.modules:
            return "FFMPEG"
        elif "av" in sys.modules:
            return "pyav"
        else:
            # Enhanced error message with installation instructions
            raise ImportError(
                "No video backend plugins are available.\n\n"
                "The bundled imageio-ffmpeg should be available by default.\n"
                "If you see this error, try reinstalling sleap-io:\n"
                "  pip install --force-reinstall sleap-io\n\n"
                "Alternative backends:\n"
                "  opencv (fastest):  pip install sleap-io[opencv]\n"
                "  pyav (balanced):   pip install sleap-io[pyav]\n\n"
                "For more information, see: https://io.sleap.ai"
            )

    @property
    def reader(self) -> object:
        """Return the reader object for the video, caching if necessary."""
        if self.keep_open:
            if self._open_reader is None:
                if self.plugin == "opencv":
                    self._open_reader = cv2.VideoCapture(self.filename)
                elif self.plugin == "pyav" or self.plugin == "FFMPEG":
                    self._open_reader = iio.imopen(
                        self.filename, "r", plugin=self.plugin
                    )
            return self._open_reader
        else:
            if self.plugin == "opencv":
                return cv2.VideoCapture(self.filename)
            elif self.plugin == "pyav" or self.plugin == "FFMPEG":
                return iio.imopen(self.filename, "r", plugin=self.plugin)

    @property
    def num_frames(self) -> int:
        """Number of frames in the video."""
        if self.plugin == "opencv":
            return int(self.reader.get(cv2.CAP_PROP_FRAME_COUNT))
        else:
            props = iio.improps(self.filename, plugin=self.plugin)
            n_frames = props.n_images
            if np.isinf(n_frames):
                legacy_reader = self.reader.legacy_get_reader()
                # Note: This might be super slow for some videos, so maybe we should
                # defer evaluation of this or give the user control over it.
                n_frames = legacy_reader.count_frames()
            return n_frames

    @property
    def fps(self) -> float | None:
        """Frames per second from video container metadata.

        Returns:
            The FPS from the video container, or None if it cannot be determined.

        Notes:
            This reads the FPS from the video file metadata using the appropriate
            method for the current plugin:
            - OpenCV: cv2.CAP_PROP_FPS
            - FFMPEG/pyav: imageio metadata
        """
        # Return cached/explicit value if set
        if self._fps is not None:
            return self._fps

        # Read from container metadata
        try:
            if self.plugin == "opencv":
                fps = self.reader.get(cv2.CAP_PROP_FPS)
                return fps if fps > 0 else None
            else:
                # Use imageio v2 API to get metadata (v3 improps doesn't include fps)
                import imageio.v2 as iio_v2

                reader = iio_v2.get_reader(self.filename, format="FFMPEG")
                meta = reader.get_meta_data()
                reader.close()
                fps = meta.get("fps")
                return float(fps) if fps is not None else None
        except Exception:
            return None

    @fps.setter
    def fps(self, value: float | None) -> None:
        """Set an explicit FPS override.

        Args:
            value: Frames per second. Must be positive if not None.

        Raises:
            ValueError: If value is not positive.

        Notes:
            Setting FPS on MediaVideo overrides the value from container metadata.
            This can be useful when the container metadata is incorrect or missing.
        """
        if value is not None and value <= 0:
            raise ValueError(f"FPS must be positive, got {value}")
        self._fps = value

    def _read_frame(self, frame_idx: int) -> np.ndarray:
        """Read a single frame from the video.

        Args:
            frame_idx: Index of frame to read.

        Returns:
            The frame as a numpy array of shape `(height, width, channels)`.

        Notes:
            This does not apply grayscale conversion. It is recommended to use the
            `get_frame` method of the `VideoBackend` class instead.
        """
        if self.plugin == "opencv":
            if self.keep_open:
                if self._open_reader is None:
                    self._open_reader = cv2.VideoCapture(self.filename)
                reader = self._open_reader
            else:
                reader = cv2.VideoCapture(self.filename)

            if reader.get(cv2.CAP_PROP_POS_FRAMES) != frame_idx:
                reader.set(cv2.CAP_PROP_POS_FRAMES, frame_idx)
            success, img = reader.read()

            if success:
                img = img[..., ::-1]  # BGR -> RGB

        elif self.plugin == "pyav" or self.plugin == "FFMPEG":
            if self.keep_open:
                img = self.reader.read(index=frame_idx)
            else:
                with iio.imopen(self.filename, "r", plugin=self.plugin) as reader:
                    img = reader.read(index=frame_idx)
            success = img is not None

        if not success:
            raise IndexError(f"Failed to read frame index {frame_idx}.")

        return img

    def _read_frames(self, frame_inds: list) -> np.ndarray:
        """Read a list of frames from the video.

        Args:
            frame_inds: List of indices of frames to read.

        Returns:
            The frame as a numpy array of shape `(frames, height, width, channels)`.

        Notes:
            This does not apply grayscale conversion. It is recommended to use the
            `get_frames` method of the `VideoBackend` class instead.
        """
        if self.plugin == "opencv":
            if self.keep_open:
                if self._open_reader is None:
                    self._open_reader = cv2.VideoCapture(self.filename)
                reader = self._open_reader
            else:
                reader = cv2.VideoCapture(self.filename)

            reader.set(cv2.CAP_PROP_POS_FRAMES, frame_inds[0])
            imgs = []
            for idx in frame_inds:
                if reader.get(cv2.CAP_PROP_POS_FRAMES) != idx:
                    reader.set(cv2.CAP_PROP_POS_FRAMES, idx)
                _, img = reader.read()
                imgs.append(img)
            imgs = np.stack(imgs, axis=0)

            imgs = imgs[..., ::-1]  # BGR -> RGB

        elif self.plugin == "pyav" or self.plugin == "FFMPEG":
            if self.keep_open:
                if self._open_reader is None:
                    self._open_reader = iio.imopen(
                        self.filename, "r", plugin=self.plugin
                    )
                reader = self._open_reader
                imgs = np.stack([reader.read(index=idx) for idx in frame_inds], axis=0)
            else:
                with iio.imopen(self.filename, "r", plugin=self.plugin) as reader:
                    imgs = np.stack(
                        [reader.read(index=idx) for idx in frame_inds], axis=0
                    )
        return imgs

EXTS = ('mp4', 'avi', 'mov', 'mj2', 'mkv') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__annotations__ = {'plugin': 'str'} class-attribute

dict() -> new empty dictionary dict(mapping) -> new dictionary initialized from a mapping object's (key, value) pairs dict(iterable) -> new dictionary initialized as if via: d = {} for k, v in iterable: d[k] = v dict(**kwargs) -> new dictionary initialized with the name=value pairs in the keyword argument list. For example: dict(one=1, two=2)

__attrs_own_setattr__ = True class-attribute

Returns True when the argument is true, False otherwise. The builtins True and False are the only two instances of the class bool. The class bool is a subclass of the class int, and cannot be subclassed.

__attrs_props__ = ClassProps(is_exception=False, is_slotted=True, has_weakref_slot=True, is_frozen=False, kw_only=<KeywordOnly.NO: 'no'>, collected_fields_by_mro=True, added_init=True, added_repr=True, added_eq=True, added_ordering=False, hashability=<Hashability.UNHASHABLE: 'unhashable'>, added_match_args=True, added_str=False, added_pickling=True, on_setattr_hook=<function pipe.<locals>.wrapped_pipe at 0x7f41e68fca40>, field_transformer=None) class-attribute

Effective class properties as derived from parameters to attr.s() or define() decorators.

This is the same data structure that attrs uses internally to decide how to construct the final class.

Warning:

This feature is currently **experimental** and is not covered by our
strict backwards-compatibility guarantees.

Attributes:

Name Type Description
is_exception bool

Whether the class is treated as an exception class.

is_slotted bool

Whether the class is slotted <slotted classes>.

has_weakref_slot bool

Whether the class has a slot for weak references.

is_frozen bool

Whether the class is frozen.

kw_only KeywordOnly

Whether / how the class enforces keyword-only arguments on the __init__ method.

collected_fields_by_mro bool

Whether the class fields were collected by method resolution order. That is, correctly but unlike dataclasses.

added_init bool

Whether the class has an attrs-generated __init__ method.

added_repr bool

Whether the class has an attrs-generated __repr__ method.

added_eq bool

Whether the class has attrs-generated equality methods.

added_ordering bool

Whether the class has attrs-generated ordering methods.

hashability Hashability

How hashable <hashing> the class is.

added_match_args bool

Whether the class supports positional match <match> over its fields.

added_str bool

Whether the class has an attrs-generated __str__ method.

added_pickling bool

Whether the class has attrs-generated __getstate__ and __setstate__ methods for pickle.

on_setattr_hook Callable[[Any, Attribute[Any], Any], Any] | None

The class's __setattr__ hook.

field_transformer Callable[[Attribute[Any]], Attribute[Any]] | None

The class's field transformers <transform-fields>.

.. versionadded:: 25.4.0

__doc__ = 'Video backend for reading videos stored as common media files.\n\nThis backend supports reading through FFMPEG (the default), pyav, or OpenCV. Here\nare their trade-offs:\n\n - "opencv": Fastest video reader, but only supports a limited number of codecs\n and may not be able to read some videos. It requires `opencv-python` to be\n installed. It is the fastest because it uses the OpenCV C++ library to read\n videos, but is limited by the version of FFMPEG that was linked into it at\n build time as well as the OpenCV version used.\n - "FFMPEG": Slowest, but most reliable. This is the default backend. It requires\n `imageio-ffmpeg` and a `ffmpeg` executable on the system path (which can be\n installed via conda). The `imageio` plugin for FFMPEG reads frames into raw\n bytes which are communicated to Python through STDOUT on a subprocess pipe,\n which can be slow. However, it is the most reliable and feature-complete. If\n you install the conda-forge version of ffmpeg, it will be compiled with\n support for many codecs, including GPU-accelerated codecs like NVDEC for\n H264 and others.\n - "pyav": Supports most codecs that FFMPEG does, but not as complete or reliable\n of an implementation in `imageio` as FFMPEG for some video types. It is\n faster than FFMPEG because it uses the `av` package to read frames directly\n into numpy arrays in memory without the need for a subprocess pipe. These\n are Python bindings for the C library libav, which is the same library that\n FFMPEG uses under the hood.\n\nAttributes:\n filename: Path to video file.\n grayscale: Whether to force grayscale. If None, autodetect on first frame load.\n keep_open: Whether to keep the video reader open between calls to read frames.\n If False, will close the reader after each call. If True (the default), it\n will keep the reader open and cache it for subsequent calls which may\n enhance the performance of reading multiple frames.\n plugin: Video plugin to use. One of "opencv", "FFMPEG", or "pyav". If `None`,\n will use the first available plugin in the order listed above.\n' class-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

__firstlineno__ = 621 class-attribute

int([x]) -> integer int(x, base=10) -> integer

Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.int(). For floating-point numbers, this truncates towards zero.

If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by '+' or '-' and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal.

int('0b100', base=0) 4

__match_args__ = ('filename', 'grayscale', 'keep_open', '_cached_shape', '_open_reader', '_fps', 'plugin') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__module__ = 'sleap_io.io.video_reading' class-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

__slots__ = ('plugin',) class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__static_attributes__ = ('_fps', '_open_reader') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

fps property

Frames per second from video container metadata.

Returns:

Type Description

The FPS from the video container, or None if it cannot be determined.

Notes

This reads the FPS from the video file metadata using the appropriate method for the current plugin: - OpenCV: cv2.CAP_PROP_FPS - FFMPEG/pyav: imageio metadata

num_frames property

Number of frames in the video.

reader property

Return the reader object for the video, caching if necessary.

__eq__(other)

Method generated by attrs for class MediaVideo.

Source code in sleap_io/io/video_reading.py
    import cv2
except ImportError:
    pass

try:
    import imageio_ffmpeg  # noqa: F401
except ImportError:
    pass

try:
    import av  # noqa: F401
except ImportError:

__init__(filename, grayscale=None, keep_open=True, cached_shape=None, open_reader=None, fps=None, plugin=NOTHING)

Method generated by attrs for class MediaVideo.

Source code in sleap_io/io/video_reading.py
    pass


# Track available backends (populated on module import)
_AVAILABLE_VIDEO_BACKENDS = {
    "opencv": "cv2" in sys.modules,
    "FFMPEG": "imageio_ffmpeg" in sys.modules,
    "pyav": "av" in sys.modules,
}

_AVAILABLE_IMAGE_BACKENDS = {
    "opencv": "cv2" in sys.modules,
    "imageio": True,  # Always available (core dependency)
}

__replace__(**changes)

Method generated by attrs for class MediaVideo.

Source code in sleap_io/io/video_reading.py
Returns:

__repr__()

Method generated by attrs for class MediaVideo.

Source code in sleap_io/io/video_reading.py
"""Backends for reading videos."""

from __future__ import annotations

import sys
from io import BytesIO
from pathlib import Path

import attrs
import h5py
import imageio.v3 as iio
import numpy as np
import simplejson as json

try:

__setattr__(name, val)

Method generated by attrs for class MediaVideo.

Source code in sleap_io/io/video_reading.py
    f = self._open_reader
else:
    f = h5py.File(self.filename, "r")

ds = f[self.dataset]
raw_bytes = ds[internal_idx]

# Handle fixed-length padding (strip trailing zeros)
is_vlen = h5py.check_vlen_dtype(ds.dtype) is not None

Video

Video class used by sleap to represent videos and data associated with them.

This class is used to store information regarding a video and its components. It is used to store the video's filename, shape, and the video's backend.

To create a Video object, use the from_filename method which will select the backend appropriately.

Attributes:

Name Type Description
filename

The filename(s) of the video. Supported extensions: "mp4", "avi", "mov", "mj2", "mkv", "h5", "hdf5", "slp", "png", "jpg", "jpeg", "tif", "tiff", "bmp". If the filename is a list, a list of image filenames are expected. If filename is a folder, it will be searched for images.

backend

An object that implements the basic methods for reading and manipulating frames of a specific video type.

backend_metadata

A dictionary of metadata specific to the backend. This is useful for storing metadata that requires an open backend (e.g., shape information) without having access to the video file itself.

source_video

The source video object if this is a proxy video. This is present when the video contains an embedded subset of frames from another video.

open_backend

Whether to open the backend when the video is available. If True (the default), the backend will be automatically opened if the video exists. Set this to False when you want to manually open the backend, or when the you know the video file does not exist and you want to avoid trying to open the file.

Notes

Instances of this class are hashed by identity, not by value. This means that two Video instances with the same attributes will NOT be considered equal in a set or dict.

Media Video Plugin Support

For media files (mp4, avi, etc.), the following plugins are supported: - "opencv": Uses OpenCV (cv2) for video reading - "FFMPEG": Uses imageio-ffmpeg for video reading - "pyav": Uses PyAV for video reading

Plugin aliases (case-insensitive): - opencv: "opencv", "cv", "cv2", "ocv" - FFMPEG: "FFMPEG", "ffmpeg", "imageio-ffmpeg", "imageio_ffmpeg" - pyav: "pyav", "av"

Plugin selection priority: 1. Explicitly specified plugin parameter 2. Backend metadata plugin value 3. Global default (set via sio.set_default_video_plugin) 4. Auto-detection based on available packages

See Also

VideoBackend: The backend interface for reading video data. sleap_io.set_default_video_plugin: Set global default plugin. sleap_io.get_default_video_plugin: Get current default plugin.

Methods:

Name Description
__attrs_post_init__

Post init syntactic sugar.

__deepcopy__

Deep copy the video object.

__getitem__

Return the frames of the video at the given indices.

__init__

Method generated by attrs for class Video.

__len__

Return the length of the video as the number of frames.

__replace__

Method generated by attrs for class Video.

__repr__

Informal string representation (for print or format).

__str__

Informal string representation (for print or format).

close

Close the video backend.

deduplicate_with

Create a new video with duplicate images removed.

exists

Check if the video file exists and is accessible.

frame_to_seconds

Convert a frame index to timestamp in seconds.

from_filename

Create a Video from a filename.

has_overlapping_images

Check if this video has overlapping images with another video.

matches_content

Check if this video has the same content as another video.

matches_path

Check if this video has the same path as another video.

matches_shape

Check if this video has the same shape as another video.

merge_with

Merge another video's images into this one.

open

Open the video backend for reading.

replace_filename

Update the filename of the video, optionally opening the backend.

save

Save video frames to a new video file.

seconds_to_frame

Convert a timestamp in seconds to frame index.

set_video_plugin

Set the video plugin and reopen the video.

Source code in sleap_io/model/video.py
@attrs.define(eq=False)
class Video:
    """`Video` class used by sleap to represent videos and data associated with them.

    This class is used to store information regarding a video and its components.
    It is used to store the video's `filename`, `shape`, and the video's `backend`.

    To create a `Video` object, use the `from_filename` method which will select the
    backend appropriately.

    Attributes:
        filename: The filename(s) of the video. Supported extensions: "mp4", "avi",
            "mov", "mj2", "mkv", "h5", "hdf5", "slp", "png", "jpg", "jpeg", "tif",
            "tiff", "bmp". If the filename is a list, a list of image filenames are
            expected. If filename is a folder, it will be searched for images.
        backend: An object that implements the basic methods for reading and
            manipulating frames of a specific video type.
        backend_metadata: A dictionary of metadata specific to the backend. This is
            useful for storing metadata that requires an open backend (e.g., shape
            information) without having access to the video file itself.
        source_video: The source video object if this is a proxy video. This is present
            when the video contains an embedded subset of frames from another video.
        open_backend: Whether to open the backend when the video is available. If `True`
            (the default), the backend will be automatically opened if the video exists.
            Set this to `False` when you want to manually open the backend, or when the
            you know the video file does not exist and you want to avoid trying to open
            the file.

    Notes:
        Instances of this class are hashed by identity, not by value. This means that
        two `Video` instances with the same attributes will NOT be considered equal in a
        set or dict.

    Media Video Plugin Support:
        For media files (mp4, avi, etc.), the following plugins are supported:
        - "opencv": Uses OpenCV (cv2) for video reading
        - "FFMPEG": Uses imageio-ffmpeg for video reading
        - "pyav": Uses PyAV for video reading

        Plugin aliases (case-insensitive):
        - opencv: "opencv", "cv", "cv2", "ocv"
        - FFMPEG: "FFMPEG", "ffmpeg", "imageio-ffmpeg", "imageio_ffmpeg"
        - pyav: "pyav", "av"

        Plugin selection priority:
        1. Explicitly specified plugin parameter
        2. Backend metadata plugin value
        3. Global default (set via sio.set_default_video_plugin)
        4. Auto-detection based on available packages

    See Also:
        VideoBackend: The backend interface for reading video data.
        sleap_io.set_default_video_plugin: Set global default plugin.
        sleap_io.get_default_video_plugin: Get current default plugin.
    """

    filename: str | list[str]
    backend: VideoBackend | None = None
    backend_metadata: dict[str, any] = attrs.field(factory=dict)
    source_video: "Video | None" = None
    open_backend: bool = True

    EXTS = MediaVideo.EXTS + HDF5Video.EXTS + ImageVideo.EXTS

    @property
    def original_video(self) -> "Video | None":
        """The root video in the provenance chain.

        For embedded videos, this returns the ultimate source video by
        traversing the source_video chain. Returns None if this video
        has no source_video (i.e., it IS an original).

        This property is computed by following the source_video chain to find
        the root. For a single-level embedding (A embeds from B), original_video
        returns B. For multi-level embedding (A <- B <- C), it returns C.
        """
        if self.source_video is None:
            return None  # This IS the original

        # Traverse to root
        v = self.source_video
        while v.source_video is not None:
            v = v.source_video
        return v

    def __attrs_post_init__(self):
        """Post init syntactic sugar."""
        if self.open_backend and self.backend is None and self.exists():
            try:
                self.open()
            except Exception:
                # If we can't open the backend, just ignore it for now so we don't
                # prevent the user from building the Video object entirely.
                pass

    def __deepcopy__(self, memo):
        """Deep copy the video object."""
        if id(self) in memo:
            return memo[id(self)]

        reopen = False
        if self.is_open:
            reopen = True
            self.close()

        new_video = Video(
            filename=self.filename,
            backend=None,
            backend_metadata=self.backend_metadata.copy(),
            source_video=self.source_video,
            open_backend=self.open_backend,
        )

        memo[id(self)] = new_video

        if reopen:
            self.open()

        return new_video

    @classmethod
    def from_filename(
        cls,
        filename: str | list[str],
        dataset: str | None = None,
        grayscale: bool | None = None,
        keep_open: bool = True,
        source_video: "Video | None" = None,
        **kwargs,
    ) -> VideoBackend:
        """Create a Video from a filename.

        Args:
            filename: The filename(s) of the video. Supported extensions: "mp4", "avi",
                "mov", "mj2", "mkv", "h5", "hdf5", "slp", "png", "jpg", "jpeg", "tif",
                "tiff", "bmp". If the filename is a list, a list of image filenames are
                expected. If filename is a folder, it will be searched for images.
            dataset: Name of dataset in HDF5 file.
            grayscale: Whether to force grayscale. If None, autodetect on first frame
                load.
            keep_open: Whether to keep the video reader open between calls to read
                frames. If False, will close the reader after each call. If True (the
                default), it will keep the reader open and cache it for subsequent calls
                which may enhance the performance of reading multiple frames.
            source_video: The source video object if this is a proxy video. This is
                present when the video contains an embedded subset of frames from
                another video.
            **kwargs: Additional backend-specific arguments passed to
                VideoBackend.from_filename. See VideoBackend.from_filename for supported
                arguments.

        Returns:
            Video instance with the appropriate backend instantiated.
        """
        backend = VideoBackend.from_filename(
            filename,
            dataset=dataset,
            grayscale=grayscale,
            keep_open=keep_open,
            **kwargs,
        )
        # If filename is a directory, VideoBackend.from_filename will expand it
        # to a list of paths to images contained within the directory. In this
        # case we want to use the expanded list as filename
        return cls(
            filename=backend.filename,
            backend=backend,
            source_video=source_video,
        )

    @property
    def shape(self) -> tuple[int, int, int, int] | None:
        """Return the shape of the video as (num_frames, height, width, channels).

        If the video backend is not set or it cannot determine the shape of the video,
        this will return None.
        """
        return self._get_shape()

    def _get_shape(self) -> tuple[int, int, int, int] | None:
        """Return the shape of the video as (num_frames, height, width, channels).

        This suppresses errors related to querying the backend for the video shape, such
        as when it has not been set or when the video file is not found.
        """
        try:
            return self.backend.shape
        except Exception:
            if "shape" in self.backend_metadata:
                return self.backend_metadata["shape"]
            return None

    @property
    def grayscale(self) -> bool | None:
        """Return whether the video is grayscale.

        If the video backend is not set or it cannot determine whether the video is
        grayscale, this will return None.
        """
        shape = self.shape
        if shape is not None:
            return shape[-1] == 1
        else:
            grayscale = None
            if "grayscale" in self.backend_metadata:
                grayscale = self.backend_metadata["grayscale"]
            return grayscale

    @grayscale.setter
    def grayscale(self, value: bool):
        """Set the grayscale value and adjust the backend."""
        if self.backend is not None:
            self.backend.grayscale = value
            self.backend._cached_shape = None

        self.backend_metadata["grayscale"] = value

    @property
    def fps(self) -> float | None:
        """Return the frames per second of the video.

        For MediaVideo backends, this reads FPS from the video container metadata.
        For other backends (ImageVideo, HDF5Video, TiffVideo), this returns the
        explicitly set value or None if not set.

        Returns:
            The FPS if known, or None if unavailable/unknown.
        """
        if self.backend is not None:
            return self.backend.fps
        return self.backend_metadata.get("fps")

    @fps.setter
    def fps(self, value: float | None):
        """Set the frames per second.

        Args:
            value: Frames per second. Must be positive if not None.

        Raises:
            ValueError: If value is not positive.

        Notes:
            For MediaVideo backends, setting FPS overrides the value from container
            metadata. For other backends, this sets the FPS directly.
        """
        if value is not None and value <= 0:
            raise ValueError(f"FPS must be positive, got {value}")

        if self.backend is not None:
            self.backend.fps = value
        self.backend_metadata["fps"] = value

    def frame_to_seconds(self, frame_idx: int) -> float | None:
        """Convert a frame index to timestamp in seconds.

        Args:
            frame_idx: Zero-indexed frame number.

        Returns:
            Time in seconds, or None if FPS is unknown.

        Notes:
            This assumes constant frame rate. For variable frame rate videos,
            the returned timestamp may be approximate.
        """
        if self.fps is None or self.fps <= 0:
            return None
        return frame_idx / self.fps

    def seconds_to_frame(self, seconds: float) -> int | None:
        """Convert a timestamp in seconds to frame index.

        Args:
            seconds: Time in seconds from video start.

        Returns:
            Zero-indexed frame number (rounded down), or None if FPS unknown.
        """
        if self.fps is None or self.fps <= 0:
            return None
        return int(seconds * self.fps)

    def __len__(self) -> int:
        """Return the length of the video as the number of frames."""
        shape = self.shape
        return 0 if shape is None else shape[0]

    def __repr__(self) -> str:
        """Informal string representation (for print or format)."""
        dataset = (
            f"dataset={self.backend.dataset}, "
            if getattr(self.backend, "dataset", "")
            else ""
        )
        return (
            "Video("
            f'filename="{self.filename}", '
            f"shape={self.shape}, "
            f"{dataset}"
            f"backend={type(self.backend).__name__}"
            ")"
        )

    def __str__(self) -> str:
        """Informal string representation (for print or format)."""
        return self.__repr__()

    def __getitem__(self, inds: int | list[int] | slice) -> np.ndarray:
        """Return the frames of the video at the given indices.

        Args:
            inds: Index or list of indices of frames to read.

        Returns:
            Frame or frames as a numpy array of shape `(height, width, channels)` if a
            scalar index is provided, or `(frames, height, width, channels)` if a list
            of indices is provided.

        See also: VideoBackend.get_frame, VideoBackend.get_frames
        """
        if not self.is_open:
            if self.open_backend:
                self.open()
            else:
                raise ValueError(
                    "Video backend is not open. Call video.open() or set "
                    "video.open_backend to True to do automatically on frame read."
                )
        return self.backend[inds]

    def exists(self, check_all: bool = False, dataset: str | None = None) -> bool:
        """Check if the video file exists and is accessible.

        Args:
            check_all: If `True`, check that all filenames in a list exist. If `False`
                (the default), check that the first filename exists.
            dataset: Name of dataset in HDF5 file. If specified, this will function will
                return `False` if the dataset does not exist.

        Returns:
            `True` if the file exists and is accessible, `False` otherwise.
        """
        if isinstance(self.filename, list):
            if check_all:
                for f in self.filename:
                    if not is_file_accessible(f):
                        return False
                return True
            else:
                return is_file_accessible(self.filename[0])

        file_is_accessible = is_file_accessible(self.filename)
        if not file_is_accessible:
            return False

        if dataset is None or dataset == "":
            dataset = self.backend_metadata.get("dataset", None)

        if dataset is not None and dataset != "":
            has_dataset = False
            if (
                self.backend is not None
                and type(self.backend) is HDF5Video
                and self.backend._open_reader is not None
            ):
                has_dataset = dataset in self.backend._open_reader
            else:
                with h5py.File(self.filename, "r") as f:
                    has_dataset = dataset in f
            return has_dataset

        return True

    @property
    def is_open(self) -> bool:
        """Check if the video backend is open."""
        return self.exists() and self.backend is not None

    def open(
        self,
        filename: str | None = None,
        dataset: str | None = None,
        grayscale: str | None = None,
        keep_open: bool = True,
        plugin: str | None = None,
    ):
        """Open the video backend for reading.

        Args:
            filename: Filename to open. If not specified, will use the filename set on
                the video object.
            dataset: Name of dataset in HDF5 file.
            grayscale: Whether to force grayscale. If None, autodetect on first frame
                load.
            keep_open: Whether to keep the video reader open between calls to read
                frames. If False, will close the reader after each call. If True (the
                default), it will keep the reader open and cache it for subsequent calls
                which may enhance the performance of reading multiple frames.
            plugin: Video plugin to use for MediaVideo files. One of "opencv",
                "FFMPEG", or "pyav". Also accepts aliases (case-insensitive).
                If not specified, uses the backend metadata, global default,
                or auto-detection in that order.

        Notes:
            This is useful for opening the video backend to read frames and then closing
            it after reading all the necessary frames.

            If the backend was already open, it will be closed before opening a new one.
            Values for the HDF5 dataset and grayscale will be remembered if not
            specified.
        """
        if filename is not None:
            self.replace_filename(filename, open=False)

        # Try to remember values from previous backend if available and not specified.
        if self.backend is not None:
            if dataset is None:
                dataset = getattr(self.backend, "dataset", None)
            if grayscale is None:
                grayscale = getattr(self.backend, "grayscale", None)

        else:
            if dataset is None and "dataset" in self.backend_metadata:
                dataset = self.backend_metadata["dataset"]
            if grayscale is None:
                if "grayscale" in self.backend_metadata:
                    grayscale = self.backend_metadata["grayscale"]
                elif "shape" in self.backend_metadata:
                    grayscale = self.backend_metadata["shape"][-1] == 1

        if not self.exists(dataset=dataset):
            msg = (
                f"Video does not exist or cannot be opened for reading: {self.filename}"
            )
            if dataset is not None:
                msg += f" (dataset: {dataset})"
            raise FileNotFoundError(msg)

        # Close previous backend if open.
        self.close()

        # Handle plugin parameter
        backend_kwargs = {}
        if plugin is not None:
            from sleap_io.io.video_reading import normalize_plugin_name

            plugin = normalize_plugin_name(plugin)
            self.backend_metadata["plugin"] = plugin

        if "plugin" in self.backend_metadata:
            backend_kwargs["plugin"] = self.backend_metadata["plugin"]

        # Create new backend.
        self.backend = VideoBackend.from_filename(
            self.filename,
            dataset=dataset,
            grayscale=grayscale,
            keep_open=keep_open,
            **backend_kwargs,
        )

    def close(self):
        """Close the video backend."""
        if self.backend is not None:
            # Try to remember values from previous backend if available and not
            # specified.
            try:
                self.backend_metadata["dataset"] = getattr(
                    self.backend, "dataset", None
                )
                self.backend_metadata["grayscale"] = getattr(
                    self.backend, "grayscale", None
                )
                self.backend_metadata["shape"] = getattr(self.backend, "shape", None)
                self.backend_metadata["fps"] = getattr(self.backend, "fps", None)
            except Exception:
                pass

            del self.backend
            self.backend = None

    def replace_filename(
        self, new_filename: str | Path | list[str] | list[Path], open: bool = True
    ):
        """Update the filename of the video, optionally opening the backend.

        Args:
            new_filename: New filename to set for the video.
            open: If `True` (the default), open the backend with the new filename. If
                the new filename does not exist, no error is raised.
        """
        if isinstance(new_filename, Path):
            new_filename = new_filename.as_posix()

        if isinstance(new_filename, list):
            new_filename = [
                p.as_posix() if isinstance(p, Path) else p for p in new_filename
            ]

        self.filename = new_filename
        self.backend_metadata["filename"] = new_filename

        if open:
            if self.exists():
                self.open()
            else:
                self.close()

    def matches_path(self, other: "Video", strict: bool = False) -> bool:
        """Check if this video has the same path as another video.

        Args:
            other: Another video to compare with.
            strict: If True, require exact path match. If False, consider videos
                with the same filename (basename) as matching.

        Returns:
            True if the videos have matching paths, False otherwise.

        Notes:
            For HDF5 video backends (e.g., embedded videos in .pkg.slp files),
            matching prioritizes the source_filename attribute since multiple
            videos can share the same HDF5 file path but reference different
            source videos. Falls back to dataset name matching if source_filename
            is not available.
        """
        # Handle HDF5 backends specially - prioritize source_filename matching
        self_is_hdf5 = isinstance(self.backend, HDF5Video)
        other_is_hdf5 = isinstance(other.backend, HDF5Video)

        if self_is_hdf5 and other_is_hdf5:
            # Both are HDF5 videos - match by source_filename first
            self_source = self.backend.source_filename
            other_source = other.backend.source_filename

            if self_source is not None and other_source is not None:
                if strict:
                    return Path(self_source).resolve() == Path(other_source).resolve()
                else:
                    return Path(self_source).name == Path(other_source).name

            # Fall back to dataset name matching if source_filename is not available
            self_dataset = self.backend.dataset
            other_dataset = other.backend.dataset

            if self_dataset is not None and other_dataset is not None:
                return self_dataset == other_dataset

            # If neither source_filename nor dataset available, cannot match
            return False

        if isinstance(self.filename, list) and isinstance(other.filename, list):
            # Both are image sequences
            if strict:
                return self.filename == other.filename
            else:
                # Compare basenames
                self_basenames = [Path(f).name for f in self.filename]
                other_basenames = [Path(f).name for f in other.filename]
                return self_basenames == other_basenames
        elif isinstance(self.filename, list) or isinstance(other.filename, list):
            # One is image sequence, other is single file
            return False
        else:
            # Both are single files
            if strict:
                return Path(self.filename).resolve() == Path(other.filename).resolve()
            else:
                return Path(self.filename).name == Path(other.filename).name

    def matches_content(self, other: "Video") -> bool:
        """Check if this video has the same content as another video.

        Args:
            other: Another video to compare with.

        Returns:
            True if the videos have the same shape and backend type.

        Notes:
            This compares metadata like shape and backend type, not actual frame data.
        """
        # Compare shapes
        self_shape = self.shape
        other_shape = other.shape

        if self_shape != other_shape:
            return False

        # Compare backend types
        if self.backend is None and other.backend is None:
            return True
        elif self.backend is None or other.backend is None:
            return False

        return type(self.backend).__name__ == type(other.backend).__name__

    def matches_shape(self, other: "Video") -> bool:
        """Check if this video has the same shape as another video.

        Args:
            other: Another video to compare with.

        Returns:
            True if the videos have the same height, width, and channels.

        Notes:
            This only compares spatial dimensions, not the number of frames.
        """
        # Try to get shape from backend metadata first if shape is not available
        if self.backend is None and "shape" in self.backend_metadata:
            self_shape = self.backend_metadata["shape"]
        else:
            self_shape = self.shape

        if other.backend is None and "shape" in other.backend_metadata:
            other_shape = other.backend_metadata["shape"]
        else:
            other_shape = other.shape

        # Handle None shapes
        if self_shape is None or other_shape is None:
            return False

        # Compare only height, width, channels (not frames)
        return self_shape[1:] == other_shape[1:]

    def has_overlapping_images(self, other: "Video") -> bool:
        """Check if this video has overlapping images with another video.

        This method is specifically for ImageVideo backends (image sequences).

        Args:
            other: Another video to compare with.

        Returns:
            True if both are ImageVideo instances with overlapping image files.
            False if either video is not an ImageVideo or no overlap exists.

        Notes:
            Only works with ImageVideo backends where filename is a list.
            Compares individual image filenames (basenames only).
        """
        # Both must be image sequences
        if not (isinstance(self.filename, list) and isinstance(other.filename, list)):
            return False

        # Get basenames for comparison
        self_basenames = set(Path(f).name for f in self.filename)
        other_basenames = set(Path(f).name for f in other.filename)

        # Check if there's any overlap
        return len(self_basenames & other_basenames) > 0

    def deduplicate_with(self, other: "Video") -> "Video":
        """Create a new video with duplicate images removed.

        This method is specifically for ImageVideo backends (image sequences).

        Args:
            other: Another video to deduplicate against. Must also be ImageVideo.

        Returns:
            A new Video object with duplicate images removed from this video,
            or None if all images were duplicates.

        Raises:
            ValueError: If either video is not an ImageVideo backend.

        Notes:
            Only works with ImageVideo backends where filename is a list.
            Images are considered duplicates if they have the same basename.
            The returned video contains only images from this video that are
            not present in the other video.
        """
        if not isinstance(self.filename, list):
            raise ValueError("deduplicate_with only works with ImageVideo backends")
        if not isinstance(other.filename, list):
            raise ValueError("Other video must also be ImageVideo backend")

        # Get basenames from other video
        other_basenames = set(Path(f).name for f in other.filename)

        # Keep only non-duplicate images
        deduplicated_paths = [
            f for f in self.filename if Path(f).name not in other_basenames
        ]

        if not deduplicated_paths:
            # All images were duplicates
            return None

        # Create new video with deduplicated images
        return Video.from_filename(deduplicated_paths, grayscale=self.grayscale)

    def merge_with(self, other: "Video") -> "Video":
        """Merge another video's images into this one.

        This method is specifically for ImageVideo backends (image sequences).

        Args:
            other: Another video to merge with. Must also be ImageVideo.

        Returns:
            A new Video object with unique images from both videos.

        Raises:
            ValueError: If either video is not an ImageVideo backend.

        Notes:
            Only works with ImageVideo backends where filename is a list.
            The merged video contains all unique images from both videos,
            with automatic deduplication based on image basename.
        """
        if not isinstance(self.filename, list):
            raise ValueError("merge_with only works with ImageVideo backends")
        if not isinstance(other.filename, list):
            raise ValueError("Other video must also be ImageVideo backend")

        # Get all unique images (by basename) preserving order
        seen_basenames = set()
        merged_paths = []

        for path in self.filename:
            basename = Path(path).name
            if basename not in seen_basenames:
                merged_paths.append(path)
                seen_basenames.add(basename)

        for path in other.filename:
            basename = Path(path).name
            if basename not in seen_basenames:
                merged_paths.append(path)
                seen_basenames.add(basename)

        # Create new video with merged images
        return Video.from_filename(merged_paths, grayscale=self.grayscale)

    def save(
        self,
        save_path: str | Path,
        frame_inds: list[int] | np.ndarray | None = None,
        fps: float | None = None,
        video_kwargs: dict[str, Any] | None = None,
    ) -> "Video":
        """Save video frames to a new video file.

        Args:
            save_path: Path to the new video file. Should end in MP4.
            frame_inds: Frame indices to save. Can be specified as a list or array of
                frame integers. If not specified, saves all video frames.
            fps: Frames per second for the output video. If not specified, uses the
                source video's FPS if available, otherwise defaults to 30.
            video_kwargs: A dictionary of keyword arguments to provide to
                `sio.save_video` for video compression.

        Returns:
            A new `Video` object pointing to the new video file.
        """
        video_kwargs = {} if video_kwargs is None else video_kwargs.copy()
        frame_inds = np.arange(len(self)) if frame_inds is None else frame_inds

        # Use source video FPS if not explicitly specified
        if fps is None:
            fps = self.fps
        if fps is not None and "fps" not in video_kwargs:
            video_kwargs["fps"] = fps

        with VideoWriter(save_path, **video_kwargs) as vw:
            for frame_ind in frame_inds:
                vw(self[frame_ind])

        new_video = Video.from_filename(save_path, grayscale=self.grayscale)
        return new_video

    def set_video_plugin(self, plugin: str) -> None:
        """Set the video plugin and reopen the video.

        Args:
            plugin: Video plugin to use. One of "opencv", "FFMPEG", or "pyav".
                Also accepts aliases (case-insensitive).

        Raises:
            ValueError: If the video is not a MediaVideo type.

        Examples:
            >>> video.set_video_plugin("opencv")
            >>> video.set_video_plugin("CV2")  # Same as "opencv"
        """
        from sleap_io.io.video_reading import MediaVideo, normalize_plugin_name

        if not self.filename.endswith(MediaVideo.EXTS):
            raise ValueError(f"Cannot set plugin for non-media video: {self.filename}")

        plugin = normalize_plugin_name(plugin)

        # Close current backend if open
        was_open = self.is_open
        if was_open:
            self.close()

        # Update backend metadata
        self.backend_metadata["plugin"] = plugin

        # Reopen with new plugin if it was open
        if was_open:
            self.open()

EXTS = ('mp4', 'avi', 'mov', 'mj2', 'mkv', 'h5', 'hdf5', 'slp', 'png', 'jpg', 'jpeg', 'tif', 'tiff', 'bmp') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__annotations__ = {'filename': 'str | list[str]', 'backend': 'VideoBackend | None', 'backend_metadata': 'dict[str, any]', 'source_video': "'Video | None'", 'open_backend': 'bool'} class-attribute

dict() -> new empty dictionary dict(mapping) -> new dictionary initialized from a mapping object's (key, value) pairs dict(iterable) -> new dictionary initialized as if via: d = {} for k, v in iterable: d[k] = v dict(**kwargs) -> new dictionary initialized with the name=value pairs in the keyword argument list. For example: dict(one=1, two=2)

__attrs_own_setattr__ = False class-attribute

Returns True when the argument is true, False otherwise. The builtins True and False are the only two instances of the class bool. The class bool is a subclass of the class int, and cannot be subclassed.

__attrs_props__ = ClassProps(is_exception=False, is_slotted=True, has_weakref_slot=True, is_frozen=False, kw_only=<KeywordOnly.NO: 'no'>, collected_fields_by_mro=True, added_init=True, added_repr=False, added_eq=False, added_ordering=False, hashability=<Hashability.LEAVE_ALONE: 'leave_alone'>, added_match_args=True, added_str=False, added_pickling=True, on_setattr_hook=<function pipe.<locals>.wrapped_pipe at 0x7f41e68fca40>, field_transformer=None) class-attribute

Effective class properties as derived from parameters to attr.s() or define() decorators.

This is the same data structure that attrs uses internally to decide how to construct the final class.

Warning:

This feature is currently **experimental** and is not covered by our
strict backwards-compatibility guarantees.

Attributes:

Name Type Description
is_exception bool

Whether the class is treated as an exception class.

is_slotted bool

Whether the class is slotted <slotted classes>.

has_weakref_slot bool

Whether the class has a slot for weak references.

is_frozen bool

Whether the class is frozen.

kw_only KeywordOnly

Whether / how the class enforces keyword-only arguments on the __init__ method.

collected_fields_by_mro bool

Whether the class fields were collected by method resolution order. That is, correctly but unlike dataclasses.

added_init bool

Whether the class has an attrs-generated __init__ method.

added_repr bool

Whether the class has an attrs-generated __repr__ method.

added_eq bool

Whether the class has attrs-generated equality methods.

added_ordering bool

Whether the class has attrs-generated ordering methods.

hashability Hashability

How hashable <hashing> the class is.

added_match_args bool

Whether the class supports positional match <match> over its fields.

added_str bool

Whether the class has an attrs-generated __str__ method.

added_pickling bool

Whether the class has attrs-generated __getstate__ and __setstate__ methods for pickle.

on_setattr_hook Callable[[Any, Attribute[Any], Any], Any] | None

The class's __setattr__ hook.

field_transformer Callable[[Attribute[Any]], Attribute[Any]] | None

The class's field transformers <transform-fields>.

.. versionadded:: 25.4.0

__doc__ = '`Video` class used by sleap to represent videos and data associated with them.\n\nThis class is used to store information regarding a video and its components.\nIt is used to store the video\'s `filename`, `shape`, and the video\'s `backend`.\n\nTo create a `Video` object, use the `from_filename` method which will select the\nbackend appropriately.\n\nAttributes:\n filename: The filename(s) of the video. Supported extensions: "mp4", "avi",\n "mov", "mj2", "mkv", "h5", "hdf5", "slp", "png", "jpg", "jpeg", "tif",\n "tiff", "bmp". If the filename is a list, a list of image filenames are\n expected. If filename is a folder, it will be searched for images.\n backend: An object that implements the basic methods for reading and\n manipulating frames of a specific video type.\n backend_metadata: A dictionary of metadata specific to the backend. This is\n useful for storing metadata that requires an open backend (e.g., shape\n information) without having access to the video file itself.\n source_video: The source video object if this is a proxy video. This is present\n when the video contains an embedded subset of frames from another video.\n open_backend: Whether to open the backend when the video is available. If `True`\n (the default), the backend will be automatically opened if the video exists.\n Set this to `False` when you want to manually open the backend, or when the\n you know the video file does not exist and you want to avoid trying to open\n the file.\n\nNotes:\n Instances of this class are hashed by identity, not by value. This means that\n two `Video` instances with the same attributes will NOT be considered equal in a\n set or dict.\n\nMedia Video Plugin Support:\n For media files (mp4, avi, etc.), the following plugins are supported:\n - "opencv": Uses OpenCV (cv2) for video reading\n - "FFMPEG": Uses imageio-ffmpeg for video reading\n - "pyav": Uses PyAV for video reading\n\n Plugin aliases (case-insensitive):\n - opencv: "opencv", "cv", "cv2", "ocv"\n - FFMPEG: "FFMPEG", "ffmpeg", "imageio-ffmpeg", "imageio_ffmpeg"\n - pyav: "pyav", "av"\n\n Plugin selection priority:\n 1. Explicitly specified plugin parameter\n 2. Backend metadata plugin value\n 3. Global default (set via sio.set_default_video_plugin)\n 4. Auto-detection based on available packages\n\nSee Also:\n VideoBackend: The backend interface for reading video data.\n sleap_io.set_default_video_plugin: Set global default plugin.\n sleap_io.get_default_video_plugin: Get current default plugin.\n' class-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

__firstlineno__ = 21 class-attribute

int([x]) -> integer int(x, base=10) -> integer

Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.int(). For floating-point numbers, this truncates towards zero.

If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by '+' or '-' and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal.

int('0b100', base=0) 4

__match_args__ = ('filename', 'backend', 'backend_metadata', 'source_video', 'open_backend') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__module__ = 'sleap_io.model.video' class-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

__slots__ = ('filename', 'backend', 'backend_metadata', 'source_video', 'open_backend', '__weakref__') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__static_attributes__ = ('backend', 'filename') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__weakref__ property

list of weak references to the object

fps property

Return the frames per second of the video.

For MediaVideo backends, this reads FPS from the video container metadata. For other backends (ImageVideo, HDF5Video, TiffVideo), this returns the explicitly set value or None if not set.

Returns:

Type Description

The FPS if known, or None if unavailable/unknown.

grayscale property

Return whether the video is grayscale.

If the video backend is not set or it cannot determine whether the video is grayscale, this will return None.

is_open property

Check if the video backend is open.

original_video property

The root video in the provenance chain.

For embedded videos, this returns the ultimate source video by traversing the source_video chain. Returns None if this video has no source_video (i.e., it IS an original).

This property is computed by following the source_video chain to find the root. For a single-level embedding (A embeds from B), original_video returns B. For multi-level embedding (A <- B <- C), it returns C.

shape property

Return the shape of the video as (num_frames, height, width, channels).

If the video backend is not set or it cannot determine the shape of the video, this will return None.

__attrs_post_init__()

Post init syntactic sugar.

Source code in sleap_io/model/video.py
def __attrs_post_init__(self):
    """Post init syntactic sugar."""
    if self.open_backend and self.backend is None and self.exists():
        try:
            self.open()
        except Exception:
            # If we can't open the backend, just ignore it for now so we don't
            # prevent the user from building the Video object entirely.
            pass

__deepcopy__(memo)

Deep copy the video object.

Source code in sleap_io/model/video.py
def __deepcopy__(self, memo):
    """Deep copy the video object."""
    if id(self) in memo:
        return memo[id(self)]

    reopen = False
    if self.is_open:
        reopen = True
        self.close()

    new_video = Video(
        filename=self.filename,
        backend=None,
        backend_metadata=self.backend_metadata.copy(),
        source_video=self.source_video,
        open_backend=self.open_backend,
    )

    memo[id(self)] = new_video

    if reopen:
        self.open()

    return new_video

__getitem__(inds)

Return the frames of the video at the given indices.

Parameters:

Name Type Description Default
inds int | list[int] | slice

Index or list of indices of frames to read.

required

Returns:

Type Description
ndarray

Frame or frames as a numpy array of shape (height, width, channels) if a scalar index is provided, or (frames, height, width, channels) if a list of indices is provided.

See also: VideoBackend.get_frame, VideoBackend.get_frames

Source code in sleap_io/model/video.py
def __getitem__(self, inds: int | list[int] | slice) -> np.ndarray:
    """Return the frames of the video at the given indices.

    Args:
        inds: Index or list of indices of frames to read.

    Returns:
        Frame or frames as a numpy array of shape `(height, width, channels)` if a
        scalar index is provided, or `(frames, height, width, channels)` if a list
        of indices is provided.

    See also: VideoBackend.get_frame, VideoBackend.get_frames
    """
    if not self.is_open:
        if self.open_backend:
            self.open()
        else:
            raise ValueError(
                "Video backend is not open. Call video.open() or set "
                "video.open_backend to True to do automatically on frame read."
            )
    return self.backend[inds]

__init__(filename, backend=None, backend_metadata=NOTHING, source_video=None, open_backend=True)

Method generated by attrs for class Video.

Source code in sleap_io/model/video.py
"""Data model for videos.

The `Video` class is a SLEAP data structure that stores information regarding
a video and its components used in SLEAP.
"""

from __future__ import annotations

from pathlib import Path
from typing import Any

__len__()

Return the length of the video as the number of frames.

Source code in sleap_io/model/video.py
def __len__(self) -> int:
    """Return the length of the video as the number of frames."""
    shape = self.shape
    return 0 if shape is None else shape[0]

__replace__(**changes)

Method generated by attrs for class Video.

__repr__()

Informal string representation (for print or format).

Source code in sleap_io/model/video.py
def __repr__(self) -> str:
    """Informal string representation (for print or format)."""
    dataset = (
        f"dataset={self.backend.dataset}, "
        if getattr(self.backend, "dataset", "")
        else ""
    )
    return (
        "Video("
        f'filename="{self.filename}", '
        f"shape={self.shape}, "
        f"{dataset}"
        f"backend={type(self.backend).__name__}"
        ")"
    )

__str__()

Informal string representation (for print or format).

Source code in sleap_io/model/video.py
def __str__(self) -> str:
    """Informal string representation (for print or format)."""
    return self.__repr__()

close()

Close the video backend.

Source code in sleap_io/model/video.py
def close(self):
    """Close the video backend."""
    if self.backend is not None:
        # Try to remember values from previous backend if available and not
        # specified.
        try:
            self.backend_metadata["dataset"] = getattr(
                self.backend, "dataset", None
            )
            self.backend_metadata["grayscale"] = getattr(
                self.backend, "grayscale", None
            )
            self.backend_metadata["shape"] = getattr(self.backend, "shape", None)
            self.backend_metadata["fps"] = getattr(self.backend, "fps", None)
        except Exception:
            pass

        del self.backend
        self.backend = None

deduplicate_with(other)

Create a new video with duplicate images removed.

This method is specifically for ImageVideo backends (image sequences).

Parameters:

Name Type Description Default
other Video

Another video to deduplicate against. Must also be ImageVideo.

required

Returns:

Type Description
Video

A new Video object with duplicate images removed from this video, or None if all images were duplicates.

Raises:

Type Description
ValueError

If either video is not an ImageVideo backend.

Notes

Only works with ImageVideo backends where filename is a list. Images are considered duplicates if they have the same basename. The returned video contains only images from this video that are not present in the other video.

Source code in sleap_io/model/video.py
def deduplicate_with(self, other: "Video") -> "Video":
    """Create a new video with duplicate images removed.

    This method is specifically for ImageVideo backends (image sequences).

    Args:
        other: Another video to deduplicate against. Must also be ImageVideo.

    Returns:
        A new Video object with duplicate images removed from this video,
        or None if all images were duplicates.

    Raises:
        ValueError: If either video is not an ImageVideo backend.

    Notes:
        Only works with ImageVideo backends where filename is a list.
        Images are considered duplicates if they have the same basename.
        The returned video contains only images from this video that are
        not present in the other video.
    """
    if not isinstance(self.filename, list):
        raise ValueError("deduplicate_with only works with ImageVideo backends")
    if not isinstance(other.filename, list):
        raise ValueError("Other video must also be ImageVideo backend")

    # Get basenames from other video
    other_basenames = set(Path(f).name for f in other.filename)

    # Keep only non-duplicate images
    deduplicated_paths = [
        f for f in self.filename if Path(f).name not in other_basenames
    ]

    if not deduplicated_paths:
        # All images were duplicates
        return None

    # Create new video with deduplicated images
    return Video.from_filename(deduplicated_paths, grayscale=self.grayscale)

exists(check_all=False, dataset=None)

Check if the video file exists and is accessible.

Parameters:

Name Type Description Default
check_all bool

If True, check that all filenames in a list exist. If False (the default), check that the first filename exists.

False
dataset str | None

Name of dataset in HDF5 file. If specified, this will function will return False if the dataset does not exist.

None

Returns:

Type Description
bool

True if the file exists and is accessible, False otherwise.

Source code in sleap_io/model/video.py
def exists(self, check_all: bool = False, dataset: str | None = None) -> bool:
    """Check if the video file exists and is accessible.

    Args:
        check_all: If `True`, check that all filenames in a list exist. If `False`
            (the default), check that the first filename exists.
        dataset: Name of dataset in HDF5 file. If specified, this will function will
            return `False` if the dataset does not exist.

    Returns:
        `True` if the file exists and is accessible, `False` otherwise.
    """
    if isinstance(self.filename, list):
        if check_all:
            for f in self.filename:
                if not is_file_accessible(f):
                    return False
            return True
        else:
            return is_file_accessible(self.filename[0])

    file_is_accessible = is_file_accessible(self.filename)
    if not file_is_accessible:
        return False

    if dataset is None or dataset == "":
        dataset = self.backend_metadata.get("dataset", None)

    if dataset is not None and dataset != "":
        has_dataset = False
        if (
            self.backend is not None
            and type(self.backend) is HDF5Video
            and self.backend._open_reader is not None
        ):
            has_dataset = dataset in self.backend._open_reader
        else:
            with h5py.File(self.filename, "r") as f:
                has_dataset = dataset in f
        return has_dataset

    return True

frame_to_seconds(frame_idx)

Convert a frame index to timestamp in seconds.

Parameters:

Name Type Description Default
frame_idx int

Zero-indexed frame number.

required

Returns:

Type Description
float | None

Time in seconds, or None if FPS is unknown.

Notes

This assumes constant frame rate. For variable frame rate videos, the returned timestamp may be approximate.

Source code in sleap_io/model/video.py
def frame_to_seconds(self, frame_idx: int) -> float | None:
    """Convert a frame index to timestamp in seconds.

    Args:
        frame_idx: Zero-indexed frame number.

    Returns:
        Time in seconds, or None if FPS is unknown.

    Notes:
        This assumes constant frame rate. For variable frame rate videos,
        the returned timestamp may be approximate.
    """
    if self.fps is None or self.fps <= 0:
        return None
    return frame_idx / self.fps

from_filename(filename, dataset=None, grayscale=None, keep_open=True, source_video=None, **kwargs) classmethod

Create a Video from a filename.

Parameters:

Name Type Description Default
filename str | list[str]

The filename(s) of the video. Supported extensions: "mp4", "avi", "mov", "mj2", "mkv", "h5", "hdf5", "slp", "png", "jpg", "jpeg", "tif", "tiff", "bmp". If the filename is a list, a list of image filenames are expected. If filename is a folder, it will be searched for images.

required
dataset str | None

Name of dataset in HDF5 file.

None
grayscale bool | None

Whether to force grayscale. If None, autodetect on first frame load.

None
keep_open bool

Whether to keep the video reader open between calls to read frames. If False, will close the reader after each call. If True (the default), it will keep the reader open and cache it for subsequent calls which may enhance the performance of reading multiple frames.

True
source_video Video | None

The source video object if this is a proxy video. This is present when the video contains an embedded subset of frames from another video.

None
**kwargs

Additional backend-specific arguments passed to VideoBackend.from_filename. See VideoBackend.from_filename for supported arguments.

required

Returns:

Type Description
VideoBackend

Video instance with the appropriate backend instantiated.

Source code in sleap_io/model/video.py
@classmethod
def from_filename(
    cls,
    filename: str | list[str],
    dataset: str | None = None,
    grayscale: bool | None = None,
    keep_open: bool = True,
    source_video: "Video | None" = None,
    **kwargs,
) -> VideoBackend:
    """Create a Video from a filename.

    Args:
        filename: The filename(s) of the video. Supported extensions: "mp4", "avi",
            "mov", "mj2", "mkv", "h5", "hdf5", "slp", "png", "jpg", "jpeg", "tif",
            "tiff", "bmp". If the filename is a list, a list of image filenames are
            expected. If filename is a folder, it will be searched for images.
        dataset: Name of dataset in HDF5 file.
        grayscale: Whether to force grayscale. If None, autodetect on first frame
            load.
        keep_open: Whether to keep the video reader open between calls to read
            frames. If False, will close the reader after each call. If True (the
            default), it will keep the reader open and cache it for subsequent calls
            which may enhance the performance of reading multiple frames.
        source_video: The source video object if this is a proxy video. This is
            present when the video contains an embedded subset of frames from
            another video.
        **kwargs: Additional backend-specific arguments passed to
            VideoBackend.from_filename. See VideoBackend.from_filename for supported
            arguments.

    Returns:
        Video instance with the appropriate backend instantiated.
    """
    backend = VideoBackend.from_filename(
        filename,
        dataset=dataset,
        grayscale=grayscale,
        keep_open=keep_open,
        **kwargs,
    )
    # If filename is a directory, VideoBackend.from_filename will expand it
    # to a list of paths to images contained within the directory. In this
    # case we want to use the expanded list as filename
    return cls(
        filename=backend.filename,
        backend=backend,
        source_video=source_video,
    )

has_overlapping_images(other)

Check if this video has overlapping images with another video.

This method is specifically for ImageVideo backends (image sequences).

Parameters:

Name Type Description Default
other Video

Another video to compare with.

required

Returns:

Type Description
bool

True if both are ImageVideo instances with overlapping image files. False if either video is not an ImageVideo or no overlap exists.

Notes

Only works with ImageVideo backends where filename is a list. Compares individual image filenames (basenames only).

Source code in sleap_io/model/video.py
def has_overlapping_images(self, other: "Video") -> bool:
    """Check if this video has overlapping images with another video.

    This method is specifically for ImageVideo backends (image sequences).

    Args:
        other: Another video to compare with.

    Returns:
        True if both are ImageVideo instances with overlapping image files.
        False if either video is not an ImageVideo or no overlap exists.

    Notes:
        Only works with ImageVideo backends where filename is a list.
        Compares individual image filenames (basenames only).
    """
    # Both must be image sequences
    if not (isinstance(self.filename, list) and isinstance(other.filename, list)):
        return False

    # Get basenames for comparison
    self_basenames = set(Path(f).name for f in self.filename)
    other_basenames = set(Path(f).name for f in other.filename)

    # Check if there's any overlap
    return len(self_basenames & other_basenames) > 0

matches_content(other)

Check if this video has the same content as another video.

Parameters:

Name Type Description Default
other Video

Another video to compare with.

required

Returns:

Type Description
bool

True if the videos have the same shape and backend type.

Notes

This compares metadata like shape and backend type, not actual frame data.

Source code in sleap_io/model/video.py
def matches_content(self, other: "Video") -> bool:
    """Check if this video has the same content as another video.

    Args:
        other: Another video to compare with.

    Returns:
        True if the videos have the same shape and backend type.

    Notes:
        This compares metadata like shape and backend type, not actual frame data.
    """
    # Compare shapes
    self_shape = self.shape
    other_shape = other.shape

    if self_shape != other_shape:
        return False

    # Compare backend types
    if self.backend is None and other.backend is None:
        return True
    elif self.backend is None or other.backend is None:
        return False

    return type(self.backend).__name__ == type(other.backend).__name__

matches_path(other, strict=False)

Check if this video has the same path as another video.

Parameters:

Name Type Description Default
other Video

Another video to compare with.

required
strict bool

If True, require exact path match. If False, consider videos with the same filename (basename) as matching.

False

Returns:

Type Description
bool

True if the videos have matching paths, False otherwise.

Notes

For HDF5 video backends (e.g., embedded videos in .pkg.slp files), matching prioritizes the source_filename attribute since multiple videos can share the same HDF5 file path but reference different source videos. Falls back to dataset name matching if source_filename is not available.

Source code in sleap_io/model/video.py
def matches_path(self, other: "Video", strict: bool = False) -> bool:
    """Check if this video has the same path as another video.

    Args:
        other: Another video to compare with.
        strict: If True, require exact path match. If False, consider videos
            with the same filename (basename) as matching.

    Returns:
        True if the videos have matching paths, False otherwise.

    Notes:
        For HDF5 video backends (e.g., embedded videos in .pkg.slp files),
        matching prioritizes the source_filename attribute since multiple
        videos can share the same HDF5 file path but reference different
        source videos. Falls back to dataset name matching if source_filename
        is not available.
    """
    # Handle HDF5 backends specially - prioritize source_filename matching
    self_is_hdf5 = isinstance(self.backend, HDF5Video)
    other_is_hdf5 = isinstance(other.backend, HDF5Video)

    if self_is_hdf5 and other_is_hdf5:
        # Both are HDF5 videos - match by source_filename first
        self_source = self.backend.source_filename
        other_source = other.backend.source_filename

        if self_source is not None and other_source is not None:
            if strict:
                return Path(self_source).resolve() == Path(other_source).resolve()
            else:
                return Path(self_source).name == Path(other_source).name

        # Fall back to dataset name matching if source_filename is not available
        self_dataset = self.backend.dataset
        other_dataset = other.backend.dataset

        if self_dataset is not None and other_dataset is not None:
            return self_dataset == other_dataset

        # If neither source_filename nor dataset available, cannot match
        return False

    if isinstance(self.filename, list) and isinstance(other.filename, list):
        # Both are image sequences
        if strict:
            return self.filename == other.filename
        else:
            # Compare basenames
            self_basenames = [Path(f).name for f in self.filename]
            other_basenames = [Path(f).name for f in other.filename]
            return self_basenames == other_basenames
    elif isinstance(self.filename, list) or isinstance(other.filename, list):
        # One is image sequence, other is single file
        return False
    else:
        # Both are single files
        if strict:
            return Path(self.filename).resolve() == Path(other.filename).resolve()
        else:
            return Path(self.filename).name == Path(other.filename).name

matches_shape(other)

Check if this video has the same shape as another video.

Parameters:

Name Type Description Default
other Video

Another video to compare with.

required

Returns:

Type Description
bool

True if the videos have the same height, width, and channels.

Notes

This only compares spatial dimensions, not the number of frames.

Source code in sleap_io/model/video.py
def matches_shape(self, other: "Video") -> bool:
    """Check if this video has the same shape as another video.

    Args:
        other: Another video to compare with.

    Returns:
        True if the videos have the same height, width, and channels.

    Notes:
        This only compares spatial dimensions, not the number of frames.
    """
    # Try to get shape from backend metadata first if shape is not available
    if self.backend is None and "shape" in self.backend_metadata:
        self_shape = self.backend_metadata["shape"]
    else:
        self_shape = self.shape

    if other.backend is None and "shape" in other.backend_metadata:
        other_shape = other.backend_metadata["shape"]
    else:
        other_shape = other.shape

    # Handle None shapes
    if self_shape is None or other_shape is None:
        return False

    # Compare only height, width, channels (not frames)
    return self_shape[1:] == other_shape[1:]

merge_with(other)

Merge another video's images into this one.

This method is specifically for ImageVideo backends (image sequences).

Parameters:

Name Type Description Default
other Video

Another video to merge with. Must also be ImageVideo.

required

Returns:

Type Description
Video

A new Video object with unique images from both videos.

Raises:

Type Description
ValueError

If either video is not an ImageVideo backend.

Notes

Only works with ImageVideo backends where filename is a list. The merged video contains all unique images from both videos, with automatic deduplication based on image basename.

Source code in sleap_io/model/video.py
def merge_with(self, other: "Video") -> "Video":
    """Merge another video's images into this one.

    This method is specifically for ImageVideo backends (image sequences).

    Args:
        other: Another video to merge with. Must also be ImageVideo.

    Returns:
        A new Video object with unique images from both videos.

    Raises:
        ValueError: If either video is not an ImageVideo backend.

    Notes:
        Only works with ImageVideo backends where filename is a list.
        The merged video contains all unique images from both videos,
        with automatic deduplication based on image basename.
    """
    if not isinstance(self.filename, list):
        raise ValueError("merge_with only works with ImageVideo backends")
    if not isinstance(other.filename, list):
        raise ValueError("Other video must also be ImageVideo backend")

    # Get all unique images (by basename) preserving order
    seen_basenames = set()
    merged_paths = []

    for path in self.filename:
        basename = Path(path).name
        if basename not in seen_basenames:
            merged_paths.append(path)
            seen_basenames.add(basename)

    for path in other.filename:
        basename = Path(path).name
        if basename not in seen_basenames:
            merged_paths.append(path)
            seen_basenames.add(basename)

    # Create new video with merged images
    return Video.from_filename(merged_paths, grayscale=self.grayscale)

open(filename=None, dataset=None, grayscale=None, keep_open=True, plugin=None)

Open the video backend for reading.

Parameters:

Name Type Description Default
filename str | None

Filename to open. If not specified, will use the filename set on the video object.

None
dataset str | None

Name of dataset in HDF5 file.

None
grayscale str | None

Whether to force grayscale. If None, autodetect on first frame load.

None
keep_open bool

Whether to keep the video reader open between calls to read frames. If False, will close the reader after each call. If True (the default), it will keep the reader open and cache it for subsequent calls which may enhance the performance of reading multiple frames.

True
plugin str | None

Video plugin to use for MediaVideo files. One of "opencv", "FFMPEG", or "pyav". Also accepts aliases (case-insensitive). If not specified, uses the backend metadata, global default, or auto-detection in that order.

None
Notes

This is useful for opening the video backend to read frames and then closing it after reading all the necessary frames.

If the backend was already open, it will be closed before opening a new one. Values for the HDF5 dataset and grayscale will be remembered if not specified.

Source code in sleap_io/model/video.py
def open(
    self,
    filename: str | None = None,
    dataset: str | None = None,
    grayscale: str | None = None,
    keep_open: bool = True,
    plugin: str | None = None,
):
    """Open the video backend for reading.

    Args:
        filename: Filename to open. If not specified, will use the filename set on
            the video object.
        dataset: Name of dataset in HDF5 file.
        grayscale: Whether to force grayscale. If None, autodetect on first frame
            load.
        keep_open: Whether to keep the video reader open between calls to read
            frames. If False, will close the reader after each call. If True (the
            default), it will keep the reader open and cache it for subsequent calls
            which may enhance the performance of reading multiple frames.
        plugin: Video plugin to use for MediaVideo files. One of "opencv",
            "FFMPEG", or "pyav". Also accepts aliases (case-insensitive).
            If not specified, uses the backend metadata, global default,
            or auto-detection in that order.

    Notes:
        This is useful for opening the video backend to read frames and then closing
        it after reading all the necessary frames.

        If the backend was already open, it will be closed before opening a new one.
        Values for the HDF5 dataset and grayscale will be remembered if not
        specified.
    """
    if filename is not None:
        self.replace_filename(filename, open=False)

    # Try to remember values from previous backend if available and not specified.
    if self.backend is not None:
        if dataset is None:
            dataset = getattr(self.backend, "dataset", None)
        if grayscale is None:
            grayscale = getattr(self.backend, "grayscale", None)

    else:
        if dataset is None and "dataset" in self.backend_metadata:
            dataset = self.backend_metadata["dataset"]
        if grayscale is None:
            if "grayscale" in self.backend_metadata:
                grayscale = self.backend_metadata["grayscale"]
            elif "shape" in self.backend_metadata:
                grayscale = self.backend_metadata["shape"][-1] == 1

    if not self.exists(dataset=dataset):
        msg = (
            f"Video does not exist or cannot be opened for reading: {self.filename}"
        )
        if dataset is not None:
            msg += f" (dataset: {dataset})"
        raise FileNotFoundError(msg)

    # Close previous backend if open.
    self.close()

    # Handle plugin parameter
    backend_kwargs = {}
    if plugin is not None:
        from sleap_io.io.video_reading import normalize_plugin_name

        plugin = normalize_plugin_name(plugin)
        self.backend_metadata["plugin"] = plugin

    if "plugin" in self.backend_metadata:
        backend_kwargs["plugin"] = self.backend_metadata["plugin"]

    # Create new backend.
    self.backend = VideoBackend.from_filename(
        self.filename,
        dataset=dataset,
        grayscale=grayscale,
        keep_open=keep_open,
        **backend_kwargs,
    )

replace_filename(new_filename, open=True)

Update the filename of the video, optionally opening the backend.

Parameters:

Name Type Description Default
new_filename str | Path | list[str] | list[Path]

New filename to set for the video.

required
open bool

If True (the default), open the backend with the new filename. If the new filename does not exist, no error is raised.

True
Source code in sleap_io/model/video.py
def replace_filename(
    self, new_filename: str | Path | list[str] | list[Path], open: bool = True
):
    """Update the filename of the video, optionally opening the backend.

    Args:
        new_filename: New filename to set for the video.
        open: If `True` (the default), open the backend with the new filename. If
            the new filename does not exist, no error is raised.
    """
    if isinstance(new_filename, Path):
        new_filename = new_filename.as_posix()

    if isinstance(new_filename, list):
        new_filename = [
            p.as_posix() if isinstance(p, Path) else p for p in new_filename
        ]

    self.filename = new_filename
    self.backend_metadata["filename"] = new_filename

    if open:
        if self.exists():
            self.open()
        else:
            self.close()

save(save_path, frame_inds=None, fps=None, video_kwargs=None)

Save video frames to a new video file.

Parameters:

Name Type Description Default
save_path str | Path

Path to the new video file. Should end in MP4.

required
frame_inds list[int] | ndarray | None

Frame indices to save. Can be specified as a list or array of frame integers. If not specified, saves all video frames.

None
fps float | None

Frames per second for the output video. If not specified, uses the source video's FPS if available, otherwise defaults to 30.

None
video_kwargs dict[str, Any] | None

A dictionary of keyword arguments to provide to sio.save_video for video compression.

None

Returns:

Type Description
Video

A new Video object pointing to the new video file.

Source code in sleap_io/model/video.py
def save(
    self,
    save_path: str | Path,
    frame_inds: list[int] | np.ndarray | None = None,
    fps: float | None = None,
    video_kwargs: dict[str, Any] | None = None,
) -> "Video":
    """Save video frames to a new video file.

    Args:
        save_path: Path to the new video file. Should end in MP4.
        frame_inds: Frame indices to save. Can be specified as a list or array of
            frame integers. If not specified, saves all video frames.
        fps: Frames per second for the output video. If not specified, uses the
            source video's FPS if available, otherwise defaults to 30.
        video_kwargs: A dictionary of keyword arguments to provide to
            `sio.save_video` for video compression.

    Returns:
        A new `Video` object pointing to the new video file.
    """
    video_kwargs = {} if video_kwargs is None else video_kwargs.copy()
    frame_inds = np.arange(len(self)) if frame_inds is None else frame_inds

    # Use source video FPS if not explicitly specified
    if fps is None:
        fps = self.fps
    if fps is not None and "fps" not in video_kwargs:
        video_kwargs["fps"] = fps

    with VideoWriter(save_path, **video_kwargs) as vw:
        for frame_ind in frame_inds:
            vw(self[frame_ind])

    new_video = Video.from_filename(save_path, grayscale=self.grayscale)
    return new_video

seconds_to_frame(seconds)

Convert a timestamp in seconds to frame index.

Parameters:

Name Type Description Default
seconds float

Time in seconds from video start.

required

Returns:

Type Description
int | None

Zero-indexed frame number (rounded down), or None if FPS unknown.

Source code in sleap_io/model/video.py
def seconds_to_frame(self, seconds: float) -> int | None:
    """Convert a timestamp in seconds to frame index.

    Args:
        seconds: Time in seconds from video start.

    Returns:
        Zero-indexed frame number (rounded down), or None if FPS unknown.
    """
    if self.fps is None or self.fps <= 0:
        return None
    return int(seconds * self.fps)

set_video_plugin(plugin)

Set the video plugin and reopen the video.

Parameters:

Name Type Description Default
plugin str

Video plugin to use. One of "opencv", "FFMPEG", or "pyav". Also accepts aliases (case-insensitive).

required

Raises:

Type Description
ValueError

If the video is not a MediaVideo type.

Examples:

>>> video.set_video_plugin("opencv")
>>> video.set_video_plugin("CV2")  # Same as "opencv"
Source code in sleap_io/model/video.py
def set_video_plugin(self, plugin: str) -> None:
    """Set the video plugin and reopen the video.

    Args:
        plugin: Video plugin to use. One of "opencv", "FFMPEG", or "pyav".
            Also accepts aliases (case-insensitive).

    Raises:
        ValueError: If the video is not a MediaVideo type.

    Examples:
        >>> video.set_video_plugin("opencv")
        >>> video.set_video_plugin("CV2")  # Same as "opencv"
    """
    from sleap_io.io.video_reading import MediaVideo, normalize_plugin_name

    if not self.filename.endswith(MediaVideo.EXTS):
        raise ValueError(f"Cannot set plugin for non-media video: {self.filename}")

    plugin = normalize_plugin_name(plugin)

    # Close current backend if open
    was_open = self.is_open
    if was_open:
        self.close()

    # Update backend metadata
    self.backend_metadata["plugin"] = plugin

    # Reopen with new plugin if it was open
    if was_open:
        self.open()

VideoBackend

Base class for video backends.

This class is not meant to be used directly. Instead, use the from_filename constructor to create a backend instance.

Attributes:

Name Type Description
filename

Path to video file(s).

grayscale

Whether to force grayscale. If None, autodetect on first frame load.

keep_open

Whether to keep the video reader open between calls to read frames. If False, will close the reader after each call. If True (the default), it will keep the reader open and cache it for subsequent calls which may enhance the performance of reading multiple frames.

fps

Frames per second of the video. For MediaVideo, this is read from container metadata. For other backends (ImageVideo, HDF5Video, TiffVideo), this must be set explicitly or will be None.

Methods:

Name Description
__eq__

Method generated by attrs for class VideoBackend.

__getitem__

Return a single frame or a list of frames from the video.

__init__

Method generated by attrs for class VideoBackend.

__len__

Return number of frames in the video.

__replace__

Method generated by attrs for class VideoBackend.

__repr__

Method generated by attrs for class VideoBackend.

detect_grayscale

Detect whether the video is grayscale.

from_filename

Create a VideoBackend from a filename.

get_frame

Read a single frame from the video.

get_frames

Read a list of frames from the video.

has_frame

Check if a frame index is contained in the video.

read_test_frame

Read a single frame from the video to test for grayscale.

Source code in sleap_io/io/video_reading.py
@attrs.define
class VideoBackend:
    """Base class for video backends.

    This class is not meant to be used directly. Instead, use the `from_filename`
    constructor to create a backend instance.

    Attributes:
        filename: Path to video file(s).
        grayscale: Whether to force grayscale. If None, autodetect on first frame load.
        keep_open: Whether to keep the video reader open between calls to read frames.
            If False, will close the reader after each call. If True (the default), it
            will keep the reader open and cache it for subsequent calls which may
            enhance the performance of reading multiple frames.
        fps: Frames per second of the video. For MediaVideo, this is read from container
            metadata. For other backends (ImageVideo, HDF5Video, TiffVideo), this must
            be set explicitly or will be None.
    """

    filename: str | Path | list[str] | list[Path]
    grayscale: bool | None = None
    keep_open: bool = True
    _cached_shape: tuple[int, int, int, int] | None = None
    _open_reader: object | None = None
    _fps: float | None = None

    @property
    def fps(self) -> float | None:
        """Frames per second of the video.

        Returns:
            The FPS if known, or None if unavailable/unknown.

        Notes:
            For MediaVideo, this is read from container metadata.
            For ImageVideo, HDF5Video, and TiffVideo, this must be set explicitly
            or inherited from source_video.
        """
        return self._fps

    @fps.setter
    def fps(self, value: float | None) -> None:
        """Set the FPS.

        Args:
            value: Frames per second. Must be positive if not None.

        Raises:
            ValueError: If value is not positive.
        """
        if value is not None and value <= 0:
            raise ValueError(f"FPS must be positive, got {value}")
        self._fps = value

    @classmethod
    def from_filename(
        cls,
        filename: str | list[str],
        dataset: str | None = None,
        grayscale: bool | None = None,
        keep_open: bool = True,
        **kwargs,
    ) -> "VideoBackend":
        """Create a VideoBackend from a filename.

        Args:
            filename: Path to video file(s).
            dataset: Name of dataset in HDF5 file.
            grayscale: Whether to force grayscale. If None, autodetect on first frame
                load.
            keep_open: Whether to keep the video reader open between calls to read
                frames. If False, will close the reader after each call. If True (the
                default), it will keep the reader open and cache it for subsequent calls
                which may enhance the performance of reading multiple frames.
            **kwargs: Additional backend-specific arguments. These are filtered to only
                include parameters that are valid for the specific backend being
                created:
                - For ImageVideo: plugin (str): Image plugin to use. One of "opencv"
                  or "imageio". Also accepts aliases (case-insensitive).
                  If None, uses global default if set, otherwise auto-detects.
                - For MediaVideo: plugin (str): Video plugin to use. One of "opencv",
                  "FFMPEG", or "pyav". Also accepts aliases (case-insensitive).
                  If None, uses global default if set, otherwise auto-detects.
                - For HDF5Video: input_format (str), frame_map (dict),
                  source_filename (str),
                  source_inds (np.ndarray), image_format (str). See HDF5Video for
                  details.

        Returns:
            VideoBackend subclass instance.
        """
        if isinstance(filename, Path):
            filename = filename.as_posix()

        if type(filename) is str and Path(filename).is_dir():
            filename = ImageVideo.find_images(filename)

        if type(filename) is list:
            filename = [Path(f).as_posix() for f in filename]
            return ImageVideo(
                filename, grayscale=grayscale, **_get_valid_kwargs(ImageVideo, kwargs)
            )
        elif filename.lower().endswith(("tif", "tiff")):
            # Detect TIFF format
            format_type, metadata = TiffVideo.detect_format(filename)

            if format_type in ("multi_page", "rank3_video", "rank4_video"):
                # Use TiffVideo for multi-page or multi-dimensional TIFFs
                tiff_kwargs = _get_valid_kwargs(TiffVideo, kwargs)
                # Add format if detected
                if format_type in ("rank3_video", "rank4_video"):
                    tiff_kwargs["format"] = metadata.get("format")
                return TiffVideo(
                    filename,
                    grayscale=grayscale,
                    keep_open=keep_open,
                    **tiff_kwargs,
                )
            else:
                # Single-page TIFF, treat as regular image
                return ImageVideo(
                    [filename],
                    grayscale=grayscale,
                    **_get_valid_kwargs(ImageVideo, kwargs),
                )
        elif filename.lower().endswith(tuple(ext.lower() for ext in ImageVideo.EXTS)):
            return ImageVideo(
                [filename], grayscale=grayscale, **_get_valid_kwargs(ImageVideo, kwargs)
            )
        elif filename.lower().endswith(tuple(ext.lower() for ext in MediaVideo.EXTS)):
            return MediaVideo(
                filename,
                grayscale=grayscale,
                keep_open=keep_open,
                **_get_valid_kwargs(MediaVideo, kwargs),
            )
        elif filename.lower().endswith(tuple(ext.lower() for ext in HDF5Video.EXTS)):
            return HDF5Video(
                filename,
                dataset=dataset,
                grayscale=grayscale,
                keep_open=keep_open,
                **_get_valid_kwargs(HDF5Video, kwargs),
            )
        else:
            raise ValueError(f"Unknown video file type: {filename}")

    def _read_frame(self, frame_idx: int) -> np.ndarray:
        """Read a single frame from the video. Must be implemented in subclasses."""
        raise NotImplementedError

    def _read_frames(self, frame_inds: list) -> np.ndarray:
        """Read a list of frames from the video."""
        return np.stack([self.get_frame(i) for i in frame_inds], axis=0)

    def read_test_frame(self) -> np.ndarray:
        """Read a single frame from the video to test for grayscale.

        Note:
            This reads the frame at index 0. This may not be appropriate if the first
            frame is not available in a given backend.
        """
        return self._read_frame(0)

    def detect_grayscale(self, test_img: np.ndarray | None = None) -> bool:
        """Detect whether the video is grayscale.

        This works by reading in a test frame and comparing the first and last channel
        for equality. It may fail in cases where, due to compression, the first and
        last channels are not exactly the same.

        Args:
            test_img: Optional test image to use. If not provided, a test image will be
                loaded via the `read_test_frame` method.

        Returns:
            Whether the video is grayscale. This value is also cached in the `grayscale`
            attribute of the class.
        """
        if test_img is None:
            test_img = self.read_test_frame()
        is_grayscale = np.array_equal(test_img[..., 0], test_img[..., -1])
        self.grayscale = is_grayscale
        return is_grayscale

    @property
    def num_frames(self) -> int:
        """Number of frames in the video. Must be implemented in subclasses."""
        raise NotImplementedError

    @property
    def img_shape(self) -> tuple[int, int, int]:
        """Shape of a single frame in the video."""
        height, width, channels = self.read_test_frame().shape
        if self.grayscale is None:
            self.detect_grayscale()
        if self.grayscale is False:
            channels = 3
        elif self.grayscale is True:
            channels = 1
        return int(height), int(width), int(channels)

    @property
    def shape(self) -> tuple[int, int, int, int]:
        """Shape of the video as a tuple of `(frames, height, width, channels)`.

        On first call, this will defer to `num_frames` and `img_shape` to determine the
        full shape. This call may be expensive for some subclasses, so the result is
        cached and returned on subsequent calls.
        """
        if self._cached_shape is not None:
            return self._cached_shape
        else:
            shape = (self.num_frames,) + self.img_shape
            self._cached_shape = shape
            return shape

    @property
    def frames(self) -> int:
        """Number of frames in the video."""
        return self.shape[0]

    def __len__(self) -> int:
        """Return number of frames in the video."""
        return self.shape[0]

    def has_frame(self, frame_idx: int) -> bool:
        """Check if a frame index is contained in the video.

        Args:
            frame_idx: Index of frame to check.

        Returns:
            `True` if the index is contained in the video, otherwise `False`.
        """
        return frame_idx < len(self)

    def get_frame(self, frame_idx: int) -> np.ndarray:
        """Read a single frame from the video.

        Args:
            frame_idx: Index of frame to read.

        Returns:
            Frame as a numpy array of shape `(height, width, channels)` where the
            `channels` dimension is 1 for grayscale videos and 3 for color videos.

        Notes:
            If the `grayscale` attribute is set to `True`, the `channels` dimension will
            be reduced to 1 if an RGB frame is loaded from the backend.

            If the `grayscale` attribute is set to `None`, the `grayscale` attribute
            will be automatically set based on the first frame read.

        See also: `get_frames`
        """
        if not self.has_frame(frame_idx):
            raise IndexError(f"Frame index {frame_idx} out of range.")

        img = self._read_frame(frame_idx)

        if self.grayscale is None:
            self.detect_grayscale(img)

        if self.grayscale:
            img = img[..., [0]]

        return img

    def get_frames(self, frame_inds: list[int]) -> np.ndarray:
        """Read a list of frames from the video.

        Depending on the backend implementation, this may be faster than reading frames
        individually using `get_frame`.

        Args:
            frame_inds: List of frame indices to read.

        Returns:
            Frames as a numpy array of shape `(frames, height, width, channels)` where
            `channels` dimension is 1 for grayscale videos and 3 for color videos.

        Notes:
            If the `grayscale` attribute is set to `True`, the `channels` dimension will
            be reduced to 1 if an RGB frame is loaded from the backend.

            If the `grayscale` attribute is set to `None`, the `grayscale` attribute
            will be automatically set based on the first frame read.

        See also: `get_frame`
        """
        imgs = self._read_frames(frame_inds)

        if self.grayscale is None:
            self.detect_grayscale(imgs[0])

        if self.grayscale:
            imgs = imgs[..., [0]]

        return imgs

    def __getitem__(self, ind: int | list[int] | slice) -> np.ndarray:
        """Return a single frame or a list of frames from the video.

        Args:
            ind: Index or list of indices of frames to read.

        Returns:
            Frame or frames as a numpy array of shape `(height, width, channels)` if a
            scalar index is provided, or `(frames, height, width, channels)` if a list
            of indices is provided.

        See also: get_frame, get_frames
        """
        if np.isscalar(ind):
            return self.get_frame(ind)
        else:
            if type(ind) is slice:
                start = (ind.start or 0) % len(self)
                stop = ind.stop or len(self)
                if stop < 0:
                    stop = len(self) + stop
                step = ind.step or 1
                ind = range(start, stop, step)
            return self.get_frames(ind)

__annotations__ = {'filename': 'str | Path | list[str] | list[Path]', 'grayscale': 'bool | None', 'keep_open': 'bool', '_cached_shape': 'tuple[int, int, int, int] | None', '_open_reader': 'object | None', '_fps': 'float | None'} class-attribute

dict() -> new empty dictionary dict(mapping) -> new dictionary initialized from a mapping object's (key, value) pairs dict(iterable) -> new dictionary initialized as if via: d = {} for k, v in iterable: d[k] = v dict(**kwargs) -> new dictionary initialized with the name=value pairs in the keyword argument list. For example: dict(one=1, two=2)

__attrs_own_setattr__ = False class-attribute

Returns True when the argument is true, False otherwise. The builtins True and False are the only two instances of the class bool. The class bool is a subclass of the class int, and cannot be subclassed.

__attrs_props__ = ClassProps(is_exception=False, is_slotted=True, has_weakref_slot=True, is_frozen=False, kw_only=<KeywordOnly.NO: 'no'>, collected_fields_by_mro=True, added_init=True, added_repr=True, added_eq=True, added_ordering=False, hashability=<Hashability.UNHASHABLE: 'unhashable'>, added_match_args=True, added_str=False, added_pickling=True, on_setattr_hook=<function pipe.<locals>.wrapped_pipe at 0x7f41e68fca40>, field_transformer=None) class-attribute

Effective class properties as derived from parameters to attr.s() or define() decorators.

This is the same data structure that attrs uses internally to decide how to construct the final class.

Warning:

This feature is currently **experimental** and is not covered by our
strict backwards-compatibility guarantees.

Attributes:

Name Type Description
is_exception bool

Whether the class is treated as an exception class.

is_slotted bool

Whether the class is slotted <slotted classes>.

has_weakref_slot bool

Whether the class has a slot for weak references.

is_frozen bool

Whether the class is frozen.

kw_only KeywordOnly

Whether / how the class enforces keyword-only arguments on the __init__ method.

collected_fields_by_mro bool

Whether the class fields were collected by method resolution order. That is, correctly but unlike dataclasses.

added_init bool

Whether the class has an attrs-generated __init__ method.

added_repr bool

Whether the class has an attrs-generated __repr__ method.

added_eq bool

Whether the class has attrs-generated equality methods.

added_ordering bool

Whether the class has attrs-generated ordering methods.

hashability Hashability

How hashable <hashing> the class is.

added_match_args bool

Whether the class supports positional match <match> over its fields.

added_str bool

Whether the class has an attrs-generated __str__ method.

added_pickling bool

Whether the class has attrs-generated __getstate__ and __setstate__ methods for pickle.

on_setattr_hook Callable[[Any, Attribute[Any], Any], Any] | None

The class's __setattr__ hook.

field_transformer Callable[[Attribute[Any]], Attribute[Any]] | None

The class's field transformers <transform-fields>.

.. versionadded:: 25.4.0

__doc__ = 'Base class for video backends.\n\nThis class is not meant to be used directly. Instead, use the `from_filename`\nconstructor to create a backend instance.\n\nAttributes:\n filename: Path to video file(s).\n grayscale: Whether to force grayscale. If None, autodetect on first frame load.\n keep_open: Whether to keep the video reader open between calls to read frames.\n If False, will close the reader after each call. If True (the default), it\n will keep the reader open and cache it for subsequent calls which may\n enhance the performance of reading multiple frames.\n fps: Frames per second of the video. For MediaVideo, this is read from container\n metadata. For other backends (ImageVideo, HDF5Video, TiffVideo), this must\n be set explicitly or will be None.\n' class-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

__firstlineno__ = 294 class-attribute

int([x]) -> integer int(x, base=10) -> integer

Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.int(). For floating-point numbers, this truncates towards zero.

If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by '+' or '-' and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal.

int('0b100', base=0) 4

__match_args__ = ('filename', 'grayscale', 'keep_open', '_cached_shape', '_open_reader', '_fps') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__module__ = 'sleap_io.io.video_reading' class-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

__slots__ = ('filename', 'grayscale', 'keep_open', '_cached_shape', '_open_reader', '_fps', '__weakref__') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__static_attributes__ = ('_cached_shape', '_fps', 'grayscale') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__weakref__ property

list of weak references to the object

fps property

Frames per second of the video.

Returns:

Type Description

The FPS if known, or None if unavailable/unknown.

Notes

For MediaVideo, this is read from container metadata. For ImageVideo, HDF5Video, and TiffVideo, this must be set explicitly or inherited from source_video.

frames property

Number of frames in the video.

img_shape property

Shape of a single frame in the video.

num_frames property

Number of frames in the video. Must be implemented in subclasses.

shape property

Shape of the video as a tuple of (frames, height, width, channels).

On first call, this will defer to num_frames and img_shape to determine the full shape. This call may be expensive for some subclasses, so the result is cached and returned on subsequent calls.

__eq__(other)

Method generated by attrs for class VideoBackend.

Source code in sleap_io/io/video_reading.py
    import cv2
except ImportError:
    pass

try:
    import imageio_ffmpeg  # noqa: F401
except ImportError:
    pass

try:
    import av  # noqa: F401

__getitem__(ind)

Return a single frame or a list of frames from the video.

Parameters:

Name Type Description Default
ind int | list[int] | slice

Index or list of indices of frames to read.

required

Returns:

Type Description
ndarray

Frame or frames as a numpy array of shape (height, width, channels) if a scalar index is provided, or (frames, height, width, channels) if a list of indices is provided.

See also: get_frame, get_frames

Source code in sleap_io/io/video_reading.py
def __getitem__(self, ind: int | list[int] | slice) -> np.ndarray:
    """Return a single frame or a list of frames from the video.

    Args:
        ind: Index or list of indices of frames to read.

    Returns:
        Frame or frames as a numpy array of shape `(height, width, channels)` if a
        scalar index is provided, or `(frames, height, width, channels)` if a list
        of indices is provided.

    See also: get_frame, get_frames
    """
    if np.isscalar(ind):
        return self.get_frame(ind)
    else:
        if type(ind) is slice:
            start = (ind.start or 0) % len(self)
            stop = ind.stop or len(self)
            if stop < 0:
                stop = len(self) + stop
            step = ind.step or 1
            ind = range(start, stop, step)
        return self.get_frames(ind)

__init__(filename, grayscale=None, keep_open=True, cached_shape=None, open_reader=None, fps=None)

Method generated by attrs for class VideoBackend.

Source code in sleap_io/io/video_reading.py
except ImportError:
    pass


# Track available backends (populated on module import)
_AVAILABLE_VIDEO_BACKENDS = {
    "opencv": "cv2" in sys.modules,

__len__()

Return number of frames in the video.

Source code in sleap_io/io/video_reading.py
def __len__(self) -> int:
    """Return number of frames in the video."""
    return self.shape[0]

__replace__(**changes)

Method generated by attrs for class VideoBackend.

Source code in sleap_io/io/video_reading.py
Returns:

__repr__()

Method generated by attrs for class VideoBackend.

Source code in sleap_io/io/video_reading.py
"""Backends for reading videos."""

from __future__ import annotations

import sys
from io import BytesIO
from pathlib import Path

import attrs
import h5py
import imageio.v3 as iio
import numpy as np
import simplejson as json

try:

detect_grayscale(test_img=None)

Detect whether the video is grayscale.

This works by reading in a test frame and comparing the first and last channel for equality. It may fail in cases where, due to compression, the first and last channels are not exactly the same.

Parameters:

Name Type Description Default
test_img ndarray | None

Optional test image to use. If not provided, a test image will be loaded via the read_test_frame method.

None

Returns:

Type Description
bool

Whether the video is grayscale. This value is also cached in the grayscale attribute of the class.

Source code in sleap_io/io/video_reading.py
def detect_grayscale(self, test_img: np.ndarray | None = None) -> bool:
    """Detect whether the video is grayscale.

    This works by reading in a test frame and comparing the first and last channel
    for equality. It may fail in cases where, due to compression, the first and
    last channels are not exactly the same.

    Args:
        test_img: Optional test image to use. If not provided, a test image will be
            loaded via the `read_test_frame` method.

    Returns:
        Whether the video is grayscale. This value is also cached in the `grayscale`
        attribute of the class.
    """
    if test_img is None:
        test_img = self.read_test_frame()
    is_grayscale = np.array_equal(test_img[..., 0], test_img[..., -1])
    self.grayscale = is_grayscale
    return is_grayscale

from_filename(filename, dataset=None, grayscale=None, keep_open=True, **kwargs) classmethod

Create a VideoBackend from a filename.

Parameters:

Name Type Description Default
filename str | list[str]

Path to video file(s).

required
dataset str | None

Name of dataset in HDF5 file.

None
grayscale bool | None

Whether to force grayscale. If None, autodetect on first frame load.

None
keep_open bool

Whether to keep the video reader open between calls to read frames. If False, will close the reader after each call. If True (the default), it will keep the reader open and cache it for subsequent calls which may enhance the performance of reading multiple frames.

True
**kwargs

Additional backend-specific arguments. These are filtered to only include parameters that are valid for the specific backend being created: - For ImageVideo: plugin (str): Image plugin to use. One of "opencv" or "imageio". Also accepts aliases (case-insensitive). If None, uses global default if set, otherwise auto-detects. - For MediaVideo: plugin (str): Video plugin to use. One of "opencv", "FFMPEG", or "pyav". Also accepts aliases (case-insensitive). If None, uses global default if set, otherwise auto-detects. - For HDF5Video: input_format (str), frame_map (dict), source_filename (str), source_inds (np.ndarray), image_format (str). See HDF5Video for details.

required

Returns:

Type Description
VideoBackend

VideoBackend subclass instance.

Source code in sleap_io/io/video_reading.py
@classmethod
def from_filename(
    cls,
    filename: str | list[str],
    dataset: str | None = None,
    grayscale: bool | None = None,
    keep_open: bool = True,
    **kwargs,
) -> "VideoBackend":
    """Create a VideoBackend from a filename.

    Args:
        filename: Path to video file(s).
        dataset: Name of dataset in HDF5 file.
        grayscale: Whether to force grayscale. If None, autodetect on first frame
            load.
        keep_open: Whether to keep the video reader open between calls to read
            frames. If False, will close the reader after each call. If True (the
            default), it will keep the reader open and cache it for subsequent calls
            which may enhance the performance of reading multiple frames.
        **kwargs: Additional backend-specific arguments. These are filtered to only
            include parameters that are valid for the specific backend being
            created:
            - For ImageVideo: plugin (str): Image plugin to use. One of "opencv"
              or "imageio". Also accepts aliases (case-insensitive).
              If None, uses global default if set, otherwise auto-detects.
            - For MediaVideo: plugin (str): Video plugin to use. One of "opencv",
              "FFMPEG", or "pyav". Also accepts aliases (case-insensitive).
              If None, uses global default if set, otherwise auto-detects.
            - For HDF5Video: input_format (str), frame_map (dict),
              source_filename (str),
              source_inds (np.ndarray), image_format (str). See HDF5Video for
              details.

    Returns:
        VideoBackend subclass instance.
    """
    if isinstance(filename, Path):
        filename = filename.as_posix()

    if type(filename) is str and Path(filename).is_dir():
        filename = ImageVideo.find_images(filename)

    if type(filename) is list:
        filename = [Path(f).as_posix() for f in filename]
        return ImageVideo(
            filename, grayscale=grayscale, **_get_valid_kwargs(ImageVideo, kwargs)
        )
    elif filename.lower().endswith(("tif", "tiff")):
        # Detect TIFF format
        format_type, metadata = TiffVideo.detect_format(filename)

        if format_type in ("multi_page", "rank3_video", "rank4_video"):
            # Use TiffVideo for multi-page or multi-dimensional TIFFs
            tiff_kwargs = _get_valid_kwargs(TiffVideo, kwargs)
            # Add format if detected
            if format_type in ("rank3_video", "rank4_video"):
                tiff_kwargs["format"] = metadata.get("format")
            return TiffVideo(
                filename,
                grayscale=grayscale,
                keep_open=keep_open,
                **tiff_kwargs,
            )
        else:
            # Single-page TIFF, treat as regular image
            return ImageVideo(
                [filename],
                grayscale=grayscale,
                **_get_valid_kwargs(ImageVideo, kwargs),
            )
    elif filename.lower().endswith(tuple(ext.lower() for ext in ImageVideo.EXTS)):
        return ImageVideo(
            [filename], grayscale=grayscale, **_get_valid_kwargs(ImageVideo, kwargs)
        )
    elif filename.lower().endswith(tuple(ext.lower() for ext in MediaVideo.EXTS)):
        return MediaVideo(
            filename,
            grayscale=grayscale,
            keep_open=keep_open,
            **_get_valid_kwargs(MediaVideo, kwargs),
        )
    elif filename.lower().endswith(tuple(ext.lower() for ext in HDF5Video.EXTS)):
        return HDF5Video(
            filename,
            dataset=dataset,
            grayscale=grayscale,
            keep_open=keep_open,
            **_get_valid_kwargs(HDF5Video, kwargs),
        )
    else:
        raise ValueError(f"Unknown video file type: {filename}")

get_frame(frame_idx)

Read a single frame from the video.

Parameters:

Name Type Description Default
frame_idx int

Index of frame to read.

required

Returns:

Type Description
ndarray

Frame as a numpy array of shape (height, width, channels) where the channels dimension is 1 for grayscale videos and 3 for color videos.

Notes

If the grayscale attribute is set to True, the channels dimension will be reduced to 1 if an RGB frame is loaded from the backend.

If the grayscale attribute is set to None, the grayscale attribute will be automatically set based on the first frame read.

See also: get_frames

Source code in sleap_io/io/video_reading.py
def get_frame(self, frame_idx: int) -> np.ndarray:
    """Read a single frame from the video.

    Args:
        frame_idx: Index of frame to read.

    Returns:
        Frame as a numpy array of shape `(height, width, channels)` where the
        `channels` dimension is 1 for grayscale videos and 3 for color videos.

    Notes:
        If the `grayscale` attribute is set to `True`, the `channels` dimension will
        be reduced to 1 if an RGB frame is loaded from the backend.

        If the `grayscale` attribute is set to `None`, the `grayscale` attribute
        will be automatically set based on the first frame read.

    See also: `get_frames`
    """
    if not self.has_frame(frame_idx):
        raise IndexError(f"Frame index {frame_idx} out of range.")

    img = self._read_frame(frame_idx)

    if self.grayscale is None:
        self.detect_grayscale(img)

    if self.grayscale:
        img = img[..., [0]]

    return img

get_frames(frame_inds)

Read a list of frames from the video.

Depending on the backend implementation, this may be faster than reading frames individually using get_frame.

Parameters:

Name Type Description Default
frame_inds list[int]

List of frame indices to read.

required

Returns:

Type Description
ndarray

Frames as a numpy array of shape (frames, height, width, channels) where channels dimension is 1 for grayscale videos and 3 for color videos.

Notes

If the grayscale attribute is set to True, the channels dimension will be reduced to 1 if an RGB frame is loaded from the backend.

If the grayscale attribute is set to None, the grayscale attribute will be automatically set based on the first frame read.

See also: get_frame

Source code in sleap_io/io/video_reading.py
def get_frames(self, frame_inds: list[int]) -> np.ndarray:
    """Read a list of frames from the video.

    Depending on the backend implementation, this may be faster than reading frames
    individually using `get_frame`.

    Args:
        frame_inds: List of frame indices to read.

    Returns:
        Frames as a numpy array of shape `(frames, height, width, channels)` where
        `channels` dimension is 1 for grayscale videos and 3 for color videos.

    Notes:
        If the `grayscale` attribute is set to `True`, the `channels` dimension will
        be reduced to 1 if an RGB frame is loaded from the backend.

        If the `grayscale` attribute is set to `None`, the `grayscale` attribute
        will be automatically set based on the first frame read.

    See also: `get_frame`
    """
    imgs = self._read_frames(frame_inds)

    if self.grayscale is None:
        self.detect_grayscale(imgs[0])

    if self.grayscale:
        imgs = imgs[..., [0]]

    return imgs

has_frame(frame_idx)

Check if a frame index is contained in the video.

Parameters:

Name Type Description Default
frame_idx int

Index of frame to check.

required

Returns:

Type Description
bool

True if the index is contained in the video, otherwise False.

Source code in sleap_io/io/video_reading.py
def has_frame(self, frame_idx: int) -> bool:
    """Check if a frame index is contained in the video.

    Args:
        frame_idx: Index of frame to check.

    Returns:
        `True` if the index is contained in the video, otherwise `False`.
    """
    return frame_idx < len(self)

read_test_frame()

Read a single frame from the video to test for grayscale.

Note

This reads the frame at index 0. This may not be appropriate if the first frame is not available in a given backend.

Source code in sleap_io/io/video_reading.py
def read_test_frame(self) -> np.ndarray:
    """Read a single frame from the video to test for grayscale.

    Note:
        This reads the frame at index 0. This may not be appropriate if the first
        frame is not available in a given backend.
    """
    return self._read_frame(0)

VideoWriter

Simple video writer using imageio and FFMPEG.

Attributes:

Name Type Description
filename

Path to output video file.

fps

Frames per second. Defaults to 30.

pixelformat

Pixel format for video. Defaults to "yuv420p".

codec

Codec to use for encoding. Defaults to "libx264".

crf

Constant rate factor to control lossiness of video. Values go from 2 to 32, with numbers in the 18 to 30 range being most common. Lower values mean less compressed/higher quality. Defaults to 25. No effect if codec is not "libx264".

preset

H264 encoding preset. Defaults to "superfast". No effect if codec is not "libx264".

keyframe_interval

Interval between keyframes in seconds. If None, uses encoder default. Lower values improve seeking but increase file size. Defaults to None.

no_audio

If True, strips audio from the output. Defaults to False.

output_params

Additional output parameters for FFMPEG. This should be a list of strings corresponding to command line arguments for FFMPEG and libx264. Use ffmpeg -h encoder=libx264 to see all options for libx264 output_params.

Notes

This class can be used as a context manager to ensure the video is properly closed after writing. For example:

with VideoWriter("output.mp4") as writer:
    for frame in frames:
        writer(frame)

Methods:

Name Description
__call__

Write a frame to the video.

__enter__

Context manager entry.

__eq__

Method generated by attrs for class VideoWriter.

__exit__

Context manager exit.

__init__

Method generated by attrs for class VideoWriter.

__replace__

Method generated by attrs for class VideoWriter.

__repr__

Method generated by attrs for class VideoWriter.

__setattr__

Method generated by attrs for class VideoWriter.

build_output_params

Build the output parameters for FFMPEG.

close

Close the video writer.

open

Open the video writer.

write_frame

Write a frame to the video.

Source code in sleap_io/io/video_writing.py
@attrs.define
class VideoWriter:
    """Simple video writer using imageio and FFMPEG.

    Attributes:
        filename: Path to output video file.
        fps: Frames per second. Defaults to 30.
        pixelformat: Pixel format for video. Defaults to "yuv420p".
        codec: Codec to use for encoding. Defaults to "libx264".
        crf: Constant rate factor to control lossiness of video. Values go from 2 to 32,
            with numbers in the 18 to 30 range being most common. Lower values mean less
            compressed/higher quality. Defaults to 25. No effect if codec is not
            "libx264".
        preset: H264 encoding preset. Defaults to "superfast". No effect if codec is not
            "libx264".
        keyframe_interval: Interval between keyframes in seconds. If None, uses encoder
            default. Lower values improve seeking but increase file size. Defaults to
            None.
        no_audio: If True, strips audio from the output. Defaults to False.
        output_params: Additional output parameters for FFMPEG. This should be a list of
            strings corresponding to command line arguments for FFMPEG and libx264. Use
            `ffmpeg -h encoder=libx264` to see all options for libx264 output_params.

    Notes:
        This class can be used as a context manager to ensure the video is properly
        closed after writing. For example:

        ```python
        with VideoWriter("output.mp4") as writer:
            for frame in frames:
                writer(frame)
        ```
    """

    filename: Path = attrs.field(converter=Path)
    fps: float = 30
    pixelformat: str = "yuv420p"
    codec: str = "libx264"
    crf: int = 25
    preset: str = "superfast"
    keyframe_interval: float | None = None
    no_audio: bool = False
    output_params: list[str] = attrs.field(factory=list)
    _writer: "imageio.plugins.ffmpeg.FfmpegFormat.Writer | None" = None

    def build_output_params(self) -> list[str]:
        """Build the output parameters for FFMPEG."""
        output_params = []
        if self.codec == "libx264":
            output_params.extend(
                [
                    "-crf",
                    str(self.crf),
                    "-preset",
                    self.preset,
                ]
            )
        # Add keyframe interval (GOP size)
        if self.keyframe_interval is not None:
            gop_size = max(1, int(self.fps * self.keyframe_interval))
            output_params.extend(["-g", str(gop_size)])
        # Strip audio if requested
        if self.no_audio:
            output_params.extend(["-an"])
        return output_params + self.output_params

    def open(self):
        """Open the video writer."""
        self.close()

        self.filename.parent.mkdir(parents=True, exist_ok=True)
        self._writer = iio_v2.get_writer(
            self.filename.as_posix(),
            format="FFMPEG",
            fps=self.fps,
            codec=self.codec,
            pixelformat=self.pixelformat,
            output_params=self.build_output_params(),
            # Disable imageio's auto-scaling for non-divisible frame sizes.
            # We handle padding manually in write_frame() to preserve coordinates.
            macro_block_size=1,
        )

    def close(self):
        """Close the video writer."""
        if self._writer is not None:
            self._writer.close()
            self._writer = None

    def write_frame(self, frame: np.ndarray):
        """Write a frame to the video.

        Args:
            frame: Frame to write to video. Should be a 2D or 3D numpy array with
                dimensions (height, width) or (height, width, channels).

        Notes:
            For libx264 codec, frames are automatically padded to dimensions divisible
            by 16 (the macro block size). Padding is only added to the bottom and right
            edges to preserve coordinate alignment.
        """
        if self._writer is None:
            self.open()

        if self.codec == "libx264":
            frame = _pad_to_macro_block(frame, macro_block_size=16)

        self._writer.append_data(frame)

    def __enter__(self):
        """Context manager entry."""
        return self

    def __exit__(
        self,
        exc_type: type[BaseException] | None,
        exc_value: BaseException | None,
        traceback: TracebackType | None,
    ) -> bool | None:
        """Context manager exit."""
        self.close()
        return False

    def __call__(self, frame: np.ndarray):
        """Write a frame to the video.

        Args:
            frame: Frame to write to video. Should be a 2D or 3D numpy array with
                dimensions (height, width) or (height, width, channels).
        """
        self.write_frame(frame)

__annotations__ = {'filename': 'Path', 'fps': 'float', 'pixelformat': 'str', 'codec': 'str', 'crf': 'int', 'preset': 'str', 'keyframe_interval': 'float | None', 'no_audio': 'bool', 'output_params': 'list[str]', '_writer': "'imageio.plugins.ffmpeg.FfmpegFormat.Writer | None'"} class-attribute

dict() -> new empty dictionary dict(mapping) -> new dictionary initialized from a mapping object's (key, value) pairs dict(iterable) -> new dictionary initialized as if via: d = {} for k, v in iterable: d[k] = v dict(**kwargs) -> new dictionary initialized with the name=value pairs in the keyword argument list. For example: dict(one=1, two=2)

__attrs_own_setattr__ = True class-attribute

Returns True when the argument is true, False otherwise. The builtins True and False are the only two instances of the class bool. The class bool is a subclass of the class int, and cannot be subclassed.

__attrs_props__ = ClassProps(is_exception=False, is_slotted=True, has_weakref_slot=True, is_frozen=False, kw_only=<KeywordOnly.NO: 'no'>, collected_fields_by_mro=True, added_init=True, added_repr=True, added_eq=True, added_ordering=False, hashability=<Hashability.UNHASHABLE: 'unhashable'>, added_match_args=True, added_str=False, added_pickling=True, on_setattr_hook=<function pipe.<locals>.wrapped_pipe at 0x7f41e68fca40>, field_transformer=None) class-attribute

Effective class properties as derived from parameters to attr.s() or define() decorators.

This is the same data structure that attrs uses internally to decide how to construct the final class.

Warning:

This feature is currently **experimental** and is not covered by our
strict backwards-compatibility guarantees.

Attributes:

Name Type Description
is_exception bool

Whether the class is treated as an exception class.

is_slotted bool

Whether the class is slotted <slotted classes>.

has_weakref_slot bool

Whether the class has a slot for weak references.

is_frozen bool

Whether the class is frozen.

kw_only KeywordOnly

Whether / how the class enforces keyword-only arguments on the __init__ method.

collected_fields_by_mro bool

Whether the class fields were collected by method resolution order. That is, correctly but unlike dataclasses.

added_init bool

Whether the class has an attrs-generated __init__ method.

added_repr bool

Whether the class has an attrs-generated __repr__ method.

added_eq bool

Whether the class has attrs-generated equality methods.

added_ordering bool

Whether the class has attrs-generated ordering methods.

hashability Hashability

How hashable <hashing> the class is.

added_match_args bool

Whether the class supports positional match <match> over its fields.

added_str bool

Whether the class has an attrs-generated __str__ method.

added_pickling bool

Whether the class has attrs-generated __getstate__ and __setstate__ methods for pickle.

on_setattr_hook Callable[[Any, Attribute[Any], Any], Any] | None

The class's __setattr__ hook.

field_transformer Callable[[Attribute[Any]], Attribute[Any]] | None

The class's field transformers <transform-fields>.

.. versionadded:: 25.4.0

__doc__ = 'Simple video writer using imageio and FFMPEG.\n\nAttributes:\n filename: Path to output video file.\n fps: Frames per second. Defaults to 30.\n pixelformat: Pixel format for video. Defaults to "yuv420p".\n codec: Codec to use for encoding. Defaults to "libx264".\n crf: Constant rate factor to control lossiness of video. Values go from 2 to 32,\n with numbers in the 18 to 30 range being most common. Lower values mean less\n compressed/higher quality. Defaults to 25. No effect if codec is not\n "libx264".\n preset: H264 encoding preset. Defaults to "superfast". No effect if codec is not\n "libx264".\n keyframe_interval: Interval between keyframes in seconds. If None, uses encoder\n default. Lower values improve seeking but increase file size. Defaults to\n None.\n no_audio: If True, strips audio from the output. Defaults to False.\n output_params: Additional output parameters for FFMPEG. This should be a list of\n strings corresponding to command line arguments for FFMPEG and libx264. Use\n `ffmpeg -h encoder=libx264` to see all options for libx264 output_params.\n\nNotes:\n This class can be used as a context manager to ensure the video is properly\n closed after writing. For example:\n\n ```python\n with VideoWriter("output.mp4") as writer:\n for frame in frames:\n writer(frame)\n ```\n' class-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

__firstlineno__ = 46 class-attribute

int([x]) -> integer int(x, base=10) -> integer

Convert a number or string to an integer, or return 0 if no arguments are given. If x is a number, return x.int(). For floating-point numbers, this truncates towards zero.

If x is not a number or if base is given, then x must be a string, bytes, or bytearray instance representing an integer literal in the given base. The literal can be preceded by '+' or '-' and be surrounded by whitespace. The base defaults to 10. Valid bases are 0 and 2-36. Base 0 means to interpret the base from the string as an integer literal.

int('0b100', base=0) 4

__match_args__ = ('filename', 'fps', 'pixelformat', 'codec', 'crf', 'preset', 'keyframe_interval', 'no_audio', 'output_params', '_writer') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__module__ = 'sleap_io.io.video_writing' class-attribute

str(object='') -> str str(bytes_or_buffer[, encoding[, errors]]) -> str

Create a new string object from the given object. If encoding or errors is specified, then the object must expose a data buffer that will be decoded using the given encoding and error handler. Otherwise, returns the result of object.str() (if defined) or repr(object). encoding defaults to 'utf-8'. errors defaults to 'strict'.

__slots__ = ('filename', 'fps', 'pixelformat', 'codec', 'crf', 'preset', 'keyframe_interval', 'no_audio', 'output_params', '_writer', '__weakref__') class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__static_attributes__ = ('_writer',) class-attribute

Built-in immutable sequence.

If no argument is given, the constructor returns an empty tuple. If iterable is specified the tuple is initialized from iterable's items.

If the argument is a tuple, the return value is the same object.

__weakref__ property

list of weak references to the object

__call__(frame)

Write a frame to the video.

Parameters:

Name Type Description Default
frame ndarray

Frame to write to video. Should be a 2D or 3D numpy array with dimensions (height, width) or (height, width, channels).

required
Source code in sleap_io/io/video_writing.py
def __call__(self, frame: np.ndarray):
    """Write a frame to the video.

    Args:
        frame: Frame to write to video. Should be a 2D or 3D numpy array with
            dimensions (height, width) or (height, width, channels).
    """
    self.write_frame(frame)

__enter__()

Context manager entry.

Source code in sleap_io/io/video_writing.py
def __enter__(self):
    """Context manager entry."""
    return self

__eq__(other)

Method generated by attrs for class VideoWriter.

Source code in sleap_io/io/video_writing.py
This preserves coordinate alignment by only adding padding to the bottom and right
edges of the frame. Without this, encoding with x264 may scale or pad symmetrically,
causing coordinate shifts.

Args:
    frame: Frame to pad. Should be a 2D or 3D numpy array with dimensions
        (height, width) or (height, width, channels).
    macro_block_size: Block size to align to. Defaults to 16 for x264.

Returns:
    Padded frame with dimensions divisible by macro_block_size, or the original
    frame if no padding is needed.
"""
h, w = frame.shape[:2]

__exit__(exc_type, exc_value, traceback)

Context manager exit.

Source code in sleap_io/io/video_writing.py
def __exit__(
    self,
    exc_type: type[BaseException] | None,
    exc_value: BaseException | None,
    traceback: TracebackType | None,
) -> bool | None:
    """Context manager exit."""
    self.close()
    return False

__init__(filename, fps=30, pixelformat='yuv420p', codec='libx264', crf=25, preset='superfast', keyframe_interval=None, no_audio=False, output_params=NOTHING, writer=None)

Method generated by attrs for class VideoWriter.

Source code in sleap_io/io/video_writing.py
# Calculate padding needed (only bottom/right)
pad_h = (macro_block_size - (h % macro_block_size)) % macro_block_size
pad_w = (macro_block_size - (w % macro_block_size)) % macro_block_size

if pad_h == 0 and pad_w == 0:
    return frame

# Pad only bottom and right
if frame.ndim == 2:
    return np.pad(frame, ((0, pad_h), (0, pad_w)), mode="constant")
else:
    return np.pad(frame, ((0, pad_h), (0, pad_w), (0, 0)), mode="constant")

__replace__(**changes)

Method generated by attrs for class VideoWriter.

__repr__()

Method generated by attrs for class VideoWriter.

Source code in sleap_io/io/video_writing.py
"""Utilities for writing videos."""

from __future__ import annotations

from pathlib import Path
from types import TracebackType

import attrs
import imageio
import imageio.v2 as iio_v2
import numpy as np


def _pad_to_macro_block(frame: np.ndarray, macro_block_size: int = 16) -> np.ndarray:
    """Pad frame to be divisible by macro_block_size, padding only bottom/right.

__setattr__(name, val)

Method generated by attrs for class VideoWriter.

build_output_params()

Build the output parameters for FFMPEG.

Source code in sleap_io/io/video_writing.py
def build_output_params(self) -> list[str]:
    """Build the output parameters for FFMPEG."""
    output_params = []
    if self.codec == "libx264":
        output_params.extend(
            [
                "-crf",
                str(self.crf),
                "-preset",
                self.preset,
            ]
        )
    # Add keyframe interval (GOP size)
    if self.keyframe_interval is not None:
        gop_size = max(1, int(self.fps * self.keyframe_interval))
        output_params.extend(["-g", str(gop_size)])
    # Strip audio if requested
    if self.no_audio:
        output_params.extend(["-an"])
    return output_params + self.output_params

close()

Close the video writer.

Source code in sleap_io/io/video_writing.py
def close(self):
    """Close the video writer."""
    if self._writer is not None:
        self._writer.close()
        self._writer = None

open()

Open the video writer.

Source code in sleap_io/io/video_writing.py
def open(self):
    """Open the video writer."""
    self.close()

    self.filename.parent.mkdir(parents=True, exist_ok=True)
    self._writer = iio_v2.get_writer(
        self.filename.as_posix(),
        format="FFMPEG",
        fps=self.fps,
        codec=self.codec,
        pixelformat=self.pixelformat,
        output_params=self.build_output_params(),
        # Disable imageio's auto-scaling for non-divisible frame sizes.
        # We handle padding manually in write_frame() to preserve coordinates.
        macro_block_size=1,
    )

write_frame(frame)

Write a frame to the video.

Parameters:

Name Type Description Default
frame ndarray

Frame to write to video. Should be a 2D or 3D numpy array with dimensions (height, width) or (height, width, channels).

required
Notes

For libx264 codec, frames are automatically padded to dimensions divisible by 16 (the macro block size). Padding is only added to the bottom and right edges to preserve coordinate alignment.

Source code in sleap_io/io/video_writing.py
def write_frame(self, frame: np.ndarray):
    """Write a frame to the video.

    Args:
        frame: Frame to write to video. Should be a 2D or 3D numpy array with
            dimensions (height, width) or (height, width, channels).

    Notes:
        For libx264 codec, frames are automatically padded to dimensions divisible
        by 16 (the macro block size). Padding is only added to the bottom and right
        edges to preserve coordinate alignment.
    """
    if self._writer is None:
        self.open()

    if self.codec == "libx264":
        frame = _pad_to_macro_block(frame, macro_block_size=16)

    self._writer.append_data(frame)

is_file_accessible(filename)

Check if a file is accessible.

Parameters:

Name Type Description Default
filename str | Path

Path to a file.

required

Returns:

Type Description
bool

True if the file is accessible, False otherwise.

Notes

This checks if the file readable by the current user by reading one byte from the file.

Source code in sleap_io/io/utils.py
def is_file_accessible(filename: str | Path) -> bool:
    """Check if a file is accessible.

    Args:
        filename: Path to a file.

    Returns:
        `True` if the file is accessible, `False` otherwise.

    Notes:
        This checks if the file readable by the current user by reading one byte from
        the file.
    """
    filename = Path(filename)
    try:
        with open(filename, "rb") as f:
            f.read(1)
        return True
    except (FileNotFoundError, PermissionError, OSError, ValueError):
        return False