Skip to content

Data model

sleap-io organizes pose tracking data into a hierarchy of containers and annotations. The core flow is: a Skeleton defines the body plan (what landmarks exist), an Instance records one animal's pose (where each landmark is), a LabeledFrame groups instances at a single video frame, and Labels ties everything together into a dataset that can be saved, loaded, and manipulated.

Overview

The data model is split into five areas, each covered on its own page:

Labels: The dataset container. Labels holds labeled frames, videos, skeletons, and tracks. LabeledFrame groups annotations for a single video frame. LabelsSet manages multiple datasets (e.g., train/val/test splits).

Video: Lazy array-like access to video data. Video wraps multiple backends (MP4, HDF5, image sequences) behind a unified interface.

Poses: The skeleton template and pose instances. A Skeleton declares landmark types (Node), connections (Edge), and symmetries. Instance and PredictedInstance store per-animal coordinates and confidence scores. Track links the same animal across frames.

3D: Multi-camera support. Camera stores calibration parameters, RecordingSession links cameras to videos, and FrameGroup/InstanceGroup pair 2D views for 3D reconstruction.

Spatial annotations: Annotation types beyond keypoints — Centroids and Boxes for detection, ROIs for vector polygons, and Segmentation (SegmentationMask, LabelImage) for pixel-level masks.

Working with annotations in frames

Spatial annotations (centroids, boxes, ROIs, masks, label images) are nested in LabeledFrame — you add them directly to a frame's annotation lists. LabeledFrame.append dispatches on the runtime type of the annotation and pushes it onto the correct per-type list — you never have to touch lf.instances, lf.bboxes, lf.centroids, lf.masks, lf.label_images, or lf.rois directly.

>>> import numpy as np
>>> import sleap_io as sio
>>> from shapely.geometry import box
>>> video = sio.Video("test.mp4", open_backend=False)
>>> lf = sio.LabeledFrame(video=video, frame_idx=0)
>>> lf.append(sio.UserBoundingBox(x1=10, y1=20, x2=50, y2=60))  # → lf.bboxes
>>> lf.append(sio.UserCentroid(x=100, y=200))                    # → lf.centroids
>>> lf.append(sio.UserSegmentationMask.from_numpy(np.zeros((8, 8), bool)))  # → lf.masks
>>> lf.append(sio.UserLabelImage.from_numpy(np.zeros((8, 8), int)))         # → lf.label_images
>>> lf.append(sio.UserROI(geometry=box(0, 0, 10, 10)))            # → lf.rois
>>> labels = sio.Labels(labeled_frames=[lf])
>>> print(len(labels.centroids), len(labels.bboxes))
1 1
>>> print(len(labels.masks), len(labels.label_images), len(labels.rois))
1 1 1

The labels.centroids, labels.bboxes, labels.masks, labels.label_images, and labels.rois properties return flattened read-only views across all frames. Static, video-level ROIs (with no frame association) live separately on Labels.static_rois — see Static vs. temporal ROIs.

Class diagram

classDiagram
    direction TB

    class Skeleton:::poses {
        +nodes
        +edges
        +symmetries
    }
    class Node:::poses {
        +str name
    }
    class Edge:::poses {
        +Node source
        +Node destination
    }
    class Symmetry:::poses {
        +Set~Node~ nodes
    }
    class Track:::poses {
        +str name
    }
    class Instance:::poses {
        +PointsArray points
        +Skeleton skeleton
        +Track track
        +Identity identity
        +float identity_score
        +Embedding identity_embedding
        +Category category
        +float category_score
        +Embedding category_embedding
    }
    class PredictedInstance:::poses {
        +float score
    }

    class Labels:::labels {
        +labeled_frames
        +videos
        +skeletons
        +tracks
    }
    class LabeledFrame:::labels {
        +Video video
        +int frame_idx
        +instances
        +centroids
        +bboxes
        +masks
        +label_images
        +rois
    }
    class SuggestionFrame:::labels {
        +Video video
        +int frame_idx
    }
    class LabelsSet:::labels {
        +labels
    }

    class Video:::video {
        +str filename
        +VideoBackend backend
    }

    class Camera:::threed {
        +ndarray matrix
        +ndarray dist
        +str name
    }
    class CameraGroup:::threed {
        +cameras
    }
    class RecordingSession:::threed {
        +CameraGroup camera_group
        +frame_groups
    }
    class FrameGroup:::threed {
        +int frame_idx
        +instance_groups
    }
    class InstanceGroup:::threed {
        +instance_by_camera
        +Instance3D instance_3d
        +Identity identity
    }
    class Identity:::threed {
        +str name
        +dict~str,str~ metadata
    }
    class Category:::labels {
        +str name
        +dict~str,str~ metadata
    }
    class Instance3D:::threed {
        +ndarray points
        +Skeleton skeleton
    }
    class PredictedInstance3D:::threed {
        +ndarray point_scores
    }

    class LabelImageWriter:::labels {
        +str filename
        +add()
    }

    class ROI:::regions {
        <<abstract>>
        +geometry
        +str name
    }
    class SegmentationMask:::regions {
        <<abstract>>
        +rle_counts
        +int height
        +int width
    }
    class BoundingBox:::regions {
        <<abstract>>
        +float x1
        +float y1
        +float x2
        +float y2
    }
    class UserBoundingBox:::regions
    class PredictedBoundingBox:::regions {
        +float score
    }

    class UserROI:::regions
    class PredictedROI:::regions {
        +float score
    }
    class UserSegmentationMask:::regions
    class PredictedSegmentationMask:::regions {
        +float score
        +ndarray score_map
    }
    class UserLabelImage:::regions
    class PredictedLabelImage:::regions {
        +float score
        +ndarray score_map
    }

    class LabelImage:::regions {
        <<abstract>>
        +ndarray data
        +dict objects
        +int n_objects
        +to_masks()
    }

    class Centroid:::regions {
        <<abstract>>
        +float x
        +float y
    }
    class UserCentroid:::regions
    class PredictedCentroid:::regions {
        +float score
    }

    Skeleton "1" *-- "1..*" Node
    Skeleton "1" *-- "0..*" Edge
    Skeleton "1" *-- "0..*" Symmetry
    Instance --> Skeleton : uses
    Instance --> Track
    Instance <|-- PredictedInstance

    Labels "1" *-- "0..*" LabeledFrame
    Labels --> Video
    Labels --> Skeleton
    Labels --> Track
    LabeledFrame "1" *-- "0..*" Instance
    LabeledFrame --> Video
    LabelsSet "1" *-- "1..*" Labels

    CameraGroup "1" *-- "0..*" Camera
    RecordingSession --> CameraGroup
    RecordingSession "1" *-- "0..*" FrameGroup
    FrameGroup "1" *-- "0..*" InstanceGroup
    InstanceGroup --> Instance
    InstanceGroup --> Camera
    InstanceGroup --> Instance3D
    InstanceGroup --> Identity
    Instance3D --> Skeleton : uses
    Instance3D <|-- PredictedInstance3D
    Labels --> Identity
    Labels --> Category
    Instance --> Category

    Centroid <|-- UserCentroid
    Centroid <|-- PredictedCentroid
    BoundingBox <|-- UserBoundingBox
    BoundingBox <|-- PredictedBoundingBox
    ROI <|-- UserROI
    ROI <|-- PredictedROI
    SegmentationMask <|-- UserSegmentationMask
    SegmentationMask <|-- PredictedSegmentationMask
    LabelImage <|-- UserLabelImage
    LabelImage <|-- PredictedLabelImage
    LabelImage --> SegmentationMask : to_masks()
    LabelImage --> BoundingBox : to_bboxes()
    LabelImageWriter --> LabelImage : streams
    LabeledFrame --> Centroid
    LabeledFrame --> BoundingBox
    LabeledFrame --> ROI
    LabeledFrame --> SegmentationMask
    LabeledFrame --> LabelImage

    classDef poses fill:#0097a7,stroke:#00796b,color:#fff
    classDef labels fill:#43a047,stroke:#2e7d32,color:#fff
    classDef video fill:#ef6c00,stroke:#e65100,color:#fff
    classDef threed fill:#7b1fa2,stroke:#6a1b9a,color:#fff
    classDef regions fill:#d32f2f,stroke:#c62828,color:#fff

Quick reference

Class Page Description
Skeleton Poses Template defining landmark types and their connections
Node Poses A single landmark type within a skeleton
Edge Poses Directed connection between two nodes
Symmetry Poses Left/right pairing between two nodes
Instance Poses One animal's pose in a single frame
PredictedInstance Poses Model-predicted pose with confidence scores
Track Poses Identity linking instances of the same animal across frames
Labels Labels Top-level dataset container
LabeledFrame Labels All instances at a specific frame of a video
SuggestionFrame Labels Frame suggested for labeling
LabelsSet Labels Named collection of Labels (e.g., train/val/test)
Video Video Video file with lazy backend loading
Camera 3D Calibrated camera with intrinsic/extrinsic parameters
CameraGroup 3D Set of cameras used together
RecordingSession 3D Multi-camera recording linking cameras to videos
FrameGroup 3D Matched labeled frames across views at one time point
InstanceGroup 3D Same animal matched across cameras, with optional 3D points
Identity 3D Cross-session persistent animal identity (distinct from per-video Track)
Category Categories Class/type a detection belongs to (e.g. female_fly), assigned by classification or re-ID
Instance3D 3D Structured triangulated 3D keypoint storage
PredictedInstance3D 3D Model-predicted 3D keypoints with per-point scores
Centroid Centroids Abstract base centroid point annotation
UserCentroid Centroids Human-annotated centroid
PredictedCentroid Centroids Model-predicted centroid with score
ROI ROIs Vector geometry annotation (polygon, etc.)
SegmentationMask Segmentation Run-length encoded pixel mask
BoundingBox Boxes Axis-aligned or rotated bounding box
UserBoundingBox Boxes Human-annotated bounding box
PredictedBoundingBox Boxes Model-predicted bounding box with score
UserROI ROIs Human-annotated region of interest
PredictedROI ROIs Model-predicted region of interest with score
UserSegmentationMask Segmentation Human-annotated segmentation mask
PredictedSegmentationMask Segmentation Model-predicted segmentation mask with score
UserLabelImage Segmentation Human-annotated label image
PredictedLabelImage Segmentation Model-predicted label image with score
LabelImage Segmentation Dense integer label image for instance segmentation
LabelImageWriter Segmentation Streaming writer for chunked label image SLP files

Hands-on examples

For practical code recipes — loading data, modifying skeletons, exporting formats, and more — see the Examples guide.