Skip to content

Changelog

Unreleased upcoming

  • sv.IconAnnotator now draws a palette icon that carries its own alpha channel (Pillow's PA mode, which a TIFF can store) instead of failing with ValueError: could not broadcast input array from shape (16,16,2) into shape (16,16,3). It reads icons with cv2.IMREAD_UNCHANGED, for which OpenCV expands any palette to BGR, but the OpenCV-free fallback backend mapped only P to color and left PA alone, returning Pillow's raw two-channel array of palette index and alpha as if it were an image. Its first channel is a lookup index rather than a color, so nothing downstream could use it; sv.draw_image rejected the same array with ValueError: Image must have 3 or 4 channels.. #2614 fixed the same two-channel PA array in sv.pillow_to_cv2, which converts an in-memory Pillow image; this is the file-reading path beside it. A palette image with alpha is now expanded to color like any other palette, matching cv2.imread. Palette images without alpha, every other mode, and every read with OpenCV installed are unchanged. (#2624)

  • sv.LineZone.trigger now ignores detections whose tracker_id is negative, the value trackers such as ByteTrackTracker from the trackers package report for tracks they have not confirmed yet. Those detections are neither counted nor given crossing state, and their entries in the returned crossed_in and crossed_out arrays are always False. Previously every unconfirmed detection was keyed under the same shared id, so several distinct unconfirmed objects on opposite sides of the line read as one track moving back and forth and silently inflated in_count and out_count: two -1 objects swapping sides each frame produced crossings without any confirmed track ever crossing. Confirmed tracks are counted exactly as before. (#2623)

  • sv.polygon_to_mask now accepts a Python list, tuple, or array-like of [x, y] vertices, returning an all-zero mask for a polygon with no vertices or fewer than MIN_POLYGON_POINT_COUNT (3). List inputs used to raise AttributeError on .astype, empty polygons failed inside OpenCV fillPoly, and a 1- or 2-vertex polygon silently drew a stray pixel or a bare line instead of an empty mask. A malformed polygon (wrong shape, non-numeric dtype, or ragged nested sequence) now raises a ValueError naming the problem instead of an opaque OpenCV or NumPy error. sv.Detections.from_sam3 draws its 2-vertex PVS fragment lines directly with OpenCV instead of through polygon_to_mask, so that behavior (see the #2625 entry below) is unaffected by the new all-zero guard. (#2622)

  • sv.Detections.from_sam3 no longer drops SAM 3's PVS-format contour fragments that have fewer than 3 vertices (a single point or a 2-point edge). from_sam3 now preserves single points directly and rasterizes 2-point edges as lines so all such fragment pixels contribute to the mask and the bounding box. Polygons with 3 or more vertices are unaffected. (#2625)

  • sv.DetectionDataset.from_yolo now loads label files whose rows carry a trailing confidence or tracker id (e.g. 1 0.5 0.5 0.2 0.4 0.87), which previously aborted the whole load with ValueError: cannot reshape array of size 5 into shape (2). Ultralytics save_txt writes that sixth column when save_conf=True or when tracking is enabled. The extra token is ignored; the box is the same as the five-column form. Segmentation lines (class id plus at least three xy pairs) and OBB four-corner lines are unchanged. (#2619)

  • sv.pillow_to_cv2, and with it every annotator, sv.crop_image, sv.resize_image, sv.letterbox_image, sv.scale_image, sv.tint_image, sv.grayscale_image and sv.plot_image handed a Pillow image, now converts every Pillow mode to the 8-bit grayscale or BGR array cv2.imread would produce for the same picture. It used to pass the raw mode bytes through: a 1-bit image came back as 0 and 1 instead of 0 and 255, a 16-bit (I;16) or 32-bit integer (I) image wrapped modulo 256 so a bright depth map drew as noise, a grayscale image with alpha (LA, PA) crashed in cvtColor with a two-channel array, and a CMYK JPEG had its cyan, magenta and yellow ink values drawn as red, green and blue. Alpha is dropped as cv2.imread drops it, 1-bit becomes 0/255, a 16-bit image keeps its high byte as OpenCV does when it reads a 16-bit PNG as 8-bit (a 32-bit integer image is clipped to the 16-bit range first), and CMYK, YCbCr, HSV and padded RGB images are converted to color through Pillow. RGB, RGBA, grayscale and palette images convert exactly as before. sv.tint_image and sv.grayscale_image also accept a single-channel (H, W) array or grayscale Image, which sv.letterbox_image already did; tint_image returns the tinted scene in color. (#2614)

  • sv.DetectionDataset.split, sv.ClassificationDataset.split and the train_test_split helper now raise ValueError for a split ratio outside the inclusive range [0, 1], including NaN and ±inf. The ratio was never validated: the helper computed int(len(data) * ratio) and sliced with it, so a finite out-of-range value looked like a successful split — for a 10-image dataset split_ratio=-0.2 returned an 8/2 partition (Python's negative slicing), and split_ratio=1.2 or an accidental percentage such as 80 returned 10/0, leaving no held-out data without any warning. NaN and ±inf already failed, but with int() conversion errors rather than a message naming the problem. The check runs before any shuffling or slicing, so an invalid ratio is rejected on an empty dataset too. 0 and 1 remain valid and keep their meaning: 0 sends every image to the test set, 1 sends every image to the train set. Splits with a ratio inside [0, 1] are unchanged. (#2611)

  • sv.ConfusionMatrix now supports MetricTarget.MASKS, so an instance segmentation model can be scored on the shapes it predicts rather than on their bounding boxes. from_detections and benchmark accept metric_target=MetricTarget.MASKS and match predictions to targets by mask IoU, computed with mask_iou_batch on dense (N, H, W) arrays or CompactMask; every other metric already offered this target, and ConfusionMatrix was the one that raised MetricTarget.MASKS is not currently supported. Two instances that share a box but not a shape are no longer counted as a match, and the validation grids written by benchmark(save_directory_path=...) fill each mask so the panels show the geometry the outcome was decided on. A non-empty Detections without mask, or predictions and targets whose masks differ in resolution, raise ValueError; Detections.empty() needs no masks. from_tensors, evaluate_detection_batch and detections_to_tensor keep rejecting masks, which have no tensor row layout, and now point at from_detections. from_detections also raises when predictions and targets differ in length instead of silently scoring the shorter list. (#2612)

  • sv.CSVSink now writes UTF-8 on every platform, preserving non-English detection labels and custom fields on Windows systems using a legacy default encoding. (#2615)

  • sv.Detections.with_nmm and sv.mask_non_max_merge now form CompactMask unions directly from run-length encoded foreground intervals. Both greedy candidate updates and final output merging avoid full-image mask allocations, including when used by InferenceSlicer(compact_masks=True). Stored foreground, exact overlap matching and detection metadata are preserved. Merging compact masks with different image shapes still raises ValueError, but the message now includes both shapes: Cannot merge CompactMask objects with different image shapes: {a} vs {b} (previously Cannot merge CompactMask objects with different image shapes.). Union cost depends on foreground column interval count (mask fragmentation and crossed columns), not logical canvas area; highly fragmented masks (a cheap O(masks) run-count check flags them, no decode needed) now union via a bbox-local dense OR instead of sorting intervals, bounded to a 16 Mi-pixel bbox so it never re-introduces a full-canvas allocation. For twelve 200×200 checkerboard masks on a 512×512 canvas (four groups of three duplicates), measured peak traced NMM allocation is 4.34 MiB (maximum of three repeats, input construction excluded), matching the pre-#2606 dense-union baseline rather than the 7.79 MiB (1.79×) a pure interval union costs there. A single mask larger than the 16 Mi-pixel bbox cap still takes the interval path regardless of fragmentation, and can still cost more time and temporary memory than a dense union would. Overlap evaluation still decodes overlapping crops. (#2606, #2616)

  • Removed, as scheduled for supervision-0.31.0: sv.ByteTrack (use ByteTrackTracker from the trackers package instead); the supervision.keypoint module (use supervision.key_points); create_tiles and overlay_image in supervision.utils.image; ensure_cv2_image_for_annotation, ensure_pil_image_for_annotation, and ensure_cv2_image_for_processing in supervision.utils.conversion; validate_keypoint_confidence and validate_keypoints_fields in supervision.validators; the normalized_xyxy argument of sv.denormalize_boxes (use xyxy); the supervision.dataset.utils import path for sv.mask_to_rle/sv.rle_to_mask (import from supervision.detection.utils.converters instead); sv.LMM and Detections.from_lmm (use sv.VLM/Detections.from_vlm); and the legacy MeanAveragePrecision in supervision.metrics.detection (use supervision.metrics.mean_average_precision.MeanAveragePrecision, exposed as sv.metrics.MeanAveragePrecision). See Deprecated for the full list. #2582

  • sv.Detections.from_vlm with sv.VLM.GOOGLE_GEMINI_2_5 or sv.VLM.GOOGLE_GEMINI_3_5 now keeps a mask pixel only where Gemini's mask scores it above 127 of 255. Gemini returns each mask as a PNG probability map with values from 0 to 255, and Google's segmentation guide binarizes it at 127, but the parser kept every pixel above 0. Any pixel the model gave even a 1/255 chance therefore joined the mask, and so did every partly covered pixel that the bilinear resize to the box blends along each edge, which grows every mask by half of one mask pixel, scaled to the box, on each side: a 4×4 mask whose middle 2×2 is set, drawn into a 40×40 box, covered 900 pixels instead of 360. Masks that are only 0 and 255 at the size of their box load as before. #2589

  • Added MetricResult abstract base class as a common parent for all metric result dataclasses, with to_pandas(), plot(), and _get_plot_details() abstract methods (#2498).

  • Added aggregate_metric_results() to combine multiple metric results into a single pd.DataFrame for model comparison (#2498).

  • Added plot_aggregate_metric_results() to visualize multiple metric results on a single grouped bar chart (#2498).

  • sv.match_detections now exposes the metrics matcher as a public primitive: given two sv.Detections, it returns one-to-one greedy IoU-matched pairs (with optional per-class filtering) as index arrays, plus the unmatched indices of each side (#2476).

  • sv.VLM.KOSMOS_2 — sv.Detections.from_vlm now parses Kosmos-2 grounding results, the (caption, entities) pair returned by the model's AutoProcessor.post_process_generation. A phrase that grounds to several regions yields one detection per region, and classes filters the result and assigns class_id by index into that list, matching the other VLM connectors. Kosmos-2 is available through sv.VLM only, not through the deprecated sv.LMM (#1903).

  • sv.Detections now validates xyxy boxes for finite numeric coordinates (no NaN/inf), raising a clear ValueError for non-finite or unsupported-dtype values instead of failing silently downstream (#2503).

  • sv.VLM.GOOGLE_GEMINI_3_6 and sv.VLM.GOOGLE_GEMINI_3_7 — sv.Detections.from_vlm now parses the structured {"boxes": [...]} detection and segmentation format, including normalized polygon masks (#2504).

  • sv.Detections.from_vlm now keeps class_id integer-typed when a classes filter removes every detection. The index array was built from an empty list, so NumPy defaulted it to float64 for sv.VLM.PALIGEMMA, sv.VLM.DEEPSEEK_VL_2 and sv.VLM.GOOGLE_GEMINI_2_0 — sv.VLM.QWEN_2_5_VL and sv.VLM.QWEN_3_VL already pinned the dtype and the rest now match them. Results with at least one surviving detection are unchanged.

0.30.5 Sep 22, 2026

  • sv.InferenceSlicer no longer raises when slices disagree on metadata (e.g. a source_image NumPy array attached per-slice by RF-DETR/inference-package connectors), which previously crashed the merge outright. A mismatched key is now dropped from the merged result with a SupervisionWarnings warning naming it; source_image is a special case, reattached afterward as the full input image rather than a single slice's tile. Also fixes two related Detections.__eq__ bugs: metadata holding NaN in a float array now compares equal to itself instead of always reading unequal, and comparing a list-valued value against an ndarray-valued one now returns False instead of raising ValueError. (#2596)

  • sv.VideoInfo.from_video_path, sv.get_video_frames_generator and sv.process_video now turn a rotated video upright when OpenCV is not installed, as they already do with OpenCV. Phones store portrait clips as landscape frames with the turn recorded in the container's display matrix; the OpenCV-free fallback ignored it, so a portrait video loaded sideways with width and height swapped — models ran on sideways frames and sv.VideoSink saved the output sideways. #2601

  • sv.ImageSink and the dataset exports that encode in-memory images now write JPEG and WebP at OpenCV's default quality (JPEG 95, lossless WebP) when OpenCV is not installed, instead of Pillow's defaults (JPEG 75, lossy WebP) — a .jpg came out with visibly stronger compression artifacts (a third of the size for one test frame), and a .webp no longer held the frame's exact pixels. A .webp written by the fallback is now larger than before, since lossless output is larger than the previous lossy default. The fallback's in-memory encoder also now accepts .jpe as a JPEG alias, and .tif/.jp2/.pgm, all four previously rejected with False, None even though file writes already accepted them. #2592

  • sv.IconAnnotator and sv.draw_image now draw CMYK JPEG and TIFF images in their real colors when OpenCV is not installed. The fallback returned the four ink channels — cyan, magenta, yellow, black — as if they were blue/green/red/alpha, so an icon drew in the wrong colors, and a CMYK image with no black ink didn't draw at all. #2602

  • sv.draw_image now scales a 16-bit PNG down to 8 bits on load instead of blending it straight into the 8-bit scene, which saturated every nonzero channel to 255: a dark gray logo drew white, and a red of 200 drew as 255. Eight-bit images, and images passed as arrays, are unaffected. #2603

  • sv.metrics.MeanAverageRecall now scores mAR@K from each image's K most confident predictions alone, instead of matching every prediction first and only then keeping the top K — the matcher pairs by highest IoU, not confidence, so a prediction ranked below K could steal a target from one ranked within it. One target with a top prediction at IoU 0.71 scored mAR@1 0.5 alone but 0.0 once a second, lower-ranked prediction at IoU 1.0 was added; mAR@1/mAR@10 now equal the recall of each image's own top-1/top-10 predictions, as documented. #2604

  • sv.KeyPoints.as_detections now leaves key points marked not visible out of the box it derives for each skeleton, as sv.KeyPoints.with_nms already does — a pose whose low-confidence joints were predicted outside the person (e.g. legs below the bottom of the frame) previously got a box reaching out to them. A skeleton with no visible key point is now dropped, as one with only missing key points already was. #2605

  • sv.IconAnnotator now draws grayscale PNG icons instead of failing with IndexError: tuple index out of range — a grayscale PNG without alpha reads as a 2-D array, which the overlay (BGR/BGRA only) can't handle; PNG optimizers like optipng/oxipng convert all-gray RGB icons to grayscale, so this hit real assets. A 16-bit icon is now scaled to 8 bits on load instead of wrapping modulo 256 (a value of 60000 previously drew as 96, near black); an icon that's neither 8-bit nor 16-bit now raises ValueError naming the file. #2591

  • sv.IconAnnotator and sv.draw_image now keep the transparency of grayscale PNGs with alpha, RGB PNGs with a transparent color, and 1-bit PNGs, when OpenCV is not installed. The fallback returned Pillow's pixel layout instead of OpenCV's for IMREAD_UNCHANGED, so a grayscale-with-alpha icon failed with ValueError: could not broadcast input array from shape (16,16,2) into shape (16,16,3), sv.draw_image failed with ValueError: Image must have 3 or 4 channels., and a transparent-color RGB icon pasted with its transparent pixels drawn black. #2588

  • sv.plot_images_grid now plots a single image in a grid_size=(1, 1) grid instead of failing with AttributeError: 'Axes' object has no attribute 'flat' — plt.subplots returns a lone Axes rather than an array for a 1x1 grid, which hit any code sizing the grid from the image count (e.g. grid_size=(1, len(images))) whenever there was only one image. #2590

  • sv.LineZone.trigger no longer counts a spurious crossing in the opposite direction when a tracker flickers to the far side of the line for fewer than minimum_crossing_threshold frames. With minimum_crossing_threshold=2, the side sequence A,A,A,B,A,A,A previously counted a crossing into A — the side the object never left — so three separated flickers gave in_count=3 instead of 0. Crossings are now measured against the last side a tracker was confirmed on rather than the oldest history entry; sustained crossings and minimum_crossing_threshold=1 (the default) are unchanged. #2600

0.30.4 Sep 17, 2026

  • sv.DetectionDataset.from_coco now rounds polygon vertices to the nearest pixel before rasterising masks, instead of truncating them to int32, as from_yolo/from_labelme/from_pascal_voc already do. Truncation shifted every mask up and to the left by up to a pixel — a polygon with corners at 2.6/7.6 covered pixels 2 to 7 from COCO but 3 to 8 from LabelMe, an IoU of 0.53 between the two masks. A non-finite vertex now raises ValueError naming the annotation id; integer vertices and RLE masks load as before. #2587

  • sv.TraceAnnotator.annotate no longer raises ValueError: Length of color lookup 3 does not match length of detections 2 when a custom_color_lookup is passed alongside pending (tracker_id == -1) tracks — the annotator skips those detections but still resolved colors against the full-length lookup, although the other annotators accept the same detections and lookup. The lookup is now filtered alongside the detections, so each confirmed track keeps its own color; frames without pending tracks, and calls without custom_color_lookup, are unchanged. #2586

  • sv.DetectionDataset.from_coco/from_labelme/from_createml/from_yolo now read JSON/YAML as UTF-8, and as_pascal_voc writes XML as UTF-8, instead of the platform default (cp1252 on Windows). A COCO category or YOLO data.yaml name outside ASCII broke on Windows: café loaded as café, 고양이 failed with UnicodeDecodeError, and as_pascal_voc either failed with UnicodeEncodeError or wrote a file from_pascal_voc rejected with ParseError: not well-formed (invalid token) — Linux/macOS, already UTF-8 by default, are unaffected. #2585

  • sv.ClassificationDataset.as_folder_structure now copies source image files unchanged instead of round-tripping them through cv2.imread/cv2.imwrite, which dropped alpha channels, downcast 16-bit PNGs to 8-bit, and recompressed JPEGs on every export — matching sv.DetectionDataset's exports (as_yolo/as_pascal_voc/as_coco), which already copied — exporting into the folder the dataset was loaded from previously rewrote its own images the same lossy way. Images held in memory are still encoded with cv2.imwrite. #2584

  • sv.get_video_frames_generator no longer reads start frames past end when iterative_seek=True — it counted start down to zero, then measured end from that zero instead of from start (start=2, end=5 yielded frames 2 to 6 instead of 2 to 4; start=4, end=6 yielded frames 4 to 9 instead of 4 to 6). A separate counter now keeps both seek modes aligned; start=0 and non-iterative calls are unchanged. #2583

  • sv.DetectionDataset.from_labelme now finds images for LabelMe files saved on Windows when loading on Linux or macOS. imagePath is written with \, but the loader took only Path(...).name, which doesn't split on \ on POSIX — the whole value became the filename and reading failed with ValueError: Could not read image from path. The filename is now taken with either separator on every system, as LabelMe itself does when reading its own files; forward-slash paths, and every file on Windows, are unchanged. #2581

  • sv.DetectionDataset.from_yolo now loads label files whose class ids are written as decimals (e.g. 1.0 0.5 0.5 0.2 0.4), which previously aborted the whole load with ValueError: invalid literal for int() — np.savetxt writes floats by default and Ultralytics tolerates them. Whole numbers load in any notation now; fractional, non-finite, or non-numeric ids still raise, naming the offending id. #2580

  • sv.DetectionDataset.from_yolo/as_coco now size EXIF-oriented images the way cv2.imread loads them (swapping width/height for orientations 5 to 8), instead of reading the un-rotated file-header size via Pillow — quarter-turned photos previously scaled boxes/polygons by the swapped dimensions and produced mismatched mask shapes. The OpenCV-free fallback backend's imread/imdecode now apply EXIF orientation too, matching OpenCV's behavior for every read except IMREAD_UNCHANGED — previously the same file loaded with a different shape depending on whether opencv-python was installed. #2577

  • sv.Detections.from_transformers now loads Transformers v5 return_binary_maps=True instance results (a (num_instances, H, W) stack), which the v5 path previously compared against each segment's id as if it were an id-map, producing a 4-D array that mask_to_xyxy rejected with ValueError: too many values to unpack (expected 3). Each segment now indexes the stack at its own id, keeping full masks for overlapping instances; segment-id-map results are unchanged. #2576

  • sv.KeyPoints.from_inference now places each key point at the slot given by its class_id (skeleton index) instead of appending in received order — Inference drops key points below keypoint_confidence and multi-skeleton models report different counts per object, so stacking as-received either raised ValueError: ... inhomogeneous shape or silently slid later key points into earlier slots, joining the wrong joints. Omitted slots now stay (0, 0) at zero confidence, already skipped by the key point annotators and as_detections; a result with every key point omitted now loads with zero key points instead of failing validation. #2575

  • sv.DetectionDataset.from_pascal_voc no longer fails on annotations with decimal coordinates (e.g. <xmin>48.5</xmin>), which aborted the whole load with ValueError: invalid literal for int() — Datumaro, which CVAT uses for its exports, writes VOC this way. Box coordinates are now read as floats and keep their precision; polygon vertices are rounded after the 1-index offset, as the YOLO and LabelMe loaders already do; non-finite values are still rejected. #2568

  • sv.Detections.from_transformers no longer crashes on a Transformers v4 panoptic result with no segments — post_process_panoptic's empty segments_info produced a (0,) mask array instead of (0, H, W), and mask_to_xyxy raised ValueError: not enough values to unpack (expected 3, got 1). The path now builds a (0, H, W) mask stack and an integer class_id, matching the v5 paths, and yields empty Detections. #2571

  • sv.KeyPoints.from_ultralytics no longer crashes on pose models whose key points carry no visibility score (kpt_shape=[K,2], where Results.keypoints.conf is None) — the connector called .cpu() on it unconditionally, raising AttributeError: 'NoneType' object has no attribute 'cpu' on every non-empty frame. Such results now load with keypoint_confidence=None; models that do report visibility are unaffected. #2570

  • sv.DetectionDataset.from_pascal_voc no longer skips .bmp, .tif, .tiff, and .webp images without a warning — the loader only listed .jpg/.jpeg/.png, even though from_yolo/from_folder_structure load those formats and as_pascal_voc writes annotations for them, so a dataset exported to Pascal VOC and read back came back smaller than it went out. It now accepts the same extensions as sv.ClassificationDataset.from_folder_structure. #2569

0.30.3 Sep 14, 2026

  • sv.Detections.from_ultralytics now assigns the placeholder class ID 0 to every mask in a masks-only result. Previously the placeholder IDs were sequential (0, 1, 2, ...) despite every mask belonging to the same image. Results that carry boxes keep their real class IDs and are unaffected (#2566).

  • sv.pad_boxes now computes integer-coordinate padding without overflow or unsigned casting errors. Integer inputs are promoted to int64, with a float64 fallback only when padded coordinates exceed its range; floating-point inputs retain their dtype. Padding can extend boxes below zero or above the input dtype's maximum without wrapping their coordinates, while ordinary integer boxes remain compatible with annotation renderers (#2565).

  • sv.scale_image and sv.resize_image(keep_aspect_ratio=True) no longer derive a zero-sized target. Both compute the output size from the input and truncate it with int(), so a small enough factor — or an aspect ratio too extreme for the target box — rounded an axis down to 0, and cv2.resize answered with error: (-215:Assertion failed) inv_scale_x > 0, an assertion that names nothing the caller passed. Each axis now keeps at least one pixel. This also reaches two callers: sv.letterbox_image could not fill the resolution it was asked for (a 1200 x 8 strip into (100, 100)), and sv.CropAnnotator with a scale_factor below 1 aborted the whole frame as soon as one detection box was a few pixels across. Sizes that did not round to zero are unchanged (#2564).

  • sv.KeyPoints.with_nms no longer stops suppressing as soon as a key point is not finite. The box it feeds to NMS is derived from the key points a skeleton considers valid, but the validity test was xy == 0 alone, and NaN is not 0 — so an undetected joint, which pose estimators report as NaN rather than dropping, stayed in the box, np.min/np.max propagated it into all four corners, and every IoU comparison against that all-NaN box was False. The duplicate skeleton therefore survived and the caller got two poses drawn on top of each other, with no error to trace it back to. Non-finite coordinates are now excluded, matching sv.KeyPoints.as_detections and the key point annotators. A skeleton the filter empties — every key point zero, non-finite, or marked invisible — now carries a zero-area box instead of the ±inf sentinels used to fill the reduction, and still passes through untouched (#2563).

  • sv.tint_image no longer tints the image it was given. It blended into its image argument via OpenCV's dst, so a NumPy input came back tinted for the caller too and np.hstack([frame, sv.tint_image(frame, ...)]) produced two tinted halves. A Pillow input was already unaffected, because the ndarray conversion shielded it, so the same call had different aliasing depending on the input type. The blend now writes to its own buffer for both, matching sv.resize_image, sv.letterbox_image, sv.scale_image and sv.grayscale_image (#2562).

  • sv.LineZone no longer consumes the triggering_anchors iterable during validation. The parameter is typed Iterable[Position], but the emptiness check ran list() over the caller's object and stored the original, so a generator or map — the natural way to build anchors from a config file — was left exhausted and the first trigger() call raised ValueError: operands could not be broadcast together with shapes (0,) (2,). The anchors are now materialized once, matching sv.PolygonZone. LineZone.triggering_anchors is therefore always a fresh list, distinct from whatever object the caller passed in — mutating the caller's own list after construction no longer affects the zone, and the default's observable type is list, not tuple (#2561).

  • The key point annotators — sv.VertexAnnotator, sv.EdgeAnnotator, sv.VertexLabelAnnotator and the sv.VertexEllipse*Annotator family — now skip key points whose coordinates are not finite instead of raising ValueError: cannot convert float NaN to integer. Pose estimators commonly report an undetected or occluded key point as NaN rather than dropping it, so a single missing joint aborted the whole frame. sv.KeyPoints.as_detections already treats non-finite coordinates as missing; the annotators now apply the same rule, drawing every other key point in the skeleton. For 3-component key points (xy shape (N, K, 3)), the check applies to the full row, so a key point with a finite, drawable (x, y) but a non-finite third component is skipped too, for parity with sv.KeyPoints.as_detections (#2560).

  • sv.ClassificationDataset.as_folder_structure now rejects images that would overwrite the same class-relative filename before writing any files. Identical basenames in different class directories remain supported (#2551).

  • sv.filter_polygons_by_area and sv.approximate_polygon now preserve local geometry for large-origin integer and float64 polygons instead of losing coordinate deltas during OpenCV conversion (#2542).

  • sv.process_video no longer hangs forever when max_frames is larger than the number of frames in the video. The reader thread used to fail on the out-of-range end before enqueuing its sentinel, leaving the main loop blocked on the read queue. max_frames is now capped at the video length so the whole video is processed, and any error raised inside the reader thread is surfaced as RuntimeError("Reader thread raised: ...") from the original exception instead of stalling the call (#2545).

  • sv.PolygonZone now rejects a polygon with fewer than three vertices, or one that is not of shape (N, 2), instead of building a zone that can never trigger. One or two vertices enclose no area, so the zone mask came out as a single pixel or a bare line and every detection tested against it read as outside — indistinguishable from a correctly configured zone that simply saw nothing. Zero vertices raised zero-size array to reduction operation maximum from NumPy rather than naming the problem. The three-vertex minimum matches MIN_POLYGON_POINT_COUNT, which sv.mask_to_polygons already enforces when producing polygons (#2554).

  • sv.Detections.from_vlm now orders each parsed box's corners, so a model that emits a corner pair backwards no longer produces an xyxy row with x_min > x_max. Every VLM parser passed such a row straight through, and nothing downstream caught it: sv.box_iou_batch clamps intersection widths at zero, so the box scored an IoU of 0.0 against itself — surviving NMS as a duplicate and counting as a total miss in mAP — while sv.Detections.box_area reported a plausible positive value, because negating both sides leaves their product positive. Correctly ordered boxes, their dtypes included, are unchanged (#2554).

  • sv.TraceAnnotator.annotate no longer raises ValueError: Thetracker_idfield is missing for a frame in which nothing was detected. An empty sv.Detections carries no tracker_id, so the documented per-frame annotate loop crashed on the first empty frame for every tracker except sv.ByteTrack, which works around it by returning an empty tracker_id array. Such a frame now draws nothing and still advances sv.Trace's frame counter, keeping trace_length a window over elapsed frames — pruning fires on the next frame that does carry detections, and only once the frames stored in the trace outnumber trace_length, so a track that had filled the window before a long gap starts a fresh trail instead of being joined to its pre-gap one. Detections that do contain boxes but no tracker_id still raise, as before (#2539).

  • sv.CSVSink no longer lets a batch with no detections fix the CSV header. Appending an empty sv.Detections — the normal result for a frame in which nothing was detected — wrote a header without the data and custom_data columns, and every later row was then silently truncated to that schema, dropping fields such as class_name for the whole file. Empty batches now write nothing and leave the header to the first batch that actually carries detections; a run in which every batch is empty still writes the header alone, and the spurious "Field names do not match the header" warning those batches logged is gone (#2539).

  • sv.scale_boxes now calculates box centers and scaled dimensions using overflow-safe arithmetic, preventing integer overflow and coordinate wrap-around for integer-coordinate bounding boxes (e.g. large int32 or uint16 coordinates) (#2540).

  • sv.scale_boxes now preserves exact integer intermediates until its final float64 conversion for 64-bit coordinates beyond 2**53, preventing scaled-corner rounding errors while retaining a vectorized fast path for exactly representable integer coordinates (#2541).

0.30.2 Sep 3, 2026

  • sv.xcycwh_to_xyxy no longer truncates coordinates for integer input arrays. Half of an odd width or height is fractional, and the previous implementation wrote those values into a copy of the integer input, silently rounding them; the converted boxes are now exact.

  • sv.denormalize_boxes no longer truncates coordinates for integer input arrays. Scaling now multiplies by a floating-point factor so integer normalized coordinates (for example VLM boxes quantized to 0..1000) map to exact absolute pixel values instead of being silently rounded down.

  • sv.Detections.box_area (and therefore sv.Detections.area for axis-aligned boxes) now computes integer-coordinate box areas in float64, preventing integer overflow for large boxes (e.g. an int32 50000 x 50000 box previously wrapped to a negative area).

  • sv.InferenceSlicer now merges slice results in source order when thread_workers > 1, restoring the ordering guarantee its docstring documents. Results were previously collected in thread-completion order, so the row order of the returned Detections — and, for tied confidences, which overlapping box survived with_nms/with_nmm — varied between runs on identical input. Both the per-slice and batch_size > 1 paths are affected (#2517).

0.30.1 Aug 24, 2026

  • RF-DETR example scripts (rfdetr_example.py) added to the count_people_in_zone, heatmap_and_track, speed_estimation, tracking, and traffic_analysis bundled examples (#2497).

  • sv.box_iou now raises TypeError for complex-valued box coordinates instead of silently discarding the imaginary part (#2485).

  • Performance: DetectionsSmoother.update_with_detections now checks active tracker IDs via set membership instead of scanning per tracked object (#2496). No output changes.

  • RF-DETR speed estimation now measures elapsed source-frame intervals, including gaps when tracked detections are temporarily missed.

  • sv.get_polygon_center now calculates polygon centroids in translated float64 coordinates, preventing integer overflow and precision loss for realistic-magnitude large-coordinate polygons.

  • sv.Detections.area and sv.oriented_box_iou_batch now translate oriented-box coordinates to local origins before floating-point area/intersection math, preserving differences representable by the input dtype and preventing self-IoU collapse for large-coordinate inputs (e.g. geospatial or stitched frames).

  • sv.Detections.with_nmm now translates oriented-box corners to a local origin with exact integer arithmetic before merging, preventing unsigned-integer wrap-around (e.g. uint16/uint64 coordinates) from corrupting both the merged extent and the winner's orientation angle.

  • DetectionsSmoother now keeps oriented-box corners aligned with smoothed xyxy geometry, including rotated tracks and mixed metadata windows.

  • sv.box_iou now calculates overlap in float64, preventing int32 area overflow for large boxes. For realistic coordinate magnitudes (below 2^53), its scalar result now matches sv.box_iou_batch; box_iou_batch still casts to float64 before subtracting, so the two can diverge above that threshold.

  • sv.list_files_with_extensions no longer includes directories when listing all files without an extension filter.

  • sv.pillow_to_cv2 now accepts RGBA images when the cv2-free fallback backend is active, matching OpenCV by dropping alpha and returning BGR channels.

  • import supervision no longer loads PyAV's native libraries when the OpenCV backend is selected. PyAV is now imported lazily on first use by the PyAV-backed video/audio fallback, preventing a duplicate libavdevice warning (and possible crash) on macOS when both av and opencv-python are installed.

0.30.0 Aug 4, 2026

Python 3.9 Support Terminated

With the upcoming supervision-0.30.0 release, we are terminating official support for Python 3.9, which reached end-of-life in October 2025. The minimum supported Python version is now 3.10.

Users on Python 3.9 should upgrade their environment before updating supervision.

  • sv.load_image_from_url — load an image from an HTTP(S) URL as an OpenCV image, with optional on-disk caching under the shared supervision cache directory ({tmpdir}/supervision/image-url/ by default, configurable via cache_dir) (#2372)

  • sv.VLM.GOOGLE_GEMINI_3_5 — sv.Detections.from_vlm now parses Google Gemini 3.5 output (detection and segmentation), reusing the Gemini 2.5 JSON format (box_2d + label, optional mask/confidence).

  • sv.get_video_frames_generator now accepts prefetch: int = 0 (#2273). When > 0, frames are decoded on a background daemon thread and buffered in a bounded queue, overlapping I/O with consumer processing. Default 0 preserves the existing synchronous behaviour.

  • cv2-free fallback — a private _cv2 backend (NumPy and Pillow, PyAV for video) reimplements the OpenCV operations supervision needs, so the library now runs on opencv-python-headless or without any OpenCV wheel installed. OpenCV remains the primary backend when available; see Changed below for the new av>=14.2 dependency this introduces.

    • Backend facade and dispatch mechanism between OpenCV and the fallback (#2430)
    • Image fallback — borders, resize, BGR↔gray conversion, and distance transform match OpenCV's fixed-point and half-pixel bilinear arithmetic exactly (#2431)
    • Geometry fallback — approxPolyDP follows OpenCV's stack traversal and closed-contour anchor selection; connected-component statistics are vectorized (#2432)
    • Drawing fallback — copyMakeBorder fills only channel 0 for a scalar border value, matching Scalar(v) semantics; addWeighted raises ValueError for a non-default dtype (#2433)
    • Text fallback — Hershey stroke fonts (#2435), later replaced by Pillow and the bundled DejaVu Sans face, dropping the 143 KB glyph table (#2440); getTextSize height/baseline padding is derived from the same stroke_width putText renders with (#2441)
    • Video fallback — file capture, writing, frame seeking, metadata, and process_video(preserve_audio=True) audio remuxing via PyAV's stream-template API (#2438)
    • Final fallback integration; removed unused raw distanceTransform, getRotationMatrix2D, and warpAffine compatibility symbols now that their only consumers use the smaller domain operations above (#2439)
  • Added: sv.ImageWindow — tkinter + Pillow desktop window that replaces cv2.imshow / cv2.waitKey, usable regardless of which OpenCV wheel (or none) is installed. Key differences from cv2:

    • wait_key() returns a tkinter keysym str (e.g. "q", "Escape") or None, not an int — update key == ord("q") to key == "q".
    • Mouse callback signature is (x: int, y: int, event_type: str) where event_type is "down", "up", or "move" — incompatible with cv2's (event, x, y, flags, param).
    • Only left-button events are captured; scroll, right-button, and modifier flags have no equivalent.
    • Requires python3-tk (not pip-installable): sudo apt-get install python3-tk on Debian/Ubuntu, brew install tcl-tk on macOS with Homebrew/pyenv.
  • KeyPoints.merge — combine a list of KeyPoints objects into one, mirroring Detections.merge. Empty inputs are ignored; all non-empty inputs must share the same number of keypoints per skeleton. Completes the merge-then-suppress workflow introduced by KeyPoints.with_nms (#2412)

  • BaseAnnotator.requires_mask — class-level bool flag on all annotators; True for MaskAnnotator, PolygonAnnotator, and HaloAnnotator; False for all others. Integrations can inspect this before materializing expensive mask payloads (#2370)

  • CompactMask.from_coco_rle — efficient COCO RLE ingestion into crop-scoped compact mask format without materializing dense (N, H, W) arrays (#2367)

  • Detections.from_inference(compact_masks=True) — opt-in compact mask representation for Roboflow/Inference segmentation results; masks are cropped to detector bounding boxes (#2367)

  • CompactMask.image_shape — new public property returning (H, W) of the full image the mask is scoped to (#2383)

  • sv.mask_to_roi — explicit exclusive mask-bound helper for NumPy slicing and crop extraction. sv.mask_to_xyxy stays inclusive for compatibility with CompactMask and current box-based adapters, so the coordinate-convention migration path is now explicit instead of implicit (#2416).

  • sv.HeatMapAnnotator now exposes a reset() method to clear accumulated heat, so a single annotator instance can be reused across independent streams without carrying over heat from a previous stream. sv.TraceAnnotator and sv.DetectionsSmoother gain the same reset() method for interface consistency, clearing their accumulated per-track history (#2418).

  • Soft-NMS: sv.box_soft_non_max_suppression, sv.mask_soft_non_max_suppression, and sv.Detections.with_soft_nms(sigma, class_agnostic, score_threshold) decay the confidence of overlapping detections instead of discarding them outright (#1624).

  • sv.PolygonZone(require_all_anchors: bool = True) — toggle between requiring every configured anchor point inside the zone (previous, default behavior) and counting a detection as soon as any anchor point is inside (#2272).

  • sv.InferenceSlicer(batch_size=...) — batches multiple image slices into a single callback invocation instead of one call per slice; batch_size=1 preserves the existing single-image callback contract (#1239).

  • sv.ConfusionMatrix.benchmark(save_directory_path=...) — optionally export an adaptive TP/FP/FN validation mosaic for each evaluated image (#2271).

  • sv.WindowedRasterDataset — public class backing sv.InferenceSlicer's windowed rasterio reads for tiled GeoTIFF inference (#2281).

  • AREA_DATA_FIELD config constant ("area") for storing per-detection area metadata in detections.data (#2428).

  • sv.denormalize_boxes and sv.xyxyxyxy_to_xyxy are now exported from the top-level supervision namespace (previously importable only from their submodules).

  • Added #2299: DetectionDataset.from_labelme and DetectionDataset.as_labelme for loading and exporting LabelMe per-image JSON annotations, following the existing COCO/YOLO/VOC convention. rectangle shapes load as boxes and polygon shapes as masks; unsupported shape types are skipped with a warning. The mask round-trip is a polygon approximation, not bit-exact.

  • Added #2284: DetectionDataset.from_createml and DetectionDataset.as_createml add load and export support for the CreateML object-detection JSON format, alongside the existing COCO, YOLO, and Pascal VOC formats.

  • Breaking: sv.JSONSink now emits native JSON types for numeric and boolean data fields instead of stringified values. Fields previously serialized as "True"/"False", "1"/"0.85", or "400.0" are now true/false, 1/0.85, 400.0. Downstream consumers that compare field values as strings (e.g. row["score"] == "1") or use strict string-typed schema validators must be updated. sv.CSVSink remains textual, but its custom-data slicing now matches sv.JSONSink: NumPy arrays, lists, and tuples are sliced per row only when their length matches the detection count; mismatched-length values are broadcast unchanged (#2400).

  • Breaking: sv.mask_non_max_merge now computes exact mask overlap at the original mask resolution and ignores the deprecated mask_dimension parameter. Code that relied on downscaled mask overlap should recalibrate thresholds. Passing overlap_metric or mask_dimension positionally is deprecated in 0.30.0 and removed in 0.33.0: the values are still honored (a positional overlap_metric still takes effect) but a DeprecationWarning is now emitted — pass both by keyword to silence it. More than five positional arguments raises TypeError (#2400).

  • Breaking: Supervision no longer installs an OpenCV distribution or provides an OpenCV extra. The default install uses the included fallback media backend; an existing compatible cv2 remains the preferred backend automatically. If your application needs OpenCV-specific behavior, install exactly one wheel family selected for that application (for example opencv-python-headless or opencv-python), then restart the process. See the OpenCV migration guide.

  • Breaking: Performance #2383: sv.Detections.merge() on mixed dense ndarray + CompactMask inputs now returns a CompactMask instead of a dense ndarray. Previously (0.29.0/0.29.1) the mixed path fell back to np.vstack, allocating a full (N, H, W) array; the new path converts dense inputs to CompactMask without materialising the full stack (~2 500× less peak memory, ~13× faster on 1080p / 40 detections). Code that checks isinstance(merged.mask, np.ndarray) or calls bare ndarray methods (.astype, .reshape, .ravel) on a mixed-merge result will need to be updated. The all-dense path is unchanged and still returns ndarray. This only affects code that explicitly merges a CompactMask-carrying Detections object with a dense-mask one via sv.Detections.merge(...) yourself — InferenceSlicer, DetectionsSmoother, and with_nms/with_nmm always merge type-homogeneous lists internally and are unaffected.

  • DetectionDataset and ClassificationDataset equality now compare the ordered classes lists directly instead of treating class labels as an unordered set. This keeps equality aligned with class_id indexing semantics, where class position is part of the dataset contract (#2388, #2408).

  • Performance: mask pixel counts now use count_nonzero (#2361), box_iou_batch_with_jaccard is vectorized (#2359), mask-annotation ROI blending is faster (#2368), the polygon annotator's square label background skips unnecessary corner circles (#2346), and compact-mask materialization is avoided inside the polygon annotator (#2369). No output changes.

  • Changed: delayed sv.ByteTrack, supervision.keypoint, normalized_xyxy for sv.denormalize_boxes, and supervision.dataset.utils RLE compatibility removals from supervision-0.30.0 to supervision-0.31.0 so the deprecated APIs keep a full transition window (#2415).

  • supervision now requires av>=14.2 as a mandatory install-time dependency for the PyAV cv2-free video fallback introduced during the OpenCV-optional transition (#2438). This doesn't change any public API — code that used supervision correctly before still behaves the same — but environments that pin exact dependency sets or vendor dependencies need to account for the new av requirement.

  • Fixed #2467 via #2468: sv.Recall now tracks classes that appear only in predictions, matching sv.Precision and sv.F1Score after #2331 and matching sklearn, which infers labels from the union of y_true and y_pred. matched_classes and recall_per_class are now aligned across the three metrics, including for samples that have predictions but no targets (background images), so per-class results can be compared row for row. matched_classes and recall_per_class gain a row for each prediction-only class under every averaging method; only the scalar MACRO recall changes value, since such a class now contributes 0.0, while the scalar MICRO and WEIGHTED aggregates are unaffected. Users relying on previous scores should re-evaluate after upgrading; no API change is required.

  • DetectionDataset.from_pascal_voc no longer raises ValueError on background images. An annotation file with no object elements produced an empty class_id array of dtype float64, which failed DetectionDataset validation, so any Pascal VOC dataset containing an unannotated image could not be loaded (#2463).

  • DetectionDataset.from_pascal_voc with force_masks=True no longer raises ValueError on background images. An annotation file with no object elements produced an empty mask of shape (0,) instead of the required (0, H, W), which failed Detections validation (#2469).

  • Reopening an existing sv.CSVSink or sv.JSONSink now starts a fresh output session: CSV files receive a new header and field schema, while JSON files no longer retain rows from the previous session (#2459).

  • Supervision now emits a UserWarning at import time when OpenCV is not installed and the cv2-free fallback backend is used, so users relying on OpenCV-specific behavior are alerted instead of silently falling back.

  • Fixed #2353: sv.Detections.from_inference no longer raises TypeError when the Inference package returns a mixed batch where only some predictions carry a tracker_id. detections.tracker_id is None for the full result in that case; fully-tracked and fully-untracked batches are unchanged.

  • sv.Detections.from_vlm with sv.VLM.GOOGLE_GEMINI_2_0, sv.VLM.GOOGLE_GEMINI_2_5, and sv.VLM.GOOGLE_GEMINI_3_5 now salvages the valid entries from a partially malformed JSON array (e.g. a single object with a syntax error) instead of discarding the whole response.

  • Geometry-aware IoU dispatch now powers the deprecated merge_inner_detections_objects, so overlapping axis-aligned envelopes no longer merge oriented boxes whose true OBB IoU is below the threshold (#2374).

  • save_coco_annotations (and therefore DetectionDataset.as_coco) now reads image sizes from file headers via lazy PIL instead of cv2-decoding every image, so labels-only COCO exports no longer decode any pixel data (#2442).

  • Fixed #2437: sv.F1Score no longer emits a spurious RuntimeWarning when true positives, false positives, and false negatives are all zero (denominator 0); the score remains 0.0.

  • Fixed #2427 via #2428: size-bucketed sv.Precision and sv.F1Score no longer count out-of-bucket detections as false positives. sv.Recall now matches only targets in the requested bucket, and all three metrics prioritize in-bucket targets during matching, matching COCO evaluation and sv.MeanAveragePrecision. A pixel-perfect detector now scores 1.0 in every bucket.

  • sv.hex_to_rgba now rejects multiple leading # characters instead of silently normalizing them, matching sv.is_valid_hex and the documented single optional prefix.

  • sv.box_iou_batch now upcasts box corners to float64 before computing areas and intersections, returning float32. This fixes integer-dtype overflow (e.g. int32 coordinates around 50_000 could previously wrap to a negative area and produce an incorrect 0.0 IoU) and gives full float64 precision to callers that pass float64/int64 coordinates directly. It does not recover precision already lost when coordinates are stored as float32 before this function is called (e.g. Detections.xyxy, which is float32 throughout the library) — such callers must upcast their own arrays to float64/int64 before calling box_iou_batch to benefit from this fix. Results for small-coordinate inputs are unchanged (#2418).

  • Legacy COCO prediction loading in sv.EvaluationDataset.load_predictions now raises ValueError for image ids absent from the ground-truth COCO set instead of relying on a bare assert, so the check is no longer silently skipped under python -O.

  • Fixed #2416: sv.process_video no longer risks hanging during shutdown; the sentinel enqueue is best-effort and worker joins are bounded.

  • Fixed #2416: COCO and CreateML dataset loaders now canonicalize resolved image paths and reject duplicate aliases for the same file.

  • Fixed #2416: DetectionDataset.as_pascal_voc() now preflights image and annotation basename collisions before writing, so exports fail fast instead of producing partial output.

  • import supervision no longer surfaces the deprecated ByteTrack warning; the top-level tracker alias now resolves lazily when accessed explicitly.

  • Fixed dataset export edge cases: DetectionDataset.split() and DetectionDataset.merge() now preserve in-memory image payloads without re-emitting the deprecation warning, and COCO/CreateML exports now reject duplicate image basenames instead of silently collapsing distinct paths into the same output key.

  • Fixed: sv.Color(...) now validates direct RGBA channel values and raises ValueError when any channel falls outside the 0-255 byte range.

  • Fixed: approximate_mask_with_polygons now defaults to no polygon simplification, matching the public dataset export methods.

  • Fixed: ImageSink.save_image() now raises OSError when cv2.imwrite() fails, and deprecation-warning control accepts the correct SUPERVISION_DEPRECATION_WARNING environment variable while still honoring the legacy misspelled alias.

  • sv.Classifications.from_timm now softmaxes model logits before exposing confidence scores, matching sv.Classifications.from_clip and keeping timm confidences on a normalized probability scale. Thresholds calibrated against raw logits may need retuning.

  • sv.download_assets now verifies MD5 hashes after fresh downloads and retries once when the downloaded payload is corrupted instead of accepting a bad file.

  • Fixed metrics scoring edge cases: legacy sv.MeanAveragePrecision now uses COCO 101-point AP averaging, sv.ConfusionMatrix rejects invalid class ids instead of wrapping them through int16/negative indexing, sv.MeanAveragePrecision preserves user-provided target ignore flags, and sv.MeanAverageRecallResult.recall_per_class now exposes per-class recall for each max-detection cutoff (#2411).

  • Fixed #2408: sv.Precision, sv.Recall, sv.F1Score, and sv.MeanAverageRecall now score size buckets by filtering targets only while leaving predictions eligible to match bucket targets. This preserves bucket matches that would otherwise be stolen by out-of-bucket filtering and keeps mAR top-K ranking intact.

  • sv.ByteTrack no longer mutates input Detections while assigning tracker IDs. It now keeps detections at the activation-threshold boundary eligible for matching, avoids impossible new-track thresholds above score 1.0, ignores invalid zero-area/non-finite tensor boxes before Kalman updates, and does not emit unconfirmed -1 IDs from first-frame tensor updates.

  • Fixed #2402: sv.KeyPoints.as_detections now accepts NumPy arrays, tuples, and generators in selected_keypoint_indices without ambiguous truth-value errors; empty index iterables select all keypoints. Valid zero-area skeletons are preserved, while all-zero and non-finite-only skeletons are filtered out.

  • Fixed #2407: sv.ColorPalette.by_idx() now raises a clear ValueError when called on an empty palette instead of leaking a ZeroDivisionError. Non-empty palettes keep the existing index-wrapping behavior.

  • Fixed #2393: sv.CropAnnotator.annotate no longer raises cv2.error when detections extend outside the scene; out-of-bounds boxes are clipped to scene bounds and zero-area results are skipped silently.

  • Fixed #2393: sv.HeatMapAnnotator.annotate no longer blanks the hottest region when the per-pixel hit count exceeds 255; the heat mask is now derived from the float32 accumulator directly, avoiding uint8 wrap-around.

  • Fixed #2393: sv.get_video_frames_generator now releases the underlying cv2.VideoCapture via try/finally, so the decoder is freed when a consumer breaks out of iteration early rather than waiting for garbage collection.

  • Fixed #2382: sv.Detections.get_anchors_coordinates now uses oriented bounding box corners (data["xyxyxyxy"]) when OBB data is present, instead of falling back to the axis-aligned envelope. Anchors on rotated detections now lie on the oriented body rather than drifting to the envelope. Non-OBB detections and Position.CENTER_OF_MASS (which requires a mask) are unaffected.

  • Fixed #2396: sv.BackgroundOverlayAnnotator.annotate no longer leaves detection regions tinted when bounding boxes have negative coordinates (extend outside the left or top scene boundary); boxes are now clipped to scene bounds before the detection region is restored.

  • Fixed: dataset IO/export edge cases now avoid mutating caller-owned Detections during DetectionDataset construction, reject non-integer and out-of-range class ids with a clear ValueError, load COCO annotations that omit optional iscrowd/area fields, expose DetectionDataset.from_coco(use_iscrowd=...) without changing the existing positional show_progress argument, export mask pixel area to COCO when no stored area is present, ignore folder-structure root clutter and non-image files inside class folders, and accept PIL-readable YOLO images such as RGBA or palette PNGs.

  • sv.Detections.from_tensorflow now scales bounding boxes by the correct image axes (#2360).

  • sv.Detections.from_inference keeps mask arrays aligned with xyxy when only some predictions in a batch carry segmentation data (#2362).

  • Replaced a deprecated 2-D np.cross call with an explicit determinant computation internally; no behavior change for callers (#2386).

  • Removed unnecessary defensive assert statements from image annotators; invalid input now surfaces through normal validation instead of being silently skipped under python -O (#2354).

  • sv.Precision, sv.Recall, and sv.F1Score now count predictions as false positives on images with an empty ground-truth set, extending the background-image handling shipped in 0.29.1 (#2397).

0.29.1 Jun 23, 2026

  • Added #2275: show_progress: bool = False parameter to all sv.DetectionDataset load and save methods — from_coco, from_yolo, from_pascal_voc, as_coco, as_yolo, as_pascal_voc, and save_dataset_images. When True, a tqdm.auto progress bar is shown (works in terminal and Jupyter). Defaults to False for full backward compatibility; no new dependencies.

  • Added #2338: sv.KeyPoints.with_nms — non-maximum suppression for keypoint detections. Derives axis-aligned bounding boxes from valid (non-zero and visible) keypoints and applies box_non_max_suppression. Requires detection_confidence; supports class-aware and class-agnostic modes via threshold, class_agnostic, and overlap_metric.

  • Fixed #2342: sv.Detections.from_vlm with sv.VLM.GOOGLE_GEMINI_2_0, sv.VLM.GOOGLE_GEMINI_2_5, and sv.VLM.QWEN_2_5_VL no longer raises when the model returns valid JSON of the wrong shape (non-list top-level or non-dict elements). A non-string or malformed "mask" value in Gemini 2.5 output no longer triggers AttributeError; invalid base64 or non-PNG mask data falls back to an empty mask, keeping xyxy, confidence, and masks arrays aligned.

  • Fixed #2341: sv.DetectionDataset.as_pascal_voc no longer mutates the source Detections.xyxy by the 1-index offset on every call. Previously, repeated exports accumulated a +1 shift in the caller's bounding boxes.

  • Fixed #2334: sv.JSONSink now serializes NumPy scalars (e.g. np.int64 frame indices) in custom_data as JSON numbers instead of raising TypeError at close time. File handle is now guaranteed to close even when serialization fails.

  • Fixed #2333: sv.DetectionsSmoother no longer raises when smoothing detections without confidence. Confidence is now averaged over the frames that carry it; when tracks in the same frame disagree on confidence presence, confidence is set to None for all smoothed detections.

  • Fixed #2332: sv.approximate_polygon now returns a polygon within the requested point-count budget (at most floor(N * (1 - percentage)) points, minimum 3). The function now also validates that epsilon_step > 0.

  • Fixed #2331: sv.Precision and sv.F1Score now count predictions on background images (empty target set) as false positives, and count predictions of classes absent from ground truth as false positives under MICRO and MACRO averaging. Previously both edge cases were silently ignored, inflating scores. WEIGHTED averaging is unchanged — absent classes retain weight 0, consistent with scikit-learn. Users relying on previous scores should re-evaluate after upgrading; no API change is required.

  • Fixed #2322: COCO export now preserves all polygon parts for multi-component masks. Previously, only the first polygon was written when a non-crowd mask had disjoint segments; all parts are now included.

  • Performance #2339: sv.HaloAnnotator now uses the same CompactMask painting path as sv.MaskAnnotator via a shared _paint_masks_by_area helper. On a 1080p frame with 30 CompactMask detections, HaloAnnotator runs approximately 4× faster; annotated output is unchanged.

  • Performance #2330: sv.mask_to_xyxy and sv.KeyPoints.as_detections are now vectorized. mask_to_xyxy uses batched occupancy-profile reductions instead of per-mask pixel scans; KeyPoints.as_detections computes all bounding boxes in a single batch operation. Both produce bit-identical results.

  • Performance #2323: Mask IoU computation now uses matrix multiplication on flattened masks instead of an explicit (N, M, H, W) intersection tensor, reducing peak memory for large mask sets. For masks larger than 4096×4096 pixels, computation automatically promotes to float64 to preserve exact pixel counts. Results are numerically identical.

0.29.0.post0 Jun 17, 2026

  • Fixed #2335: sv.KeyPoints(confidence=...) now works again. The 0.29.0 refactor accidentally dropped the deprecated confidence constructor kwarg; it is now accepted and mapped to keypoint_confidence with a deprecation warning.

0.29.0 Jun 15, 2026

  • Added #2314: new cookbook Oriented Bounding Boxes showing how an oriented box differs from an axis-aligned one on a marina of boats: DOTA-pretrained detection, the effect on with_nms and Detections.area, and YOLO OBB dataset export.

  • Fixed #2306: sv.Detections.area now returns the rotated body's area for detections carrying data["xyxyxyxy"] (oriented box corners) instead of the area of the derived axis-aligned bounding box, which overestimates by up to ~2x at 45° rotation. Affects annotator z-ordering inside MaskAnnotator and HaloAnnotator, and any user code that filters or sorts OBB detections by area. The mask path and the non-OBB AABB fallback are unchanged.

  • Added #2277, #2286: sv.VertexEllipseAreaAnnotator, sv.VertexEllipseOutlineAnnotator, and sv.VertexEllipseHaloAnnotator for visualizing keypoint uncertainty as covariance ellipses. Requires models that output keypoint uncertainty (e.g. RF-DETR keypoint models).

  • Added #2303: sv.oriented_box_non_max_suppression and sv.oriented_box_non_max_merge for performing NMS and NMM directly on oriented bounding boxes using oriented-box IoU instead of axis-aligned IoU.

  • Added #2247: sv.ConfusionMatrix now supports MetricTarget.ORIENTED_BOUNDING_BOXES, computing IoU via oriented_box_iou_batch on xyxyxyxy corners. Previously, OBB inputs silently fell back to axis-aligned bounding-box IoU, producing incorrect match scores for rotated detections.

  • Added #2252: sv.process_video gains a preserve_audio parameter. When enabled, the audio stream from the source video is muxed into the output using ffmpeg.

  • Added #2302, #2289: sv.DetectionDataset.as_yolo gains an is_obb parameter for exporting oriented bounding box annotations in the YOLO OBB format (9-token lines with 4 corner coordinates).

  • Added #2312: sv.xyxyxyxy_to_xyxy — vectorised utility that converts oriented bounding box corners (N, 4, 2) to axis-aligned bounding boxes (N, 4).

  • Changed #2286: sv.KeyPoints now separates keypoint-level and detection-level confidence into distinct fields: keypoint_confidence (shape (n, m)) and detection_confidence (shape (n,)). A new visible mask (shape (n, m)) controls per-keypoint visibility. The legacy KeyPoints.confidence property still works but is deprecated.

  • Changed #2286: sv.EdgeAnnotator and sv.VertexAnnotator now respect the visible mask. Invisible keypoints and their edges are skipped during rendering.

  • Changed #2286: sv.EdgeAnnotator and sv.VertexLabelAnnotator now support per-class skeleton definitions, enabling correct rendering when multiple skeleton topologies (e.g. person + animal) coexist in one frame.

  • Changed #2303: sv.Detections.with_nms and sv.Detections.with_nmm now use oriented-box IoU when data["xyxyxyxy"] coordinates are present, instead of axis-aligned box IoU. Callers relying on the previous axis-aligned behaviour should remove data["xyxyxyxy"] before calling, or recalibrate any IoU thresholds.

  • Changed #2312: sv.Detections.with_nmm now computes the merged oriented bounding box as the tightest rectangle at the winner's orientation enclosing all corners from every detection in a merge group.

  • Changed #2325: sv.VertexEllipseAreaAnnotator, sv.VertexEllipseOutlineAnnotator, and sv.VertexEllipseHaloAnnotator now draw sigma levels level-by-level (outermost first) across all points, ensuring correct visual layering when ellipses overlap.

  • Changed #2256: sv.InferenceSlicer now detects OBB outputs from callbacks and automatically falls back to sequential processing to avoid thread-safety issues when thread_workers > 1.

  • Changed #2324: Project-wide deprecation policy unified to a minimum 3-minor-release window. All current deprecations (including KeyPoints.confidence and validate_* helpers) are scheduled for removal in 0.32.0.

  • Fixed #2252: sv.process_video audio muxing path now correctly creates temp files on the same filesystem, decodes ffmpeg errors, and avoids muxing incomplete output.

  • Fixed #2282, #2317: sv.oriented_box_iou_batch now computes exact IoU via convex polygon intersection (cv2.intersectConvexConvex) and uses an axis-aligned bounding box envelope gate to skip pairs that cannot overlap, improving both accuracy and performance. Previously, rasterization on a discrete grid was used, which assumed square dimensions and introduced quantisation noise.

  • Fixed #2239: sv.Detections.from_vlm no longer returns None for class_id on empty VLM parses; now returns an empty int ndarray.

  • Fixed #2270: sv.Detections.from_inference now preserves class_name as a string-dtype array when predictions are empty.

  • Fixed #2269: sv.HeatMapAnnotator no longer crashes with a divide-by-zero when called with empty detections.

  • Fixed #2276: COCO export now emits 1-indexed category_id values as required by the COCO specification.

  • Fixed #2267: COCO annotation and image IDs are now sequential across train/val/test splits via starting_image_id and starting_annotation_id parameters.

  • Fixed #2289: sv.DetectionDataset.as_yolo no longer loses OBB rotation when exporting oriented bounding boxes.

  • Fixed #2296: YOLO dataset loading now sorts class names by numeric keys when data.yaml uses integer class IDs.

  • Fixed #2297: Letterbox utility now supports grayscale images.

  • Fixed #2298: File extension filters now normalize casing (e.g. .JPG matches .jpg).

  • Fixed #2321: sv.DetectionDataset.as_coco() now round-trips polygon and RLE segmentation data. Segmentations loaded from COCO annotations are preserved in detections.data["coco_raw_segmentation"] and written back on export, preventing data loss in train/val/test split workflows.

  • Deprecated: KeyPoints.confidence (use KeyPoints.keypoint_confidence), merge_inner_detection_object_pair, merge_inner_detections_objects, merge_inner_detections_objects_without_iou, validate_detections_fields, validate_vlm_parameters, validate_fields_both_defined_or_none, validate_xyxy, validate_mask, validate_class_id, validate_confidence, validate_tracker_id, validate_data, validate_xy, validate_key_point_confidence, validate_key_points_fields, validate_resolution, validate_custom_values, validate_input_tensors, and validate_labels are deprecated in 0.29.0 and will be removed in 0.32.0.

0.28.0 Apr 30, 2026

  • Added #2159: sv.CompactMask for memory-efficient mask storage. Masks are stored as crop-region bounding boxes plus RLE-encoded data instead of full-resolution bitmaps, reducing memory by up to 240× for sparse masks. Integrates transparently with sv.Detections.mask — filtering, merging, and area all work without materialising the full array.

  • Added #2227: sv.CompactMask.resize(new_image_shape) rescales all stored crops to match a new image resolution, enabling use across frames or after image resizing pipelines.

  • Added #2178: sv.Detections.from_inference now supports compressed COCO RLE masks. Inference responses with rle or rle_mask fields containing a compressed counts string (as produced by pycocotools) are decoded directly into binary masks, avoiding a lossy polygon round-trip.

  • Added #2004: sv.Color.from_hex now accepts 8-digit hexadecimal RGBA codes (e.g. #ff00ff80). Color.as_hex() serialises back, including alpha when not fully opaque. New utility functions sv.hex_to_rgba, sv.rgba_to_hex, and sv.is_valid_hex are exported at the top level.

  • Added #709: sv.BlurAnnotator and sv.PixelateAnnotator now support dynamic sizing. When kernel_size=None or pixel_size=None (the new default), the size is computed per detection as a fraction of the shorter bounding-box dimension, producing consistent visual results across objects of different sizes.

  • Added #2186: sv.InferenceSlicer now emits a warning when detections returned by the callback fall outside the tile boundaries, helping catch coordinate-system bugs in custom callbacks.

  • Added #2103, #2152: New sv.Detections.from_sam3() classmethod parses SAM3 PCS (text-prompted) and PVS (visual-prompted video segmentation) response formats into a standard sv.Detections, both from the local inference package and from Roboflow-hosted server responses.

  • Added #2154: The library now uses Python's logging module instead of print for diagnostic output. Messages are emitted under the supervision logger so applications can capture, filter, or silence them through standard logging configuration.

  • Added #932: sv.ImageAssets for downloading sample images alongside existing video assets, useful for examples and tutorials.

  • Changed #2169: sv.MeanAveragePrecisionResult and related metric arrays (mAP_scores, ap_per_class, iou_thresholds, precision/recall) are now float32 instead of float64. Reduces memory and speeds up computation; numerical results may differ in the last few digits.

  • Changed #2178: sv.rle_to_mask and sv.mask_to_rle moved to supervision.detection.utils.converters. The old import path supervision.dataset.utils continues to work but is deprecated.

  • Fixed #2178: sv.rle_to_mask now returns NDArray[bool] as declared in its signature. Previously the implementation returned uint8 despite the bool annotation; code that relied on the undocumented uint8 output (e.g. mask * 255 producing uint8) should wrap the result with .astype(np.uint8).

  • Fixed #2210: sv.VideoInfo.fps now returns a float instead of a truncated int. Previously, frame rates like 23.976, 29.97, and 59.94 were silently truncated, causing frame-timing drift that accumulates over long videos. The type of VideoInfo.fps has changed from int to float; callers that pass fps to APIs requiring an integer (such as deque(maxlen=...) or TraceAnnotator(trace_length=...)) should wrap the value with int().

  • Fixed #2209: sv.Detections.is_empty() now returns True for detections filtered down to zero rows, even when tracker_id is an empty array. Previously this case incorrectly returned False.

  • Fixed #2199: sv.CSVSink now correctly slices numpy array values in custom_data per row. Previously the full array was written for every detection.

  • Fixed #2216: sv.CSVSink and sv.JSONSink now slice plain Python list and tuple values in custom_data per detection row. Lists and tuples matching the detection count are indexed per row, consistent with np.ndarray behavior.

  • Fixed #2217: sv.TraceAnnotator no longer crashes in smooth mode when a tracker remains stationary. Duplicate consecutive points caused splprep to fail; the annotator now deduplicates anchor points and falls back to a raw polyline when fewer than 4 unique points are available.

  • Fixed #2218: load_coco_annotations now rejects COCO annotations whose file_name escapes the images directory via ../ traversal or absolute paths, preventing path-traversal attacks from malicious annotation files.

  • Fixed #2187: Extreme memory usage when loading OBB (oriented bounding box) datasets, caused by allocating full-image masks for each rotated box, has been resolved.

  • Fixed #2188: sv.KeyPoints boolean mask indexing now works correctly when all instances have the same keypoint count (uniform-count selection).

  • Fixed #2185: sv.DetectionDataset.as_coco() now preserves area and iscrowd fields instead of silently dropping them in the round-trip.

  • Fixed #1746: Precision loss when converting annotations with force_mask=True in dataset format converters.

  • Fixed #1991: sv.PolygonZone no longer double-counts the same object when multiple zones overlap. Detection bounding boxes were incorrectly clipped to each zone's ROI before anchor computation, causing the same detection to appear at a different anchor point in each zone; anchor is now computed from the original bounding box so containment is independent per zone.

  • Fixed #1868: sv.LineZone no longer mis-attributes crossings when a tracker reuses the same tracker_id across different classes. Class-aware bookkeeping prevents a new object from inheriting another class's prior crossing state.

  • Fixed #2022: sv.process_video now raises immediately when the user callback throws, instead of silently swallowing the exception and hanging until the writer is flushed.

  • Fixed #2156: sv.DetectionDataset now populates data["class_name"] on every loaded annotation, matching what model connectors produce. Downstream code can rely on class_name being present whether detections come from a dataset or a model.

  • Fixed #1364: sv.ByteTrack now preserves externally assigned tracker_id values instead of overwriting them with internal ids on the first update.

  • Fixed #1853: sv.ConfusionMatrix evaluate_detection_batch now matches predictions to ground truth correctly when multiple detections fall on the same target. Previously, double-counting inflated false-positive and false-negative counts.

  • Fixed #2136: sv.MeanAverageRecall now computes mAR@K using the top-K detections per image, matching the COCO definition. Previous values were inflated relative to pycocotools.

  • Fixed #1086, #265: COCO export and force_masks behaviour are now consistent across dataset formats. Empty polygons no longer raise during as_coco, and force_masks=True produces masks regardless of source format.

  • Deprecated #2215: sv.ByteTrack is deprecated in favour of ByteTrackTracker from the external trackers package (pip install trackers). The update method is renamed from update_with_detections() to update(). Removal is now planned for supervision-0.31.0.

  • Deprecated #2214: supervision.keypoint module is deprecated; use supervision.key_points instead. create_tiles in supervision.utils.image, ensure_cv2_image_for_processing in supervision.utils.conversion, and keypoint validation utilities in supervision.validators are deprecated. The LMM enum (use VLM) and from_lmm method (use from_vlm) were deprecated in 0.26.0; this release migrates their deprecation mechanism to pydeprecate.

  • Deprecated: normalized_xyxy argument in sv.denormalize_boxes renamed to xyxy. Passing normalized_xyxy= now emits a FutureWarning; support will be removed in supervision-0.31.0.

0.27.0 Nov 16, 2025

  • Added #2008: sv.filter_segments_by_distance to keep the largest connected component and nearby components within an absolute or relative distance threshold, for cleaning segmentation predictions from models such as SAM, SAM2, YOLO segmentation, and RF-DETR segmentation.

  • Added #2006: sv.xyxy_to_mask to convert bounding boxes into 2D boolean masks, where each mask corresponds to a single box.

  • Added #1943: sv.tint_image to apply a solid color overlay to an image at a given opacity. Works with both NumPy and PIL inputs.

  • Added #1943: sv.grayscale_image to convert an image to 3 channel grayscale for compatibility with color based drawing utilities.

  • Added #2014: sv.get_image_resolution_wh as a unified way to read image width and height from NumPy and PIL inputs.

  • Added #1912: sv.edit_distance for Levenshtein distance between two strings. Supports insert, delete, and substitute operations.

  • Added #1912: sv.fuzzy_match_index to find the first close match in a list using edit distance.

  • Changed #2015: sv.Detections.from_vlm and legacy from_lmm now support Qwen3 VL via vlm=sv.VLM.QWEN_3_VL.

  • Changed #1884: sv.Detections.from_vlm and legacy from_lmm now support DeepSeek VL 2 via vlm=sv.VLM.DEEPSEEK_VL_2.

  • Changed #2015: sv.Detections.from_vlm now parses Qwen 2.5 VL outputs more robustly and handles incomplete or truncated JSON responses.

  • Changed #2014: sv.InferenceSlicer now uses a new offset generation logic that removes redundant tiles and aligns borders cleanly, shortening inference time without hurting detection quality.

  • Changed #2016: sv.Detections now includes a box_aspect_ratio property for vectorized aspect ratio computation, useful for filtering detections based on box shape.

  • Changed #2001: Improved the performance of sv.box_iou_batch, approximately 2x to 5x faster on internal benchmarks.

  • Changed #1997: sv.process_video now uses a threaded reader, processor, and writer pipeline, removing I/O stalls and improving throughput while keeping the callback single threaded and safe for stateful models.

  • Changed: sv.denormalize_boxes now accepts arrays of shape (N, 4) and returns a batch of absolute pixel coordinates.

  • Changed #1917: sv.LabelAnnotator and sv.RichLabelAnnotator now accept text_offset=(x, y) to shift the label relative to text_position. Works with smart label position and line wrapping.

Removed

Removed the deprecated overlap_ratio_wh argument from sv.InferenceSlicer. Use the pixel based overlap_wh argument to control slice overlap.

Tip

Convert your old ratio based overlap to pixel based overlap by multiplying each ratio by the slice dimensions.

# before

slice_wh = (640, 640)
overlap_ratio_wh = (0.25, 0.25)

slicer = sv.InferenceSlicer(
    callback=callback,
    slice_wh=slice_wh,
    overlap_ratio_wh=overlap_ratio_wh,
    overlap_filter=sv.OverlapFilter.NON_MAX_SUPPRESSION,
)

# after

overlap_wh = (
    int(overlap_ratio_wh[0] * slice_wh[0]),
    int(overlap_ratio_wh[1] * slice_wh[1]),
)

slicer = sv.InferenceSlicer(
    callback=callback,
    slice_wh=slice_wh,
    overlap_wh=overlap_wh,
    overlap_filter=sv.OverlapFilter.NON_MAX_SUPPRESSION,
)

0.26.1 Jul 22, 2025

0.26.0 Jul 16, 2025

Removed

supervision-0.26.0 drops python3.8 support and upgrade all codes to python3.9 syntax style.

Tip

Supervision’s documentation theme now has a fresh look that is consistent with the documentations of all Roboflow open-source projects. (#1858)

  • Added #1774: Support for the IOS (Intersection over Smallest) overlap metric that measures how much of the smaller object is covered by the larger one in sv.Detections.with_nms, sv.Detections.with_nmm, sv.box_iou_batch, and sv.mask_iou_batch.

    import numpy as np
    import supervision as sv
    
    boxes_true = np.array([[100, 100, 200, 200], [300, 300, 400, 400]])
    boxes_detection = np.array([[150, 150, 250, 250], [320, 320, 420, 420]])
    
    sv.box_iou_batch(
        boxes_true=boxes_true,
        boxes_detection=boxes_detection,
        overlap_metric=sv.OverlapMetric.IOU,
    )
    
    # array([[0.14285714, 0.        ],
    #        [0.        , 0.47058824]])
    
    sv.box_iou_batch(
        boxes_true=boxes_true,
        boxes_detection=boxes_detection,
        overlap_metric=sv.OverlapMetric.IOS,
    )
    
    # array([[0.25, 0.  ],
    #        [0.  , 0.64]])
    
  • Added #1874: sv.box_iou that efficiently computes the Intersection over Union (IoU) between two individual bounding boxes.

  • Added #1816: Support for frame limitations and progress bar in sv.process_video.

  • Added #1788: Support for creating sv.KeyPoints objects from ViTPose and ViTPose++ inference results via sv.KeyPoints.from_transformers.

  • Added #1823: sv.xyxy_to_xcycarh function to convert bounding box coordinates from (x_min, y_min, x_max, y_max) into measurement space to format (center x, center y, aspect ratio, height), where the aspect ratio is width / height.

  • Added #1788: sv.xyxy_to_xywh function to convert bounding box coordinates from (x_min, y_min, x_max, y_max) format to (x, y, width, height) format.

  • Changed #1820: sv.LabelAnnotator now supports the smart_position parameter to automatically keep labels within frame boundaries, and the max_line_length parameter to control text wrapping for long or multi-line labels.

  • Changed #1825: sv.LabelAnnotator now supports non-string labels.

  • Changed #1792: sv.Detections.from_vlm now supports parsing bounding boxes and segmentation masks from responses generated by Google Gemini models.

    import supervision as sv
    
    gemini_response_text = (
        "```json\n"
        "    [\n"
        '        {"box_2d": [543, 40, 728, 200], "label": "cat", "id": 1},\n'
        '        {"box_2d": [653, 352, 820, 522], "label": "dog", "id": 2}\n'
        "    ]\n"
        "```"
    )
    
    detections = sv.Detections.from_vlm(
        sv.VLM.GOOGLE_GEMINI_2_5,
        gemini_response_text,
        resolution_wh=(1000, 1000),
        classes=["cat", "dog"],
    )
    
    detections.xyxy
    # array([[543., 40., 728., 200.], [653., 352., 820., 522.]])
    
    detections.data
    # {'class_name': array(['cat', 'dog'], dtype='<U26')}
    
    detections.class_id
    # array([0, 1])
    
  • Changed #1878: sv.Detections.from_vlm now supports parsing bounding boxes from responses generated by Moondream.

    import supervision as sv
    
    moondream_result = {
        "objects": [
            {
                "x_min": 0.5704046934843063,
                "y_min": 0.20069346576929092,
                "x_max": 0.7049859315156937,
                "y_max": 0.3012596592307091,
            },
            {
                "x_min": 0.6210969910025597,
                "y_min": 0.3300672620534897,
                "x_max": 0.8417936339974403,
                "y_max": 0.4961046129465103,
            },
        ]
    }
    
    detections = sv.Detections.from_vlm(
        sv.VLM.MOONDREAM,
        moondream_result,
        resolution_wh=(1000, 1000),
    )
    
    detections.xyxy
    # array([[1752.28,  818.82, 2165.72, 1229.14],
    #        [1908.01, 1346.67, 2585.99, 2024.11]])
    
  • Changed #1709: sv.Detections.from_vlm now supports parsing bounding boxes from responses generated by Qwen-2.5 VL.

    import supervision as sv
    
    qwen_2_5_vl_result = (
        "```json\n"
        "[\n"
        '    {"bbox_2d": [139, 768, 315, 954], "label": "cat"},\n'
        '    {"bbox_2d": [366, 679, 536, 849], "label": "dog"}\n'
        "]\n"
        "```"
    )
    
    detections = sv.Detections.from_vlm(
        sv.VLM.QWEN_2_5_VL,
        qwen_2_5_vl_result,
        input_wh=(1000, 1000),
        resolution_wh=(1000, 1000),
        classes=["cat", "dog"],
    )
    
    detections.xyxy
    # array([[139., 768., 315., 954.], [366., 679., 536., 849.]])
    
    detections.class_id
    # array([0, 1])
    
    detections.data
    # {'class_name': array(['cat', 'dog'], dtype='<U10')}
    
    detections.class_id
    # array([0, 1])
    
  • Changed #1786: Improved the speed of HSV color mapping in sv.HeatMapAnnotator, approximately 28x faster on 1920x1080 frames.

  • Fixed #1834: Supervision’s sv.MeanAveragePrecision is now fully aligned with pycocotools, the official COCO evaluation tool. This update enabled us to launch a new version of the Computer Vision Model Leaderboard.

    import supervision as sv
    from supervision.metrics import MeanAveragePrecision
    
    predictions = sv.Detections(...)
    targets = sv.Detections(...)
    
    map_metric = MeanAveragePrecision()
    map_metric.update(predictions, targets).compute()
    
    # Average Precision (AP) @[ IoU=0.50:0.95 | area=   all | maxDets=100 ] = 0.464
    # Average Precision (AP) @[ IoU=0.50      | area=   all | maxDets=100 ] = 0.637
    # Average Precision (AP) @[ IoU=0.75      | area=   all | maxDets=100 ] = 0.203
    # Average Precision (AP) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.284
    # Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.497
    # Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.629
    
  • Fixed #1767: Fixed losing sv.Detections.data when detections filtering.

0.25.0 Nov 12, 2024

  • No removals or deprecations in this release!

  • Added minimum_crossing_threshold to LineZone to stop jittering detections from being counted twice or more at a crossing. Setting it to 2 or more requires extra confirmation frames, significantly improving accuracy. (#1540)

  • Objects detected as KeyPoints can now be tracked; see the step-by-step guide in the Object Tracking Guide. (#1658)

    import numpy as np
    import supervision as sv
    from ultralytics import YOLO
    
    model = YOLO("yolov8m-pose.pt")
    tracker = sv.ByteTrack()
    trace_annotator = sv.TraceAnnotator()
    
    
    def callback(frame: np.ndarray, _: int) -> np.ndarray:
        results = model(frame)[0]
        key_points = sv.KeyPoints.from_ultralytics(results)
    
        detections = key_points.as_detections()
        detections = tracker.update_with_detections(detections)
    
        annotated_image = trace_annotator.annotate(frame.copy(), detections)
        return annotated_image
    
    
    sv.process_video(
        source_path="input_video.mp4", target_path="output_video.mp4", callback=callback
    )
    
  • Added is_empty method to KeyPoints to check if there are any keypoints in the object. (#1658)

  • Added as_detections method to KeyPoints that converts KeyPoints to Detections. (#1658)

  • Added a new video to the supervision.assets download catalog. (#1657)

    from supervision.assets import download_assets, VideoAssets
    
    path_to_video = download_assets(VideoAssets.SKIING)
    
  • Supervision can now be used with Python 3.13, most notably its ability to run without the Global Interpreter Lock (GIL). Dependency support for this is expected to be inconsistent, but if you try it, let us know the results! (#1595)

  • Added Mean Average Recall mAR metric, which returns a recall score, averaged over IoU thresholds, detected object classes, and limits imposed on maximum considered detections. (#1661)

    import supervision as sv
    from supervision.metrics import MeanAverageRecall
    
    predictions = sv.Detections(...)
    targets = sv.Detections(...)
    
    map_metric = MeanAverageRecall()
    map_result = map_metric.update(predictions, targets).compute()
    
    map_result.plot()
    
  • Added Precision and Recall metrics, providing a baseline for comparing model outputs to ground truth or another model (#1609)

    import supervision as sv
    from supervision.metrics import Recall
    
    predictions = sv.Detections(...)
    targets = sv.Detections(...)
    
    recall_metric = Recall()
    recall_result = recall_metric.update(predictions, targets).compute()
    
    recall_result.plot()
    
  • All Metrics now support Oriented Bounding Boxes (OBB) (#1593)

    import supervision as sv
    from supervision.metrics import F1_Score
    
    predictions = sv.Detections(...)
    targets = sv.Detections(...)
    
    f1_metric = MeanAverageRecall(metric_target=sv.MetricTarget.ORIENTED_BOUNDING_BOXES)
    f1_result = f1_metric.update(predictions, targets).compute()
    
  • Added Smart Labels: when smart_position is set for LabelAnnotator, RichLabelAnnotator or VertexLabelAnnotator, labels move to avoid overlapping others. (#1625)

    import supervision as sv
    from ultralytics import YOLO
    
    image = cv2.imread("image.jpg")
    
    label_annotator = sv.LabelAnnotator(smart_position=True)
    
    model = YOLO("yolo11m.pt")
    results = model(image)[0]
    detections = sv.Detections.from_ultralytics(results)
    
    annotated_frame = label_annotator.annotate(first_frame.copy(), detections)
    sv.plot_image(annotated_frame)
    
  • Added the metadata variable to Detections, storing custom data per-image rather than per-detected-object as data does, for example the source video path, camera model, or camera parameters. (#1589)

    import supervision as sv
    from ultralytics import YOLO
    
    model = YOLO("yolov8m")
    
    result = model("image.png")[0]
    detections = sv.Detections.from_ultralytics(result)
    
    # Items in `data` must match length of detections
    object_ids = [num for num in range(len(detections))]
    detections.data["object_number"] = object_ids
    
    # Items in `metadata` can be of any length.
    detections.metadata["camera_model"] = "Luxonis OAK-D"
    
  • Added a py.typed type hints metafile, signaling to type annotators and IDEs that type support is available. (#1586)

  • ByteTrack no longer requires detections to have a class_id (#1637)

  • draw_line, draw_rectangle, draw_filled_rectangle, draw_polygon, draw_filled_polygon and PolygonZoneAnnotator now comes with a default color (#1591)

  • Dataset classes are treated as case-sensitive when merging multiple datasets. (#1643)

  • Expanded metrics documentation with example plots and printed results (#1660)

  • Added usage example for polygon zone (#1608)

  • Small improvements to error handling in polygons: (#1602)

  • Updated ByteTrack to remove shared variables between instances, previously requiring liberal use of tracker.reset(). (#1603), (#1528)

  • Fixed a bug where class_agnostic setting in MeanAveragePrecision would not work. (#1577) hacktoberfest

  • Large refactor of ByteTrack: STrack moved to separate class, removed superfluous BaseTrack class, removed unused variables (#1603)

  • Large refactor of RichLabelAnnotator, matching its contents with LabelAnnotator. (#1625)

0.24.0 Oct 4, 2024

  • Added F1 score as a new metric for detection and segmentation. #1521

    import supervision as sv
    from supervision.metrics import F1Score
    
    predictions = sv.Detections(...)
    targets = sv.Detections(...)
    
    f1_metric = F1Score()
    f1_result = f1_metric.update(predictions, targets).compute()
    
    print(f1_result)
    print(f1_result.f1_50)
    print(f1_result.small_objects.f1_50)
    
  • Added new cookbook: Small Object Detection with SAHI, a guide to using InferenceSlicer for small object detection. #1483

  • Added an Embedded Workflow, which allows you to preview annotators. #1533

  • Enhanced LineZoneAnnotator: labels now align with the line even when it's not horizontal. You can also disable the text background and draw labels off-center to minimize overlap for multiple LineZone labels. #854

    import supervision as sv
    import cv2
    
    image = cv2.imread("<SOURCE_IMAGE_PATH>")
    
    line_zone = sv.LineZone(start=sv.Point(0, 100), end=sv.Point(50, 200))
    line_zone_annotator = sv.LineZoneAnnotator(
        text_orient_to_line=True, display_text_box=False, text_centered=False
    )
    
    annotated_frame = line_zone_annotator.annotate(
        frame=image.copy(), line_counter=line_zone
    )
    
    sv.plot_image(frame)
    
  • Added per-class counting to LineZone and introduced LineZoneAnnotatorMulticlass to visualize per-class counts — tracks individual classes crossing a line, useful for traffic monitoring or crowd analysis. #1555

    import supervision as sv
    import cv2
    
    image = cv2.imread("<SOURCE_IMAGE_PATH>")
    
    line_zone = sv.LineZone(start=sv.Point(0, 100), end=sv.Point(50, 200))
    line_zone_annotator = sv.LineZoneAnnotatorMulticlass()
    
    frame = line_zone_annotator.annotate(frame=frame, line_zones=[line_zone])
    
    sv.plot_image(frame)
    
  • Added from_easyocr, integrating results from EasyOCR, an open-source OCR library, into the supervision framework. #1515

    import supervision as sv
    import easyocr
    import cv2
    
    image = cv2.imread("<SOURCE_IMAGE_PATH>")
    
    reader = easyocr.Reader(["en"])
    result = reader.readtext("<SOURCE_IMAGE_PATH>", paragraph=True)
    detections = sv.Detections.from_easyocr(result)
    
    box_annotator = sv.BoxAnnotator(color_lookup=sv.ColorLookup.INDEX)
    label_annotator = sv.LabelAnnotator(color_lookup=sv.ColorLookup.INDEX)
    
    annotated_image = image.copy()
    annotated_image = box_annotator.annotate(scene=annotated_image, detections=detections)
    annotated_image = label_annotator.annotate(scene=annotated_image, detections=detections)
    
    sv.plot_image(annotated_image)
    
  • Added oriented_box_iou_batch to detection.utils — computes Intersection over Union (IoU) for oriented or rotated bounding boxes (OBB). #1502

    import numpy as np
    
    boxes_true = np.array([[[1, 0], [0, 1], [3, 4], [4, 3]]])
    boxes_detection = np.array([[[1, 1], [2, 0], [4, 2], [3, 3]]])
    ious = sv.oriented_box_iou_batch(boxes_true, boxes_detection)
    print("IoU between true and detected boxes:", ious)
    
  • Extended PolygonZoneAnnotator with an opacity option for zone fill (adjustable transparency). #1527

    import cv2
    from ncnn.model_zoo import get_model
    import supervision as sv
    
    image = cv2.imread("<SOURCE_IMAGE_PATH>")
    model = get_model(
        "yolov8s",
        target_size=640,
        prob_threshold=0.5,
        nms_threshold=0.45,
        num_threads=4,
        use_gpu=True,
    )
    result = model(image)
    detections = sv.Detections.from_ncnn(result)
    

Removed

The frame_resolution_wh parameter in PolygonZone has been removed.

Removed

Supervision installation methods "headless" and "desktop" were removed, as they are no longer needed. pip install supervision[headless] will install the base library and harmlessly warn of non-existent extras.

  • Supervision now depends on opencv-python rather than opencv-python-headless. #1530

  • Fixed the COCO 101-point Average Precision algorithm to correctly interpolate precision, avoiding averaged-out intermediate values. #1500

0.23.0 Aug 28, 2024

  • Added #930: IconAnnotator, a new annotator that draws an icon on each detection — useful for per-class icons.

    import supervision as sv
    from inference import get_model
    
    image = "<SOURCE_IMAGE_PATH>"
    icon_dog = "<DOG_PNG_PATH>"
    icon_cat = "<CAT_PNG_PATH>"
    
    model = get_model(model_id="yolov8n-640")
    results = model.infer(image)[0]
    detections = sv.Detections.from_inference(results)
    
    icon_paths = []
    for class_name in detections.data["class_name"]:
        if class_name == "dog":
            icon_paths.append(icon_dog)
        elif class_name == "cat":
            icon_paths.append(icon_cat)
        else:
            icon_paths.append("")
    
    icon_annotator = sv.IconAnnotator()
    annotated_frame = icon_annotator.annotate(
        scene=image.copy(), detections=detections, icon_path=icon_paths
    )
    
  • Added #1385: BackgroundColorAnnotator, that draws an overlay on the background images of the detections.

    import supervision as sv
    from inference import get_model
    
    image = "<SOURCE_IMAGE_PATH>"
    
    model = get_model(model_id="yolov8n-640")
    results = model.infer(image)[0]
    detections = sv.Detections.from_inference(results)
    
    background_overlay_annotator = sv.BackgroundOverlayAnnotator()
    annotated_frame = background_overlay_annotator.annotate(
        scene=image.copy(), detections=detections
    )
    
  • Added #1386: Support for Transformers v5 functions in sv.Detections.from_transformers. This includes the DetrImageProcessor methods post_process_object_detection, post_process_panoptic_segmentation, post_process_semantic_segmentation, and post_process_instance_segmentation.

    import torch
    import supervision as sv
    from PIL import Image
    from transformers import DetrImageProcessor, DetrForObjectDetection
    
    processor = DetrImageProcessor.from_pretrained("facebook/detr-resnet-50")
    model = DetrForObjectDetection.from_pretrained("facebook/detr-resnet-50")
    
    image = Image.open("<SOURCE_IMAGE_PATH>")
    inputs = processor(images=image, return_tensors="pt")
    
    with torch.no_grad():
        outputs = model(**inputs)
    
    width, height = image.size
    target_size = torch.tensor([[height, width]])
    results = processor.post_process_object_detection(
        outputs=outputs, target_sizes=target_size
    )[0]
    detections = sv.Detections.from_transformers(
        transformers_results=results, id2label=model.config.id2label
    )
    
  • Added #1354: Ultralytics SAM (Segment Anything Model) support in sv.Detections.from_ultralytics. SAM2 was released during this update, and is already supported via sv.Detections.from_sam.

    import supervision as sv
    from segment_anything import sam_model_registry, SamAutomaticMaskGenerator
    
    sam_model_reg = sam_model_registry[MODEL_TYPE]
    sam = sam_model_reg(checkpoint=CHECKPOINT_PATH).to(device=DEVICE)
    mask_generator = SamAutomaticMaskGenerator(sam)
    sam_result = mask_generator.generate(IMAGE)
    detections = sv.Detections.from_sam(sam_result=sam_result)
    
  • Added #1458: outline_color options for TriangleAnnotator and DotAnnotator.

  • Added #1409: text_color option for VertexLabelAnnotator keypoint annotator.

  • Changed #1434: InferenceSlicer now features an overlap_wh parameter, making it easier to compute slice sizes when handling overlapping slices.

  • Fixed #1448: Various annotator type issues have been resolved, supporting expanded error handling.

  • Fixed #1348: Added a new method for seeking to a specific video frame, for cases where traditional seeking fails. Enable with iterative_seek=True.

    import supervision as sv
    
    for frame in sv.get_video_frames_generator(
        source_path="<SOURCE_VIDEO_PATH>", start=60, iterative_seek=True
    ):
        # process frame
        pass
    
  • Fixed #1424: plot_image function now clearly indicates that the size is in inches.

Removed

The track_buffer, track_thresh, and match_thresh parameters in ByteTrack are deprecated and were removed as of supervision-0.23.0. Use lost_track_buffer, track_activation_threshold, and minimum_matching_threshold instead.

Removed

The triggering_position parameter in sv.PolygonZone was removed as of supervision-0.23.0. Use triggering_anchors instead.

Deprecated

overlap_filter_strategy in InferenceSlicer.__init__ is deprecated and will be removed in supervision-0.27.0. Use overlap_strategy instead.

Deprecated

overlap_ratio_wh in InferenceSlicer.__init__ is deprecated and will be removed in supervision-0.27.0. Use overlap_wh instead.

0.22.0 Jul 12, 2024

Deprecated

Constructing DetectionDataset with parameter images as Dict[str, np.ndarray] is deprecated and will be removed in supervision-0.26.0. Please pass a list of paths List[str] instead.

Deprecated

The DetectionDataset.images property is deprecated and will be removed in supervision-0.26.0. Please loop over images with for path, image, annotation in dataset:, as that does not require loading all images into memory.

import roboflow
from roboflow import Roboflow
import supervision as sv

roboflow.login()
rf = Roboflow()

project = rf.workspace("<WORKSPACE_ID>").project("<PROJECT_ID>")
dataset = project.version("<PROJECT_VERSION>").download("coco")

ds_train = sv.DetectionDataset.from_coco(
    images_directory_path=f"{dataset.location}/train",
    annotations_path=f"{dataset.location}/train/_annotations.coco.json",
)

path, image, annotation = ds_train[0]
# loads image on demand
# iterate to inspect all entries

for path, image, annotation in ds_train:
    # loads image on demand
    pass
  • Added #1296: sv.Detections.from_lmm now supports parsing results from the Florence 2 model, a Large Multimodal Model (LMM) — includes detailed object detection, OCR with region proposals, segmentation, and more. Find out more in our Colab notebook.

  • Added #1232: keypoint detection support with Mediapipe — both legacy and modern pipelines. See sv.KeyPoints.from_mediapipe for more.

  • Added #1316: sv.KeyPoints.from_mediapipe extended to support FaceMesh — processes face landmarks from both FaceLandmarker and legacy FaceMesh.

  • Added #1310: sv.KeyPoints.from_detectron2, a new KeyPoints method for extracting keypoints from the popular Detectron 2 platform.

  • Added #1300: sv.Detections.from_detectron2 now supports Detectron2 segmentation models — resulting masks work with sv.MaskAnnotator.

    import supervision as sv
    from detectron2 import model_zoo
    from detectron2.engine import DefaultPredictor
    from detectron2.config import get_cfg
    import cv2
    
    image = cv2.imread("<SOURCE_IMAGE_PATH>")
    cfg = get_cfg()
    cfg.merge_from_file(
        model_zoo.get_config_file("COCO-InstanceSegmentation/mask_rcnn_R_50_FPN_3x.yaml")
    )
    cfg.MODEL.WEIGHTS = model_zoo.get_checkpoint_url(
        "COCO-InstanceSegmentation/mask_rcnn_R_50_FPN_3x.yaml"
    )
    predictor = DefaultPredictor(cfg)
    
    result = predictor(image)
    detections = sv.Detections.from_detectron2(result)
    
    mask_annotator = sv.MaskAnnotator()
    annotated_frame = mask_annotator.annotate(scene=image.copy(), detections=detections)
    
  • Added #1277: if you provide a font that supports symbols of a language, sv.RichLabelAnnotator will draw them on your images.

    • Various annotators revised for correct in-place behavior with numpy arrays; fixed a bug where sv.ColorAnnotator filled boxes with solid color when used in-place.

      import cv2
      import supervision as sv
      from inference import get_model
      
      image = cv2.imread("<SOURCE_IMAGE_PATH>")
      
      model = get_model(model_id="yolov8n-640")
      results = model.infer(image)[0]
      detections = sv.Detections.from_inference(results)
      
      rich_label_annotator = sv.RichLabelAnnotator(font_path="<TTF_FONT_PATH>")
      annotated_image = rich_label_annotator.annotate(
          scene=image.copy(), detections=detections
      )
      
  • Added #1227: support for loading Oriented Bounding Box datasets in YOLO format.

    import supervision as sv
    
    train_ds = sv.DetectionDataset.from_yolo(
        images_directory_path="/content/dataset/train/images",
        annotations_directory_path="/content/dataset/train/labels",
        data_yaml_path="/content/dataset/data.yaml",
        is_obb=True,
    )
    
    _, image, detections = train_ds[0]
    
    obb_annotator = OrientedBoxAnnotator()
    annotated_image = obb_annotator.annotate(scene=image.copy(), detections=detections)
    
  • Fixed #1312: CropAnnotator.

Removed

BoxAnnotator was removed, however BoundingBoxAnnotator has been renamed to BoxAnnotator. Use a combination of BoxAnnotator and LabelAnnotator to simulate old BoundingBox behavior.

Deprecated

The name BoundingBoxAnnotator has been deprecated and will be removed in supervision-0.26.0. It has been renamed to BoxAnnotator.

  • Added #975 📝 New Cookbooks: serialize detections into json and csv.

  • Added #1290: internal change — file utility functions now support both str and pathlib paths.

  • Added #1340: Two new methods for converting between bounding box formats - xywh_to_xyxy and xcycwh_to_xyxy

Removed

from_roboflow method has been removed due to deprecation. Use from_inference instead.

Removed

Color.white() has been removed due to deprecation. Use color.WHITE instead.

Removed

Color.black() has been removed due to deprecation. Use color.BLACK instead.

Removed

Color.red() has been removed due to deprecation. Use color.RED instead.

Removed

Color.green() has been removed due to deprecation. Use color.GREEN instead.

Removed

Color.blue() has been removed due to deprecation. Use color.BLUE instead.

Removed

ColorPalette.default() has been removed due to deprecation. Use ColorPalette.DEFAULT instead.

Removed

FPSMonitor.__call__ has been removed due to deprecation. Use the attribute FPSMonitor.fps instead.

0.21.0 Jun 5, 2024

  • Added #500: sv.Detections.with_nmm to perform non-maximum merging on the current set of object detections.

  • Added #1221: sv.Detections.from_lmm to parse Large Multimodal Model (LMM) text results into sv.Detections. Currently supports only PaliGemma result parsing.

    import supervision as sv
    
    paligemma_result = "<loc0256><loc0256><loc0768><loc0768> cat"
    detections = sv.Detections.from_lmm(
        sv.LMM.PALIGEMMA,
        paligemma_result,
        resolution_wh=(1000, 1000),
        classes=["cat", "dog"],
    )
    detections.xyxy
    # array([[250., 250., 750., 750.]])
    
    detections.class_id
    # array([0])
    
  • Added #1236: sv.VertexLabelAnnotator allowing to annotate every vertex of a keypoint skeleton with custom text and color.

    import supervision as sv
    
    image = ...
    key_points = sv.KeyPoints(...)
    
    edge_annotator = sv.EdgeAnnotator(color=sv.Color.GREEN, thickness=5)
    annotated_frame = edge_annotator.annotate(scene=image.copy(), key_points=key_points)
    
  • Added #1147: sv.KeyPoints.from_inference allowing to create sv.KeyPoints from Inference result.

  • Added #1138: sv.KeyPoints.from_yolo_nas allowing to create sv.KeyPoints from YOLO-NAS result.

  • Added #1163: sv.mask_to_rle and sv.rle_to_mask allowing for easy conversion between mask and rle formats.

  • Changed #1236: sv.InferenceSlicer allowing to select overlap filtering strategy (NONE, NON_MAX_SUPPRESSION and NON_MAX_MERGE).

  • Changed #1178: sv.InferenceSlicer adding instance segmentation model support.

    import cv2
    import numpy as np
    import supervision as sv
    from inference import get_model
    
    model = get_model(model_id="yolov8x-seg-640")
    image = cv2.imread("<SOURCE_IMAGE_PATH>")
    
    
    def callback(image_slice: np.ndarray) -> sv.Detections:
        results = model.infer(image_slice)[0]
        return sv.Detections.from_inference(results)
    
    
    slicer = sv.InferenceSlicer(callback=callback)
    detections = slicer(image)
    
    mask_annotator = sv.MaskAnnotator()
    label_annotator = sv.LabelAnnotator()
    
    annotated_image = mask_annotator.annotate(scene=image, detections=detections)
    annotated_image = label_annotator.annotate(scene=annotated_image, detections=detections)
    
  • Changed #1228: sv.LineZone making it 10-20 times faster, depending on the use case.

  • Changed #1163: sv.DetectionDataset.from_coco and sv.DetectionDataset.as_coco adding support for run-length encoding (RLE) mask format.

0.20.0 April 24, 2024

  • Added #1128: sv.KeyPoints to provide initial support for pose estimation and broader keypoint detection models.

  • Added #1128: sv.EdgeAnnotator and sv.VertexAnnotator to enable rendering of results from keypoint detection models.

    import cv2
    import supervision as sv
    from ultralytics import YOLO
    
    image = cv2.imread("<SOURCE_IMAGE_PATH>")
    model = YOLO("yolov8l-pose")
    
    result = model(image, verbose=False)[0]
    keypoints = sv.KeyPoints.from_ultralytics(result)
    
    edge_annotators = sv.EdgeAnnotator(color=sv.Color.GREEN, thickness=5)
    annotated_image = edge_annotators.annotate(image.copy(), keypoints)
    
  • Changed #1037: sv.LabelAnnotator by adding a corner_radius argument to round the corners of the bounding box.

  • Changed #1109: sv.PolygonZone so the frame_resolution_wh argument is no longer required to initialize it.

Deprecated

The frame_resolution_wh parameter in sv.PolygonZone is deprecated and will be removed in supervision-0.24.0.

  • Changed #1084: sv.get_polygon_center to calculate a more accurate polygon centroid.

  • Changed #1069: sv.Detections.from_transformers by adding support for Transformers segmentation models and extract class names values.

    import torch
    import supervision as sv
    from PIL import Image
    from transformers import DetrImageProcessor, DetrForSegmentation
    
    processor = DetrImageProcessor.from_pretrained("facebook/detr-resnet-50-panoptic")
    model = DetrForSegmentation.from_pretrained("facebook/detr-resnet-50-panoptic")
    
    image = Image.open("<SOURCE_IMAGE_PATH>")
    inputs = processor(images=image, return_tensors="pt")
    
    with torch.no_grad():
        outputs = model(**inputs)
    
    width, height = image.size
    target_size = torch.tensor([[height, width]])
    results = processor.post_process_segmentation(
        outputs=outputs, target_sizes=target_size
    )[0]
    detections = sv.Detections.from_transformers(results, id2label=model.config.id2label)
    
    mask_annotator = sv.MaskAnnotator()
    label_annotator = sv.LabelAnnotator(text_position=sv.Position.CENTER)
    
    annotated_image = mask_annotator.annotate(scene=image, detections=detections)
    annotated_image = label_annotator.annotate(scene=annotated_image, detections=detections)
    
  • Fixed #787: sv.ByteTrack.update_with_detections which removed segmentation masks while tracking; ByteTrack now works alongside segmentation models.

0.19.0 March 15, 2024

  • Added #818: sv.CSVSink allowing for the straightforward saving of image, video, or stream inference results in a .csv file.

    import supervision as sv
    from ultralytics import YOLO
    
    model = YOLO("<SOURCE_MODEL_PATH>")
    csv_sink = sv.CSVSink("<RESULT_CSV_FILE_PATH>")
    frames_generator = sv.get_video_frames_generator("<SOURCE_VIDEO_PATH>")
    
    with csv_sink:
        for frame in frames_generator:
            result = model(frame)[0]
            detections = sv.Detections.from_ultralytics(result)
            csv_sink.append(detections, custom_data={"<CUSTOM_LABEL>": "<CUSTOM_DATA>"})
    
  • Added #819: sv.JSONSink allowing for the straightforward saving of image, video, or stream inference results in a .json file.

    import supervision as sv
    from ultralytics import YOLO
    
    model = YOLO("<SOURCE_MODEL_PATH>")
    json_sink = sv.JSONSink("<RESULT_JSON_FILE_PATH>")
    frames_generator = sv.get_video_frames_generator("<SOURCE_VIDEO_PATH>")
    
    with json_sink:
        for frame in frames_generator:
            result = model(frame)[0]
            detections = sv.Detections.from_ultralytics(result)
            json_sink.append(detections, custom_data={"<CUSTOM_LABEL>": "<CUSTOM_DATA>"})
    
  • Added #847: sv.mask_iou_batch allowing to compute Intersection over Union (IoU) of two sets of masks.

  • Added #847: sv.mask_non_max_suppression allowing to perform Non-Maximum Suppression (NMS) on segmentation predictions.

  • Added #888: sv.CropAnnotator allowing users to annotate the scene with scaled-up crops of detections.

    import cv2
    import supervision as sv
    from inference import get_model
    
    image = cv2.imread("<SOURCE_IMAGE_PATH>")
    model = get_model(model_id="yolov8n-640")
    
    result = model.infer(image)[0]
    detections = sv.Detections.from_inference(result)
    
    crop_annotator = sv.CropAnnotator()
    annotated_frame = crop_annotator.annotate(scene=image.copy(), detections=detections)
    
  • Changed #827: sv.ByteTrack.reset allowing users to clear trackers state, enabling the processing of multiple video files in sequence.

  • Changed #802: sv.LineZoneAnnotator allowing to hide in/out count using display_in_count and display_out_count properties.

  • Changed #787: sv.ByteTrack input arguments and docstrings updated to improve readability and ease of use.

Deprecated

The track_buffer, track_thresh, and match_thresh parameters in sv.ByteTrack are deprecated and will be removed in supervision-0.23.0. Use lost_track_buffer, track_activation_threshold, and minimum_matching_threshold instead.

  • Changed #910: sv.PolygonZone to now accept a list of specific box anchors that must be in zone for a detection to be counted.

Deprecated

The triggering_position parameter in sv.PolygonZone is deprecated and will be removed in supervision-0.23.0. Use triggering_anchors instead.

  • Changed #875: annotators adding support for Pillow images — all supervision Annotators now accept an image as either a numpy array or a Pillow Image, auto-detect its type, and return the output in the same format as the input.

  • Fixed #944: sv.DetectionsSmoother removing tracking_id from sv.Detections.

0.18.0 January 25, 2024

  • Added #720: sv.PercentageBarAnnotator allowing to annotate images and videos with percentage values representing confidence or other custom property.

    import supervision as sv
    
    image = ...
    detections = sv.Detections(...)
    
    percentage_bar_annotator = sv.PercentageBarAnnotator()
    annotated_frame = percentage_bar_annotator.annotate(
        scene=image.copy(), detections=detections
    )
    
  • Added #702: sv.RoundBoxAnnotator allowing to annotate images and videos with rounded corners bounding boxes.

  • Added #770: sv.OrientedBoxAnnotator allowing to annotate images and videos with OBB (Oriented Bounding Boxes).

    import cv2
    import supervision as sv
    from ultralytics import YOLO
    
    image = cv2.imread("<SOURCE_IMAGE_PATH>")
    model = YOLO("yolov8n-obb.pt")
    
    result = model(image)[0]
    detections = sv.Detections.from_ultralytics(result)
    
    oriented_box_annotator = sv.OrientedBoxAnnotator()
    annotated_frame = oriented_box_annotator.annotate(
        scene=image.copy(), detections=detections
    )
    
  • Added #696: sv.DetectionsSmoother allowing for smoothing detections over multiple frames in video tracking.

  • Added #769: sv.ColorPalette.from_matplotlib allowing users to create a sv.ColorPalette instance from a Matplotlib color palette.

    import supervision as sv
    
    sv.ColorPalette.from_matplotlib("viridis", 5)
    # ColorPalette(colors=[Color(r=68, g=1, b=84), Color(r=59, g=82, b=139), ...])
    
  • Changed #770: sv.Detections.from_ultralytics adding support for OBB (Oriented Bounding Boxes).

  • Changed #735: sv.LineZone to now accept a list of specific box anchors that must cross the line for a detection to be counted — a single anchor such as sv.Position.BOTTOM_CENTER, or any combination defined as List[sv.Position], instead of requiring all four box corners.

  • Changed #756: sv.Color's and sv.ColorPalette's method of accessing predefined colors, transitioning from a function-based approach (sv.Color.red()) to a property-based method (sv.Color.RED).

Deprecated

sv.ColorPalette.default() is deprecated and will be removed in supervision-0.22.0. Use sv.ColorPalette.DEFAULT instead.

Deprecated

Detections.from_roboflow() is deprecated and will be removed in supervision-0.22.0. Use Detections.from_inference instead.

  • Fixed #735: sv.LineZone functionality to accurately update the counter when an object crosses a line from any direction, including from the side — enabling more precise tracking and analytics, such as calculating individual in/out counts for each lane on the road.

0.17.0 December 06, 2023

0.16.0 October 19, 2023

  • Added #422: sv.BoxMaskAnnotator allowing to annotate images and videos with mox masks.

  • Added #433: sv.HaloAnnotator allowing to annotate images and videos with halo effect.

    import supervision as sv
    
    image = ...
    detections = sv.Detections(...)
    
    halo_annotator = sv.HaloAnnotator()
    annotated_frame = halo_annotator.annotate(scene=image.copy(), detections=detections)
    
  • Added #466: sv.HeatMapAnnotator allowing to annotate videos with heat maps.

  • Added #492: sv.DotAnnotator allowing to annotate images and videos with dots.

  • Added #449: sv.draw_image allowing to draw an image onto a given scene with specified opacity and dimensions.

  • Added #280: sv.FPSMonitor for monitoring frames per second (FPS) to benchmark latency.

  • Added #454: 🤗 Hugging Face Annotators space.

  • Changed #482: sv.LineZone.trigger now returns Tuple[np.ndarray, np.ndarray] — the first array indicates detections that crossed the line from outside to inside, the second from inside to outside.

  • Changed #465: Annotator argument name from color_map: str to color_lookup: ColorLookup enum to increase type safety.

  • Changed #426: sv.MaskAnnotator allowing 2x faster annotation.

  • Fixed #477: Poetry env definition allowing proper local installation.

  • Fixed #430: sv.ByteTrack to return np.array([], dtype=int) when svDetections is empty.

Deprecated

sv.Detections.from_yolov8 and sv.Classifications.from_yolov8 as those are now replaced by sv.Detections.from_ultralytics and sv.Classifications.from_ultralytics.

0.15.0 October 5, 2023

0.14.0 August 31, 2023

  • Added #282: support for SAHI inference technique with sv.InferenceSlicer.

    import cv2
    import supervision as sv
    from ultralytics import YOLO
    
    image = cv2.imread(SOURCE_IMAGE_PATH)
    model = YOLO(...)
    
    
    def callback(image_slice: np.ndarray) -> sv.Detections:
        result = model(image_slice)[0]
        return sv.Detections.from_ultralytics(result)
    
    
    slicer = sv.InferenceSlicer(callback=callback)
    
    detections = slicer(image)
    
  • Added #297: Detections.from_deepsparse to enable seamless integration with DeepSparse framework.

  • Added #281: sv.Classifications.from_ultralytics to enable seamless integration with Ultralytics framework. This will enable you to use supervision with all models that Ultralytics supports.

Deprecated

sv.Detections.from_yolov8 and sv.Classifications.from_yolov8 are now deprecated and will be removed with supervision-0.16.0 release.

0.13.0 August 8, 2023

  • Added #236: support for mean average precision (mAP) for object detection models with sv.MeanAveragePrecision.

    import supervision as sv
    from ultralytics import YOLO
    
    dataset = sv.DetectionDataset.from_yolo(...)
    
    model = YOLO(...)
    
    
    def callback(image: np.ndarray) -> sv.Detections:
        result = model(image)[0]
        return sv.Detections.from_yolov8(result)
    
    
    mean_average_precision = sv.MeanAveragePrecision.benchmark(
        dataset=dataset, callback=callback
    )
    
    mean_average_precision.map50_95
    # 0.433
    
  • Added #256: support for ByteTrack for object tracking with sv.ByteTrack.

  • Added #222: sv.Detections.from_ultralytics to enable seamless integration with Ultralytics framework. This will enable you to use supervision with all models that Ultralytics supports.

Deprecated

sv.Detections.from_yolov8 is now deprecated and will be removed with supervision-0.15.0 release.

0.12.0 July 24, 2023

Python 3.7. Support Terminated

With the supervision-0.12.0 release, we are terminating official support for Python 3.7.

  • Added #177: initial support for object detection model benchmarking with sv.ConfusionMatrix.

    import supervision as sv
    from ultralytics import YOLO
    
    dataset = sv.DetectionDataset.from_yolo(...)
    
    model = YOLO(...)
    
    
    def callback(image: np.ndarray) -> sv.Detections:
        result = model(image)[0]
        return sv.Detections.from_yolov8(result)
    
    
    confusion_matrix = sv.ConfusionMatrix.benchmark(dataset=dataset, callback=callback)
    
    confusion_matrix.matrix
    # array([
    #     [0., 0., 0., 0.],
    #     [0., 1., 0., 1.],
    #     [0., 1., 1., 0.],
    #     [1., 1., 0., 0.]
    # ])
    
  • Added #173: Detections.from_mmdetection to enable seamless integration with MMDetection framework.

  • Added #130: ability to install package in headless or desktop mode.

  • Changed #180: packing method from setup.py to pyproject.toml.

  • Fixed #188: sv.DetectionDataset.from_cooc can't be loaded when there are images without annotations.

  • Fixed #226: sv.DetectionDataset.from_yolo can't load background instances.

0.11.1 June 29, 2023

0.11.0 June 28, 2023

  • Added #150: ability to load and save sv.DetectionDataset in COCO format using as_coco and from_coco methods.

    import supervision as sv
    
    ds = sv.DetectionDataset.from_coco(images_directory_path="...", annotations_path="...")
    
    ds.as_coco(images_directory_path="...", annotations_path="...")
    
  • Added #158: ability to merge multiple sv.DetectionDataset together using merge method.

    import supervision as sv
    
    ds_1 = sv.DetectionDataset(...)
    len(ds_1)
    # 100
    ds_1.classes
    # ['dog', 'person']
    
    ds_2 = sv.DetectionDataset(...)
    len(ds_2)
    # 200
    ds_2.classes
    # ['cat']
    
    ds_merged = sv.DetectionDataset.merge([ds_1, ds_2])
    len(ds_merged)
    # 300
    ds_merged.classes
    # ['cat', 'dog', 'person']
    
  • Added #162: additional start and end arguments to sv.get_video_frames_generator allowing to generate frames only for a selected part of the video.

  • Fixed #157: incorrect loading of YOLO dataset class names from data.yaml.

0.10.0 June 14, 2023

  • Added #125: ability to load and save sv.ClassificationDataset in a folder structure format.

    import supervision as sv
    
    cs = sv.ClassificationDataset.from_folder_structure(root_directory_path="...")
    
    cs.as_folder_structure(root_directory_path="...")
    
  • Added #125: support for sv.ClassificationDataset.split, dividing sv.ClassificationDataset into two parts.

  • Added #110: ability to extract masks from Roboflow API results using sv.Detections.from_roboflow.

  • Added commit hash: Supervision Quickstart notebook covering Detection, Dataset and Video APIs.

  • Changed #135: sv.get_video_frames_generator documentation to better describe actual behavior.

0.9.0 June 7, 2023

  • Added #118: ability to select sv.Detections by index, list of indexes or slice.

    import supervision as sv
    
    detections = sv.Detections(...)
    len(detections[0])
    # 1
    len(detections[[0, 1]])
    # 2
    len(detections[0:2])
    # 2
    
  • Added #101: ability to extract masks from YOLOv8 result using sv.Detections.from_yolov8.

  • Added #122: ability to crop image using sv.crop.

  • Added #120: ability to conveniently save multiple images into directory using sv.ImageSink.

    import supervision as sv
    
    with sv.ImageSink(target_dir_path="target/directory/path") as sink:
        for image in sv.get_video_frames_generator(
            source_path="source_video.mp4", stride=10
        ):
            sink.save_image(image=image)
    
  • Fixed #106: inconvenient handling of sv.PolygonZone coordinates. Now sv.PolygonZone accepts coordinates in the form of [[x1, y1], [x2, y2], ...] that can be both integers and floats.

0.8.0 May 17, 2023

  • Added #100: support for dataset inheritance — Dataset renamed to DetectionDataset, now inheriting from BaseDataset, to keep future computer vision dataset APIs consistent.

  • Added #100: ability to save datasets in YOLO format using DetectionDataset.as_yolo.

    import roboflow
    from roboflow import Roboflow
    import supervision as sv
    
    roboflow.login()
    
    rf = Roboflow()
    
    project = rf.workspace(WORKSPACE_ID).project(PROJECT_ID)
    dataset = project.version(PROJECT_VERSION).download("yolov5")
    
    ds = sv.DetectionDataset.from_yolo(
        images_directory_path=f"{dataset.location}/train/images",
        annotations_directory_path=f"{dataset.location}/train/labels",
        data_yaml_path=f"{dataset.location}/data.yaml",
    )
    
    ds.classes
    # ['dog', 'person']
    
  • Added #103: support for DetectionDataset.split, dividing DetectionDataset into two parts.

    import supervision as sv
    
    ds = sv.DetectionDataset(...)
    train_ds, test_ds = ds.split(split_ratio=0.7, random_state=42, shuffle=True)
    
    len(train_ds), len(test_ds)
    # (700, 300)
    
  • Changed #100: default value of approximation_percentage parameter from 0.75 to 0.0 in DetectionDataset.as_yolo and DetectionDataset.as_pascal_voc.

0.7.0 May 11, 2023

  • Added #91: Detections.from_yolo_nas to enable seamless integration with YOLO-NAS model.
  • Added #86: ability to load datasets in YOLO format using Dataset.from_yolo.
  • Added #84: Detections.merge to merge multiple Detections objects together.
  • Fixed #81: LineZoneAnnotator.annotate does not return annotated frame.
  • Changed #44: LineZoneAnnotator.annotate to allow for custom text for the in and out tags.

0.6.0 April 19, 2023

  • Added #71: initial Dataset support and ability to save Detections in Pascal VOC XML format.
  • Added #71: new mask_to_polygons, filter_polygons_by_area, polygon_to_xyxy and approximate_polygon utilities.
  • Added #72: ability to load Pascal VOC XML object detections dataset as Dataset.
  • Changed #70: order of Detections attributes to make it consistent with order of objects in __iter__ tuple.
  • Changed #71: generate_2d_mask to polygon_to_mask.

0.5.2 April 13, 2023

  • Fixed #63: LineZone.trigger function expects 4 values instead of 5.

0.5.1 April 12, 2023

  • Fixed Detections.__getitem__ method did not return mask for selected item.
  • Fixed Detections.area crashed for mask detections.

0.5.0 April 10, 2023

  • Added #58: Detections.mask to enable segmentation support.
  • Added #58: MaskAnnotator to allow easy Detections.mask annotation.
  • Added #58: Detections.from_sam to enable native Segment Anything Model (SAM) support.
  • Changed #58: Detections.area behaviour to work not only with boxes but also with masks.

0.4.0 April 5, 2023

  • Added #46: Detections.empty to allow easy creation of empty Detections objects.
  • Added #56: Detections.from_roboflow to allow easy creation of Detections objects from Roboflow API inference results.
  • Added #56: plot_images_grid to allow easy plotting of multiple images on single plot.
  • Added #56: initial support for Pascal VOC XML format with detections_to_voc_xml method.
  • Changed #56: show_frame_in_notebook refactored and renamed to plot_image.

0.3.2 March 23, 2023

  • Changed #50: Allow Detections.class_id to be None.

0.3.1 March 6, 2023

  • Fixed #41: PolygonZone throws an exception when the object touches the bottom edge of the image.
  • Fixed #42: Detections.wth_nms method throws an exception when Detections is empty.
  • Changed #36: Detections.wth_nms support class agnostic and non-class agnostic case.

0.3.0 March 6, 2023

  • Changed: Allow Detections.confidence to be None.
  • Added: Detections.from_transformers and Detections.from_detectron2 to enable seamless integration with Transformers and Detectron2 models.
  • Added: Detections.area to dynamically calculate bounding box area.
  • Added: Detections.wth_nms to filter out double detections with NMS. Initial - only class agnostic - implementation.

0.2.0 February 2, 2023

  • Added: Advanced Detections filtering with pandas-like API.
  • Added: Detections.from_yolov5 and Detections.from_yolov8 to enable seamless integration with YOLOv5 and YOLOv8 models.

0.1.0 January 19, 2023

Say hello to Supervision 👋