Skip to content

Conversion Utils

supervision.utils.conversion.cv2_to_pillow(image: npt.NDArray[np.uint8]) -> Image.Image

Converts an OpenCV image into a Pillow image, reordering channels from OpenCV's BGR(A) convention to Pillow's RGB(A).

Parameters:

Name Type Description Default

image

NDArray[uint8]

OpenCV image. Accepted shapes: - (H, W) — grayscale, passed through unchanged. - (H, W, 3) — BGR, converted to RGB. - (H, W, 4) — BGRA, converted to RGBA.

required

Returns:

Type Description
Image

Input image converted to Pillow format.

Raises:

Type Description
ValueError

If image is not 2-D or 3-D with 3 or 4 channels.

Examples:

>>> import numpy as np
>>> from supervision.utils.conversion import cv2_to_pillow
>>> scene = np.zeros((10, 10, 3), dtype=np.uint8)
>>> scene[:, :, 2] = 255
>>> image = cv2_to_pillow(scene)
>>> image.size
(10, 10)
>>> image.getpixel((0, 0))
(255, 0, 0)
Source code in src/supervision/utils/conversion.py
def cv2_to_pillow(image: npt.NDArray[np.uint8]) -> Image.Image:
    """Converts an OpenCV image into a Pillow image, reordering channels from OpenCV's
    BGR(A) convention to Pillow's RGB(A).

    Args:
        image: OpenCV image. Accepted shapes:
            - `(H, W)` — grayscale, passed through unchanged.
            - `(H, W, 3)` — BGR, converted to RGB.
            - `(H, W, 4)` — BGRA, converted to RGBA.

    Returns:
        Input image converted to Pillow format.

    Raises:
        ValueError: If `image` is not 2-D or 3-D with 3 or 4 channels.

    Examples:
        ```pycon
        >>> import numpy as np
        >>> from supervision.utils.conversion import cv2_to_pillow
        >>> scene = np.zeros((10, 10, 3), dtype=np.uint8)
        >>> scene[:, :, 2] = 255
        >>> image = cv2_to_pillow(scene)
        >>> image.size
        (10, 10)
        >>> image.getpixel((0, 0))
        (255, 0, 0)

        ```
    """
    if image.ndim == 2:
        return Image.fromarray(np.ascontiguousarray(image))
    if image.ndim == 3 and image.shape[2] == 3:
        rgb_image = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)
        return Image.fromarray(rgb_image)
    if image.ndim == 3 and image.shape[2] == 4:
        return Image.fromarray(np.ascontiguousarray(image[..., [2, 1, 0, 3]]))
    raise ValueError(f"Expected shape (H,W), (H,W,3), or (H,W,4), got {image.shape}.")

supervision.utils.conversion.pillow_to_cv2(image: Image.Image) -> npt.NDArray[np.uint8]

Converts Pillow image into OpenCV image, handling RGB -> BGR conversion.

Every Pillow mode is reduced to the 8-bit layout OpenCV would hand back for the same picture: an (H, W) grayscale array for single-channel modes and an (H, W, 3) BGR array for everything else. Palette images are expanded to RGB so palette indices are resolved to their actual colors, and CMYK ink values are converted to color instead of being read as RGB plus an extra channel. Alpha is dropped, matching cv2.imread with its default flags, so RGBA becomes BGR and LA becomes grayscale. A 1-bit image becomes 0 and 255. An integer image deeper than 8 bits (I;16 and its endian variants, or the signed 32-bit I) is clipped to the 16-bit range and keeps its high byte, as cv2.imread does when it reads a 16-bit PNG as 8-bit. A 32-bit float image is clipped to 0-255 the way Pillow's own convert("L") clips it.

Parameters:

Name Type Description Default

image

Image

Pillow image in any mode.

required

Returns:

Type Description
NDArray[uint8]

Input image converted to OpenCV format: uint8 of shape (H, W) for

NDArray[uint8]

single-channel modes or (H, W, 3) BGR otherwise.

Examples:

>>> from PIL import Image
>>> from supervision.utils.conversion import pillow_to_cv2
>>> image = Image.new("RGB", (10, 10), color=(255, 0, 0))
>>> scene = pillow_to_cv2(image)
>>> scene.shape
(10, 10, 3)
>>> scene[0, 0].tolist()
[0, 0, 255]
>>> pillow_to_cv2(Image.new("1", (2, 2), color=1)).tolist()
[[255, 255], [255, 255]]
Source code in src/supervision/utils/conversion.py
def pillow_to_cv2(image: Image.Image) -> npt.NDArray[np.uint8]:
    """Converts Pillow image into OpenCV image, handling RGB -> BGR conversion.

    Every Pillow mode is reduced to the 8-bit layout OpenCV would hand back for the
    same picture: an `(H, W)` grayscale array for single-channel modes and an
    `(H, W, 3)` BGR array for everything else. Palette images are expanded to RGB so
    palette indices are resolved to their actual colors, and CMYK ink values are
    converted to color instead of being read as RGB plus an extra channel. Alpha is
    dropped, matching `cv2.imread` with its default flags, so RGBA becomes BGR and
    LA becomes grayscale. A 1-bit image becomes `0` and `255`. An integer image
    deeper than 8 bits (`I;16` and its endian variants, or the signed 32-bit `I`) is
    clipped to the 16-bit range and keeps its high byte, as `cv2.imread` does when
    it reads a 16-bit PNG as 8-bit. A 32-bit float image is clipped to `0`-`255`
    the way Pillow's own `convert("L")` clips it.

    Args:
        image: Pillow image in any mode.

    Returns:
        Input image converted to OpenCV format: `uint8` of shape `(H, W)` for
        single-channel modes or `(H, W, 3)` BGR otherwise.

    Examples:
        ```pycon
        >>> from PIL import Image
        >>> from supervision.utils.conversion import pillow_to_cv2
        >>> image = Image.new("RGB", (10, 10), color=(255, 0, 0))
        >>> scene = pillow_to_cv2(image)
        >>> scene.shape
        (10, 10, 3)
        >>> scene[0, 0].tolist()
        [0, 0, 255]
        >>> pillow_to_cv2(Image.new("1", (2, 2), color=1)).tolist()
        [[255, 255], [255, 255]]

        ```
    """
    values = np.asarray(image)
    if values.dtype.kind in "iu" and values.dtype.itemsize > 1:
        # Any integer mode deeper than 8 bits, signed or not. Keep the high byte: a
        # 16-bit value cast to uint8 wraps modulo 256 and redraws a bright pixel as
        # a dark one, and a signed or 32-bit value must be clipped before the cast.
        clipped = np.clip(values, 0, np.iinfo(np.uint16).max).astype(np.uint16)
        return cast(npt.NDArray[np.uint8], (clipped >> 8).astype(np.uint8))

    if image.mode in _SINGLE_CHANNEL_MODES:
        # Annotators draw into the returned array, so hand back a writable copy
        # rather than the read-only view `np.asarray` makes of a Pillow buffer.
        if image.mode != "L":
            return np.array(image.convert("L"), dtype=np.uint8)
        return np.array(values, dtype=np.uint8)

    if image.mode != "RGB":
        values = np.asarray(image.convert("RGB"))

    scene = cv2.cvtColor(values, cv2.COLOR_RGB2BGR)
    # cvtColor already returns uint8 here, so astype is a no-op other than the
    # full-image copy it forces; copy=False keeps the dtype guard without it.
    return cast(npt.NDArray[np.uint8], scene.astype(np.uint8, copy=False))

supervision.utils.conversion.images_to_cv2(images: list[npt.NDArray[np.uint8] | Image.Image]) -> list[npt.NDArray[np.uint8]]

Converts images provided either as Pillow images or OpenCV images into OpenCV format.

Parameters:

Name Type Description Default

images

list[NDArray[uint8] | Image]

Images to be converted

required

Returns:

Type Description
list[NDArray[uint8]]

List of input images in OpenCV format (with order preserved).

Source code in src/supervision/utils/conversion.py
def images_to_cv2(
    images: list[npt.NDArray[np.uint8] | Image.Image],
) -> list[npt.NDArray[np.uint8]]:
    """Converts images provided either as Pillow images or OpenCV images into OpenCV
    format.

    Args:
        images: Images to be converted

    Returns:
        List of input images in OpenCV format
            (with order preserved).
    """
    result: list[npt.NDArray[np.uint8]] = []
    for image in images:
        if isinstance(image, Image.Image):
            result.append(pillow_to_cv2(image))
        else:
            result.append(image)
    return result

Comments