Range extension for object detection in machine vision systems

By using image slicing technology in machine vision systems, intelligently selecting cameras and determining the number of image slicing, the challenge of detecting small objects in large spaces is solved, and efficient object detection and recognition is achieved.

CN120298640APending Publication Date: 2025-07-11NOKIA NETWORKS OY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510027391.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-10
Filing Date
2025-01-08
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing machine vision systems face challenges in detecting and identifying small objects in large spaces, especially in environments covered by sparse cameras, where resources are consumed and the probability of detection is low.

Method used

By using image slicing technology in machine vision systems, intelligently select cameras and determine the number of image slicing, perform image slicing and select processing image slicing, and combine machine learning models for object detection and tracking.

Benefits of technology

Efficient detection and recognition of small objects in large spaces is achieved, reducing computational complexity and eliminating the need to retrain machine learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298640A_ABST
    Figure CN120298640A_ABST
Patent Text Reader

Abstract

Various example embodiments are presented herein for supporting object detection in a machine vision system. Various example embodiments directed to supporting object detection in a machine vision system may be configured to support range extension for object detection in a machine vision system. Various example embodiments directed to supporting object detection in a machine vision system may be configured to support flexible and efficient range extension for object detection in a machine vision system. Various example embodiments for flexible and efficient range extension for supporting object detection in a machine vision system may be configured to support flexible and efficient range extension for object detection in a machine vision system based on use of image slices.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Each exemplary embodiment generally relates to a machine vision system, and more specifically but not exclusively to object detection in a machine vision system. Background Art

[0002] Machine vision systems can use camera systems for object detection in various contexts such as industrial automation, autonomous vehicle tracking, etc. Machine vision systems are configured to perform object tracking for tracking various types of objects in such contexts. There are various challenges associated with the implementation of machine vision systems, including object detection for various types of objects, management of resources required to support object detection, etc. Summary of the Invention

[0003] In at least some example embodiments, the apparatus includes: at least one processor, and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: detect an object within an environment monitored by a set of cameras; based on a determination that the object is no longer detected, determine from the set of cameras a set of cameras for which image slices will be activated, where the set of cameras for which image slices will be activated includes: a first camera on which the object was last detected before the object was no longer detected; and a second camera that is determined to be: at the location closest to the object when the object was last detected; obtain a set of images, where the set of images includes: corresponding images for the respective cameras from each camera in the set of cameras for which image slices will be activated; obtain a set of selected image slices from the set of images based on application of image slices to each image in the set of images; and determine the location of the object within the environment based on processing of the set of selected image slices. In at least some example embodiments, to detect the object, the instructions, when executed by the at least one processor, cause the apparatus to at least: determine a distance between the location of the camera and the location of the object for cameras in the set of cameras on which the object is detected; and detect the object based on a determination that the distance between the location of the camera and the location of the object satisfies a threshold. In at least some example embodiments, to detect the object, the instructions, when executed by the at least one processor, cause the apparatus to at least: based on a determination that the layout of the set of cameras is unknown, periodically activate image slices for each camera in the set of cameras; and detect the object based on the image slices for each camera in the set of cameras. In at least some example embodiments, the set of cameras for which image slices will be activated is determined based on a length of time since the object was last detected satisfying a threshold. In at least some example embodiments, the second camera from the set of cameras is determined to be at the location closest to the object when the object was last detected based on a determination that the second camera has a field of view closest to the location of the object when the object was last detected. In at least some example embodiments, to select the second camera, the instructions, when executed by the at least one processor, cause the apparatus to at least: calculate a respective field of view for each camera in the set of cameras; determine a set of distance values that includes: for each camera in the set of cameras, a respective distance between the location of the object when the object was last detected and the respective field of view of the camera; and select the second camera from the set of cameras based on the set of distance values, where the second camera is determined to be at the location closest to the object when the object was last detected.In at least some example embodiments, to calculate the corresponding field of view of a camera, the instructions, when executed by at least one processor, cause the apparatus to at least: receive a homography matrix calculated using a reference image resolution that matches the image resolution of the corresponding image from the corresponding camera; obtain a set of points by calculating, for each pixel in a set of pixels of the corresponding image and based on the homography matrix, a corresponding point specifying the physical location of the region captured by the corresponding pixel; filter any points in the set of points that are beyond a predefined distance threshold in the set of points to form a filtered set of points; and calculate the corresponding field of view of the camera based on a convex structure and using the filtered set of points. In at least some example embodiments, to obtain a set of selected image slices, the instructions, when executed by at least one processor, cause the apparatus to at least: obtain, for each image in a set of images, a corresponding set of image slices for the corresponding image based on the application of image slices to the corresponding image in the set of images; and obtain the set of selected image slices by selecting at least a portion of the image slices in the corresponding set of image slices from each image slice in the set of image slices formed by applying the image slices to the corresponding image in the set of images. In at least some example embodiments, to obtain a corresponding set of image slices for a corresponding image based on the application of image slices to the corresponding image in a set of images, the instructions, when executed by at least one processor, cause the apparatus to at least: determine the number of slices into which the corresponding image in the set of images is to be sliced; and perform slicing of the corresponding image in the set of images based on the number of slices into which the corresponding image in the set of images is to be sliced to form a corresponding set of image slices for the corresponding image in the set of images. In at least some example embodiments, the number of slices into which the corresponding image in the set of images is to be sliced is determined based on at least one of: the distance between the camera capturing the corresponding image and the object, or the number of pixels in the bounding box of the detection of the object. In at least some example embodiments, at least a portion of the image slices in the corresponding set of image slices for the corresponding image in the set of images is selected based on a slice selection algorithm for a region of interest (ROI).In at least some example embodiments, to select at least a portion of the image slices in a corresponding set of image slices to form a corresponding set of selected image slices, the instructions, when executed by at least one processor, cause the apparatus to at least: collect a set of bounding boxes for an object; combine the set of bounding boxes for the object into a union bounding box for the object; slice an image from a camera into N image slices and, for each of the N image slices of the corresponding image, obtain the corresponding coordinates of the corresponding image slice; and for each of the image slices, calculate the intersection over union (IoU) between the corresponding image slice and the union bounding box, and based on a determination that the corresponding IoU for the corresponding image slice meets a threshold, then select the corresponding image slice for inclusion in the corresponding set of selected image slices. In at least some example embodiments, based on a cropped slice selection algorithm, at least a portion of the image slices in a corresponding set of image slices for a corresponding image in a set of images is selected. In at least some example embodiments, to select at least a portion of the image slices in a corresponding set of image slices to form a corresponding set of selected image slices, the instructions, when executed by at least one processor, cause the apparatus to at least: collect a set of bounding boxes for an object; combine the set of bounding boxes for the object into a union bounding box for the object; expand the union bounding box based on a defined image size of a machine learning (ML) model to form a region of interest; based on a determination that the size of the region of interest is greater than the defined image size of the ML model in at least one of a height parameter or a width parameter, crop a bounding box region from the image to form a cropped image; and slice the cropped image to obtain a corresponding set of selected image slices. In at least some example embodiments, the cropped image is sliced based on the size of the region of interest and the defined image size of the ML model to obtain a corresponding set of image slices.

[0004] In at least some example embodiments, a non-transitory computer-readable medium stores computer program instructions that, when executed by a device, cause the device to at least: detect an object within an environment monitored by a set of cameras; based on a determination that the object is no longer detected, determine from the set of cameras a set of cameras for which image slices will be activated, wherein the set of cameras for which image slices will be activated includes: a first camera on which the object was last detected before the object was no longer detected; and a second camera that is determined to be: at the location closest to the object when the object was last detected; obtain a set of images, wherein the set of images includes: corresponding images for each camera in the set of cameras for which image slices will be activated, for the respective cameras; obtain a set of selected image slices from the set of images based on application of image slices to each image in the set of images; and determine the location of the object within the environment based on processing of the set of selected image slices. In at least some example embodiments, to detect an object, the computer program instructions, when executed by the device, cause the device to at least: determine a distance between the location of a camera and the location of the object for each camera in the set of cameras on which the object is detected; and detect the object based on a determination that the distance between the location of the camera and the location of the object satisfies a threshold. In at least some example embodiments, to detect an object, the computer program instructions, when executed by the device, cause the device to at least: based on a determination that the layout of the set of cameras is unknown, periodically activate image slices for each camera in the set of cameras; and detect the object based on the image slices for each camera in the set of cameras. In at least some example embodiments, the set of cameras for which image slices will be activated is determined based on a length of time since the object was last detected satisfying a threshold. In at least some example embodiments, the second camera from the set of cameras is determined to be at the location closest to the object when the object was last detected based on a determination that the second camera has a field of view closest to the location of the object when the object was last detected. In at least some example embodiments, to select the second camera, the computer program instructions, when executed by the device, cause the device to at least: calculate a respective field of view for each camera in the set of cameras; determine a set of distance values that includes: for each camera in the set of cameras, a respective distance between the location of the object when the object was last detected and the respective field of view of the camera; and select the second camera from the set of cameras based on the set of distance values, the second camera being determined to be at the location closest to the object when the object was last detected.In at least some example embodiments, to calculate the corresponding field of view of a camera, the computer program instructions, when executed by a device, cause the device to at least: receive a homography matrix calculated using a reference image resolution that matches the image resolution of a corresponding image from the corresponding camera; obtain a set of points by calculating, for each pixel in a set of pixels of the corresponding image, based on the homography matrix, a corresponding point that specifies the physical location of the region captured by the corresponding pixel; filter any points in the set of points that are beyond a predefined distance threshold in the set of points to form a filtered set of points; and calculate the corresponding field of view of the camera based on a convex structure and using the filtered set of points. In at least some example embodiments, to obtain a selected set of image slices, the computer program instructions, when executed by a device, cause the device to at least: obtain, for each image in a set of images, a corresponding set of image slices for the corresponding image based on the application of image slices to the corresponding image in the set of images; and obtain the selected set of image slices by selecting at least a portion of the image slices in the corresponding set of image slices from each image slice in the set of image slices formed by applying the image slices to the corresponding image in the set of images. In at least some example embodiments, to obtain a corresponding set of image slices for a corresponding image based on the application of image slices to the corresponding image in a set of images, the computer program instructions, when executed by a device, cause the device to at least: determine the number of slices into which the corresponding image in the set of images is to be sliced; and perform slicing of the corresponding image in the set of images based on the number of slices into which the corresponding image in the set of images is to be sliced to form a corresponding set of image slices for the corresponding image in the set of images. In at least some example embodiments, the number of slices into which the corresponding image in the set of images is to be sliced is determined based on at least one of the following: the distance between the camera that captured the corresponding image and an object, or the number of pixels in the bounding box of the detection of the object. In at least some example embodiments, at least a portion of the image slices in the corresponding set of image slices for a corresponding image in the set of images is selected based on a slice selection algorithm for a region of interest (ROI).In at least some example embodiments, to select at least a portion of the image slices in a corresponding set of image slices to form a corresponding set of selected image slices, the computer program instructions, when executed by a device, cause the device to at least: collect a set of bounding boxes for an object; combine the set of bounding boxes for the object into a union bounding box for the object; slice an image from a camera into N image slices, and for each of the N image slices of the corresponding image, obtain the corresponding coordinates of the corresponding image slice; and for each of the image slices, calculate the intersection over union (IoU) between the corresponding image slice and the union bounding box, and based on a determination that the corresponding IoU for the corresponding image slice meets a threshold, then select the corresponding image slice for inclusion in the corresponding set of selected image slices. In at least some example embodiments, based on a cropped slice selection algorithm, at least a portion of the image slices in a corresponding set of image slices for a corresponding image in a set of images is selected. In at least some example embodiments, to select at least a portion of the image slices in a corresponding set of image slices to form a corresponding set of selected image slices, the computer program instructions, when executed by a device, cause the device to at least: collect a set of bounding boxes for an object; combine the set of bounding boxes for the object into a union bounding box for the object; expand the union bounding box based on a defined image size of a machine learning (ML) model to form a region of interest; based on a determination that the size of the region of interest is greater than the defined image size of the ML model in at least one of a height parameter or a width parameter, crop a bounding box region from the image to form a cropped image; and slice the cropped image to obtain a corresponding set of selected image slices. In at least some example embodiments, the cropped image is sliced based on the size of the region of interest and the defined image size of the ML model to obtain a corresponding set of image slices.

[0005] In at least some example embodiments, the method includes: detecting an object within an environment monitored by a set of cameras; based on a determination that the object is no longer detected, determining from the set of cameras a set of cameras for which image slices will be activated, wherein the set of cameras for which image slices will be activated includes: a first camera on which the object was last detected before the object was no longer detected; and a second camera that is determined to be the location closest to the object when the object was last detected; obtaining a set of images, wherein the set of images includes: corresponding images for each camera in the set of cameras for which image slices will be activated, for the respective cameras; obtaining a set of selected image slices from the set of images based on application of image slices to each image in the set of images; and determining the location of the object within the environment based on processing of the set of selected image slices. In at least some example embodiments, detecting the object includes: determining a distance between the location of a camera on which the object is detected from the set of cameras and the location of the object; and detecting the object based on a determination that the distance between the location of the camera and the location of the object satisfies a threshold. In at least some example embodiments, detecting the object includes: based on a determination that the layout of the set of cameras is unknown, periodically activating image slices for each camera in the set of cameras; and detecting the object based on the image slices for each camera in the set of cameras. In at least some example embodiments, the set of cameras for which image slices will be activated is determined based on a length of time since the object was last detected satisfying a threshold. In at least some example embodiments, the second camera from the set of cameras is determined to be the location closest to the object when the object was last detected based on a determination that the second camera has a field of view closest to the location of the object when the object was last detected. In at least some example embodiments, selecting the second camera includes: for each camera in the set of cameras, calculating the respective field of view of the camera; determining a set of distance values, the set of distance values including: for each camera in the set of cameras, a respective distance between the location of the object when the object was last detected and the respective field of view of the camera; and selecting the second camera from the set of cameras based on the set of distance values, the second camera being determined to be the location closest to the object when the object was last detected.In at least some example embodiments, the corresponding field of view of a computing camera includes: receiving a homography matrix calculated using a reference image resolution that matches the image resolution of the corresponding image from the corresponding camera; obtaining a set of points by calculating, for each pixel in a set of pixels of the corresponding image, a corresponding point that specifies the physical location of the region captured by the corresponding pixel, based on the homography matrix; filtering any points in the set of points that are beyond a predefined distance threshold in the set of points to form a filtered set of points; and calculating the corresponding field of view of the camera based on a convex structure and using the filtered set of points. In at least some example embodiments, obtaining a set of selected image slices includes: obtaining, for each image in a set of images, a corresponding set of image slices for the corresponding image based on the application of image slices to the corresponding image in the set of images; and obtaining the set of selected image slices by selecting at least a portion of the image slices in the set of image slices corresponding to the corresponding image from each image slice in the set of image slices formed by applying image slices to the corresponding image in the set of images. In at least some example embodiments, obtaining the corresponding set of image slices for the corresponding image based on the application of image slices to the corresponding image in the set of images includes: determining the number of slices into which the corresponding image in the set of images is to be sliced; and performing slicing of the corresponding image in the set of images based on the number of slices into which the corresponding image in the set of images is to be sliced to form a corresponding set of image slices for the corresponding image in the set of images. In at least some example embodiments, the number of slices into which the corresponding image in the set of images is to be sliced is determined based on at least one of: the distance between the camera capturing the corresponding image and the object, or the number of pixels in the bounding box of the detection of the object. In at least some example embodiments, at least a portion of the image slices in the corresponding set of image slices for the corresponding image in the set of images is selected based on a slice selection algorithm for a region of interest (ROI). In at least some example embodiments, selecting at least a portion of the image slices in the corresponding set of image slices to form a corresponding set of selected image slices includes: collecting a set of bounding boxes for the object; combining the set of bounding boxes for the object into a union bounding box for the object; slicing the image from the camera into N image slices and obtaining the corresponding coordinates of the corresponding image slices for each of the N image slices of the corresponding image; and calculating the intersection over union (IoU) between the corresponding image slice and the union bounding box for each image slice, and based on a determination that the corresponding IoU for the corresponding image slice satisfies a threshold, then selecting the corresponding image slice for inclusion in the corresponding set of selected image slices.In at least some example embodiments, based on a cropping-based slice selection algorithm, for a respective image in a set of images, at least a portion of the image slices in the set of respective image slices is selected. In at least some example embodiments, selecting at least a portion of the image slices in the set of respective image slices to form a set of respective selected image slices includes: collecting a set of bounding boxes for an object; combining the set of bounding boxes for the object into a union bounding box for the object; expanding the union bounding box based on a defined image size of a machine learning (ML) model to form a region of interest; cropping a bounding box region from the image based on a determination that a size of the region of interest is greater than the defined image size of the ML model in at least one of a height parameter or a width parameter to form a cropped image; and slicing the cropped image to obtain the set of respective selected image slices. In at least some example embodiments, the cropped image is sliced based on the size of the region of interest and the defined image size of the ML model to obtain the set of respective image slices.

[0006] In at least some example embodiments, the apparatus includes: components for detecting an object within an environment monitored by a set of cameras; components for determining, based on a determination that the object is no longer detected, a set of cameras from the set of cameras for which image slices will be activated, wherein the set of cameras for which image slices will be activated includes: a first camera on which the object was last detected before the object was no longer detected; and a second camera that is determined to be the location closest to the object when the object was last detected; components for obtaining a set of images, wherein the set of images includes: corresponding images for corresponding cameras, the corresponding cameras being each from the set of cameras for which image slices will be activated; components for obtaining a set of selected image slices from the set of images based on the application of image slices to each image in the set of images; and components for determining the location of the object within the environment based on the processing of the set of selected image slices. In at least some example embodiments, the components for detecting an object include: components for determining a distance between the location of a camera from the set of cameras on which the object was detected and the location of the object; and components for detecting the object based on a determination that the distance between the location of the camera and the location of the object meets a threshold. In at least some example embodiments, the components for detecting an object include: components for periodically activating image slices based on a determination that the layout of the set of cameras is unknown, the image slices being for each camera in the set of cameras; and components for detecting the object based on the image slices, the image slices being for each camera in the set of cameras. In at least some example embodiments, the set of cameras for which image slices will be activated is determined based on a length of time since the object was last detected meeting a threshold. In at least some example embodiments, the second camera from the set of cameras is determined to be the location closest to the object when the object was last detected based on a determination that the second camera has a field of view closest to the location of the object when the object was last detected. In at least some example embodiments, the components for selecting the second camera include: components for calculating a corresponding field of view for each camera in the set of cameras; components for determining a set of distance values, the set of distance values including: for each camera in the set of cameras, a corresponding distance between the location of the object when the object was last detected and the corresponding field of view of the camera; and components for selecting the second camera from the set of cameras based on the set of distance values, the second camera being determined to be the location closest to the object when the object was last detected.In at least some example embodiments, the components for calculating the respective fields of view of the cameras include: a component for receiving a homography matrix calculated using a reference image resolution that matches the image resolution of the respective images from the respective cameras; a component for obtaining a set of points by calculating, for each pixel in a set of pixels of the respective image and based on the homography matrix, a respective point specifying the physical location of the region captured by the respective pixel; a component for filtering any points in the set of points that are beyond a predefined distance threshold in the set of points to form a filtered set of points; and a component for calculating the respective fields of view of the cameras based on a convex structure and using the filtered set of points. In at least some example embodiments, the components for obtaining a selected set of image slices include: a component for obtaining, for each image in the set of images, a respective set of image slices for the respective image based on the application of image slices to the respective image in the set of images; and a component for obtaining the selected set of image slices by selecting at least a portion of the image slices in the respective set of image slices from each image slice in the set of image slices, the set of image slices being formed from the application of image slices to the respective image in the set of images. In at least some example embodiments, the component for obtaining, based on the application of image slices to the respective image in the set of images, a respective set of image slices for the respective image includes: a component for determining the number of slices into which the respective image in the set of images is to be sliced; and a component for performing the slicing of the respective image in the set of images based on the number of slices into which the respective image in the set of images is to be sliced to form a respective set of image slices for the respective image in the set of images. In at least some example embodiments, the number of slices into which the respective image in the set of images is to be sliced is determined based on at least one of the following: the distance between the camera capturing the respective image and the object, or the number of pixels in the bounding box of the detection of the object. In at least some example embodiments, at least a portion of the image slices in the respective set of image slices for the respective image in the set of images is selected based on a slice selection algorithm for a region of interest (ROI).In at least some example embodiments, the components for selecting at least a portion of the image slices in a corresponding set of image slices to form a corresponding set of selected image slices include: components for collecting a set of bounding boxes for an object; components for combining the set of bounding boxes for the object into a union bounding box for the object; components for slicing an image from a camera into N image slices, and components for obtaining corresponding coordinates of the corresponding image slices for each of the N image slices of the corresponding image; and components for, for each image slice in the image slices, calculating the intersection over union (IoU) between the corresponding image slice and the union bounding box, and then selecting the corresponding image slice based on the determination that the corresponding IoU for the corresponding image slice meets a threshold, for inclusion in the corresponding set of selected image slices. In at least some example embodiments, based on a cropped slice selection algorithm, at least a portion of the image slices in a corresponding set of image slices for a corresponding image in a set of images is selected. In at least some example embodiments, the components for selecting at least a portion of the image slices in a corresponding set of image slices to form a corresponding set of selected image slices include: components for collecting a set of bounding boxes for an object; components for combining the set of bounding boxes for the object into a union bounding box for the object; components for expanding the union bounding box based on a defined image size of a machine learning (ML) model to form a region of interest; components for cropping a bounding box region from an image based on the determination that at least one of a height parameter or a width parameter of the size of the region of interest is greater than the defined image size of the ML model, to form a cropped image; and components for slicing the cropped image to obtain a corresponding set of selected image slices. In at least some example embodiments, the cropped image is sliced based on the size of the region of interest and the defined image size of the ML model to obtain a corresponding set of image slices. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] The teachings herein can be readily understood by considering the following detailed description in conjunction with the accompanying drawings, in which:

[0008] Figure 1 An example embodiment of an environment including an object detection and tracking system is depicted, the object detection and tracking system being configured to: perform object detection and tracking for objects within the environment;

[0009] Figure 2 An example embodiment of a method for performing object detection and tracking in a multi-camera system is depicted;

[0010] Figure 3 An example embodiment of a method for calculating the field of view of a camera based on a homography matrix and a distance limit is depicted;

[0011] Figure 4 Illustrates an example embodiment of a slice activation pipeline for supporting image slice activation in a multi-camera system;

[0012] Figure 5 Illustrates an example embodiment of a method for implementing a slice activation pipeline;

[0013] Figure 6 Illustrates an example embodiment of a method for camera selection for image slices for object detection;

[0014] Figure 7 Illustrates an example embodiment of a method for closest camera selection for image slices for object detection;

[0015] Figure 8 Illustrates an example embodiment of a method for image slices for object detection that combines periodic monitoring and camera selection;

[0016] Figure 9 Illustrates an example embodiment of image slice activation and image slice deactivation based on the position of a detected object of interest relative to a camera;

[0017] Figure 10 Illustrates an example embodiment of a method for selecting a camera for image slices for object detection;

[0018] Figure 11 Illustrates an example embodiment of obtaining a set of image slices by slicing an image via a set of horizontal splits and a set of vertical splits;

[0019] Figure 12 Illustrates an example embodiment of changing the number of image slices used in a slice based on the distance of an object to a camera position;

[0020] Figure 13 Illustrates an example embodiment of changing the number of image slices used in a slice based on the number of pixels in a detected bounding box or contour;

[0021] Figure 14 Illustrates an example embodiment of a method for determining the number of image slices to be used when slicing an image based on distance or number of pixels;

[0022] Figure 15 Illustrates an example embodiment of a method for determining the number of image slices to be used when slicing an image in object detection;

[0023] Figure 16Depicts the selection of a subset of image slices of an image based on a region of interest (ROI) slice selection algorithm;

[0024] Figure 17 Depicts an example embodiment of an ROI-based slice selection algorithm configured to select a subset of image slices of an image for processing;

[0025] Figure 18 Depicts the selection of a subset of image slices of an image based on a cropping-based slice selection algorithm;

[0026] Figure 19 Depicts an example embodiment of a cropping-based slice selection algorithm configured to select a subset of image slices of an image for processing;

[0027] Figure 20 Depicts an example embodiment of an object tracking system configured to perform object tracking for a set of objects in an environment; and

[0028] Figure 21 Depicts an example embodiment of a computer suitable for performing the various functions presented herein.

[0029] For ease of understanding, the same reference numerals are used throughout this document as much as possible to denote the same elements common to the various figures. Detailed Description

[0030] Various example embodiments for supporting object detection in a machine vision system are presented. Various example embodiments for supporting object detection in a machine vision system can be configured to: support range extension for object detection in a machine vision system. Various example embodiments for supporting object detection in a machine vision system can be configured to: support flexible and efficient range extension for object detection in a machine vision system. Various example embodiments for supporting flexible and efficient range extension for object detection in a machine vision system can be configured to: support flexible and efficient range extension for object detection in a machine vision system by: selecting a set of cameras for which image slices will be used, capturing images from the set of cameras for which image slices will be used, and processing images from the set of cameras for which image slices will be used for object detection (e.g., determining the number of slices to be used, performing image slicing, selecting which image slices to process for object detection, processing the selected image slices for object detection, combining detections across image slices, etc., and various combinations thereof), and optionally, for other functions related to object detection (e.g., object localization, object tracking, object control, etc., and various combinations thereof). Various example embodiments for supporting flexible and efficient range extension for object detection in a machine vision system can be configured to: support flexible and efficient range extension for object detection in a machine vision system in various environments that support object detection for various object types (e.g., automatic object detection in an automated factory, automated mining, automated shipping fulfillment, an automated vehicle tracking system, etc., and various combinations thereof).

[0031] Various example embodiments for supporting object detection in a machine vision system can be configured to support range extension for object detection in a machine vision system. Since a camera sensor provides a rich source of information about the physical space in which the camera is installed, a machine vision system is a core component of various types of systems that perform vision-based object tracking, such as automation systems in Industry 4.0 / 5.0, vehicle tracking in autonomous vehicle systems, etc. Object detection continues to play an important role in a machine vision system. Since excellent performance is usually associated with the use of such ML models, many machine vision systems utilize machine learning (ML) models for object detection; however, since such ML models typically consume a large amount of resources during both training and inference, such ML models are typically kept at a reasonable size by continuously training on images with a relatively small image size. However, this leads to the following problem: Detecting and recognizing objects that appear small due to their distance or their size is a greater challenge because the relatively small image size used during model training means fewer pixels in the image plane capture the object of interest, thus affecting the probability of detection. This problem is further exacerbated when an object moves (e.g., a robot moves in a long corridor with relatively sparse fixed cameras) in a large space (e.g., a factory, a warehouse, etc.) with little coverage area by the camera sensor. Various example embodiments for supporting object detection in a machine vision system can be configured to use image slicing to support range extension for object detection in a machine vision system without retraining the ML model and while keeping the overall computational complexity relatively low, based on: intelligent selection of the camera for which slices will be activated, intelligent determination of the number of slices to be used, intelligent slicing of the image to form image slices, intelligent selection of the image slices to be processed for object detection, intelligent combination of object detections across image slices, etc.

[0032] It should be understood that, as Figure 1 shown, these, and various other example embodiments for flexible and efficient range extension for supporting object detection in a machine vision system, can be further understood by considering a machine vision system in the form of a multi-camera system, which is configured to perform object detection and tracking for a set of objects in an environment.

[0033] Figure 1 An example embodiment depicting an environment including an object detection and tracking system, which is configured to perform object detection and tracking for objects within the environment.

[0034] As Figure 1As depicted, environment 100 includes: a set of objects 110 to be detected and tracked within environment 100 (schematically, two objects 110 are shown as: object 110-1 and object 110-2), a set of cameras 120 configured to capture images within environment 100 (schematically, three cameras 120 are shown as: camera 120-1, camera 120-2, and camera 120-3), and a controller 130 configured to perform processing of the images captured by cameras 120 to support detection and tracking of objects 110 within environment 100. It should be noted that cameras 120 and controller 130 can cooperate to provide an object detection and tracking system configured to support detection and tracking of objects 110 within environment 100.

[0035] Environment 100 can include a physical space within which object detection and tracking can be performed. The physical space can include the physical space within which objects 110 are located, where objects 110 can move within the physical space for various purposes, and where cameras 120 can be deployed at various locations within the physical space to capture images of the physical space to facilitate detection and tracking of objects 110 within environment 100. For example, environment 100 can be a building that houses a factory in the case of an automated factory, a mining site in the case of an automated mining operation, a geographical area associated with a road network in the case of an automated vehicle tracking service, etc., and various combinations thereof. It should be understood that environment 100 can include various other types of indoor and / or outdoor locations. It should be understood that although primarily presented as having a specific arrangement of the physical space within environment 100, the physical space within environment 100 can be implemented in various other ways.

[0036] The set of objects 110 can include any objects that can be located within the physical space of environment 100 and can move within the physical space of environment 100, which may depend on the type of environment of environment 100. For example, when environment 100 is a factory (as shown in the example of Figure 1 ), objects 110 can include factory robots, autonomous vehicles, etc., and various combinations thereof. For example, when environment 100 is a mining site (not shown in the example of Figure 1 ), objects 110 can include mining equipment, autonomous vehicles, etc., and various combinations thereof. It should be understood that various other types of objects can be detected and tracked in various other environments. It should be understood that although primarily presented with respect to a set of objects 110 including a specific number of objects to be detected and tracked within environment 100, the set of objects 110 can include fewer or more objects to be detected and tracked within environment 100.

[0037] The set of cameras 120 can include: any camera configured to capture images within the physical space of the environment 100. For example, the camera 120 can be configured to: capture images including pictures, videos, etc., and various combinations thereof. The camera 120 can include: local processing resources configured to: perform various processing functions to support object detection and tracking for the set of objects 110 in the environment 100. The camera 120 can include: communication resources configured to: support communication between the camera 120 and the controller 130 to support object detection and tracking for the set of objects 110 in the environment 100. It should be understood that although the use of the set of cameras 120 including a specific number of cameras, with the specific number of cameras having specific positions within the physical space of the environment 100, the set of cameras 120 can include fewer or more cameras, can include cameras having different positions within the physical space of the environment 100, etc., and various combinations thereof.

[0038] The controller 130 is configured to: support the detection and tracking of the object 110 based on the images captured by the set of cameras 120. The controller 130 is configured to: support the detection and tracking of the object 110 based on the use of image slices, the use of image slices including: selecting one of the cameras 120 for which the image slices will be used, capturing an image from one of the cameras 120 for which the image slices are to be used, and processing the image from one of the cameras 120 for which the image slices will be used for object detection (e.g., determining the number of image slices to be used, performing image slicing, selecting which image slices to process for object detection, processing the selected image slices for object detection, combining detections across the image slices, etc., and various combinations thereof). It should be understood that based on the application of the image slices of the images captured by the set of cameras 120, the operation of the controller 130 in performing the detection and tracking of the set of objects 110 can be further understood to be performed by referring to Figure 2 to perform. It should be understood that although the use of a single controller 130 located within the physical space of the environment 100 is mainly proposed, the controller 130 can be implemented in various other ways (e.g., using multiple local controllers located within the environment 100 or otherwise associated with the environment 100, using one or more edge computing devices or edge computing resources accessible from the environment 100, using one or more cloud computing devices or cloud computing resources accessible from the environment 100, etc., and using various combinations thereof).

[0039] It should be understood that the environment 100 can be implemented in various other ways to support object detection and tracking.

[0040] Figure 2 Illustrates an example embodiment of a method for performing object detection and tracking in a multi-camera system. It should be understood that although method 200 is primarily presented as being performed sequentially, at least some of the functions of method 200 may be performed simultaneously or in a different order than that shown by reference Figure 2 as shown.

[0041] At block 201, method 200 begins.

[0042] At block 210 (also referred to herein as step 1), prior information for performing object detection is collected. Prior information collection may include: calculating or otherwise obtaining a homography matrix, calculating the camera FOV of a camera based on the homography matrix, etc., and various combinations thereof. It should be understood that aspects of prior information collection will be further discussed below, and at least some example embodiments of prior information collection may be further understood by reference Figure 3 thereto.

[0043] At block 220 (also referred to herein as step 2), image slices for one or more cameras are activated. Activation of the image slices may include: selection of one or more cameras for which image slices are to be activated, initiation of slice activation for one or more cameras for which image slices are to be activated, etc., and various combinations thereof. Selection of the (multiple) cameras for which image slices are activated may be based on the use of a combination of: periodic monitoring, and distance metrics related to the position of the object relative to the camera position and field of view. It should be noted that the use of image slices enables an expansion of the range of object detection and an improvement in the ability to detect small objects. It should be understood that aspects of image slice activation will be further discussed below, and at least some example embodiments of image slice activation may be further understood by reference Figures 4 to 10 thereto.

[0044] At block 230 (also referred to herein as step 3), the number of image slices to be used for performing image slicing is determined. The determination of the number of image slices to be used for performing image slicing may be based on one or both of the following two items: changing the number of image slices according to the distance of the object (based on the last detection) to the camera position, or changing the number of image slices according to the number of pixels in the detection bounding box or contour. It should be understood that aspects of the number of image slices to be used for performing image slicing will be further discussed below, and at least some example embodiments of determining the number of image slices to be used for performing image slicing may be further understood by reference Figures 11 to 14 thereto.

[0045] At block 240 (also referred to herein as step 4), it is determined which image slices are to be processed. The determination of which image slices are to be processed can be based on: computational constraints associated with performing slice processing, which can vary in different situations. The determination of which image slices are to be processed can be based on: a slice selection algorithm based on a region of interest (ROI), or a slice selection algorithm based on cropping. It should be understood that the various aspects of determining which image slices are to be processed will be further discussed below, and at least some example embodiments of determining which image slices are to be processed can be further understood by reference to Figures 15 to 18 herein.

[0046] At block 250 (also referred to herein as step 5), image slicing is performed to slice the original image captured by the camera and selected for activation of the image slice. Image slicing is performed to obtain the image slices that have been selected for object detection processing. The result of the image slicing of the image is a set of image slices, the set of image slices including: the image slices that will be processed for object detection.

[0047] At block 260 (also referred to herein as step 6), object detection is performed for each of the image slices to be processed. Object detection for the image slices can be performed based on a pre-trained ML model detector. The result of the object detection performed for each image slice is a set of object detections, the set of object detections for each image slice in the set of image slices to be processed.

[0048] At block 270 (also referred to herein as step 7), object detection and tracking are performed. Object detection and tracking can be performed by combining the set of object detections for each image slice in the set of image slices to be processed. The result of object detection and tracking is: the set of object detections, and the localization information for the object (e.g., the object has been detected, and the position of the object within the environment has been determined).

[0049] At block 299, method 200 ends. It should be understood that although mainly presented in an ending manner (for clarity purposes), method 200 can include: one or more additional blocks for additional functions that can be performed after the object has been detected, such as, performing tracking of the object, performing object recognition for the object, initiating one or more control functions to control one or more aspects of the object (e.g., speed, trajectory, the function(s) being performed, etc., and various combinations thereof), etc., and various combinations thereof. It should be understood that although mainly presented in an ending manner (for clarity purposes), method 200 can continue to execute and / or re-execute for continued tracking of objects in the environment.

[0050] In step 1 regarding Figure 2 as represented by box 210 in Figure 2 , prior information collection is performed. To develop a camera selection algorithm and potentially select the next camera for image slice activation based on the position of an object in space, having information about the field of view of the camera is an important step. Although the field of view of a camera can be determined using the camera information of the camera (i.e., the focal length of the camera and the size of the camera sensor of the camera), and the distance from the camera to the subject, in at least some example embodiments, the field of view of the camera can be calculated using a homography matrix. The homography matrix transforms pixel image points to physical space. Given the image size, the homography matrix, and the defined viewing limits, the field of view of the camera can be constructed by: performing some preliminary processing to calculate the homography matrix and determine the distance limit, and then calculating the field of view of the camera based on the homography matrix and the distance limit.

[0051] The homography matrix (Φ) can be calculated by collecting ground truth data. This involves identifying reference points in the field of view of the camera and, for each reference point, finding the pixel coordinates w = (u, v) and the corresponding real-world coordinates r = (x, y). This data is used to calculate the transformation matrix (referred to as the homography matrix) that predicts the real-world coordinates with the minimum error relative to the ground truth world coordinates when the input pixel coordinates are provided.

[0052] The distance limit (D L ) and the image resolution depend on: the camera model used and the configuration parameters of the camera. The distance limit is the range of detection for a given camera and the ML model used for detection. Typically, this corresponds to the farthest distance of the smallest-sized object (from the class of objects for which the model was trained) that can be detected within it. This is determined by applying the trained ML to the video output without performing image slicing. For example, using cameras in a factory with an image resolution of 1920x1080, the smallest robots detected in most cameras are up to 20 meters, which is considered the distance limit for these cameras.

[0053] Given a homography matrix and a distance limit, the field of view of a camera can be calculated as follows: (1) obtain an image having the same resolution as that used to calculate the homography matrix, (2) for each pixel of the image, apply the calculated homography matrix to predict the real-world coordinates of the pixel (i.e., the physical location of the area captured by the pixel), (3) filter out points that exceed a predefined distance limit, and (4) use the calculated points to construct a continuous shape of the view (e.g., based on a convex structure, such as a convex hull structure or other suitable type of convex structure). It should be understood that these steps are general enough to be applied to any deployed camera, rather than being specific to a particular location. This process can be further understood by referring to Figure 3 for further understanding.

[0054] Figure 3 An example embodiment of a method for calculating the field of view of a camera based on a homography matrix and a distance limit is depicted. It should be understood that Figure 3 method 300 of Figure 2 can be used to implement Figure 3 block 210 of Figure 3 . It should be understood that although omitted for clarity,

[0055] method 300 of Figure 2 may include: one or more initial blocks for calculating the homography matrix and / or the distance limit. It should be understood that although presented primarily as being executed sequentially, at least part of the functionality of method 300 can be executed simultaneously, or in a different order than that shown with respect to Figure 2 Figure 3 . At block 301, method 300 begins. At block 310, a homography matrix calculated using a reference image resolution that matches the image resolution of the corresponding image from the corresponding camera is received. At block 320, a set of points is obtained by calculating, for each pixel in a set of pixels of the corresponding image, a corresponding point specifying the physical location of the area captured by the corresponding pixel, based on the homography matrix. The result for the set of cameras is illustrated as block 321. At block 330, any points in the set of points that exceed a predefined distance threshold in the set of points are filtered out to form a filtered set of points. At block 340, the corresponding field of view of the camera is calculated based on a convex structure and using the filtered set of points. The result for the set of cameras is illustrated as block 341. At block 399, method 300 ends.

[0055] In step 2 as discussed with respect to Figure 2 (in Figure 2In (represented as box 220), image slices are activated for one or more cameras. Activation of an image slice can include: selection of the (multiple) cameras for which its image slice will be activated (i.e., determining where and when the slice pipeline should be activated). Activation of an image slice can be based on a camera selection algorithm that is configured to: select the (multiple) cameras for which its image slice will be activated. The camera selection algorithm can be configured to: support image slice activation based on the use of a camera selection pipeline that is coupled to a periodic sensing of an environment with full image slice activation, and thus, the camera selection algorithm can be divided into two parts: periodic camera monitoring and the camera selection pipeline. The camera selection algorithm can be configured to: support image slice activation in a way that reduces system resource usage. The camera selection algorithm can be further understood by referring to Figures 4 to 10 for further understanding.

[0056] Figure 4 Depicts an example embodiment of a slice activation pipeline for supporting image slice activation in a multi-camera system.

[0057] As Figure 4 shown, the slice activation pipeline 400 includes: a box 410 for initiating image slice activation in a multi-camera system, and then proceeding to boxes 420 and 421 for periodic camera monitoring, or proceeding to boxes 430 and 431 for the camera selection pipeline.

[0058] As shown in box 420 and box 421, periodic camera monitoring activates image slices for all cameras to check for moving objects in a long corridor. As shown in box 421, periodic camera monitoring can be performed by: every N seconds, turning on the image slices for all cameras for T seconds. It should be understood that periodic camera monitoring may be useful in cases where an object first appears in physical space and is not detected by any camera. In other words, periodic camera monitoring helps to detect distant objects in the space first.

[0059] As shown in block 430 and block 431, a camera selection is made to select a camera and activate image slices for the selected camera and the cameras closest to the selected camera. As shown in block 431, the camera selection pipeline may be performed by: if an object disappears from the current camera, activating the image slices for that camera and the closest camera for M seconds. It should be understood that the camera selection pipeline may depend on events such as the last position where the object was detected and the use of the layout of the camera network. Typically, camera networks are deployed to address specific use cases such as monitoring or tracking, and thus the layout of the camera network is known; however, if information about the camera network layout or camera positions is not available, the camera selection pipeline may rely on periodic monitoring to first detect distant objects, and after this occurs, the image slices will remain active (such that computational resource usage is still minimized).

[0060] The slice activation pipeline 400 for controlling the intelligent use of periodic camera monitoring and the camera selection pipeline may be supported using the following two distance metrics: (1) the distance between the position of the object and the field of view of the camera (denoted as D cam-fov,obj ), and (2) the distance between the position of the object and the position of the camera (denoted as D cam-location,obj ). In the case of D cam-fov,obj , the position of the object may be calculated by applying the computed homography matrix (from step 1) to the predicted pixel position of the detected object : And since the field of view of the camera has been calculated (in step 1), the field of view of the camera is known. In the case of D cam-location,obj , the position of the object is known from the calculation of the object position as discussed with respect to D cam-location,obj (or may be calculated in the same manner as discussed above with respect to D cam-location,obj ), and assuming the position of the camera is known (e.g., based on CAD, a map of the region of interest, etc.).

[0061] In the slice activation pipeline 400, where the camera selection pipeline is used and the closest camera needs to be determined, the closest camera may be defined as: the camera with the field of view closest to the current position of the object. The field of view is a shape formed by lines. The distance between the object and the FoV may be calculated using an algorithm configured to: calculate the distance between a point and a line segment, where the position of the object is considered the point and each line forming the FoV is considered. In the FoV, the minimum distance between the point and any line is used to determine the closest camera.

[0062] It should be understood that the camera with the field of view closest to the current position of the object may not necessarily be the camera that is physically closest to the current position of the object, because there may be one or more cameras that, although physically closer to the current position of the object, may have an occluded field of view of the object. Therefore, they may not be selected as the closest camera for object detection because the camera will not be able to capture an object image sufficient for image processing purposes. This can be seen, for example, Figure 1 in the environment 100, where it can be seen that, since the field of view of camera 120-2 to object 110-1 is occluded by the wall therebetween, camera 120-2 may not be considered as a candidate for the closest camera for the detection of object 110-1. It should be understood that determining whether a camera has an occluded field of view of an object such that the camera may not be selected as the closest camera for object detection can be performed in various ways. It should be understood that, based on a determination that the degree of occlusion is within an acceptable degree of occlusion (e.g., not occluded at all, degree of occlusion less than 1%, degree of occlusion less than 5%, degree of occlusion less than 10%, etc.), a camera determined to have an occluded field of view of an object can still be selected as the closest camera for that object, and the known or expected degree of occlusion can still achieve acceptable results in terms of object detection and tracking, where the acceptable degree of occlusion and / or acceptable results can be different for different object types, different environmental types, etc., and various combinations thereof.

[0063] Figure 5 An example embodiment of a method for implementing a slice activation pipeline is depicted. It should be understood that Figure 5 the method 500 of Figure 4 can be used to implement Figure 5 the slice activation pipeline 400 of Figure 5 . It should be understood that, although omitted for clarity, Figure 5 the method 500 of cam-fov,obj and D LIf the object is detected (determined by, for example, the object being within a certain distance from the camera), then the object is detected. If the object is far from the camera, then periodic image slices for each camera are used to increase the probability of detection. At block 530, after the object disappears (i.e., the object is detected and then stops being detected), the camera continues to monitor the object and tracks the length of time the object has not been detected. At block 540, it is determined whether the length of time the object has not been detected meets a threshold (denoted as last_seen_limit). If the length of time the object has not been detected does not meet the threshold, then method 500 returns to block 530 such that the camera continues to monitor the object and will continue to track the length of time the object has not been detected (i.e., the object detection timer continues to run as long as the object is not detected). If the length of time the object has not been detected meets the threshold, then method 500 proceeds to block 550. At block 550, the camera where the object was last detected and the closest camera use image slices to attempt to detect the object. Here, based on the determination that the length of time the object has not been detected meets the threshold, the image slices are activated for the camera where the object was last detected and the closest camera, and then the camera where the object was last detected and the closest camera use the image slices to attempt to detect the object. At block 560, it is determined whether the object has been detected. If the object has been detected, then method 500 proceeds to block 599 where method 500 ends. If the object has not been detected, then method 500 proceeds to block 570. At block 570, it is determined whether the length of time the image slices for the camera and the closest camera have been activated exceeds a threshold. If the length of time the image slices for the camera have been activated does not exceed the threshold, then method 500 returns to block 560 such that the image slices will continue to be used to monitor the object and the length of time the image slices have been activated will continue to be tracked (i.e., the image slice activation timer continues to run as long as the object is not detected). If the length of time the image slices for the camera have been activated does not exceed the threshold, then method 500 proceeds to block 580. At block 580, the camera used for the image slices is deselected and method 500 proceeds to block 599 where method 500 ends. At block 599, method 500 ends.

[0064] Figure 6 An example embodiment of a method for using a combination of periodic monitoring and camera selection for object detection using image slices is depicted. It should be understood that Figure 6 method 600 can be used to implement Figure 4 parts of the slice activation pipeline 400, and / or Figure 5 parts of method 500. It should be understood that, although omitted for clarity, Figure 6 method 600 can include: one or more initial blocks to perform processing to obtain atFigure 6 the information used in the context of method 600. It should be understood that although mainly presented as being executed continuously, at least some of the functions of method 600 may be executed simultaneously or in an order different from the order shown with respect to Figure 6 Shown. At block 601, method 600 begins. At block 610, messages from the positioning socket are listened for. At block 620, every monitoring_frequency, slices for all configured cameras are activated, where monitoring_frequency indicates the frequency of the slices activated when all cameras perform periodic detection. At block 630, based on the data in the object_per_camera_info dictionary, which is a dictionary including information identifying the objects detected in the field of view of the cameras, the necessary cameras are activated. In at least some example embodiments, block 630 may use Figure 7 The method to implement. At block 640, a message is received from the positioning socket, the message including: information indicating the (multiple) positions of the object and the current (multiple) cameras (represented as active_camera) for which image slices have been activated for the object. At block 650, it is determined whether the current (multiple) cameras for which image slices have been activated for the object are configured. If the current (multiple) cameras for which image slices have been activated for the object are not configured, method 600 returns to block 610 to continue listening for messages from the positioning socket. If the current (multiple) cameras for which image slices have been activated for the object are configured, method 600 proceeds to block 660. At block 660, the object information is saved. At block 670, the object information on the closest camera is updated. In at least some example embodiments, block 670 may use Figure 8 The method to implement. At block 699, method 600 ends.

[0065] Figure 7 An example embodiment of a method for camera selection for image slices for object detection is depicted. It should be understood that Figure 7 Method 700 can be used to implement Figure 4 Parts of the slice activation pipeline 400 of Figure 5 And / or Figure 7 Parts of method 500 of Figure 6 It should be understood that Figure 7 Method 700 can be used to implement Figure 7The information used in the context of method 700. It should be understood that although mainly presented as being executed sequentially, at least portions of the functionality of method 700 may be executed simultaneously, or in an order different from the order shown with respect to Figure 7 At block 701, method 700 begins. At block 710, it is determined whether the object for which camera selection is being performed exists in the object_per_camera_info dictionary, which is a dictionary that includes information identifying the objects detected in the field of view of the cameras. If the object for which camera selection is being performed does not exist in the object_per_camera_info dictionary, method 700 proceeds to block 799 where method 700 ends. If the object for which camera selection is being performed exists in the object_per_camera_info dictionary, method 700 proceeds to block 720. At block 720, it is determined whether a threshold time length (denoted as last_seen_limit, and which represents the time length after the closest camera where the object was last seen is activated for slicing) has been exceeded since the object was last detected by a camera. If the threshold time length (last_seen_limit) has not been exceeded since the object was last detected by a camera, method 700 returns to block 710. If the threshold time length (last_seen_limit) has been exceeded since the object was last detected by a camera, method 700 proceeds to block 730. At block 730, the camera on which the object was last detected is activated for image slicing. This camera is saved to camera_collection in the database, where camera_collection is a list of cameras currently activated for image slicing. At block 740, the closest camera is selected as the camera closest to the camera on which the object was last detected, and this closest camera is activated for image slicing. This closest camera is saved to camera_collection in the database, where camera_collection is a list of cameras currently activated for image slicing. At block 750, the object is removed from the object_per_camera_info dictionary, which is a dictionary that includes information identifying the objects detected in the field of view of the cameras. Method 700 returns from block 750 to block 710.

[0066] Figure 8 Depicts an example embodiment of a method for closest camera selection for image slicing for object detection. It should be understood that Figure 8 Method 800 can be used to implement Figure 4 Portions of the slicing activation pipeline 400, and / orFigure 5 Part of method 500. It should be understood that Figure 8 Method 800 can be used to implement Figure 6 Box 670 of method 600. It should be understood that, although omitted for clarity, Figure 8 Method 800 may include: one or more initial boxes to perform processing to obtain information used in the context of Figure 8 Method 800. It should be understood that, although mainly presented as being executed continuously, at least part of the functions of method 800 can be executed simultaneously, or in an order different from the order shown with respect to Figure 8 At box 801, method 800 begins. At box 810, it is determined whether the object for which camera selection is being performed exists in the object_per_camera_info dictionary, which is a dictionary including information identifying objects detected in the field of view of the camera. If the object for which camera selection is being performed does not exist in the object_per_camera_info dictionary, method 800 proceeds to box 899, where method 800 ends. If the object for which camera selection is being performed exists in the object_per_camera_info dictionary, method 800 proceeds to box 820. At box 820, the field of view of the closest camera (denoted as closest_camera) is determined. Based on the current predicted position of the object and a priori calculations of the camera FoV, the closest camera (closest_camera) is the camera with the FoV closest to that of the given object. It should be noted that the camera with the FoV closest to that of the given object can be the camera with the closest unobstructed view of the object (or, if the view is partially obstructed, with an acceptable level of object occlusion). At box 830, the distance between the position of the object and the closest camera (closest_camera) is determined. This distance is denoted as object_to_camera_dist (which can also be denoted as D cam-location,obj)。At block 840, it is determined whether the distance (object_to_camera_dist) between the position of the object and the closest camera is greater than a threshold (represented as distance_limit). If the distance between the position of the object and the closest camera is not greater than the threshold (object_to_camera_dist < distance_limit), method 800 returns to block 810. If the distance between the position of the object and the closest camera is greater than the threshold (object_to_camera_dist > distance_limit), method 800 proceeds to block 850. At block 850, the object_per_camera_info dictionary (i.e., the dictionary that includes information identifying the objects detected in the field of view of the camera) is updated to include: the time when the object was last detected, and an indication of the closest camera.

[0067] Figure 9 Depicts an example embodiment in accordance with the position of the detected object of interest relative to the camera, image slice activation, and image slice deactivation.

[0068] For example, if the distance between the object and the position of the closest camera (represented as D cam-location,obj ) is relatively small (e.g., less than D L ), it is assumed that the object will appear in front of the camera, and thus no image slice is required (i.e., the object will be large enough to be detected). This is shown in Figure 9 for object positions "1", "2", "3", and "4", for which the indicated image slice is "OFF".

[0069] For example, if the object appears at the edge of the field of view of the closest camera, in which case the distance between it and the camera position is large (e.g., greater than D L ), therefore, the image slice should be activated. This is shown in Figure 9 for object position "5", for which the indicated image slice is "ON".

[0070] For example, after the object disappears from the current scene, the current camera and the closest camera will be activated for image slicing. This is illustrated in Figure 9 for object position "6", for which the indicated image slice is "ON". It should be noted that the current camera will be activated for image slicing because the object may still be in the same corridor, but the detection may be lost.

[0071] It should be understood that these are just a few scenarios where intelligent activation and deactivation of image slices can be used to support range extension for improved object detection in a camera system.

[0072] Figure 10 Illustrates an example embodiment of a method for a camera to select image slices for object detection. It should be understood that although mainly presented as being executed continuously, at least some of the functions of method 1000 may be executed simultaneously, or in an order different from the order shown with respect to Figure 10 the order shown.

[0073] At block 1001, method 1000 begins.

[0074] At block 1010, within the environment monitored by a set of cameras, an object is detected. In at least some example embodiments, the detection of the object may include: determining the distance between the position of the camera and the position of the object for the camera in the set of cameras on which the object is detected; and detecting the object based on the determination that the distance between the position of the camera and the position of the object satisfies a threshold. In at least some example embodiments, the detection of the object may include: periodically activating image slices for each camera in the set of cameras based on the determination that the layout of the set of cameras is unknown; and detecting the object based on the image slices for each camera in the set of cameras.

[0075] At block 1020, based on the determination that the object is no longer detected, a set of cameras is determined from the set of cameras for which image slices will be activated, where the set of cameras for which image slices will be activated includes: a first camera on which the object was last detected before the object was no longer detected; and a second camera from the set of cameras that is determined to be: the location closest to the object when the object was last detected. In at least some example embodiments, based on the determination that the length of time since the object was last detected meets a threshold, the set of cameras for which image slices will be activated is determined. In at least some example embodiments, based on the determination that the second camera has a field of view closest to the location of the object when the object was last detected, the second camera from the set of cameras can be determined to be: closest to the location of the object when the object was last detected. In at least some example embodiments, the selection of the second camera can include: for each camera in the set of cameras, calculating the corresponding field of view of the camera; determining a set of distance values that includes: for each camera in the set of cameras, the corresponding distance between the location of the object and the corresponding field of view of the camera when the object was last detected; and based on the set of distance values, selecting the second camera from the set of cameras determined to be closest to the location of the object when the object was last detected. In at least some example embodiments, calculating the corresponding field of view of the camera can include: receiving a homography matrix calculated using a reference image resolution that matches the image resolution of the corresponding image from the corresponding camera; obtaining a set of points by calculating, for each pixel in a set of pixels in the corresponding image and based on the homography matrix, the corresponding point that specifies the physical location of the region captured by the corresponding pixel; filtering any points in the set of points that are beyond a predefined distance threshold in the set of points to form a filtered set of points; and calculating the corresponding field of view of the camera based on a convex structure and using the filtered set of points.

[0076] At block 1030, a set of images is obtained, where the set of images includes: the corresponding images for the corresponding cameras from each camera in the set of cameras for which image slices will be activated. In at least some example embodiments, the set of images can include: at least one image from each camera in the set of cameras for which image slices will be activated (e.g., at least one image captured by the first camera and at least one image captured by the second camera).

[0077] At block 1040, a set of selected image slices is obtained from the set of images based on the application of image slicing to each image in the set of images. In at least some example embodiments, obtaining the set of selected image slices may include: for each image in the set of images, obtaining a set of corresponding image slices for the corresponding image based on applying image slicing to the corresponding image in the set of images (e.g., by determining the number of slices into which the corresponding image in the set of images is to be sliced (such as based on at least one of: the distance between the camera capturing the corresponding image and the object, or the number of pixels in the bounding box of object detection), and based on the number of slices into which the corresponding image in the set of images is to be sliced, performing slicing on the corresponding image in the set of images to form a set of corresponding image slices for the corresponding image in the set of images); and obtaining the set of selected image slices by selecting at least a portion of the image slices in the set of corresponding image slices from each image slice in the set of image slices formed by applying image slicing to the corresponding image in the set of images (e.g., based on at least one of: a slice selection algorithm based on a region of interest (ROI), or a slice selection algorithm based on cropping).

[0078] At block 1050, based on the processing of the set of selected image slices, the position of the object within the environment is determined. The determination of the position of the object within the environment can be performed using various image slicing processing capabilities. The determination of the position of the object within the environment can include: determining the actual position of the object within the environment, determining an estimate of the position of the object within the environment, etc.

[0079] At block 1099, method 1000 ends.

[0080] In connection with Figure 2 step 3 discussed above (represented as block 230 in Figure 2 ), after the camera for which image slicing has been activated has been selected, the number of image slices for performing image slicing is determined. It should be noted that if image slicing has been activated, this is likely triggered by an object approaching the camera from a distance or a small object being detected in the image plane during periodic monitoring (although it should be understood that image slicing can be activated under other conditions). Also note that there may be other objects in the scene for which image slicing is not required, and the camera images are being processed in the normal manner.

[0081] By first considering multiple definitions of the various variables that can be applied within the context of the process for determining the number of image slices for performing image slicing, the determination of the number of image slices can be better understood. The ML model training image size can be represented as (ML H , MLW ) where ML H represents the height for the ML model training image size, and ML W is represented as the width for the ML model training image size. The overlap factor between image slices can be represented as O f . The size of the sliced image can be represented as (H, W), where H represents the height of the sliced image and W represents the width of the sliced image. The maximum number of image slices can be represented as N max . The minimum bounding box size that allows the recognition / detection of objects in the object classes for which the ML model is trained can be represented as ML pixels . The number of image slices is represented as N. The number used to implement the horizontal and vertical splitting of the image slices can be represented as (S h , S v ), where S h is the number used to implement the horizontal splitting of the image slices and S v is the number used to implement the vertical splitting of the image slices.

[0082] The number of image slices N can be given by N = (S h + 1) * (S v + 1). For example, slicing the image via two horizontal splits and two vertical splits results in nine image slices. For example, slicing the image via one horizontal split and one vertical split results in four image slices. For example, slicing the image via three horizontal splits and three vertical splits results in sixteen image slices. Figure 11 The case of using vertical splitting to slice the image to form a set of image slices is depicted in Figure 11 . As shown in

[0083] , slicing the image via two horizontal splits and two vertical splits results in nine image slices, which are labeled "Slice 1" to "Slice 9". It should be understood that applying image slicing to an image may result in slicing the image into fewer or more slices. max The maximum number of image slices N max can be determined as: N f = [H / ((1 - O H ) * ML f )] * [W / ((1 - O W ) * ML f = 0.2, and the sliced image is (1080, 1920), then the maximum number of image slices N max is 9.

[0084] It should be noted that after the image slicing is activated for the camera, there are two independent criteria that can be used to change the number of image slices: (1) the distance of the object (based on the last detection) to the camera position, or (2) the number of pixels in the detected bounding box or contour.

[0085] As described above, after the image slicing is activated for the camera, one criterion that can be used to change the number of image slices is: the distance of the object (based on the last detection) to the camera position. In short, at a specific distance threshold, as the distance increases, the number of image slices also increases. Figure 12 The example embodiments are illustrated therein, which depict the example embodiments for changing the number of image slices used in the image slicing according to the distance between the object and the camera position. As Figure 12 depicted therein, as the distance from the camera 1201 increases, the number of image slices of the original image 1202 also increases. Within the first distance range 1211, no image slicing is performed (i.e., the original image 1202 is scaled down without image slicing). Within the second distance range 1212, image slicing is performed, and the original image is sliced into four image slices (i.e., scaled down and image sliced). Within the third distance range 1213, image slicing is performed, and the original image is sliced into nine image slices (i.e., full resolution and image sliced). It should be understood that fewer or more distance ranges and associated distance thresholds can be defined for changing the number of image slices according to the distance.

[0086] As described above, after the image slicing is activated for the camera, one criterion that can be used to change the number of image slices is: the number of pixels in the detected bounding box or contour. This can be determined when the detection occurs and the bounding box information is available. In order to be able to use this criterion, it is assumed that information about the minimum bounding box size that allows the recognition / detection of the object is available, and the object is in the object category for which the ML model is trained. This is called ML pixels 。 Figure 13 The example embodiments are depicted therein. As Figure 13As depicted, camera 1301 captures the original image 1302, and the number of image slices of the original image 1302 varies based on the number of pixels in the detected bounding box or contour. In the first range, no image slicing is performed (i.e., the original image 1302 is scaled down without image slicing). In the second range, image slicing is performed, and the original image is sliced into four image slices (i.e., scaled down and image sliced). In the third range, image slicing is performed, and the original image is sliced into nine image slices (i.e., full resolution and image sliced). It should be understood that fewer or more ranges and associated pixel number thresholds can be defined for varying the number of image slices according to the number of pixels in the detected bounding box or contour.

[0087] It should be noted that, given that there may be multiple objects in the scene and only one object can trigger image slicing, there are at least three cases to consider: (1) Image slicing has been activated and no other detections have been observed in the past T seconds based on the image without image slicing, (2) Slicing has been activated and detections have been observed in the past T seconds based on the image without image slicing, (3) Slicing has not been activated and detections have been observed in the past T seconds based on the image without image slicing. Additionally, if the object is moving away from the camera, switch to more and higher resolution image slices, and alternatively, if the object is moving towards the camera, switch to fewer and lower resolution image slices. It should be understood that applying these scenarios within the context of the overall algorithm for selecting the number N of image slices can be further understood by referring to Figure 14 for further understanding.

[0088] Figure 14 An example embodiment of a method for determining the number of image slices to be used when slicing an image according to distance or the number of pixels is depicted. It should be understood that, although omitted for clarity, Figure 14 method 1400 may include: one or more initial boxes to perform processing to obtain information used in the context of Figure 14 method 1400. It should be understood that, although mainly presented as being executed sequentially, at least part of the functions of method 1400 can be executed simultaneously or in an order different from the order shown with respect to Figure 14 shown.

[0089] At block 1401, method 1400 begins.

[0090] At block 1405, it is determined whether the image slice is activated. If the image slice is not activated, method 1400 proceeds to block 1410 (where the number of slices is set to be equal to one (i.e., N = 1)), and then method 1400 proceeds to block 1499, where method 1400 ends. If the image slice is activated, method 1400 proceeds to block 1415.

[0091] At block 1415, it is determined whether there is a detection from the image slice. If there is no detection from the image slice, method 1400 proceeds to block 1499, where method 1400 ends. If there is a detection from the image slice, method 1400 proceeds to block 1420.

[0092] At block 1420, it is determined whether information on the minimum bounding box size that allows the recognition / detection of an object is available, where the object is in the object class for which the ML model (ML pixels ) is trained. If ML pixels is available, method 1400 proceeds to the first branch of method 1400, where the number of pixels is used to determine the number of image slices (schematically, method 1400 proceeds to block 1425, and this branch of method 1400 includes blocks 1425 to 1450). If ML pixels is not available, method 1400 proceeds to the second branch of method 1400, where the distance to the camera is used to determine the number of image slices (schematically, method 1400 proceeds to block 1455, and this branch of method 1400 includes blocks 1455 to 1480).

[0093] At block 1425, it is determined whether the number of pixels (N pixels ) is greater than the minimum bounding box size (ML pixels ) adjusted by a factor (γ0). That is, it is determined whether N pixels > (γ0)ML pixels . If N pixels is not greater than (γ0)ML pixels , method 1400 proceeds to block 1430 (where the number of image slices N is set to be equal to N max (i.e., N = N max = (S h + 1)*(S v + 1)), and then method 1400 proceeds to block 1499, where method 1400 ends. If N pixels is greater than (γ0)ML pixels , method 1400 proceeds to block 1435.

[0094] At block 1435, it is determined whether the number of pixels (Npixels ) Is it greater than the minimum bounding box size (ML) adjusted by the factor (γ1) pixels ). That is, determine whether N pixels >(γ1)ML pixels . If N pixels is not greater than (γ1)ML pixels , then method 1400 proceeds to block 1440 (where the number N of image slices is set to be equal to S h and S v 's product (i.e., N = S h *S v ), and then method 1400 proceeds to block 1499, where method 1400 ends. If N pixels is greater than (γ1)ML pixels , then method 1400 proceeds to block 1445.

[0095] At block 1445, determine whether the number of pixels (N pixels ) is greater than the minimum bounding box size (ML k ) adjusted by the factor (γ pixels ). That is, determine whether N pixels >(γ k )ML pixels . If N pixels is not greater than (γ k )ML pixels , then method 1400 proceeds to block 1440 (where the number N of image slices is set based on a minimization function (i.e., N = min(1, [(S h –k + 1)*(S v –k + 1)])), and then method 1400 proceeds to block 1499, where method 1400 ends. If N pixels is greater than (γ k )ML pixels , then method 1400 proceeds to block 1499, where method 1400 ends.

[0096] At block 1455, determine whether the distance (D) is greater than the maximum distance (D max ) adjusted by the factor (β0). That is, determine whether D > (β0)D max . If D is greater than (β0)D max , then method 1400 proceeds to block 1460 (where the number N of image slices is set to be equal to N max (i.e., N = N max =(S h +1)*(S v+1)), and then method 1400 proceeds to block 1499, where method 1400 ends. If D is not greater than (β0)D max , then method 1400 proceeds to block 1465.

[0097] At block 1465, it is determined whether the distance (D) is greater than the maximum distance (D max ) adjusted by the factor (β1). That is, it is determined whether D > (β1)D max . If D is greater than (β1)D max , then method 1400 proceeds to block 1470 (where the number N of image slices is set equal to the product of S h and S v (i.e., N = S h *S v ), and then method 1400 proceeds to block 1499, where method 1400 ends. If D is not greater than (β1)D max , then method 1400 proceeds to block 1475.

[0098] At block 1475, it is determined whether the distance (D) is greater than the maximum distance (D k ) adjusted by the factor (β max ). That is, it is determined whether D > (β k )D max . If D is greater than (β k )D max , then method 1400 proceeds to block 1480 (where the number N of image slices is set based on a minimization function (i.e., N = min(1, [(S h –k + 1)*(S v –k + 1)])), and then method 1400 proceeds to block 1499, where method 1400 ends. If D is not greater than (β1)D max , then method 1400 proceeds to block 1499, where method 1400 ends.

[0099] At block 1499, method 1400 ends.

[0100] Figure 15 Illustrates an exemplary embodiment of a method for determining the number of image slices to be used when slicing an image in object detection. It should be understood that although presented primarily as being executed sequentially, at least portions of the functions of method 1500 may be executed simultaneously, or in an order different from the order shown with respect to Figure 15 .

[0101] At block 1501, method 1500 begins.

[0102] At block 1510, an image captured by a camera is received. The received image depicts an object.

[0103] At block 1520, based on at least one of the following, determine the number of image slices into which the image is to be sliced: the distance between the camera and the object, or the number of pixels in the bounding box of the object detection.

[0104] The number of image slices into which the image is sliced can be changed in different ways. In at least some example embodiments, based on a set of distance thresholds, the number of image slices into which the image is sliced increases as the distance between the camera and the object increases. In at least some example embodiments, the number of image slices into which the image is sliced increases as the number of pixels in the bounding box of the object detection increases. The number of image slices into which the image is sliced can be changed in other ways.

[0105] The number of image slices into which the image is sliced can be determined in a variety of ways.

[0106] In at least some example embodiments, determining the number of image slices into which the image is sliced can include: determining an ML model trained for the object class based on the object class of the object; determining whether information indicating a minimum bounding box size is available, the information allowing detection of an object in the object class for which the ML model is trained; and based on whether the information indicating the minimum bounding box size is available, determining whether to determine the number of image slices into which the image is sliced based on the distance between the camera and the object, or based on the number of pixels in the bounding box of the object detection.

[0107] In at least some example embodiments, determining the number of image slices into which the image is sliced can include: determining a minimum bounding box size that allows object detection, the object being in the object class for which the ML model is trained; and based on comparing the number of pixels in the bounding box of the object detection with a set of thresholds determined based on the minimum bounding box size, determining the number of slices into which the image is sliced. In at least some example embodiments, based on a determination that the number of pixels in the bounding box of the object detection is less than a first threshold, the number of image slices into which the image is sliced is set to be equal to a first value, where the first value is the maximum number of image slices, based on a determination that the number of pixels in the bounding box of the object detection is greater than the first threshold and less than a second threshold, the number of image slices into which the image is sliced is set to be equal to a second value (the second value being less than the first value), or based on a determination that the number of pixels in the bounding box of the object detection is greater than the second threshold and less than a third threshold, the number of image slices into which the image is sliced is set to be equal to a third value (the third value being less than the second value).

[0108] In at least some example embodiments, determining the number of image slices into which an image is sliced may include: determining a maximum distance between a camera and an object; and determining the number of image slices into which the image is sliced based on comparing the distance between the camera and the object with a set of thresholds, the set of thresholds being determined based on the maximum distance between the camera and the object. In at least some example embodiments, based on a determination that the distance between the camera and the object is less than a first threshold, the number of image slices into which the image is sliced is set to be equal to a first value, where the first value is the maximum number of image slices, based on a determination that the distance between the camera and the object is greater than the first threshold and less than a second threshold, the number of image slices into which the image is sliced is set to be equal to a second value (the second value being less than the first value), or based on a determination that the distance between the camera and the object is greater than the second threshold and less than a third threshold, the number of image slices into which the image is sliced is set to be equal to a third value (the third value being less than the second value).

[0109] At block 1530, based on the number of image slices into which the image is sliced, the image is sliced into a set of image slices.

[0110] At block 1540, based on the set of image slices, object detection is performed.

[0111] At block 1599, method 1500 ends.

[0112] In connection with Figure 2 step 4 discussed above (represented as block 240 in Figure 2 ), after the number of image slices into which the image is sliced for image slicing is determined, which image slice to process is determined. The selection of the image slice to be processed may be based on computational constraints associated with performing the slicing process, which may vary in different contexts. For example, in a factory environment, it is known that a robot follows a specific path, so the robot is likely to be in a specific region of the image and detected in a specific region of the image, and this can be used to select a subset of the image slices of the image for processing (e.g., the number of image slices that need to be processed can be reduced to save computational resources). For example, in an autonomous driving environment, the path taken by the vehicle may not be easily predictable and may be more random, so the selection of the image slices of the image for processing can be performed based on other factors to attempt to reduce the number of image slices of the image selected for processing (e.g., the number of image slices that need to be processed can be controlled to save computational resources). It should be understood that a variety of algorithms can be used to determine which image slices to process (including slice selection algorithms based on regions of interest (ROIs) and slice selection algorithms based on cropping, each of which will be discussed in detail below).

[0113] In at least some example embodiments, the selection of image slices of an image to be processed can be performed using an ROI-based slice selection algorithm. The ROI-based slice selection algorithm can be performed as follows. First, subscribe to a service with detection results (e.g., bounding boxes and image frames) associated therewith. Second, within a certain time limit, collect the necessary data (including collecting the bounding boxes for all detected objects and combining them into a union box), and save an image for each camera. Third, for each image from each camera, slice the image into N slices and obtain the positions of the image slices (i.e., the (X, Y) coordinates of the image slices). Fourth, calculate the intersection over union (IoU) between each image slice and the union of the bounding boxes. Finally, if the IoU is greater than a certain threshold, use these image slices for processing. It should be noted that the ROI-based slice selection algorithm can be further understood by referring to Figure 16 and Figure 17 for further understanding.

[0114] Figure 16 depicts the selection of a subset of image slices of an image using an ROI-based slice selection algorithm. As Figure 16 illustrated, the robot moves along a path within the field of view of the camera, thereby capturing an image as depicted in Figure 16 . In a three-by-three grid, the image is sliced into nine image slices, labeled using slice identifiers, referred to as "slice 1" to "slice 9". In this example, the union of the bounding boxes is depicted. In this example, assuming no occlusion, slice 4, slice 5, and slice 6 would be selected for processing. At this time, only three out of the nine image slices of the image need to be processed, thereby reducing the computational resources consumed for detecting and tracking the robot compared to other solutions that would need to process the entire image for detecting and tracking the robot.

[0115] Figure 17 depicts an example embodiment of an ROI-based slice selection algorithm configured to select a subset of image slices of an image for processing. It should be understood that, although omitted for clarity purposes, the method 1700 of Figure 17 may include: one or more initial boxes to perform processing to obtain information used in the context of the method 1700 of Figure 17 . It should be understood that, although mainly presented as being executed sequentially, at least some of the functions of the method 1700 can be executed simultaneously, or in a manner related to Figure 17Execute in a different order than shown. At block 1701, method 1700 begins. At block 1710, the IoU limit is initialized (denoted as iou_limit). At block 1720, all bounding boxes from the subscribed socket are collected. At block 1730, the union of the bounding boxes (denoted as bbox_union_config) is determined. At block 1740, iterate over the bounding boxes to determine if a bounding box exists in bbox_union_config. If the bounding box exists in bbox_union_config, method 1700 proceeds to block 1750, otherwise method 1700 proceeds to block 1799 where method 1700 ends. At block 1750, the image is sliced into N image slices. At block 1760, determine whether to iterate over the image slices. If it is determined not to iterate over the image slices, method 1700 returns to block 1740, otherwise method 1700 proceeds to block 1760. At block 1770, the IoU between the bounding box and the image slice is calculated. At block 1780, determine whether the IoU between the bounding box and the image slice is greater than the IoU limit (iou_limit). If the IoU between the bounding box and the image slice is not greater than the IoU limit, method 1700 returns to block 1760 to continue iterating over the image slices, otherwise method 1700 proceeds to block 1790. At block 1790, based on the determination that the IoU between the bounding box and the image slice is greater than the IoU limit, the selected image slices are saved to the configuration, and method 1700 returns to block 1760 to continue iterating over the slices. At block 1799, method 1700 ends.

[0116] In at least some example embodiments, the selection of image slices of an image to be processed can be performed using a cropping-based slice selection algorithm. The cropping-based slice selection algorithm can be performed as follows. First, subscribe to a service with detection results (e.g., bounding boxes and image frames) associated therewith. Second, within a certain time limit: (a) collect the bounding boxes for all detected objects and combine them into a combined bounding box, (2) if the extrapolation parameter is activated (since it helps to expand the region of interest when the system is running without performing image slicing during data collection), the bounding boxes will collect the positions of the objects together and extrapolate the positions of the objects as splines, convert the received end points to pixel coordinates, and combine the pixels with the union of the bounding boxes, and (3) save an image for each camera. Third, expand the collected ROI to the input height and width of the model. Finally, if the height or width is greater than the height or width of the model, slice the ROI into several image slices (depending on the ROI size and the input size of the model). It should be noted that the ROI-based slice selection algorithm can be performed by referring toFigure 18 and Figure 19 for further understanding.

[0117] Figure 18 depicts the selection of a subset of image slices of an image based on a cropping-based slice selection algorithm. As Figure 18 illustrated, the robot moves along a path within the field of view of the camera, thereby capturing an image as Figure 18 depicted. In this example, a bounding box union is depicted. In this example, the bounding box union is expanded to provide a ROI selected for processing. In this example, since the height or width of the expanded ROI is greater than the height or width of the model, the expanded ROI is sliced into multiple slices (denoted as "Slice 1" and "Slice 2"). At this time, only a subset of the complete image needs to be processed, thereby reducing the computational resources consumed for detecting and tracking the robot compared to other solutions, where other solutions need to process the entire image for detecting and tracking the robot.

[0118] Figure 19 depicts an example embodiment of a cropping-based slice selection algorithm configured to: select a subset of image slices of an image for processing. It should be understood that, although omitted for clarity, Figure 19 Method 1900 may include: one or more initial boxes to perform processing to obtain information used in the context of Figure 19 Method 1900. It should be understood that, although presented primarily as being executed sequentially, at least portions of the functions of Method 1900 may be executed simultaneously, or in an order different from that regarding Figure 19Performed in a different order than shown. At block 1901, method 1900 begins. At block 1905, parameters indicating the model size (in terms of height and width), denoted as model_width and model_height, are initialized. At block 1910, all bounding boxes from the subscribed socket are collected. At block 1915, the union of the bounding boxes, denoted as bbox_union_config, is determined. At block 1920, the bounding boxes are iterated over to determine whether a bounding box exists in bbox_union_config. If the bounding box exists in bbox_union_config, method 1900 proceeds to block 1925; otherwise, method 1900 proceeds to block 1960. At block 1925, it is determined whether the width of the bounding box is less than the defined width of the model (model_width). If it is determined that the width of the bounding box is less than the defined width of the model (model_width), method 1900 proceeds to block 1930; otherwise, method 1900 skips block 1930 and proceeds to block 1935. At block 1930, based on the determination that the width of the bounding box is less than the defined width of the model, the width of the bounding box is extended to the defined width of the model, and then method 1900 proceeds to block 1935. At block 1935, it is determined whether the height of the bounding box is less than the defined height of the model (model_height). If it is determined that the height of the bounding box is less than the defined height of the model (model_height), method 1900 proceeds to block 1940; otherwise, method 1900 skips block 1940 and proceeds to block 1945. At block 1940, based on the determination that the height of the bounding box is less than the defined height of the model, the height of the bounding box is extended to the defined height of the model, and then method 1900 proceeds to block 1945. At block 1945, it is determined whether the height of the bounding box is greater than the defined height of the model (model_height) or the width of the bounding box is greater than the defined width of the model (model_width). If the height of the bounding box is not greater than the defined height of the model and the width of the bounding box is not greater than the defined width of the model, method 1900 returns to block 1920. If the height of the bounding box is greater than the defined height of the model or the width of the bounding box is greater than the defined width of the model, method 1900 proceeds to block 1950. At block 1950, the bounding box region is cropped from the image to obtain the cropped image. At block 1955, slices are applied to the cropped image, and then method 1900 returns to block 1920. At block 1960, based on the determination that the bounding box does not exist in bbox_union_config, the slice configuration is saved, and method 1900 proceeds to block 1999, where method 1900 ends.At block 1999, method 1900 ends.

[0119] Figure 20 An example embodiment of an object tracking system is depicted. The object tracking system is configured to perform object tracking for a set of objects in an environment. As Figure 20 illustrated, the object tracking system 2000 includes: a camera 2001, a database 2002, a flexible slicing algorithm 2010, an image slicer 2020, a learning model-based detector 2030, a slice processor 2040, a tracker 2050, a locator 2060, a motion model processor 2070, and a spatial analysis processor 2080. The camera 2001 captures an image, which may include one or more objects to be detected and tracked. The database 2002 stores: a homography matrix, camera position information, and calculated FoV information. The flexible slicing algorithm 2010 determines a slicing configuration for use by the image slicer 2020 to slice the image from the camera 2001, based on the homography matrix from the database 2002 and feedback from the motion model processor 2070. The flexible slicing algorithm 2010 may determine the slicing configuration based on steps 1 to 4 presented herein (e.g., Figure 2 blocks 210 to 240). The image slicer 2020 receives the image from the camera 2001 and the slicing configuration from the flexible slicing algorithm 2020, and performs image slicing to form image slices, which are provided to the learning model-based detector 2030. The image slicer 2020 may perform image slicing based on step 5 presented herein (e.g., Figure 2 block 250). The learning model-based detector 2030 receives the image slices from the image slicer 2020 and performs model-based object detection on the image slices to detect objects. The learning model-based detector 2030 may perform object detection based on step 6 presented herein (e.g., Figure 2 block 260). The learning model-based detector 2030 provides object detection data related to the detection of objects in the image slices to the slice processor 2040. The slice processor 2040 receives the object detection data for the image slices, processes the image detection data to combine the detections of objects in the image slices and obtain object detection / location information for the objects, and provides the object detection / location information for the objects to the tracker 2050. The slice processor 2040 may process the image detection data to based on step 7 presented herein (e.g., Figure 2The frame 270), obtaining object detection / localization information for the object. The tracker 2050 receives the object detection / localization information for the object from the slicing processor 2040, performs object tracking based on the object detection / localization information to generate object tracking information, and provides the object tracking information to the locator 2060. The locator 2060 receives the object tracking information for the object from the tracker 2050, performs object localization based on the object tracking information to generate object localization information, and provides the object localization information to the motion model processor 2070. The motion model processor 2070 receives the object localization information from the locator 2060 and learns motion parameters (such as speed, acceleration, etc.) based on the motion of each object in the environment to construct a corresponding model, which can then be used to fill in missed positions (or detections), predict future positions, etc., and various combinations thereof. The spatial analysis processor 2080 analyzes the data received from multiple cameras in the environment or space and combines them to generate a single position estimate for the object.

[0120] It should be understood that the various elements of the object tracking system can be implemented in various ways. In at least some example embodiments, for example, the various functions of the object tracking system can be implemented using various combinations of local computing resources, edge computing resources, cloud computing resources, etc., and various combinations thereof. In at least some example embodiments, for example, in order to maintain relatively low latency and network bandwidth consumption, the image preprocessing / perception steps (such as slicing, downscaling, detection) can be kept close to the source (such as local computing and / or edge computing), so that the image does not need to be sent to the cloud computing. This is also useful in protecting the privacy of sensitive data. In this case, the slice configuration information sent from the cloud computing to the edge computing can be checked to check different slice configurations. It should be understood that the various elements of the object tracking system can be implemented in various other ways.

[0121] It should be understood that although the detection and tracking of a single object within the scene are mainly presented, object detection and tracking can be performed for multiple objects in the scene. In at least some such example embodiments, the object detection and tracking algorithms can be configured to decide when and how to activate the image slices. In at least some example embodiments, for example, the image slices can be processed together with the full-resolution image, and the detections in all the images are combined, thus allowing the detector to identify the objects close to the camera and those moving from a farther distance (although it should be noted that in this case, there is always an additional image to be processed). In at least some example embodiments, for example, the image slices can be activated for the camera based on the camera selection pipeline, and if at least one robot is moving from a farther distance, the image slice algorithm can be applied to identify the object.

[0122] Various example embodiments for supporting object detection in a machine vision system can provide various advantages or potential advantages. For example, various example embodiments for supporting object detection in a machine vision system can be configured to: support object detection based on an intelligent selection of a set of cameras for which image slices will be used; capture images from the set of cameras for which image slices will be used; and process images from the set of cameras for which image slices will be used for object detection (e.g., determine the number of image slices to be used, perform image slicing, select which image slices to process for object detection, process the selected image slices for object detection, combine detections across image slices, etc., and various combinations thereof). For example, various example embodiments for supporting object detection in a machine vision system can be configured to: use image slices to improve object detection while alleviating the typically increased linear computations that result from using image slices for object detection. For example, various example embodiments for supporting object detection in a machine vision system can be configured to: support object detection based on an intelligent calculation of the FOV of a camera, where the FOV of the camera is for a camera involved in object detection based on the use of a homography matrix, rather than simply relying on the camera parameter of the camera FOV for calculation, and the FOV of the camera is for a camera involved in object detection. For example, various example embodiments for supporting object detection in a machine vision system can be configured to: support object detection based on an intelligent selection of a set of cameras for which image slices will be used, where the set of cameras is based on the combined use of periodic monitoring and a distance metric related to the position of the object relative to the camera position and field of view, without a map of the space in which the cameras are deployed and the object is located (although it should be understood that such a map can also be used if available). For example, various example embodiments for supporting object detection in a machine vision system can be configured to: support object detection based on an intelligent selection of the number of image slices to be used for the image slices, where the number of image slices to be used for the image slices is a function of the distance (between the object and the camera) and the number of pixels in the bounding box of the contour (if segmentation is used), while considering the computational impact from using image slices for object detection, and without the need to manipulate a processing module (e.g., no need to retrain an ML detection model, no need to use image patches to train a convolutional network, no need to build and train a custom neural network, etc.). For example, various example embodiments for supporting object detection in a machine vision system can be configured to: support object detection based on the effective processing of image slices using a model that learns the active regions in the image plane for each camera and matches the slices to be processed to the active regions (e.g., based on a preconfigured number of image slices for the entire image, or sliced in the active regions of the image).For example, various example embodiments for supporting object detection in a machine vision system can be configured to support object detection without increasing the image resolution input to a deep neural network (DNN), which may otherwise result in a significant increase in the complexity of the DNN, thus affecting the computational cost during the training phase and inference. For example, various example embodiments for supporting object detection in a machine vision system can be configured to support object detection without using a smaller anchor size in an ML model, which may otherwise result in the need to collect samples with small objects for training and retraining and may not work well when the objects are not small. It should be understood that various example embodiments for supporting object detection in a machine vision system can provide various other advantages or potential advantages.

[0123] Figure 21 Depicts an example embodiment of a computer suitable for performing the various functions presented herein.

[0124] The computer 2100 includes a processor 2102 (e.g., a central processing unit (CPU), a processor, a processor having a set of processor cores, a processor core of a processor, etc.) and a memory 2104 (e.g., a random access memory (RAM), a read only memory (ROM), etc.). In at least some example embodiments, the computer 2100 may include at least one processor and at least one memory storing instructions that, when executed by the at least one processor, cause the computer 2100 to perform the various functions presented herein.

[0125] The computer 2100 may further include a cooperation element 2105. The cooperation element 2105 may be a hardware device. The cooperation element 2105 may be a process that can be loaded into the memory 2104 and executed by the processor 2102 to implement the various functions presented herein (in which case, for example, the cooperation element 2105 (including associated data structures) may be stored on a non-transitory computer-readable medium, such as a storage device, or other suitable type of storage element (e.g., a magnetic drive, an optical drive, etc.)).

[0126] The computer 2100 may further include one or more input / output devices 2106. The input / output devices 2106 may include one or more of the following: user input devices (e.g., a keyboard, a keypad, a mouse, a microphone, a camera, etc.), user output devices (e.g., a display, a speaker, etc.), one or more network communication devices or elements (e.g., an input port, an output port, a receiver, a transmitter, a transceiver, etc.), one or more storage devices (e.g., a tape drive, a floppy drive, a hard drive, an optical drive, etc.), etc., and various combinations thereof.

[0127] It should be understood that the computer 2100 can represent a general architecture and functionality suitable for implementing the following: the functional elements described herein, parts of the functional elements described herein, etc., and various combinations thereof. For example, the computer 2100 can provide a general architecture and functionality suitable for implementing one or more elements presented herein. For example, the computer 2100 can provide a general architecture and functionality suitable for implementing at least one of the following: a camera or a part thereof, an object detection processing element or a part thereof, etc., and various combinations thereof.

[0128] It should be understood that at least some of the functions presented herein can be implemented in software (e.g., via software on one or more processors, for execution on a general-purpose computer (e.g., by execution by one or more processors) to provide a special-purpose computer, etc.) and / or can be implemented in hardware (e.g., using a general-purpose computer, one or more application-specific integrated circuits, and / or any other hardware equivalent).

[0129] It should be understood that at least some of the functions presented herein can be implemented within hardware, e.g., as circuitry that cooperates with a processor to perform various functions. Portions of the functions / elements described herein can be implemented as a computer program product, where computer instructions, when processed by a computer, adjust the operation of the computer such that the methods and / or techniques described herein are invoked or otherwise provided. Instructions for invoking various methods can be stored on a fixed or removable medium (e.g., a non-transitory computer-readable medium), transmitted via a data stream in a broadcast or other signal-bearing medium, and / or stored in a memory within a computing device that operates according to the instructions.

[0130] It should be understood that the term "non-transitory" as used herein is a limitation of the medium itself (i.e., tangible, rather than a signal), rather than a limitation on data storage persistence (e.g., RAM vs. ROM).

[0131] It should be understood that as used herein, "at least one of " and "at least one of the following: " and similar phrases, where the list of two or more elements is joined by "and" or "or", means at least any one element, or at least any two or more elements, or at least all elements.

[0132] It should be understood that as used herein, the term "or" refers to an inclusive "or" unless otherwise specified (e.g., using "otherwise" or "alternatively").

[0133] It should be understood that, although various embodiments incorporating the teachings presented herein have been shown and described in detail herein, those skilled in the art can readily devise many other different embodiments that still incorporate these teachings.

Claims

1. An apparatus for detecting an object, comprising components configured to perform the following: Detect the object within an environment monitored by a set of cameras; Based on the determination that the object is no longer detected, determine, from the set of cameras, the set of cameras for which image slices will be activated, where the set of cameras for which image slices will be activated includes: A first camera on which the object was last detected before the object is no longer detected; And a second camera that is determined to be the position closest to the object when the object was last detected; Obtain a set of images, wherein the set of images includes: corresponding images for each camera from among the set of cameras for which image slices will be activated, for the respective cameras; Based on the application of image slices to each image in the set of images, obtain a set of selected image slices from the set of images; and Based on the processing of the set of selected image slices, determine the position of the object within the environment.

2. The apparatus according to claim 1, wherein, in order to detect the object, the components are configured to perform: Determine the distance between the position of the camera on which the object was detected and the position of the object, for the cameras from the set of cameras on which the object was detected; and Detect the object based on the determination that the distance between the position of the camera and the position of the object satisfies a threshold.

3. The apparatus according to any one of claims 1 to 2, wherein, in order to detect the object, the components are configured to: Based on the determination that the layout of the set of cameras is unknown, periodically activate image slices for each camera in the set of cameras; and Detect the object based on the image slices for each camera in the set of cameras.

4. The apparatus according to any one of claims 1 to 3, wherein the set of cameras for which image slices will be activated is determined based on the length of time since the object was last detected satisfying a threshold.

5. The apparatus according to any one of claims 1 to 4, wherein the second camera from the set of cameras is determined to be the position closest to the object when the object was last detected, based on the determination that the second camera has the field of view closest to the position of the object when the object was last detected.

6. The apparatus according to any one of claims 1 to 5, wherein, in order to select the second camera, the components are configured to perform: For each camera in the set of cameras, calculate the respective field of view of the camera; Determine a set of distance values, the set of distance values including: For each camera in the set of cameras, the respective distance between the position of the object when the object was last detected and the respective field of view of the camera; And Based on the set of distance values, select the second camera from the set of cameras, the second camera being determined to be the position closest to the object when the object was last detected.

7. The apparatus according to claim 6, wherein, in order to calculate the respective field of view of the camera, the component is configured to perform: Receiving a homography matrix calculated using a reference image resolution that matches the image resolution of the respective image from the respective camera; Obtaining a set of points by calculating, for each pixel in a set of pixels of the respective image, a respective point specifying a physical location of a region captured by the respective pixel, based on the homography matrix; Filtering any points in the set of points that are beyond a predefined distance threshold in the set of points to form a filtered set of points; And Calculating the respective field of view of the camera based on a convex structure and using the filtered set of points.

8. The apparatus according to any one of claims 1 to 7, wherein, in order to obtain the set of selected image slices, the component is configured to perform: Obtaining, for each image in the set of images, a respective set of image slices for the respective image based on an application of image slices to the respective image in the set of images; and Obtaining the set of selected image slices by selecting at least a portion of the image slices in the respective set of image slices from each image slice in the set of image slices formed by applying image slices to the respective image in the set of images.

9. The apparatus according to claim 8, wherein, in order to obtain the respective set of image slices for the respective image based on an application of image slices to the respective image in the set of images, the component is configured to perform: Determining the number of slices into which the respective image in the set of images is to be sliced; and Performing slicing of the respective image in the set of images based on the number of slices into which the respective image in the set of images is to be sliced to form the respective set of image slices for the respective image in the set of images.

10. The apparatus according to claim 9, wherein the number of slices into which the respective image in the set of images is to be sliced is determined based on at least one of: the distance between the camera capturing the respective image and the object, or the number of pixels in the bounding box of the detection of the object.

11. The apparatus according to claim 8, wherein at least a portion of the image slices in the respective set of image slices for the respective image in the set of images is selected based on a slice selection algorithm for a region of interest (ROI).

12. The apparatus according to claim 8, wherein, in order to select at least a portion of the image slices in the respective set of image slices to form a respective set of selected image slices, the component is configured to perform: Collecting a set of bounding boxes for the object; Combining the set of bounding boxes for the object into a combined bounding box for the object; Slice the image from the camera into N image slices, and for each of the N image slices of the corresponding image, obtain the corresponding coordinates of the corresponding image slice; and For each of the image slices, calculate the intersection over union (IoU) between the corresponding image slice and the union bounding box, and based on the determination that the corresponding IoU for the corresponding image slice meets a threshold, then select the corresponding image slice for inclusion in the set of the corresponding selected image slices.

13. The apparatus according to claim 8, wherein based on a cropped slice selection algorithm, at least a portion of the image slices in the set of image slices for the corresponding image in the set of images is selected.

14. The apparatus according to claim 8, wherein in order to select at least a portion of the image slices in the set of image slices for the corresponding image to form the set of the corresponding selected image slices, the component is configured to perform: Collect a set of bounding boxes for the object; Combine the set of bounding boxes for the object into a union bounding box for the object; Based on the defined image size of a machine learning (ML) model, expand the union bounding box to form a region of interest; Based on the determination that the size of the region of interest is greater than the defined image size of the ML model in at least one of a height parameter or a width parameter, crop the bounding box region from the image to form a cropped image; and Slice the cropped image to obtain the set of the corresponding selected image slices.

15. A computer-implemented method, comprising: Detect an object within an environment monitored by a set of cameras; Based on the determination that the object is no longer detected, determine, from the set of cameras, a set of cameras for which image slices will be activated, wherein the set of cameras for which image slices will be activated includes: a first camera on which the object was last detected before the object was no longer detected; and a second camera that is determined to be the position closest to the object when the object was last detected; Obtain a set of images, wherein the set of images includes: corresponding images for the corresponding cameras from each camera in the set of cameras for which image slices will be activated; Based on the application of image slices to each of the images in the set of images, obtain a set of selected image slices from the set of images; and Based on the processing of the set of selected image slices, determine the position of the object within the environment.

Citation Information

Cited By

  • Map-Anchored Object Detection

    US20250131591A1