Object counting using monocular three-dimensional (3D) perception
Through monocular 3D perception technology, a monocular RGB camera is used to perform implicit space of interest calculation, which solves the problem of expensive sensors and high-cost 3D model construction in existing technologies and achieves low-cost and efficient object counting.
Patent Information
- Application Number
- CN202480012979.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-22
- Filing Date
- 2024-01-10
- Publication Date
- 2025-09-12
AI Technical Summary
Existing 3D scene understanding or perception solutions require additional sensors such as stereo cameras, LIDAR, ToF sensors, etc., and the cost of building a 3D model of the environment is high, making accurate object counting difficult.
Using monocular 3D perception technology, a single RGB camera is used to implicitly calculate the space of interest through the relative depth changes between scenes, avoiding expensive camera calibration and 3D reconstruction algorithms and reducing computing resource requirements.
It achieves low-cost accurate object counting, adapts to various scenarios and applications, and reduces processing speed and power consumption.
Smart Images

Figure CN120641949A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to image processing. For example, aspects of the present disclosure relate to providing accurate object counting using monocular three-dimensional (3D) perception. Background Art
[0002] The increasing versatility of digital camera products has allowed them to be integrated into a wide variety of devices and has expanded their use across diverse applications. For example, phones, drones, cars, computers, televisions, and many other devices today are often equipped with camera devices. Camera devices allow users to capture images and / or video from any system equipped with a camera device. Images and / or video can be captured for recreational use, professional photography, surveillance, automation, and other applications. Furthermore, camera devices are increasingly equipped with specialized features for modifying images or creating artistic effects on images. For example, many camera devices are equipped with image processing capabilities for generating various effects on captured images.
[0003] In some applications, images and / or video frames may be processed to obtain object counts. Accurate object counting can be important and has many real-life application scenarios. Various different types of objects can be counted, including but not limited to people, animals, tangible items, and / or electronic devices. Artificial intelligence (AI)-based methods can enhance object counting compared to traditional methods. However, erroneous object counts may still occur when objects appear in mirrors or glass in the image. In addition, sometimes it is only desirable to count objects (e.g., people) that are within a given 3D space, such as counting people waiting for an elevator in a lounge area, rather than counting people in the elevator itself.
[0004] This type of accurate people counting requires 3D scene understanding, particularly for verifying whether objects are located within the space of interest. Currently, existing 3D scene understanding or perception solutions require additional sensors such as stereo cameras, light detection and ranging (LIDAR), and time-of-flight (ToF) sensors. Often, a 3D model of the environment is also required, which can be expensive to construct. Therefore, an improved approach to obtaining accurate object counts would be beneficial. Summary of the Invention
[0005] The following presents a simplified summary of one or more aspects disclosed herein. Therefore, the following summary should neither be considered an exhaustive overview of all contemplated aspects nor be considered to identify key or critical elements related to all contemplated aspects or to delineate the scope associated with any particular aspect. Therefore, the sole purpose of the following summary is to present certain concepts related to one or more aspects of the mechanisms disclosed herein in a simplified form prior to the detailed description presented below.
[0006] Systems and techniques for accurate object counting using monocular 3D perception are described. According to at least one illustrative example, a method for processing one or more frames includes: generating a reference depth map based on a reference frame depicting a space of interest; generating a current depth map based on a current frame depicting the space of interest and one or more objects; comparing the current depth map with the reference depth map to determine a corresponding depth change for each of the one or more objects; comparing the corresponding depth change for each of the one or more objects to a threshold; and determining whether each of the one or more objects is located within the space of interest based on comparing the corresponding depth change for each of the one or more objects to the threshold.
[0007] In another illustrative example, an apparatus for processing one or more frames is provided. The apparatus includes at least one memory and at least one processor, the at least one processor being coupled to the at least one memory and configured to: generate a reference depth map based on a reference frame depicting a space of interest; generate a current depth map based on a current frame depicting the space of interest and one or more objects; compare the current depth map with the reference depth map to determine a corresponding depth change for each of the one or more objects; compare the corresponding depth change for each of the one or more objects with a threshold; and determine whether each of the one or more objects is located within the space of interest based on comparing the corresponding depth change for each of the one or more objects with the threshold.
[0008] In another illustrative example, a non-transitory computer-readable medium is provided, having instructions stored thereon that, when executed by at least one processor, cause the at least one processor to: generate a reference depth map based on a reference frame depicting a space of interest; generate a current depth map based on a current frame depicting the space of interest and one or more objects; compare the current depth map with the reference depth map to determine a corresponding depth change for each of the one or more objects; compare the corresponding depth change for each of the one or more objects with a threshold; and determine whether each of the one or more objects is located within the space of interest based on comparing the corresponding depth change for each of the one or more objects with the threshold.
[0009] In another illustrative example, an apparatus for processing one or more frames is provided. The apparatus includes: means for generating a reference depth map based on a reference frame depicting a space of interest; means for generating a current depth map based on a current frame depicting the space of interest and one or more objects; means for comparing the current depth map with the reference depth map to determine a corresponding depth change for each of the one or more objects; means for comparing the corresponding depth change for each of the one or more objects with a threshold; and means for determining whether each of the one or more objects is located within the space of interest based on comparing the corresponding depth change for each of the one or more objects with the threshold.
[0010] Aspects generally include methods, apparatus, systems, computer program products, non-transitory computer-readable media, user devices, user equipment, wireless communication devices, and / or processing systems as substantially described with reference to and as illustrated in the accompanying drawings and description.
[0011] In some aspects, each of the aforementioned devices is, or may be part of, or include, a mobile device, a smart or connected device, a camera system, and / or an extended reality (XR) device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device). In some examples, the device may include, or be part of, a vehicle, a mobile device (e.g., a mobile phone or so-called "smartphone" or other mobile device), a wearable device, a personal computer, a laptop computer, a tablet computer, a server computer, a robotic device or system, an aviation system, or other device. In some aspects, the device includes an image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, the device includes one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, the device includes one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, the devices described above may include one or more sensors. In some cases, the one or more sensors may be used to determine the device's location, the device's status (e.g., tracking status, operational status, temperature, humidity level, and / or other status), and / or for other purposes.
[0012] Some aspects include a device having a processor configured to perform one or more operations of any of the methods outlined above. Further aspects include a processing device for use in a device, the processing device configured with processor-executable instructions to perform the operations of any of the methods outlined above. Further aspects include a non-transitory processor-readable storage medium having stored thereon processor-executable instructions configured to cause the processor of the device to perform the operations of any of the methods outlined above. Further aspects include a device having components for performing the functions of any of the methods outlined above.
[0013] The features and technical advantages of the examples according to the present disclosure have been outlined quite broadly above so that the detailed description that follows may be better understood. Additional features and advantages will be described below. The concepts and specific examples disclosed may be readily used as a basis for modifying or designing other structures for achieving the same purposes of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. The characteristics of the concepts disclosed herein (both their organization and method of operation) and the associated advantages will be better understood from the following description when considered in conjunction with the accompanying drawings. Each of the figures in the drawings is provided for the purpose of illustration and description and not as a definition of limitations on the claims. The foregoing and other features and aspects will become more apparent upon reference to the following description, claims, and drawings.
[0014] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. This subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.
[0015] The foregoing and other features and embodiments will become more fully apparent upon reference to the following description, claims, and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings are presented to assist in describing various aspects of the present disclosure and are provided for illustration only and not limitation of the various aspects. To enable a detailed understanding of the aforementioned features of the present disclosure, a more detailed description, briefly summarized above, may be obtained by reference to various aspects, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only certain typical aspects of the present disclosure and are not to be considered limiting of its scope, as the description may admit to other equally effective aspects. The same reference numerals in different drawings may identify the same or similar elements.
[0017] Figure 1 is a block diagram illustrating an example image processing system according to some examples of the present disclosure.
[0018] Figure 2 is a diagram illustrating frames of an example scene with multiple objects (eg, people) located outside a space of interest (eg, an elevator waiting room) according to some examples of the present disclosure.
[0019] Figure 3 is a diagram illustrating frames of an example scene with multiple objects (eg, people) located within a space of interest (eg, an elevator waiting room) according to some examples of the present disclosure.
[0020] Figure 4 is a diagram illustrating examples of frames and depth maps according to some examples of the present disclosure.
[0021] Figure 5 is a flow chart illustrating an example computational flow for accurate people counting according to some examples of the present disclosure.
[0022] Figure 6 is a flowchart illustrating an example process for accurate object counting using monocular 3D perception according to some examples of the present disclosure.
[0023] Figure 7 An example computing device architecture according to some examples of the present disclosure is illustrated. DETAILED DESCRIPTION
[0024] Provide certain aspects and embodiments of the present disclosure below. Some of these aspects and embodiments can be applied independently, and some of them can be applied in combination, which will be apparent to those skilled in the art. In the following description, specific details are set forth for explanation purposes to provide a thorough understanding of each embodiment of the application. However, it will be apparent that each embodiment can be put into practice without these specific details. Each drawing and description are not intended to be restrictive.
[0025] The following description provides only example embodiments and is not intended to limit the scope, applicability or configuration of the present disclosure. On the contrary, the following description of the exemplary embodiments will provide an enabling description for implementing the exemplary embodiments to those skilled in the art. It should be understood that various changes may be made to the function and arrangement of elements without departing from the spirit and scope of the present application as set forth in the appended claims.
[0026] As previously mentioned, computing devices are increasingly equipped with the ability to capture images, perform various image processing tasks, generate various image effects, and the like. Images and / or video frames can be processed to obtain object counts. Accurate object counting can be important and has many real-life applications. Various different types of objects can be counted, including but not limited to people, animals, tangible items, and / or electronic devices. Artificial intelligence (AI)-based methods can enhance object counting compared to traditional methods. However, erroneous object counts may still occur (for example, when an object appears in a mirror or glass in the image). In addition, sometimes it is only desirable to count objects (for example, people) located within a given 3D space (for example, such as counting people waiting for an elevator in a lounge area, rather than counting people in the elevator itself).
[0027] This type of accurate people counting may require 3D scene understanding, such as to verify whether objects are located within a space of interest. Existing 3D scene understanding or perception solutions require additional sensors (e.g., stereo cameras, LIDAR, Time of Flight sensors, etc.). These also typically require building a 3D model of the environment, which can be expensive to construct. Therefore, an improved method for obtaining accurate object counts is needed.
[0028] In the following disclosure, this document describes systems, apparatuses, methods (also referred to as processes), and computer-readable media (collectively referred to herein as "systems and techniques") for providing accurate object counting using monocular 3D perception. In some examples, the systems and techniques described herein can provide a low-cost object counting solution based on monocular depth estimation. The methods of the present invention can benefit the Internet of Things (IoT), safety and security monitoring systems, robotic systems, smart home systems, mapping systems, detection systems, positioning systems, entertainment systems, alternative reality (AR), extended reality (XR), virtual reality (VR), and mobile applications. These systems and techniques can be used in a variety of different scenarios that utilize spaces of interest.
[0029] In one or more aspects, these systems and techniques for providing accurate object counting can utilize a single red, green, and blue (RGB) camera, eliminating the need for additional depth sensors. These systems and techniques perform implicit space-of-interest calculations using relative depth changes between scenes (e.g., current frame versus reference frame). These systems and techniques do not require expensive camera calibration and, therefore, are flexible and adaptable to a variety of scenarios and applications. These systems and techniques do not require 3D reconstruction algorithms and / or implementations, resulting in reduced demands on computing resources, which can lead to faster processing and lower power consumption.
[0030] Additional aspects of the disclosure are described in more detail below.
[0031] Figure 1 is a diagram illustrating an example image processing system 100 according to some examples. As described herein, the image processing system 100 can perform object counting techniques. Furthermore, as described herein, the image processing system 100 can perform various image processing tasks and generate various image processing effects. For example, the image processing system 100 can perform object counting, image segmentation, foreground prediction, background replacement, depth of field effects, chroma keying effects, feature extraction, object detection, image recognition, machine vision, and / or any other image processing and computer vision tasks.
[0032] exist Figure 1In the example shown, image processing system 100 includes an image capture device 102, storage 108, a computing component 110, an image processing engine 120, one or more neural networks 122, and a rendering engine 124. Image processing system 100 can also optionally include one or more additional image capture devices 104; one or more sensors 106, such as a light detection and ranging (LIDAR) sensor, a radio detection and ranging (RADAR) sensor, an accelerometer, a gyroscope, a light sensor, an inertial measurement unit (IMU), a proximity sensor, and the like. In some cases, image processing system 100 may include multiple image capture devices capable of capturing images with different FOVs. For example, in a dual-camera or image sensor application, image processing system 100 may include image capture devices with different types of lenses (e.g., wide-angle, telephoto, standard, zoom, etc.) capable of capturing images with different FOVs (e.g., different viewing angles, different depths of field, etc.).
[0033] The image processing system 100 can be part of a computing device or multiple computing devices. In some examples, the image processing system 100 can be part of an electronic device (or multiple electronic devices), such as a camera system (e.g., an RGB camera, a digital camera, an IP camera, a video camera, a security camera, etc.), a phone system (e.g., a smartphone, a cellular phone, a conferencing system, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a display device, a digital media player, a game console, a video streaming device, a drone, an in-vehicle computer, an IoT (Internet of Things) device, a smart wearable device, an extended reality (XR) device (e.g., a head-mounted display, smart glasses, etc.), or any other suitable electronic device.
[0034] In some implementations, the image capture device 102, the image capture device 104, the other sensor 106, the storage 108, the computing component 110, the image processing engine 120, the neural network 122, and the rendering engine 124 can be part of the same computing device. For example, in some cases, the image capture device 102, the image capture device 104, the other sensor 106, the storage 108, the computing component 110, the image processing engine 120, the neural network 122, and the rendering engine 124 can be integrated into a smartphone, a laptop, a tablet, a smart wearable device, a gaming system, an XR device, and / or any other computing device. However, in some implementations, the image capture device 102, the image capture device 104, the other sensor 106, the storage 108, the computing component 110, the image processing engine 120, the neural network 122, and / or the rendering engine 124 can be part of two or more separate computing devices.
[0035] In some examples, image capture devices 102, 104 can be any image and / or video capture device, such as a digital camera, a video camera, a smartphone camera, a camera device on an electronic device (such as a television or computer), a camera system, etc. In some cases, image capture devices 102, 104 can be part of a camera or computing device (such as a digital camera, a video camera, an IP camera, a smartphone, a smart TV, a gaming system, etc. In some examples, image capture devices 102, 104 can be part of a dual-camera assembly. Image capture devices 102, 104 can capture image and / or video content (e.g., raw image and / or video data), which can then be processed by computing component 110, image processing engine 120, neural network 122, and / or rendering engine 124, as described herein.
[0036] In some cases, image capture devices 102 and 104 may include image sensors and / or lenses for capturing image data (e.g., still images, video frames, etc.). Image capture devices 102 and 104 may be capable of capturing image data with different or the same FOVs, including different or the same viewing angles, different or the same depths of field, different or the same size, etc. For example, in some cases, image capture devices 102 and 104 may include different image sensors with different FOVs. In other examples, image capture devices 102 and 104 may include different types of lenses with different FOVs, such as wide-angle lenses, telephoto lenses (e.g., short telephoto lenses, medium telephoto lenses, etc.), standard lenses, zoom lenses, etc. In some examples, image capture device 102 may include one type of lens and image capture device 104 may include different types of lenses. In some cases, image capture devices 102 and 104 may respond to different types of light. For example, in some cases, image capture device 102 may respond to visible light, and image capture device 104 may respond to infrared light.
[0037] Another sensor 106 can be any sensor for detecting and measuring information such as distance, motion, position, depth, speed, etc. Non-limiting examples of sensors include LIDAR, ultrasonic sensors, gyroscopes, accelerometers, magnetometers, RADAR, IMUs, audio sensors, light sensors, etc. In one illustrative example, the sensor 106 can be a LIDAR configured to sense or measure distance and / or depth information that may be used in calculating depth of field and other effects. In some cases, the image processing system 100 may include other sensors such as machine vision sensors, intelligent scene sensors, voice recognition sensors, impact sensors, position sensors, tilt sensors, light sensors, etc.
[0038] Storage 108 may include any storage device for storing data, such as, for example, image data. Storage 108 may store data from any of the components of image processing system 100. For example, storage 108 may store data or measurements (e.g., processing parameters, output, video, image, segmentation map, depth map, filtering results, calculation results, etc.) from any of image capture devices 102, 104, another sensor 106, computing component 110, and / or data or measurements (e.g., output images, processing results, parameters, etc.) from any of image processing engine 120, neural network 122, and / or rendering engine 124. In some examples, storage 108 may include a buffer for storing data (e.g., image data) processed by computing component 110.
[0039] In some implementations, the computing component 110 may include a central processing unit (CPU) 112, a graphics processing unit (GPU) 114, a digital signal processor (DSP) 116, and / or an image signal processor (ISP) 118. The computing component 110 may perform various operations such as image enhancement, feature extraction, object or image segmentation, depth estimation, computer vision, graphics rendering, XR (e.g., augmented reality, virtual reality, mixed reality, etc.), image / video processing, sensor processing, recognition (e.g., text recognition, object recognition, feature recognition, facial recognition, pattern recognition, scene recognition, etc.), foreground prediction, machine learning, filtering, depth effect calculation or rendering, tracking, localization, and / or any of the various operations described herein. In some examples, the computing component 110 may implement an image processing engine 120, a neural network 122, and a rendering engine 124. In other examples, the computing component 110 may also implement one or more other processing engines.
[0040] The operations of image processing engine 120, neural network 122, and rendering engine 124 may be implemented by one or more computing components in computing components 110. In one illustrative example, image processing engine 120 and neural network 122 (and associated operations) may be implemented by CPU 112, DSP 116, and / or ISP 118, and rendering engine 124 (and associated operations) may be implemented by GPU 114. In some cases, computing component 110 may include other electronic circuitry or hardware, computer software, firmware, or any combination thereof to perform any of the various operations described herein.
[0041] In some cases, computing component 110 may receive data captured by image capture device 102 and / or image capture device 104 and process the data to generate an output image or video with certain visual and / or image processing effects (such as, for example, depth of field effects, background replacement, tracking, object detection, etc.). For example, computing component 110 may receive image data (e.g., one or more still images or video frames, etc.) captured by image capture devices 102, 104, perform depth estimation, image segmentation, and depth filtering, and generate output segmentation results, as described herein. The image (or frame) may be: an RGB image having red, green, and blue color components per pixel; a luminance, redness, blueness (YCbCr) image having a luminance component and two chrominance (color) components (redness and blueness) per pixel; or any other suitable type of color or monochrome picture.
[0042] The computing component 110 may implement an image processing engine 120 and a neural network 122 to perform various image processing operations and generate image effects. For example, the computing component 110 may implement an image processing engine 120 and a neural network 122 to perform feature extraction, superpixel detection, foreground prediction, spatial mapping, saliency detection, segmentation, depth estimation, depth filtering, pixel classification, cropping, upsampling / downsampling, blurring, modeling, filtering, color correction, noise reduction, scaling, sorting, adaptive Gaussian thresholding, and / or other image processing tasks. The computing component 110 may process: image data captured by the image capture devices 102 and / or 104; image data in the storage device 108; image data received from a remote source (such as a remote camera, server, or content provider); image data obtained from a combination of multiple sources; and the like.
[0043] In some examples, the computing component 110 may generate a depth map based on a monocular image captured by the image capture device 102; generate a segmentation map based on the monocular image; generate a refined or updated segmentation map based on depth filtering performed by comparing the depth map to the segmentation map to filter pixels / regions having at least a threshold depth; and generate a segmentation output. In some cases, the computing component 110 may use spatial information (e.g., a center prior map), a probability map, disparity information (e.g., a disparity map), an image query, a saliency map, etc. to segment objects and / or regions in one or more images and generate an output image with an image effect (such as a depth effect). In other cases, the computing component 110 may also use other information such as facial detection information, sensor measurements (e.g., depth measurements), depth measurements, etc.
[0044] In some examples, computing component 110 can perform segmentation (e.g., foreground-background segmentation, object segmentation, etc.) with pixel-level or region-level accuracy (or near such accuracy). In some cases, computing component 110 can perform segmentation using images with different FOVs. For example, computing component 110 can perform segmentation using an image with a first FOV captured by image capture device 102 and an image with a second FOV captured by image capture device 104. The segmentation can also enable (or be used in conjunction with) other image adjustments or image processing operations, such as, for example, but not limited to, depth enhancement and object-aware auto-exposure, auto-white balance, auto-focus, tone mapping, and the like.
[0045] Although the image processing system 100 is shown as including certain components, one of ordinary skill in the art will appreciate that the image processing system 100 may include more than Figure 1 For example, in some instances, the image processing system 100 may further include one or more memory devices (e.g., RAM, ROM, cache, etc.), one or more networking interfaces (e.g., wired and / or wireless communication interfaces, etc.), one or more display devices, and / or Figure 1 Other hardware or processing devices not shown. Figure 7 Illustrative examples of computing devices and hardware components that may be implemented utilizing image processing system 100 are described.
[0046] As previously mentioned, computing devices are increasingly equipped with the ability to capture images, perform various image processing tasks, generate various image effects, and more. Images and / or video frames can be processed to obtain object counts. Accurate object counting can be important and has numerous real-life applications. Different types of objects can be counted, including but not limited to people, animals, tangible items, and / or electronic devices. AI-based methods can be used to enhance object counting compared to traditional methods. However, erroneous object counts can still occur (for example, when objects appear in a mirror or glass in the image).
[0047] In one or more examples, you may want to count objects (e.g., people) within a given 3D space. For example, in an elevator scenario, such as a smart elevator, you may want to count people waiting in a rest area, but not in the elevator itself. In another example, in a video game console (such as an XR or AR device), you may want to count people within a space of interest (e.g., a specific area defined for the gaming experience), while excluding people outside of that space from the count to provide a better user gaming experience. In other examples, you may want to count people in a variety of other scenarios where a space of interest needs to be defined.
[0048] Accurate people counting requires 3D scene understanding, especially for verifying whether an object is located within the space of interest. Currently, existing 3D scene understanding or perception solutions require the use of additional sensors (e.g., stereo cameras, LIDAR, ToF sensors, etc.). Generally speaking, it is also necessary to build a 3D model of the environment. Existing solutions determine whether an object is located within the space of interest by determining coordinates related to the boundaries of the 3D model of the environment (e.g., coordinates that model the space of interest) and the coordinates of objects located in the environment, and performing a comparison (e.g., determining the distance of the object relative to the space of interest). The cost of building an accurate 3D model of the environment can be high, and equipment calibration may be required.
[0049] These systems and techniques use monocular 3D perception (e.g., based on monocular depth estimation) to provide accurate object counting. In one or more aspects, these systems and techniques can use a single RGB camera to provide accurate object counting (e.g., no additional depth sensor is required). These systems and techniques perform implicit space of interest calculations by using relative depth changes between scenes (e.g., current frame compared to a reference frame). These systems and techniques do not require camera calibration and, therefore, are easily adaptable to a variety of different scenarios and applications. These systems and techniques do not require 3D reconstruction algorithms and / or specific implementations, which can result in reduced demands on computing resources, thereby allowing for faster processing speeds and lower power consumption.
[0050] Figure 2 An example of a frame (e.g., a still image or a frame from a video) of scene 200 is shown. Specifically, Figure 2 is a diagram illustrating frames of an example of a scene 200 containing multiple objects (e.g., persons, such as person A 205a and person B 205b), wherein the objects (e.g., persons, such as person A 205a and person B 205b) are located outside a space of interest (e.g., a seating area 260 that may be used as an elevator waiting room). In one or more examples, Figure 2 The frame may be an RGB image with red, green, and blue color components per pixel.
[0051] Figure 2 The scene 200 shows an elevator scene, where the elevator 220 may or may not be a smart elevator. Figure 2In the example, a frame (e.g., captured from a still image or a video frame) showing scene 200 may be captured by a security camera located in rest area 260. In one or more examples, the security camera may be an RGB camera. Rest area 260 may be used by people waiting for elevator 220. In scene 200, rest area 260 is formed by wall 240, two glass doors 250a, 250b, and an elevator door of elevator 220 (e.g., shown as Figure 4 In the scene 200, the two glass doors 250a and 250b are shown as being located on opposite sides of the rest area 260.
[0052] In scene 200, the elevator doors of elevator 220 (e.g., shown as Figure 4 Elevator door 430 (e.g., elevator door 430) is shown open, and elevator threshold 230 is visible in scene 200. Also shown in scene 200 are two people (e.g., person A 205a and person B 205b). These two people (e.g., person A 205a and person B 205b) are shown inside elevator 220, such that neither person (e.g., person A 205a or person B 205b) has crossed elevator threshold 230 to enter rest area 260.
[0053] In one or more examples, such as in a smart elevator scenario, it may be desirable to count people waiting for elevator 220 in rest area 260, rather than people within elevator 220 itself. This requires identifying a space of interest (e.g., a defined 3D space) and objects of interest (e.g., people) for counting. In these examples, rest area 260 can be identified as the space of interest, and people (e.g., person A 205a and person B 205b) can be identified as objects of interest for counting. In one or more examples, image processing system 100 can perform the counting. Objects (e.g., people) within the space of interest (e.g., rest area 260) can be included in the count. Objects (e.g., people) not within the space of interest (e.g., rest area 260) and other objects (e.g., objects other than people) can be excluded from the count. Other objects (e.g., non-people) within the space of interest (e.g., rest area 260) also can be excluded from the count.
[0054] In some examples, image processing system 100 may detect persons (eg, person A 205a and person B 205b) in scene 200 of a frame. Figure 2As shown, the image processing system may generate bounding boxes 210a, 210b, and 210c (e.g., which may be segmentation masks) to indicate detections of people in the scene 200 of the frame. For example, bounding box 210a in the scene 200 may indicate detection of person A 205a, and bounding box 210c in the scene 200 may indicate detection of person B 205b. However, bounding box 210b is a false detection of a person because bounding box 210b results from detection of person B 205b's reflection in a mirror (e.g., located at the back of the elevator 220).
[0055] exist Figure 2 In the illustrated scene 200, the detected persons (e.g., person A 205a and person B 205b) in bounding boxes 210a and 210c (or the falsely detected persons in bounding box 210b) are not shown crossing the elevator threshold 230 to enter the rest area 230 (e.g., the space of interest). Therefore, the detected persons (e.g., person A 205a and person B 205b) (or the falsely detected persons in bounding box 210b) are not considered to be within the space of interest (e.g., the rest area 260). Because the detected persons (e.g., person A 205a and person B 205b) (or the falsely detected persons in bounding box 210b) are not considered to be within the space of interest (e.g., the rest area 230), the bounding boxes 210a, 210b, and 210c are depicted with bold solid lines.
[0056] Since none of the detected persons (e.g., person A 205a and person B 205b) (or the falsely detected persons in the bounding box 210b) are considered to be located within the space of interest (e.g., rest area 230), none of the detected persons (e.g., person A 205a and person B 205b) (or the falsely detected persons in the bounding box 210b) may be included in the count. Therefore, for the scene 200 of the frame, the image processing system 100 may count zero objects (e.g., persons) located within the space of interest (e.g., rest area 260).
[0057] Figure 3 An example of a frame (eg, a still image or a frame from a video) of a scene 300 is shown, the scene including the seating area 260, captured by a camera at Figure 2 200 frames of the scene are captured within a short time. Specifically, Figure 3 is a diagram illustrating a frame of an example of a scene 300 having multiple objects (e.g., persons, such as person A 205a and person B 205b), wherein the objects (e.g., persons, such as person A 205a and person B 205b) are located within a space of interest (e.g., rest area 260). In one or more examples, Figure 3The frame may be an RGB image with red, green, and blue color components per pixel.
[0058] exist Figure 3 In scene 300, a frame showing scene 300 (e.g., captured from a still image or video frame) may be captured by a security camera located within rest area 260. The frame showing scene 300 may be captured by the security camera a short time after the frame showing scene 200 is captured by the security camera. The security camera may be an RGB camera.
[0059] In some examples, the image processing system 100 may detect persons (eg, person A 205a and person B 205b) in the scene 300 of the frame. Figure 3 As shown, the image processing system may generate bounding boxes 310a, 310b, 320a, 320b (e.g., which may be segmentation masks) to indicate detections of people in a scene 300 of a frame. For example, bounding box 320a in scene 300 may indicate detection of person A 205a, and bounding box 320b in scene 300 may indicate detection of person B 205b. Bounding boxes 310a and 310b are false detections of people because bounding box 310a results from detection of person A 205a reflected in a mirror (e.g., located at the rear of elevator 220), and bounding box 310b results from detection of person B 205b reflected in a mirror (e.g., located at the rear of elevator 220).
[0060] exist Figure 3 In the scene 300, the elevator door of the elevator 220 (for example, shown as Figure 4 Elevator doors 430 are shown as being open. A person (e.g., person B 205b) within bounding box 320b is shown as being completely outside elevator 220 and completely within rest area 260 (e.g., a space of interest). Because this person (e.g., person B 205b) is shown as being completely outside elevator 220 and completely within rest area 260, this person (e.g., person B 205b) is considered to be within the space of interest (e.g., rest area 260), and bounding box 320b is depicted as a bold dashed line.
[0061] Another person (e.g., person A 205a) in bounding box 320a is shown in scene 300 as crossing (e.g., using their feet) over elevator threshold 230 and entering rest area 260. Because the person (e.g., person A 205a) is shown as being at least partially within rest area 260 (e.g., the area of interest), the person (e.g., person A 205a) is considered to be within the space of interest (e.g., rest area 260), and bounding box 320a is depicted as a bold dashed line.
[0062] The misdetected persons indicated by bounding boxes 310a and 310b are shown as being located within the elevator 220. Since the misdetected persons in the bounding boxes 310a and 310b are not at least partially located within the rest area 260 (e.g., the space of interest), the misdetected persons in the bounding boxes 310a and 310b are not considered to be located within the space of interest (e.g., the rest area 260), and the bounding boxes 310a and 310b are depicted in bold solid lines.
[0063] Since the detected persons (e.g., person A 205a and person B 205b) are considered to be located within the space of interest (e.g., rest area 230), the detected persons (e.g., person A 205a and person B 205b) can be included in the count. Since the falsely detected persons in bounding boxes 310a and 310b are not considered to be located within the space of interest (e.g., rest area 260), the falsely detected persons in bounding boxes 310a and 310b can be excluded from the count. Therefore, for the scene 300 of the frame, the image processing system 100 can count two objects (e.g., persons, person A 205a and person B 205b) located within the space of interest (e.g., rest area 260).
[0064] As described above, some existing solutions (e.g., AI-based or traditional methods) for counting objects (e.g., people) within a space of interest may incorrectly include falsely detected objects (e.g., reflections of objects (such as people) in a mirror) in their object counts. These systems and techniques provide a solution for counting objects within a space of interest that can distinguish real objects from falsely detected objects (e.g., reflections of objects) by using monocular depth estimation. In one or more aspects, these systems and techniques utilize a single RGB camera to provide depth estimation. These systems and techniques can perform implicit space of interest calculations using relative depth changes between scenes (e.g., between a current frame and a reference frame). By examining the relative depth changes in a two-dimensional region of interest (e.g., an object's bounding box or an object's segmentation mask), these systems and techniques can determine whether an object (e.g., a person) is within or outside the space of interest.
[0065] Figure 4 An example of a frame and associated depth map that can be used to count objects (e.g., people) using monocular depth based on an implicit space of interest (e.g., the lounge area 260 of the elevator 220) is shown. Specifically, Figure 4 is a diagram 400 illustrating examples of frames (eg, reference frame 410a and current frame 410b) and depth maps (eg, reference depth map 420a and current depth map 420b).
[0066] exist Figure 4In the image processing system 100, reference frame 410a is a frame (e.g., a still image or a frame from a video) that can be captured by a camera (such as an RGB camera) at a first time (e.g., a reference time). Image processing system 100 can use reference frame 410a to determine (e.g., define) a space of interest (e.g., which can be rest area 260). In reference frame 410a, elevator door 420 of elevator 220 is shown as closed, and therefore, the space of interest can be determined as (e.g., defined as) rest area 260 (e.g., the 3D space of rest area 260) enclosed by elevator door 230 of elevator 220.
[0067] Image processing system 100 may process (e.g., using image processing engine 120 of image processing system 100) reference frame 410a to generate reference depth map 420a (e.g., a monocular depth map). Reference depth map 420a illustrates a mapping corresponding to the depths of objects within the scene in reference frame 410a.
[0068] In addition Figure 4 In the example, current frame 410b is a frame (e.g., a still image or a frame from a video) that may be captured by a camera (e.g., an RGB camera) at a second time (e.g., the current time), which is subsequent to a first time (e.g., the reference time). Image processing system 100 may use current frame 410b to determine a count of objects (e.g., people) located within a determined (e.g., defined) space of interest (e.g., rest area 260). In current frame 410b, the elevator doors are open, and two people (e.g., person A 205a and person B 205b) are shown as being inside the elevator and not within rest area 260 (e.g., the space of interest). Since the people (e.g., person A 205a and person B 205b) are not within the space of interest (e.g., rest area 260), they should not be counted.
[0069] Image processing system 100 may process (e.g., using image processing engine 120 of image processing system 100) current frame 410b to generate current depth map 420b (e.g., a monocular depth map). Current depth map 420b shows a mapping corresponding to the depths of objects within the scene in reference frame 410b.
[0070] After the image processing system 100 has generated the reference depth map 420a and the current depth map 420b, the image processing system 100 may use the monocular depths in the reference depth map 420a and the current depth map 420b to calculate (e.g., compute) a depth variance from the reference depth map 420a and the current depth map 420b. The image processing system 100 may determine the depth variance from the reference depth map 420a and the current depth map 420b by subtracting the depth of the object in the current depth map 420b from the depth of the object in the reference depth map 420a. The image processing system 100 may then use the determined depth variance of the object of interest (e.g., indicated by the object's bounding box or segmentation mask) to determine whether the object (e.g., a person) is located within the space of interest (e.g., the rest area 260). When the image processing system 100 determines that the object of interest (e.g., indicated by the object's bounding box or segmentation mask) is located within the space of interest (e.g., the rest area 260), the image processing system 100 may include the object (e.g., the person) associated with the object of interest in the count. However, when the image processing system 100 determines that the target of interest (e.g., indicated by the object's bounding box or segmentation mask) is not located within the space of interest (e.g., rest area 260), the image processing system 100 may not include the object (e.g., person) associated with the target of interest in the count.
[0071] Figure 5 An example of a process that can be used (eg, by image processing system 100) to count objects (eg, targets of interest, such as people) within a space of interest (eg, rest area 260) is shown. Specifically, Figure 5 is a flow chart illustrating an example computational process 500 for accurate people counting. Figure 5 In the present invention, a reference frame 410a (e.g., a still image or a frame from a video) including a scene (e.g., a reference scene) may be captured by a camera (e.g., an RGB camera) at a first time (e.g., a reference time). The scene (e.g., the reference scene) of the reference frame 410a may not include any target of interest (e.g., an object such as a person). The reference frame 410a may be pre-processed (e.g., by the image processing system 100) during a system setup phase.
[0072] Image processing system 100 may process (e.g., using image processing engine 120 of image processing system 100) reference frame 410a to generate reference depth map 420a (e.g., a monocular depth map). In one or more examples, image processing system 100 may use depth estimator 510a to generate reference depth map 420a. Depth estimator 510a may be a machine learning model (e.g., a deep neural network trained using monocular depth-based deep learning training) or may implement such a machine learning model to generate reference depth map 420a. In one or more examples, depth estimator 510a may use self-supervised training, semi-self-supervised training, and / or fully supervised training for deep learning training. The generated reference depth map 420a shows a map corresponding to the depths of objects within the scene in reference frame 410a.
[0073] Then, at a second time (e.g., the current time) later than the first time (e.g., the reference time), a current frame 410b (e.g., a still image or a frame from a video) including a scene (e.g., the current scene) may be captured by a camera (e.g., an RGB camera). The image processing system 100 may process the current frame 410b in real time (e.g., using the image processing engine 120 of the image processing system 100) to generate a current depth map 420b (e.g., a monocular depth map). In some examples, the image processing system 100 may use a depth estimator 510b to generate the current depth map 420a. The depth estimator 510b may be a machine learning model (e.g., a deep neural network trained using deep learning based on monocular depth (such as self-supervised training, semi-self-supervised training, and / or fully supervised training)) or may implement the machine learning model to generate the current depth map 420b. The current reference depth map 420b shows a mapping corresponding to the depths of objects within the scene in the current frame 410b.
[0074] After a camera captures a current frame 410b (e.g., a still image or a frame from a video), the image processing system 100 may process the current frame 410b to detect an object 520 (e.g., a person) within the scene of the current frame 410b. After the image processing system 100 has detected the object 520 (e.g., a person) within the scene, the image processing system 100 may generate a bounding box (e.g., which may be an object segmentation mask) 520 around the detected object (e.g., a person) to indicate the detection of the object (e.g., a person) in the scene of the current frame 410b. For example, one bounding box may be generated within the scene of the current frame 410b to indicate the detection of a first person (e.g., person A), and a second bounding box may be generated within the scene of the current frame 410b to indicate the detection of a second person (e.g., person B).
[0075] Based on the generated bounding box (e.g., or segmentation mask), the image processing system 100 may generate 530 (e.g., determine and / or calculate) a relative change in depth 540 (e.g., a depth difference) of the detected object (e.g., surrounded by the bounding box or segmentation mask) to determine whether the object (e.g., a person) is present (e.g., located) within or outside the space of interest (e.g., the rest area 260). In one or more examples, the relative change in depth 540 (e.g., a depth difference) of the detected object (e.g., surrounded by the bounding box or segmentation mask) may be determined by comparing the depth of the reference depth map 420a with the depth of the current depth map 420b on a per-pixel basis.
[0076] After the image processing system 100 has determined the change in depth 540 (e.g., depth difference), the image processing system 100 may compare the change in depth 540 (e.g., depth difference) of each detected object to a threshold 550 (e.g., which may be a predetermined threshold, such as a distance value in meters) to determine whether the detected object is located within the space of interest (e.g., rest area 260). When the image processing system 100 determines that the change in depth 540 (e.g., depth difference) of the detected object is greater than 560 the threshold 550, the image processing system 100 may determine that the detected object (e.g., person) is located 570 within the space of interest (e.g., rest area 260). Conversely, when the image processing system 100 determines that the change in depth 540 (e.g., depth difference) of the detected object is less than 580 the threshold 550, the image processing system 100 may determine that the detected object (e.g., person) is not located 590 within the space of interest (e.g., rest area 260). The image processing system 100 may then count 570 the detected objects (eg, people) determined to be located within the space of interest (eg, the rest area 260 ) for object counting.
[0077] In one or more examples, to cope with changes in scene environment (eg, lighting conditions, newly added furniture, etc.), the reference frame 410 a may be updated when no target of interest (eg, an object such as a person) appears.
[0078] Figure 6 is a flow chart illustrating an example of a process 600 for accurate object counting using monocular 3D perception. The process 600 may be performed by a computing device (or system), or by a component or system (e.g., a chipset) of the computing device or system. In one illustrative example, the process 600 may be performed by Figure 1 The operations of process 600 may be implemented as a single process on one or more processors (e.g., Figure 1 One or more computing components in computing components 110, Figure 1 Image processor engine 120, Figure 1Rendering engine 124, Figure 7 software components executed and running on the processor 710, any combination thereof, and / or other processors).
[0079] At block 610, the computing device (or a component thereof) may generate a graph based on a representation of a space of interest (e.g., Figures 2 to 4 The rest area 260) of the reference frame (eg, Figure 4 410a) to generate a reference depth map (e.g., Figure 4 At block 620, the computing device (or a component thereof) may generate a reference depth map 420a for the space of interest and one or more objects based on a current frame (e.g., Figure 4 410b) to generate a current depth map (eg, Figure 4 In some cases, the one or more objects include people (e.g., Figures 2 to 4 In some aspects, the computing device (or its components) may obtain a reference scene (e.g., Figure 4 A reference frame (e.g., a scene depicted in the reference frame 410a) may be obtained that captures a current scene including the space of interest and one or more objects (e.g., Figure 4 410b). For example, the computing device (or a component thereof) may obtain a reference frame and a current frame from a camera. In some examples, the reference frame is a monocular frame captured by a monocular camera device (e.g., image capture device 102). Additionally or alternatively, in some examples, the current frame is a monocular frame captured by a monocular camera device (e.g., image capture device 102). In some cases, the reference frame and the current frame are each one of an image or a video frame.
[0080] In some aspects, a computing device (or a component thereof) may use a machine learning model (e.g., a neural network or other type of machine learning model) to generate the reference depth map and the current depth map. In some cases, the machine learning model is trained using self-supervised training, semi-self-supervised training, fully supervised training, any combination thereof, and / or other training processes.
[0081] In some aspects, the computing device (or a component thereof) may detect one or more objects in the current frame (e.g., Figure 5Detected objects / object segmentation mask 520). In some cases, the computing device (or a component thereof) may generate a corresponding bounding box for each of the one or more objects based on detecting the one or more objects. In some examples, the computing device (or a component thereof) may generate a segmentation mask (e.g., detected objects / object segmentation mask 520) for the one or more objects based on performing instance segmentation on the current frame.
[0082] At block 630, the computing device (or a component thereof) may compare the current depth map to the reference depth map to determine a corresponding depth change for each of the one or more objects. Figure 5 As described in the computational flow 500 , the computing device may compare the depth of the reference depth map 420 a with the depth of the current depth map 420 b (e.g., on a per-pixel basis) to determine a relative change (e.g., a depth difference) in the depth 540 of a detected object (e.g., surrounded by a bounding box or segmentation mask).
[0083] At block 640, the computing device (or a component thereof) may compare the corresponding depth change of each of the one or more objects to a threshold value. Figure 5 As described in the computational flow 500 , the computing device may compare a change in depth 540 (eg, a depth difference) of each detected object with a threshold 550 (eg, a predetermined threshold, such as a distance value in meters).
[0084] At block 650, the computing device (or a component thereof) may determine whether each of the one or more objects is located within the space of interest based on comparing the corresponding depth change of each of the one or more objects to a threshold value. Figure 5 As described in the computational flow 500 of FIG. 5 , the computing device may compare the change in depth 540 of each detected object to a threshold 550 to determine whether the detected object is located within the space of interest (e.g., the rest area 260). In some aspects, the computing device (or a component thereof) may count at least one of the one or more objects (e.g., a target of interest, such as a person) located within the space of interest. For example, as described with respect to FIG. Figure 4 and Figure 5 As described, a computing device (or a component thereof) may determine that one or more objects of interest (eg, indicated by a bounding box or segmentation mask of the object) are located in a space of interest (eg, Figures 2 to 4If the computing device (or a component thereof) determines that the target of interest is not located within the space of interest (e.g., rest area 260), in which case the computing device (or a component thereof) may include the object (e.g., person) associated with the target of interest in the count. However, if the computing device (or a component thereof) determines that the target of interest is not located within the space of interest (e.g., rest area 260), the computing device (or a component thereof) may not include the object (e.g., person) associated with the target of interest in the count.
[0085] In some examples, process 600 may be performed by one or more computing devices or apparatuses. In one illustrative example, process 600 may be performed by Figure 1 The image processing system 100 shown and / or having Figure 7 The computing device architecture 700 shown is executed by one or more computing devices. In some cases, such computing devices or devices may include a processor, a microprocessor, a microcomputer, or other components of a device configured to perform the steps of process 600. In some examples, such computing devices or devices may include one or more sensors configured to capture image data. For example, the computing device may include a smartphone, a head-mounted display, a mobile device, a camera, a tablet computer, or other suitable device. In some examples, such computing devices or devices may include a camera configured to capture one or more images or videos. In some cases, such computing devices may include a display for displaying images. In some examples, one or more sensors and / or cameras are separate from the computing device, in which case the computing device receives the sensed data. Such computing devices may further include a network interface configured to communicate data.
[0086] Components of a computing device can be implemented in circuitry. For example, a component may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein. A computing device may also include a display (as an example of an output device or in addition to an output device), a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive data based on the Internet Protocol (IP) or other types of data.
[0087] Process 600 is illustrated as a logical flow diagram, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, each operation represents computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally speaking, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform specific functions or implement specific data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the process.
[0088] Additionally, process 600 may be executed under the control of one or more computer systems configured with executable instructions and may be implemented via hardware or a combination thereof as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is collectively executed on one or more processors. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions that can be executed by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0089] Figure 7 An example computing device architecture 700 illustrates an example computing device that can implement the various techniques described herein. For example, the computing device architecture 700 can implement Figure 1 1 and 2. The example computing device architecture 700 is shown as a computer system that is electrically connected to the computer system 100. The components of the computing device architecture 700 are shown as being in electrical communication with each other using connections 705, such as a bus. The example computing device architecture 700 includes a processing unit (CPU or processor) 710 and computing device connections 705 that couple various computing device components, including computing device memory 715, such as read-only memory (ROM) 720 and random access memory (RAM) 725, to the processor 710.
[0090] The computing device architecture 700 may include a cache of high-speed memory directly connected to, near, or integrated as part of the processor 710. The computing device architecture 700 may copy data from memory 715 and / or storage device 730 to cache 712 for quick access by the processor 710. In this way, the cache may provide a performance boost by preventing the processor 710 from being delayed while waiting for data. These and other modules may control or be configured to control the processor 710 to perform various actions. Other computing device memory 715 may also be available for use. The memory 715 may include a variety of different types of memory with different performance characteristics.
[0091] Processor 710 may include any general-purpose processor and hardware or software services configured to control processor 710 (such as service 1 732, service 2 734, and service 3 736 stored in storage device 730), as well as specialized processors that incorporate software instructions into the processor design. Processor 710 may be a self-contained system containing multiple cores or processors, a bus, a memory controller, a cache, etc. Multi-core processors may be symmetric or asymmetric.
[0092] To enable user interaction with the computing device architecture 700, the input device 745 can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphic input, a keyboard, a mouse, motion input, voice, etc. The output device 735 can also be one or more of a variety of output mechanisms known to those skilled in the art, such as a display, a projector, a television, a speaker device. In some instances, a multimodal computing device can enable a user to provide multiple types of input to communicate with the computing device architecture 700. The communication interface 740 can generally control and manage user input and computing device output. There is no limitation on operating on any particular hardware arrangement, and therefore the basic features herein can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.
[0093] Storage device 730 is a non-volatile memory and can be a hard disk or other type of computer-readable medium that can store computer-accessible data, such as a magnetic tape cartridge, a flash memory card, a solid-state memory device, a digital versatile disk, a magnetic cassette, random access memory (RAM) 725, read-only memory (ROM) 720, or a combination thereof. Storage device 730 may include services 732, 734, and 736 for controlling processor 710. Other hardware or software modules are contemplated. Storage device 730 may be connected to computing device connector 705. In one aspect, a hardware module that performs a particular function may include a software component stored in a computer-readable medium connected to the necessary hardware components (such as processor 710, connector 705, output device 735, etc.) to perform that function.
[0094] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media that can store data and does not include carrier waves and / or transient electronic signals propagating wirelessly or over wired connections. Examples of non-transitory media include, but are not limited to, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memory, or storage devices. A computer-readable medium may have stored thereon code and / or machine-executable instructions, which may represent a procedure, function, subroutine, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means, including memory sharing, message passing, token passing, network transmission, and the like.
[0095] In some embodiments, computer-readable storage devices, media, and memories may include wired or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media specifically excludes media such as power consumption, carrier signals, electromagnetic waves, and signals themselves.
[0096] Specific details are provided in the description above to provide a thorough understanding of the embodiments and examples provided herein. However, it will be understood by those skilled in the art that embodiments can be put into practice without these specific details. For clarity of explanation, in some instances, the present technology can be presented as including a separate functional block, which includes a device, device assembly, step or routine in a method embodied in a combination of software or hardware and software. Additional components other than those components shown in the accompanying drawings and / or described herein can be used. For example, circuits, systems, networks, processes and other components can be shown as components in block diagram form to avoid these embodiments from becoming difficult to understand in unnecessary details. In other instances, known circuits, processes, algorithms, structures and techniques can be shown without necessary details to avoid making each embodiment difficult to understand.
[0097] Individual embodiments may be described above as processes or methods depicted as flowcharts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams. Although a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or concurrently. Furthermore, the order of the operations may be rearranged. A process is terminated when its operations are completed, but a process may have additional steps not included in the accompanying drawings. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, termination of the process may correspond to the function returning to the calling function or main function.
[0098] The processes and methods according to the examples described above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtained from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions. Portions of the computer resources used may be accessible over a network. The computer-executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, source code. Examples of computer-readable media that can be used to store instructions, information used, and / or information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, and the like.
[0099] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., a computer program product) for performing the necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein may also be embodied in peripheral devices or add-in cards. By way of further example, such functionality may also be implemented on circuit boards in different chips or different processes executed on a single device.
[0100] Instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.
[0101] In the foregoing description, various aspects of the present application have been described with reference to the specific embodiments of the present application, but those skilled in the art will recognize that the present application is not limited thereto. Thus, although the exemplary embodiments of the present application have been described in detail herein, it is to be understood that the inventive concept can be embodied and adopted in various other ways, and the appended claims are intended to be interpreted as including such variations, unless limited by the prior art. The various features and aspects of the above-mentioned applications may be used individually or in combination. In addition, without departing from the broader essence and scope of this specification, the embodiments may be used in any number of environments and applications beyond the environments and applications described herein. Therefore, the description and the accompanying drawings should be considered as illustrative rather than restrictive. For illustrative purposes, each method is described in a specific order. It should be understood that in an alternative embodiment, each method may be performed in a different order than described.
[0102] Those of ordinary skill in the art will appreciate that the less than ("<") and greater than (">") symbols or terms used herein can be replaced by less than or equal to (" ") and greater than or equal to (" )” symbol without departing from the scope of this description.
[0103] Where a component is described as being “configured to” perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., a microprocessor or other suitable electronic circuits) to perform the operations, or any combination thereof.
[0104] The phrase “coupled to” refers to any component being directly or indirectly physically connected to another component, and / or any component being in direct or indirect communication with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0105] Claim language or other language reciting "at least one of" a set and / or "one or more of" a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, claim language reciting "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language reciting "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A, B, and C. The language "at least one of" a set and / or "one or more of" a set does not limit the set to the items listed in the set. For example, claim language reciting "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.
[0106] The various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the examples disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the entire system. Technicians may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be interpreted as departing from the scope of this application.
[0107] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices, or integrated circuit devices with multiple uses, including applications in wireless communication devices and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be implemented at least in part by a computer-readable data storage medium containing program code, which, when executed, includes instructions that perform one or more of the methods, algorithms, and / or operations described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0108] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such processors may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Thus, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein.
[0109] Illustrative aspects of the present disclosure include:
[0110] Aspect 1. A method for processing one or more frames, the method comprising: generating a reference depth map based on a reference frame depicting a space of interest; generating a current depth map based on a current frame depicting the space of interest and one or more objects; comparing the current depth map with the reference depth map to determine a corresponding depth change for each of the one or more objects; comparing the corresponding depth change for each of the one or more objects with a threshold; and determining whether each of the one or more objects is located within the space of interest based on comparing the corresponding depth change for each of the one or more objects with the threshold.
[0111] Aspect 2. The method according to aspect 1 further includes: obtaining the reference frame that captures the reference scene including the space of interest; and obtaining the current frame that captures the current scene including the space of interest and the one or more objects.
[0112] Aspect 3. The method according to aspect 2, wherein obtaining the reference frame and obtaining the current frame are performed using a camera.
[0113] Aspect 4. The method according to any one of aspects 1 to 3, wherein the reference frame and the current frame are each one of an image or a video frame.
[0114] Aspect 5. The method according to any one of aspects 1 to 4 further includes detecting the one or more objects in the current frame.
[0115] Aspect 6. The method according to aspect 5, further comprising generating a corresponding bounding box for each of the one or more objects based on detecting the one or more objects.
[0116] Aspect 7. The method according to any one of aspects 1 to 6, further comprising generating a segmentation mask for the one or more objects based on performing instance segmentation on the current frame.
[0117] Aspect 8. The method according to any one of aspects 1 to 7, further comprising counting at least one object among the one or more objects located within the space of interest.
[0118] Aspect 9. The method according to any one of aspects 1 to 8, wherein the one or more objects include at least one of a person, an animal, a tangible item, or an electronic device.
[0119] Aspect 10. The method according to any one of aspects 1 to 9, wherein the reference depth map and the current depth map are generated using a machine learning model.
[0120] Aspect 11. A method according to aspect 10, wherein the machine learning model is trained using at least one of self-supervised training, semi-self-supervised training, or fully supervised training.
[0121] Aspect 12. A device for processing one or more frames, the device comprising: at least one memory; and at least one processor, the at least one processor being coupled to the at least one memory and configured to: generate a reference depth map based on a reference frame depicting a space of interest; generate a current depth map based on a current frame depicting the space of interest and one or more objects; compare the current depth map with the reference depth map to determine a corresponding depth change for each of the one or more objects; compare the corresponding depth change for each of the one or more objects with a threshold; and determine whether each of the one or more objects is located within the space of interest based on comparing the corresponding depth change for each of the one or more objects with the threshold.
[0122] Aspect 13. An apparatus according to Aspect 12, wherein the at least one processor is configured to: obtain the reference frame that captures the reference scene including the space of interest; and obtain the current frame that captures the current scene including the space of interest and the one or more objects.
[0123] Clause 14. The apparatus of clause 13, wherein the at least one processor is configured to obtain the reference frame and the current frame from a camera.
[0124] Clause 15. The apparatus of any one of clauses 12 to 14, wherein the reference frame and the current frame are each one of an image or a video frame.
[0125] Aspect 16. The apparatus according to any one of aspects 12 to 15, wherein the at least one processor is configured to detect the one or more objects in the current frame.
[0126] Aspect 17. The apparatus of aspect 16, wherein the at least one processor is configured to generate a corresponding bounding box for each of the one or more objects based on detecting the one or more objects.
[0127] Clause 18. The apparatus according to any one of clauses 12 to 17, wherein the at least one processor is configured to generate a segmentation mask for the one or more objects based on performing instance segmentation on the current frame.
[0128] Aspect 19. The apparatus according to any one of aspects 12 to 18, wherein the at least one processor is configured to count at least one object of the one or more objects located within the space of interest.
[0129] Aspect 20. The apparatus according to any one of aspects 12 to 19, wherein the one or more objects include at least one of a person, an animal, a tangible item, or an electronic device.
[0130] Aspect 21. An apparatus according to any one of aspects 12 to 20, wherein the at least one processor is configured to generate the reference depth map and the current depth map using a machine learning model.
[0131] Aspect 22. An apparatus according to Aspect 21, wherein the machine learning model is trained using at least one of self-supervised training, semi-self-supervised training, or fully supervised training.
[0132] Aspect 23. A non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by at least one processor, causing the at least one processor to perform the operations according to any one of aspects 1 to 11.
[0133] Aspect 24. An apparatus for wireless communication, the apparatus comprising one or more means for performing the operations of any one of aspects 1 to 11.
[0134] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the aspects shown herein, but are to be accorded the full scope consistent with the language of the claims, wherein reference to an element in the singular is not intended to mean "one and only one" unless specifically stated otherwise, but rather "one or more."
Claims
1. A method for processing one or more frames, the method comprising: generating a reference depth map based on a reference frame depicting the space of interest; generating a current depth map based on a current frame depicting the space of interest and the one or more objects; comparing the current depth map to the reference depth map to determine a respective depth change for each of the one or more objects; comparing the corresponding depth change of each of the one or more objects to a threshold; as well as Whether each of the one or more objects is located within the space of interest is determined based on comparing the corresponding depth change of each of the one or more objects to the threshold.
2. The method according to claim 1, further comprising: obtaining the reference frame capturing a reference scene including the space of interest; as well as The current frame capturing a current scene including the space of interest and the one or more objects is obtained. The method of claim 2 , wherein obtaining the reference frame and obtaining the current frame are performed using a camera. The method of claim 1 , wherein the reference frame and the current frame are each one of an image or a video frame. The method of claim 1 , further comprising detecting the one or more objects in the current frame. The method of claim 5 , further comprising generating a corresponding bounding box for each of the one or more objects based on detecting the one or more objects. 7 . The method of claim 1 , further comprising generating a segmentation mask for the one or more objects based on performing instance segmentation on the current frame. 8 . The method of claim 1 , further comprising counting at least one of the one or more objects located within the space of interest.
9. The method of claim 1, wherein the one or more objects include at least one of a person, an animal, a tangible item, or an electronic device.
10. The method of claim 1, wherein the reference depth map and the current depth map are generated using a machine learning model.
11. The method of claim 10, wherein the machine learning model is trained using at least one of self-supervised training, semi-self-supervised training, or fully supervised training.
12. An apparatus for processing one or more frames, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: generating a reference depth map based on a reference frame depicting the space of interest; generating a current depth map based on a current frame depicting the space of interest and the one or more objects; comparing the current depth map to the reference depth map to determine a respective depth change for each of the one or more objects; comparing the corresponding depth change of each of the one or more objects to a threshold; as well as Whether each of the one or more objects is located within the space of interest is determined based on comparing the corresponding depth change of each of the one or more objects to the threshold.
13. The apparatus of claim 12, wherein the at least one processor is configured to: obtaining the reference frame capturing a reference scene including the space of interest; and The current frame capturing a current scene including the space of interest and the one or more objects is obtained. The apparatus of claim 13 , wherein the at least one processor is configured to obtain the reference frame and the current frame from a camera.
15. The device of claim 12, wherein the reference frame and the current frame are each one of an image or a video frame.
16. The apparatus of claim 12, wherein the at least one processor is configured to detect the one or more objects in the current frame. 17 . The apparatus of claim 16 , wherein the at least one processor is configured to generate a corresponding bounding box for each of the one or more objects based on detecting the one or more objects.
18. The apparatus of claim 12, wherein the at least one processor is configured to generate a segmentation mask for the one or more objects based on performing instance segmentation on the current frame.
19. The apparatus of claim 12, wherein the at least one processor is configured to count at least one of the one or more objects located within the space of interest.
20. The apparatus of claim 12, wherein the one or more objects comprise at least one of a person, an animal, a tangible item, or an electronic device.
21. The apparatus of claim 12, wherein the at least one processor is configured to generate the reference depth map and the current depth map using a machine learning model.
22. The apparatus of claim 21, wherein the machine learning model is trained using at least one of self-supervised training, semi-self-supervised training, or fully supervised training.
23. A non-transitory computer-readable medium having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to: generating a reference depth map based on a reference frame depicting the space of interest; generating a current depth map based on a current frame depicting the space of interest and the one or more objects; comparing the current depth map to the reference depth map to determine a respective depth change for each of the one or more objects; comparing the corresponding depth change of each of the one or more objects to a threshold; as well as Whether each of the one or more objects is located within the space of interest is determined based on comparing the corresponding depth change of each of the one or more objects to the threshold.
24. The non-transitory computer-readable medium of claim 23, wherein the instructions, when executed by the at least one processor, cause the at least one processor to: obtaining the reference frame capturing a reference scene including the space of interest; and The current frame capturing a current scene including the space of interest and the one or more objects is obtained.
25. The non-transitory computer-readable medium of claim 23, wherein the instructions, when executed by the at least one processor, cause the at least one processor to detect the one or more objects in the current frame.
26. The non-transitory computer-readable medium of claim 25, wherein the instructions, when executed by the at least one processor, cause the at least one processor to generate a respective bounding box for each of the one or more objects based on detecting the one or more objects.
27. The non-transitory computer-readable medium of claim 23, wherein the instructions, when executed by the at least one processor, cause the at least one processor to generate a segmentation mask for the one or more objects based on performing instance segmentation on the current frame.
28. The non-transitory computer-readable medium of claim 23, wherein the instructions, when executed by the at least one processor, cause the at least one processor to count at least one of the one or more objects located within the space of interest.
29. The non-transitory computer-readable medium of claim 23, wherein the instructions, when executed by the at least one processor, cause the at least one processor to generate the reference depth map and the current depth map using a machine learning model.
30. The non-transitory computer-readable medium of claim 29, wherein the machine learning model is trained using at least one of self-supervised training, semi-self-supervised training, or fully supervised training.