Multi-task object detection system for detecting occluded objects in image
By generating feature maps and using Hungarian algorithms, the problem of inaccurate object detection in convolutional neural networks is solved, and accurate correlation and tracking of objects and parts in the image is achieved.
Patent Information
- Application Number
- CN202380084807.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-21
- Filing Date
- 2023-10-20
- Publication Date
- 2025-07-18
AI Technical Summary
When existing convolutional neural networks detect objects that are blocked in images, it is difficult to accurately identify the relationships and parts of the objects, resulting in inaccurate detection.
A multi-task object detection system is adopted to identify and associate objects and their parts in the image by generating feature maps and using Hungarian algorithms to construct cost functions, combining depth information, and maintain accuracy especially in occlusion.
It improves the detection accuracy and tracking ability of occluded objects, can correctly associate the parts of the objects under complex visual conditions, and reduces false detection and missed detection.
Smart Images

Figure CN120345007A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to object detection via neural networks. In some examples, aspects of the present disclosure relate to a multi-task object detection system for detecting occluded objects in an image. Background Art
[0002] Deep neural networks, specifically convolutional neural networks, are used for object detection. A convolutional neural network is capable of extracting high-level features (such as facial shapes) from an input image and using these high-level features to output, for example, the probability that the input image includes a dog, a cat, a boat, or a bird. An artificial neural network attempts to use computer technology to replicate the logical reasoning performed by a biological neural network.
[0003] Although convolutional neural networks can identify objects within a scene, many of those objects may have relationships. However, the identification of relationships within the input image may be difficult, and objects may be misdetected due to occlusion or other reasons. For example, some parts of a person may be missed due to occlusion or other reasons, and body part detection is not as accurate as face detection. Summary of the Invention
[0004] In some examples, systems and techniques for detecting occluded objects in an image are described. The systems and techniques may improve the identification of objects and sub-features of those objects when the objects are at least partially occluded.
[0005] According to at least one example, a method for processing an image is provided. The method includes: obtaining an image including at least a first object; generating a feature map based on providing the image to a neural network; identifying a plurality of objects based on the feature map, the plurality of objects including a first part of the first object; and identifying a first set of object parts corresponding to the first object within the plurality of objects.
[0006] In another example, an apparatus for processing an image is provided, the apparatus including at least one memory and at least one processor, the at least one processor being coupled to the at least one memory. The at least one processor is configured to: obtain an image including at least a first object; generate a feature map based on providing the image to a neural network; identify a plurality of objects based on the feature map, the plurality of objects including a first part of the first object; and identify a first set of object parts corresponding to the first object within the plurality of objects.
[0007] In another example, a non-transitory computer-readable medium is provided that has instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: obtain an image that includes at least a first object; generate a feature map based on providing the image to a neural network; identify a plurality of objects based on the feature map, the plurality of objects including a first part of the first object; and identify a first set of object parts within the plurality of objects that correspond to the first object.
[0008] In another example, an apparatus for processing an image is provided. The apparatus includes: means for obtaining an image that includes at least a first object; means for generating a feature map based on providing the image to a neural network; means for identifying a plurality of objects based on the feature map, the plurality of objects including a first part of the first object; and means for identifying a first set of object parts within the plurality of objects that correspond to the first object.
[0009] In some aspects, one or more of the apparatuses described herein are the following devices, are part of the following devices, and / or include the following devices: a mobile device (e.g., a mobile phone and / or a mobile handset and / or a so-called "smartphone" or other mobile device), an extended reality (XR) device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device, a head-mounted device (HMD) device), a vehicle or a computing system, device, or component of a vehicle, a wearable device (e.g., a network-connected watch or other wearable device), a wireless communication device, a camera, a personal computer, a laptop computer, a server computer, another device, or a combination thereof. In some aspects, the apparatus includes one camera or a plurality of cameras for capturing one or more images. In some aspects, the apparatus further includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the aforementioned apparatus may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyro testers, one or more accelerometers, any combination thereof, and / or other sensors).
[0010] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to the appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.
[0011] The foregoing and other features and aspects will become more apparent when referring to the following specification, claims, and appended drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Exemplary aspects of the present application are described in detail below with reference to the following figures:
[0013] Figure 1 is a block diagram illustrating the architecture of an image capture and processing system according to some examples;
[0014] Figure 2 is a diagram illustrating an example of a model for a convolutional neural network.
[0015] Figure 3 is a conceptual diagram illustrating the architecture of a computer vision system 300 according to various aspects of the present disclosure.
[0016] Figure 4 illustrates a block diagram of a computer vision system 400 for performing object part association according to various aspects of the present disclosure.
[0017] Figure 5A and Figure 5B illustrates an example result of a bounding box identified from an image according to various aspects of the present disclosure.
[0018] Figure 6 illustrates an example joint attention module (JAM) configured to identify relevant channels similar to squeeze and excitation blocks.
[0019] Figure 7A illustrates an image of the sigmoid operation performed in a squeeze and excitation block.
[0020] Figure 7B illustrates an image of the clipped rectified linear unit (ReLU) operation performed in a JAM according to various aspects of the present disclosure.
[0021] Figure 8 is a flowchart illustrating an example of a method for performing object part association according to certain aspects of the present disclosure.
[0022] Figure 9 is a flowchart illustrating an example of a method for performing object part association according to certain aspects of the present disclosure.
[0023] Figure 10 is a diagram illustrating an example of a system for implementing certain aspects described herein. Detailed Description
[0024] Certain aspects of the present disclosure are provided below. Some of these aspects can be applied independently, and some of them can be applied in combination, which will be obvious to those skilled in the art. In the following description, specific details are set forth for purposes of explanation to provide a thorough understanding of the various aspects of the present application. However, it will be apparent that the various aspects can be practiced without these specific details. The accompanying drawings and description are not intended to be restrictive.
[0025] The following description provides only example aspects and is not intended to limit the scope, applicability, or configuration of the present disclosure. Instead, the following description of example aspects will provide those skilled in the art with a description that can be used to implement the example aspects. It should be understood that various changes can be made to the function and arrangement of elements without departing from the spirit and scope of the present application as set forth in the appended claims.
[0026] The term "exemplary" and / or "example" is used herein to mean "serving as an example, instance, or illustration". Any aspect described herein as "exemplary" and / or "example" need not be construed as superior or better than other aspects. Similarly, the term "aspects of the present disclosure" does not require that all aspects of the present disclosure include the discussed features, advantages, or modes of operation.
[0027] A camera is a device that uses an image sensor to receive light and capture image frames (such as still images or video frames). The terms "image", "image frame", and "frame" can be used interchangeably herein. A camera can be configured with various image capture and image processing settings. Different settings produce images with different appearances. Some camera settings are determined and applied before or during the capture of one or more image frames, such as ISO, exposure time, aperture size, aperture value, shutter speed, focus, and gain. For example, settings or parameters can be applied to the image sensor used to capture one or more image frames. Other camera settings can configure the post-processing of one or more image frames, such as changes in contrast, brightness, saturation, sharpness, levels, curves, or color. For example, settings or parameters can be applied to a processor (e.g., an image signal processor or ISP) used to process one or more image frames captured by the image sensor.
[0028] Computer vision is an important task that can be performed by computer systems to carry out various tasks such as detecting human occupancy, identifying a person, tracking the movement of a person, etc. One type of computer vision is face recognition based on facial landmark detection, which involves locating key (e.g., important or relevant) facial points in an image frame. Key facial points can include facial points such as the corners of the eyes, the corners of the mouth, and the tip of the nose. The positions of the detected facial landmarks can characterize and / or indicate the shape of the face (and the shape of one or more facial features such as the nose or mouth). Computer vision can also identify other parts of a person such as the hands, the body center, and the feet to track the person. Computer vision can be used for many purposes such as identifying the movement of objects within an environment to enable an autonomous device to safely navigate the scene.
[0029] For devices and systems that implement computer vision, understanding complex object activities is a challenging and critical task. For example, an autonomous vehicle (AV) tracks a person to prevent the AV from hitting a person crossing the street. Computer vision typically requires further understanding after detecting different objects (such as after separately detecting the bodies, faces, and hands of multiple people). Due to various reasons such as body detection (which is not as accurate as face detection), parts of an object (such as a person) may be incorrectly detected. Occlusions within an image often result in incorrect identification of parts of an object (such as a person) and incorrect association with incorrect objects.
[0030] This disclosure describes systems, devices, methods, and computer-readable media for performing human body part association in images (collectively referred to as "systems and techniques"). For example, the systems and techniques can identify body parts of a person and associate different body parts with each person. Aspects of this disclosure include generating a feature map of an object and identifying different features from the feature map. Examples of the object are a person, but the object can also be another mammal or another moving object such as a vehicle, a robot, etc. In the case of a person, the systems and techniques can identify the face, hands, body center, feet, and other features of the person. When body parts of a person can be identified from the feature map, the systems and techniques can associate each object with the person by constructing a cost function. The systems and techniques can associate different objects with body parts based on the cost function (e.g., using the Hungarian algorithm). The systems and techniques can also generate a region corresponding to the person based on the identification of the body parts of the person to enable tracking of the person in a next image.
[0031] When an object is occluded and challenging visual conditions are created, the systems and techniques can also perform object association in an image. For example, a person in the foreground may have body parts that at least partially overlap with a person in the background. The systems and techniques can be configured to use body part association to build person containers to enable detection of at least one body part, such that the entire person can be detected for tracking, activity understanding, and other tasks. In one aspect, the systems and techniques determine the depth of the body parts to disambiguate the foreground person and their corresponding body parts from the background person and their corresponding body parts. The systems and techniques can determine that the identity of at least one body part may be incorrectly identified. In such a case, the systems and techniques can also determine the depth of each object and construct a cost function that incorporates the depth. The systems and techniques can associate different objects with body parts based on the cost function (e.g., using the Hungarian algorithm). The systems and techniques can also generate regions corresponding to the foreground person and the background person based on the identity of the body parts of each person.
[0032] Additional details and aspects of the present disclosure are described in more detail below with respect to the drawings.
[0033] Figure 1 is a block diagram illustrating the architecture of an image capture and processing system 100. The image capture and processing system 100 includes various components for capturing and processing images of a scene (e.g., an image of scene 110). The image capture and processing system 100 can capture individual images (or photos), and / or can capture video including multiple images (or video frames) in a particular sequence. In some cases, the lens 115 and the image sensor 130 can be associated with an optical axis. In one illustrative example, the photosensitive area of the image sensor 130 (e.g., a photodiode) and the lens 115 can both be centered on the optical axis. The lens 115 of the image capture and processing system 100 faces the scene 110 and receives light from the scene 110. The lens 115 bends the incoming light from the scene towards the image sensor 130. The light received by the lens 115 passes through an aperture. In some cases, the aperture (e.g., the aperture size) is controlled by one or more control mechanisms 120 and is received by the image sensor 130. In some cases, the aperture can have a fixed size.
[0034] One or more control mechanisms 120 may control exposure, focus, and / or zoom based on information from the image sensor 130 and / or based on information from the image processor 150. The one or more control mechanisms 120 may include multiple mechanisms and components; for example, the control mechanism 120 may include one or more exposure control mechanisms 125A, one or more focus control mechanisms 125B, and / or one or more zoom control mechanisms 125C. The one or more control mechanisms 120 may also include additional control mechanisms other than those illustrated, such as control mechanisms for controlling analog gain, flash, HDR, depth of field, and / or other image capture attributes.
[0035] The focus control mechanism 125B of the control mechanism 120 may obtain a focus setting. In some examples, the focus control mechanism 125B stores the focus setting in a memory register. Based on the focus setting, the focus control mechanism 125B may adjust the positioning of the lens 115 relative to the positioning of the image sensor 130. For example, based on the focus setting, the focus control mechanism 125B may move the lens 115 closer to or farther from the image sensor 130 by actuating a motor or a servo system (or other lens mechanism) to adjust the focus. In some cases, additional lenses may be included in the image capture and processing system 100, such as one or more microlenses located above each photodiode of the image sensor 130, each of the one or more microlenses bending the light received from the lens 115 toward the corresponding photodiode before the light reaches the corresponding photodiode. The focus setting may be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), hybrid autofocus (HAF), or some combination thereof. The focus setting may be determined using the control mechanism 120, the image sensor 130, and / or the image processor 150. The focus setting may be referred to as an image capture setting and / or an image processing setting. In some cases, the lens 115 may be fixed relative to the image sensor, and the focus control mechanism 125B may be omitted without departing from the scope of the present disclosure.
[0036] The exposure control mechanism 125A of the control mechanism 120 may obtain an exposure setting. In some cases, the exposure control mechanism 125A stores the exposure setting in a memory register. Based on the exposure setting, the exposure control mechanism 125A may control the size of the aperture (e.g., aperture size or f-number), the duration for which the aperture is open (e.g., exposure time or shutter speed), the duration for which the sensor collects light (e.g., exposure time or electronic shutter speed), the sensitivity of the image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 130, or any combination thereof. The exposure setting may be referred to as an image capture setting and / or an image processing setting.
[0037] The zoom control mechanism 125C of the control mechanism 120 can obtain a zoom setting. In some examples, the zoom control mechanism 125C stores the zoom setting in a memory register. Based on the zoom setting, the zoom control mechanism 125C can control the focal length of an assembly of lens elements (lens assembly) including the lens 115 and one or more additional lenses. For example, the zoom control mechanism 125C can control the focal length of the lens assembly by actuating one or more motors or servo systems (or other lens mechanisms) to move one or more lenses relative to each other. The zoom setting can be referred to as an image capture setting and / or an image processing setting. In some examples, the lens assembly can include a parfocal zoom lens or a variable focal length zoom lens. In some examples, the lens assembly can include a focusing lens (in some cases, the focusing lens can be the lens 115), which first receives light from the scene 110, and the light then passes through an afocal zoom system between the focusing lens (e.g., the lens 115) and the image sensor 130 before reaching the image sensor 130. In some cases, the afocal zoom system can include two positive (e.g., converging, convex) lenses with equal or similar focal lengths (e.g., within a threshold difference from each other), with a negative (e.g., diverging, concave) lens between them. In some cases, the zoom control mechanism 125C moves one or more of the lenses in the afocal zoom system, such as the negative lens and one or both of the positive lenses. In some cases, the zoom control mechanism 125C can control the zoom by capturing an image from an image sensor (e.g., including the image sensor 130) among a plurality of image sensors with a zoom corresponding to the zoom setting. For example, the image capture and processing system 100 can include a wide-angle image sensor with a relatively low zoom and a telephoto image sensor with a larger zoom. In some cases, based on the selected zoom setting, the zoom control mechanism 125C can capture an image from the corresponding sensor.
[0038] The image sensor 130 includes one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image generated by the image sensor 130. In some cases, different photodiodes may be covered by different filters. In some cases, different photodiodes may be covered in different color filters and may thus measure light that matches the color of the filter covering the photodiode. Various color filter arrays can be used, including a Bayer color filter array, a four-color filter array (also known as a four-color Bayer filter array or QCFA), and / or any other color filter array. For example, a Bayer color filter includes a red color filter, a blue color filter, and a green color filter, where each pixel of the image is generated based on red light data from at least one photodiode covered in the red color filter, blue light data from at least one photodiode covered in the blue color filter, and green light data from at least one photodiode covered in the green color filter.
[0039] Return to Figure 1 , yellow, magenta, and / or cyan (also known as "emerald") filters can be used to replace or supplement the red, blue, and / or green filters. In some cases, some photodiodes may be configured to measure infrared (IR) light. In some specific implementations, the photodiodes that measure IR light may not be covered by any filter, thus allowing the IR photodiodes to measure both visible light (e.g., color) and IR light. In some examples, the IR photodiodes may be covered by an IR filter, thus allowing IR light to pass through and blocking light from other parts of the spectrum (e.g., visible light, color). Some image sensors (e.g., the image sensor 130) may lack filters altogether (e.g., color, IR, or any other part of the spectrum) and may instead use different photodiodes (vertically stacked in some cases) throughout the pixel array. The different photodiodes throughout the pixel array may have different spectral sensitivity curves, thereby responding to light of different wavelengths. Monochromatic image sensors may also lack filters and thus lack color depth.
[0040] In some cases, the image sensor 130 may alternatively or additionally include opaque and / or reflective masks that block light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles. In some cases, the opaque and / or reflective masks may be used for PDAF. In some cases, the opaque and / or reflective masks may be used to block portions of the electromagnetic spectrum from reaching the photodiodes of the image sensor (e.g., IR cut-off filter, ultraviolet (UV) cut-off filter, band-pass filter, low-pass filter, high-pass filter, etc.). The image sensor 130 may also include an analog gain amplifier for amplifying the analog signal output by the photodiodes and / or an analog-to-digital converter (ADC) for converting the analog signal output by the photodiodes (and / or the analog signal amplified by the analog gain amplifier) into a digital signal. In some cases, certain components or functions discussed with respect to one or more of the control mechanisms 120 may alternatively or additionally be included in the image sensor 130. The image sensor 130 may be a charge-coupled device (CCD) sensor, an electron-multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), complementary metal-oxide semiconductor (CMOS), N-type metal-oxide semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.
[0041] The image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 154), one or more host processors (including host processor 152), and / or with respect to Figure 10One or more processors of any other type of processor 1010 discussed in the computing system 1000. The host processor 152 can be a digital signal processor (DSP) and / or other types of processors. In some specific embodiments, the image processor 150 is a single integrated circuit or chip (e.g., referred to as a system-on-chip or SoC) that includes the host processor 152 and the ISP 154. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 156), a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a broadband modem (e.g., 3G, 4G or LTE, 5G, etc.), memory, connection components (e.g., BluetoothTM, Global Positioning System (GPS), etc.), any combination thereof and / or other components. The I / O port 156 can include any suitable input / output port or interface according to one or more protocols or specifications, such as an inter-integrated circuit 1 (I2C) interface, an inter-integrated circuit 3 (I3C) interface, a serial peripheral interface (SPI) interface, a serial general-purpose input / output (GPIO) interface, a Mobile Industry Processor Interface (MIPI) (such as a MIPI CSI-2 physical (PHY) layer port or interface, an Advanced High-Performance Bus (AHB) bus, any combination thereof and / or other input / output ports. In an illustrative example, the host processor 152 can communicate with the image sensor 130 using the I2C port, and the ISP 154 can communicate with the image sensor 130 using the MIPI port.
[0042] The image processor 150 can perform multiple tasks, such as demosaicking, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form an HDR image, image recognition, object recognition, feature recognition, receiving inputs, managing outputs, managing memory, or some combination thereof. The image processor 150 can store the image frames and / or the processed images in the random access memory (RAM) 140 / 1025, read-only memory (ROM) 145 / 1020, cache, memory cells, another storage device, or some combination thereof.
[0043] Various input / output (I / O) devices 160 may be connected to the image processor 150. The I / O devices 160 may include a display screen, a keyboard, a keypad, a touch screen, a touchpad, a touch-sensitive surface, a printer, any other output device, any other input device, or any combination thereof. In some cases, captions may be input into the image processing device 105B via the physical keyboard or keypad of the I / O device 160, or via the virtual keyboard or keypad of the touch screen of the I / O device 160. The I / O 160 may include one or more ports, jacks, or other connectors that implement a wired connection between the image capture and processing system 100 and one or more peripheral devices, through which the image capture and processing system 100 may receive data from and / or send data to one or more peripheral devices. The I / O 160 may include one or more wireless transceivers that implement a wireless connection between the image capture and processing system 100 and one or more peripheral devices, through which the image capture and processing system 100 may receive data from and / or send data to one or more peripheral devices. The peripheral devices may include any type of I / O device 160 discussed previously, and once they are coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector, they may themselves be considered I / O devices 160.
[0044] In some cases, the image capture and processing system 100 may be a single device. In some cases, the image capture and processing system 100 may be two or more separate devices, including an image capture device 105A (e.g., a camera) and an image processing device 105B (e.g., a computing device coupled to the camera). In some specific implementations, the image capture device 105A and the image processing device 105B may be coupled together, for example, via one or more wires, cables, or other electrical connectors, and / or wirelessly via one or more wireless transceivers. In some specific implementations, the image capture device 105A and the image processing device 105B may be disconnected from each other.
[0045] As Figure 1 shown, the vertical dashed line divides Figure 1 the image capture and processing system 100 into two parts, representing the image capture device 105A and the image processing device 105B, respectively. The image capture device 105A includes a lens 115, a control mechanism 120, and an image sensor 130. The image processing device 105B includes an image processor 150 (including an ISP 154 and a host processor 152), a RAM 140, a ROM 145, and an I / O 160. In some cases, certain components illustrated in the image capture device 105A (such as the ISP 154 and / or the host processor 152) may be included in the image capture device 105A.
[0046] The image capture and processing system 100 may include an electronic device, such as a mobile or fixed telephone handset (e.g., a smart phone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the image capture and processing system 100 may include one or more wireless transceivers for wireless communication (such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof). In some specific implementations, the image capture device 105A and the image processing device 105B may be different devices. For example, the image capture device 105A may include a camera device, and the image processing device 105B may include a computing device, such as a mobile phone, a desktop computer, or other computing device.
[0047] Although the image capture and processing system 100 is shown as including certain components, those of ordinary skill in the art should understand that the image capture and processing system 100 may include more components than Figure 1 those shown therein. The components of the image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some specific implementations, the components of the image capture and processing system 100 may include electronic circuits or other electronic hardware, and / or may be implemented using electronic circuits or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a GPU, a DSP, a CPU, and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof, and / or may be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of the electronic device implementing the image capture and processing system 100.
[0048] In some examples, Figure 3 the computer vision system (described below) may include the image capture and processing system 100, the image capture device 105A, the image processing device 105B, or a combination thereof.
[0049] In various aspects, systems, methods, and computer-readable media are provided for a neural network architecture that uses deconvolution to preserve spatial information. In various aspects, the neural network architecture includes one or more deconvolution operations performed in a deconvolution layer. The deconvolution operations in the deconvolution layer can be used to upsample or magnify a higher-level feature map that may have a lower resolution than a lower-level feature map. Then, the higher-resolution feature map generated by the deconvolution layer can be combined with the feature map generated by a convolution layer. Then, the combined feature map can be used to determine an output prediction.
[0050] A convolutional neural network that uses deconvolution to preserve spatial information can have better performance than a convolutional neural network that does not include deconvolution. For example, better performance can be measured according to the accuracy with which the network identifies the content of image data. For example, when the object to be identified is small compared to the size of the input image, a convolutional neural network with deconvolution can have better accuracy.
[0051] In particular, adding deconvolution to small neural networks can improve the accuracy of these neural networks. Small neural networks have fewer layers and / or perform depthwise convolutions, and thus require fewer computations. Deep neural networks can perform a very large number of computations and can generate a large amount of intermediate data. Therefore, deep neural networks have been run using systems with a large amount of processing power and storage space. However, due to factors such as device size, available battery power, and the need for device lightweighting, mobile devices (such as smartphones, tablet computers, laptops, and other devices designed to be easily transportable) may have less powerful processors and may have less storage space. Smaller neural networks can be used for resource-constrained applications. Smaller neural networks may not be as accurate as deep neural networks, and thus techniques such as those discussed herein can be applied to improve the prediction accuracy of small neural networks.
[0052] Artificial neural networks attempt to use computer technology to replicate the logical reasoning performed by the biological neural networks that make up an animal's brain. Neural networks belong to a subfield of artificial intelligence called machine learning. Machine learning is a field of study that studies the ability to endow a computer with the ability to learn without being explicitly programmed. Explicitly programmed software programs must consider all possible inputs, scenarios, and outcomes. In contrast, software programs that use machine learning algorithms learn by being provided with inputs and receiving feedback on the correctness of the outputs produced by the program. The feedback is incorporated into the program so that the program can produce better results for the same or similar inputs.
[0053] Neural networks are inspired by the operating mechanisms of the human brain to the extent of understanding these operations. According to various models of the brain, the main computational elements of the brain are neurons. Neurons are connected to multiple elements, where the elements entering the neuron are called dendrites, and the elements leaving the neuron are called axons. Neurons receive signals via dendrites, perform calculations on these signals, and output signals on the axons. The input and output signals are called activations. The axon of one neuron can branch out and connect to the dendrites of multiple neurons. The connection between the branches of the axon and the dendrites is called a synapse.
[0054] Synapses can scale the signals passing through the synapse. The scaling factor is called the weight and is considered the way the brain can learn, where different weights are generated by different responses to the input. Learning can change the weights, but the organization of neurons and synapses does not need to change to achieve this learning. The static structure of the neural network is used as a model of the program, and the weights can reflect one or more tasks that the program has learned to perform.
[0055] The operating concept of a neural network is that the calculation of a neuron involves a weighted sum of input values. These weighted sums correspond to the value scaling performed by the synapses and the combination of those values in the neuron. A function operation is performed on the combined inputs in the neuron. In the brain model, the operation seems to be a non - linear function that causes the neuron to generate an output only when the input crosses a certain threshold. In some aspects, by analogy, the nodes of a neural network can apply a non - linear function to the weighted sum of the values input to the node.
[0056] Figure 2 FIG. 200 is a diagram illustrating a model 200 of a convolutional neural network. Model 200 illustrates the operations that can be included in a convolutional neural network: convolution, activation, pooling or subsampling, batch normalization, and output generation (e.g., fully - connected layer). Any given convolutional network includes at least one convolutional layer and can have dozens of convolutional layers. Additionally, a pooling layer does not need to follow each convolutional layer. In some examples, a pooling layer can occur after multiple convolutional layers or may not occur at all. Figure 2 The illustrated example convolutional network classifies the input image 220 into one of the following three categories: book, coffee, or laptop computer. In the illustrated example, when receiving the image as input, the example neural network outputs the highest probability (0.94) for "laptop computer" among the output predictions 214.
[0057] To produce the illustrated output prediction 214, an example convolutional neural network performs a first convolution 202 using a rectified linear unit (ReLU), pooling 204, a second convolution 206 using ReLU, additional pooling 208, and then uses two fully connected layers for classification. In the first convolution 202 step using ReLU, the input image 220 is convolved to produce one or more output feature maps 222. The first pooling 204 operation produces additional feature maps that serve as the input feature maps 224 for the second convolution and ReLU 206 operations. The second convolution and ReLU 206 operations produce a second set of output feature maps 226. The second pooling 208 step also produces a feature map 228 that is input into the first fully connected 210 layer. The output of the first fully connected 210 layer is input into the second fully connected 212 layer. The output of the second fully connected 212 layer is the output prediction 214. In a convolutional neural network, the terms "higher level" and "higher layers" refer to layers that are further from the input image (e.g., in the example model 200, the second fully connected 212 layer is the highest layer).
[0058] Other aspects may include additional or fewer convolution operations, ReLU operations, pooling operations, and / or fully connected layers. Convolution, non-linearity (e.g., ReLU), pooling or subsampling, and classification operations will be explained in more detail below.
[0059] When performing image recognition, a convolutional neural network operates on a numerical or digital representation of an image. An image can be represented in a computer as a matrix of pixel values. For example, a video frame captured at 1080p includes a pixel array that is 1920 pixels wide and 1080 pixels high. Certain components of an image can be referred to as channels. For example, a color image has three channels: red, green, and blue. In this example, a color image can be represented as three two-dimensional matrices, one two-dimensional matrix for each color, where the horizontal and vertical axes indicate the position of the pixels in the image, and values between 0 and 255 indicate the color intensity of the pixel. As another example, a grayscale image has only one channel and can thus be represented as a single two-dimensional matrix of pixel values. In this example, the pixel values can also be between 0 and 255, where for example 0 indicates black and 255 indicates white. In these examples, an upper limit value of 255 is assumed for pixels represented by 8-bit values. In other examples, more bits (e.g., 16, 32, or more bits) can be used to represent pixels, and the pixels can thus have a higher upper limit value.
[0060] Convolution is a mathematical operation that can be used to extract features from an input image. Features that can be extracted include, for example, edges, curves, corners, blobs, and ridges. Convolution learns image features by using small squares of the input data, thus preserving the spatial relationships between pixels.
[0061] In some aspects, MobileNet is an example of a convolutional neural network optimized for mobile vision applications and other environments with low computational power, such as embedded devices. In this regard, convolutional units from MobileNet and some additional convolutional units are used to generate feature maps. Then, the highest-level (and smallest) feature map is upsampled by a deconvolution layer and combined with a feature map of the same size using a lateral concatenation operation. Feature maps of several combinations with different resolutions can also be generated. The combined feature map and the highest-level feature map can be used for object detection.
[0062] Figure 3 FIG. 4 is a conceptual diagram illustrating the architecture of a computer vision system 300 according to various aspects of the present disclosure. In some aspects, an input image 302 is provided to a multi-task detection system 304 of the computer vision system 300 for detecting various objects within the input image 302. The multi-task detection system 304 can operate as a feature extraction backbone for extracting features from one or more images including the input image 302. The extracted features can be output from the multi-task detection system 304 to a Feature Pyramid Network (FPN) 308. In one illustrative aspect, the multi-task detection system 304 includes a neural network, such as a convolutional neural network, that includes one or more neural networks configured to detect objects, perform depth estimation, and / or perform other tasks. Non-limiting examples of objects include faces, body parts of a person, different mammals (dogs, cats, etc.), moving objects (e.g., cars, bicycles, etc.), and stationary objects (coffee, food, etc.).
[0063] As described above, according to some aspects, an input image (including the input image 302) is processed by the multi-task detection system 304 to extract features from the input image. The extracted features are provided to the FPN 308. The FPN 308 is configured to use different features from the multi-task detection system 304 and use multiple detection engines to detect objects and / or predict depth from the image. In one aspect, the FPN 308 can be configured to detect an object, such as a person. In this case, the FPN 308 can include a person detection engine 310 for detecting the body of a person, a face detection engine 320 for the face of the person, a hand detection engine 330 for detecting the hand of the person, another detection engine 340 for detecting other objects related to the image, and a depth estimation engine 350 for estimating the depth of the objects in the image.
[0064] In some aspects, the person detection engine 310 is configured to detect the body of a person. If a body is detected, the person detection engine 310 provides a bounding region 312 and key points 314 that identify the region of the image corresponding to the body and the landmark features corresponding to the body.
[0065] In one aspect, the face detection engine 320 is configured to detect a person's face. If a face is detected, the face detection engine 320 is configured to provide a bounding region 322 and key points 324, where the bounding region identifies the region of the image corresponding to the person's face, and the key points identify landmark features corresponding to the face. For example, the key points 324 can identify features such as the nose, eyes, and mouth.
[0066] The hand detection engine 330 is configured to detect a person's hand. If one or more hands are detected, the hand detection engine 330 provides at least one bounding region 332 corresponding to the detected hand. In some cases, the hand detection engine 330 can also be key points that identify landmark features of the hand.
[0067] The detection engine 340 is configured to detect other objects related to the image. For example, the detection engine 340 can detect an object that is close to or held by a person, such as a cup of coffee. For example, the computer vision system 300 can be implemented in a security system, and the detection engine 340 can be configured to detect a person holding a badge with authentication information (such as a picture of a person). The security system can authenticate a person based on the identity of another object, based on the picture on their badge. In this aspect, when the detection engine 340 detects an object, the detection engine 340 is configured to provide at least one bounding box 342 related to the object detected in the image.
[0068] The depth estimation engine 350 is configured to estimate the depth of features within the image and provide depth information 352. In one illustrative example, the depth information is a bitmap that identifies the depth of various features within the image.
[0069] An example of a neural network that can be used as a feature extraction backbone (e.g., for extracting features from input images, including input image 302) by the multi-task detection system 304, various detection engines 310, 320, 330, 340, and / or depth estimation engine 350 is the Mobilenet neural network. In some aspects, the Mobilenet neural network uses squeeze-and-excitation blocks to adaptively recalibrate per-channel feature responses by modeling the interdependencies between channels, which can improve the speed of inference. In some cases, standard convolutions can be used in the Mobilenet neural network instead of depthwise separable convolutional layers. For example, depthwise separable convolutions can provide high efficiency on the CPU, but they may be less accurate and may not have the same speed advantage on processing units with highly parallel architectures (such as GPUs, DSPs, NPUs, etc., which are commonly used to implement neural networks). In some cases, the number of convolutional channels can be adjusted to a multiple of 32 to maximize the use of the hardware memory layout. In some cases, Mobilenet includes squeeze-and-excitation blocks. The squeeze-and-excitation block takes an input such as a three-dimensional (3D) tensor vector and squeezes the tensor into a one-dimensional (1D) vector, and the 1D vector is restored to its original value. In one aspect, the squeeze-and-excitation block may not be supported in hardware, or a library including software support for the squeeze-and-excitation block may not be available to developers. In some aspects, the multi-task detection system 304, various detection engines 310, 320, 330, 340, and / or depth estimation engine 350 may include a joint attention module (JAM) (e.g., instead of the squeeze-and-excitation block), which is configured to perform functions similar to those of the squeeze-and-excitation block and is discussed in Figure 6 below.
[0070] In an illustrative aspect, the JAM maintains the weight and height dimensions and instead uses 1x1 convolutions to reduce the number of channels. For example, using 1x1 convolutions, the JAM can perform per-channel attention similar to the cross-attention between different channels in the squeeze-and-excitation block. Another convolution can be performed to return the information to the original dimensions. In another aspect, the JAM can implement clipped ReLU instead of the sigmoid operation. The sigmoid operation is a non-linear operation, and the ReLU clipping is a linear operation that clips the input values between 0 and 1 to approximate the normalization function performed by the sigmoid operation. However, clipped ReLU is much faster than the sigmoid operation. In some aspects, the functions performed by the JAM can be executed in hardware, which can improve its efficiency relative to software computations.
[0071] The FPN 308 includes an object association engine 360 that is configured to associate various detected objects within an image. For example, the object association engine 360 is configured to associate different detected body parts of a person (e.g., hands, feet, etc.). In some aspects, the object association engine 360 is configured to construct containers (e.g., bounding boxes) that enclose regions corresponding to the object and its associated objects. For example, in the case of detecting a person, the object association engine 360 can identify various objects (e.g., hands, face, body, etc.) and determine the region that includes the person. In some aspects, these containers can overlap. For example, a person in the foreground may stand in front of a person in the background, and the foreground person and the background person at least partially overlap.
[0072] Figure 4 FIG. illustrates a block diagram of a computer vision system 400 for performing object part association in accordance with various aspects of the present disclosure. In some aspects, the computer vision system 400 is configured to identify an object and a part of the object. For example, the computer vision system 400 can be configured to identify a person and at least one body part of the object.
[0073] In one aspect, the computer vision system 400 includes a neural network 410 that is configured to perform various computer vision tasks, such as generating various images at different layers of the neural network 410. For example, an image 402 is provided into the neural network 410, and the neural network 410 performs various operations, such as downsampling the image 402 to remove features from the image 402. As Figure 4 illustrated, the neural network 410 generates a feature map represented by C and provides the feature map to the FPN 420. In some aspects, the number after C (e.g., 4, 8, etc.) indicates the amount of downsampling provided by the corresponding layer of the neural network 410. In this case, the C4 feature map corresponds to an image that is reduced by a factor of four, and the C4 feature map has 1 / 4 of the original image size. Convolution reduces the number of features available in the C4 feature map. Similarly, the C8 feature map corresponds to an image that is reduced by a factor of 8 (e.g., 1 / 8 of the original size), the C16 feature map corresponds to an image that is reduced by a factor of 16 (e.g., 1 / 16 of the original size), and the P32 feature map corresponds to an image that is reduced by a factor of 32 (e.g., 1 / 32 of the original size). In this case, the C8 feature map follows the C4 feature map in the neural network 410, the C16 feature map follows the C8 feature map in the neural network 410, and the P32 feature map follows the C16 feature map. In some aspects, upsampling can be achieved through a deconvolution or bilinear sampling process.
[0074] In some aspects, the C4, C8, C16, and P32 feature maps are provided to the FPN 420. The FPN 420 fuses features of different scales from different output layers. For example, the P32 feature map is upsampled and merged with the C16 feature map to produce the P16 feature map. The P16 feature map is upsampled and merged with the C8 feature map, and the P8 feature map is upsampled and merged with the C4 feature map. In some aspects, the number of features being searched corresponds to the number of feature maps extracted from the neural network 410. For example, the computer vision system 400 may be configured to search for people, faces, hands, and feet, and will have four feature maps from the neural network 410.
[0075] The FPN 308 provides the P4 feature map to the single-stage headless (SSH) 430 component. The SSH 430 includes a number of convolutions, batch normalizations, and activations to enhance the visual representation of various features within the P4 feature map. In one aspect, the convolutions, batch normalizations, and activations correspond to ReLU blocks, and additional parameters for the SSH 430 are provided that can be learned to improve the visual representation of the features. In an illustrative example, the SSH 430 can perform the functions of 5x5 and 7x7 convolutions by using multiple serial 3x3 convolutions. In this aspect, the SSH 430 produces a feature map from the image 402, which is provided to the classifier 440, the bounding box identifier 442, the keypoint detector 444, and the depth estimator 446.
[0076] In some aspects, the classifier 440 is configured to classify objects within the image, such as parts of a human body (e.g., a person, a face, a hand, etc.). For example, the classifier 440 may include at least a portion of the person detection engine 310, the face detection engine 320, the hand detection engine 330, and the detection engine 340. The classifier 440 is configured to produce a heatmap that identifies points associated with the object.
[0077] In some aspects, the bounding box identifier 442 is configured to identify bounding regions, such as bounding boxes that enclose the regions corresponding to each classified object in the image. For example, the bounding box identifier 442 may be configured to regress four bounding box coordinates (e.g., top, bottom, left, right) to identify the corresponding regions. In this case, if the heatmap provided by the classifier 440 includes one pixel with a response, that pixel indicates the detection of a person or an object, and the pixel can be sampled at the location of the corresponding bounding box.
[0078] In some aspects, the keypoint detector 444 is configured to identify relevant keypoints associated with at least some of the classified objects in the image, and the depth estimator 446 is configured to determine the depth of each classified object in the image. In such cases, the keypoint detector 444 may have multiple regression targets based on the features to be identified (such as fingers, hands, wrists, faces, feet, etc.). The keypoints from the keypoint detector 444 can be sampled from the class center (e.g., from the classifier 440) to identify features (such as the mouth, eyes, lips) or body joint positions (such as the wrists, shoulders, etc.).
[0079] The object association engine 450 is configured to identify the different objects detected by the classifier 440 and associate them into a complete object. For example, in the case where the object is a person, the object association engine 450 is configured to associate body parts (such as the face, hands, body) with the person, and the object association engine 450 is configured to generate a container corresponding to the person. In some aspects, the object association engine 450 is configured to generate multiple overlapping containers, such as a person in the foreground and a person in the background that at least partially overlap.
[0080] In one aspect, the object association engine 450 is configured to associate parts of an object with the object. Illustrative examples of a person are further described, but the object association engine 450 can be configured to associate other parts of other objects (such as animals or moving objects). In one aspect, the object association engine 450 uses the keypoints of a person's body and, based on the distances from these keypoints to other detected objects (such as bounding boxes and keypoints), associates other parts of the person. In the case of a person's face, the object association engine 450 uses the keypoints associated with the head (e.g., the center point on the head or other points on the head) to determine the distance from the head to the body (e.g., the center point on the body or other points on the body). In the case of a hand, the object association engine 450 can determine the center point of the hand (e.g., the average of the keypoints identifying the hand boundary) to associate the hand with the body using the body keypoints, which include the positions of both hands. In some aspects, the hand bounding box including the center point of the hand and the body keypoints including two hand keypoints can be separately output from the neural network. For example, the body keypoints can be regressed or sampled from the body center, the association between the keypoints, the body center point, and the body boundary can be known, and the object association engine 450 can be configured to determine the association between the hand and the body based on the center of the hand and the body keypoints.
[0081] After constructing the distances, two cost functions illustrated in Equation 1 are constructed for the face and the hand, and the cost function identifies the cost of associating different features.
[0082]
[0083] Where B and P represent the bounding box and body key points respectively. X(*) and Y(*) are the coordinates of the center point of the boundary (referred to as the box center) and the key points. The superscript c represents the body part category (e.g., face, hand, etc.). The subscripts i, j represent the pairs of the bounding box and key points to be associated.
[0084] The object association engine 450 can use a solution such as the Hungarian algorithm to identify the one-to-one alignment of the person's face bounding box and the corresponding head key points, as well as the hand bounding box and the corresponding wrist key points. After the object association engine 450 identifies the correspondence between the face bounding box, hand bounding box, and body points, it also finds the relationship between the person bounding box and their face and hands. The object association engine 450 can also construct a container for tracking the person, so that if one or more body parts are not detected due to occlusion or other challenging visual conditions, the object association engine 450 will still not lose track of the person.
[0085] In some aspects, two-dimensional association is sufficient. However, in the case of occlusion, such as when the foreground person is partially in front of the background person, the bounding box may be occluded. In this regard, the object association engine 450 can determine that object identification and association may fail. For example, the object association engine 450 can determine that the first object may be incorrectly assigned to the second object because the key points are close to each other.
[0086] In some aspects, 3D information (such as depth) can be used in the cost function. Equation 2 below identifies the cost functions generated for the hand and face:
[0087]
[0088] Where B and P represent the bounding box and body key points respectively, X(*) and Y(*) are the coordinates of the box center and the key points, Z(*) indicates the average depth of the surrounding area of the bounding box or key points, the superscript c represents the body part category (e.g., face or hand), the subscripts i, j represent the pairs of the bounding box and key points to be associated, and λ is a hyperparameter of the depth cost weight. For example, if the classifier 440 outputs m people and n hand objects, Equation 2 will generate an m×n matrix of cost values input to the Hungarian algorithm.
[0089] As Figure 3 shown, in a typical in-cabin occupancy monitoring scenario where two people mostly overlap, the two-dimensional (2D) association method depicted in the previous section may fail because the body key points of one person are closer in 2D to the body part bounding box of another person. Such ambiguities occur more frequently under extreme visual conditions because key point detection tends to be less accurate. In such cases, if the depth information of the bounding box and key points is known, depth can simply be added to the distance cost calculation, which will penalize the association between the front person container and the rear person container.
[0090] Figure 5A Illustrates example results of bounding boxes identified from an image according to aspects of the present disclosure. In this aspect, 2D detection is used to identify four people. In this case, the first person 610 overlaps with the second person 620, where the second person 620 is behind the first person. In this case, the hand 630 of the second person 620 is incorrectly identified as the hand of the first person 610.
[0091] Figure 5B Illustrates another example result of bounding boxes identified from an image according to aspects of the present disclosure. In this aspect, the same image from Figure 5A is used, and 3D detection is used to identify four people. As described above, the first person 610 overlaps with the second person 620, where the second person 620 is behind the first person. However, the hand 630 of the second person 620 is correctly identified as the hand of the second person 620 because the depth of the hand 630 is correctly associated with the second person 620.
[0092] Figure 6 Illustrates an example JAM 600 that can be used by Figure 3 the multi-task detection system 304, various detection engines 310, 320, 330, 340, and / or the depth estimation engine 350. In some aspects, the squeeze-and-excitation block adaptively recalibrates the per-channel feature responses by modeling the interdependencies between channels to improve the detection of features. In some aspects, the squeeze-and-excitation block may not be handled by hardware features, and available libraries may not support the squeeze-and-excitation block.
[0093] In an illustrative aspect, JAM 600 maintains the weight and height dimensions and instead uses 1x1 convolutions to reduce the number of channels. Different from the squeeze-and-excitation block, JAM 600 does not squeeze the spatial locations of the input feature map. By using 1x1 convolutions, JAM performs per-channel attention similar to the cross-attention between different channels in the squeeze-and-excitation block. Another convolution is performed to return the information to the original dimensions. In another aspect, JAM can implement clipped ReLU to replace the sigmoid operation. The sigmoid operation is a non-linear operation, and the ReLU clipping is a linear operation that clips the input values between 0 and 1 to approximate the normalization function performed by the sigmoid operation. However, the clipped ReLU is much faster than the sigmoid operation. In some aspects, the functions performed by JAM can be executed in hardware, which has higher efficiency compared to software calculations.
[0094] As Figure 6As shown, JAM 600 receives a feature map 602 as input, which may include multiple features in a tensor or other representation. The feature map 602 can be of any shape and is shown as having dimensions C*H*W (corresponding to multiple channels (C), each channel having a height (H) and a width (W)). JAM 600 first applies a 1x1 convolution to generate a tensor (or other representation) with a reduced number of channels compared to the input feature map 602 (e.g., reducing the channels from C to C', where C' < C). Then, JAM 600 can apply another 1x1 convolution to reshape the tensor (or other representation) into a feature map 606 with the original dimensions C*H*W. The output feature map 606 from the second 1x1 convolutional layer can be passed into an activation function (e.g., a ReLU layer) for activation, and then using a clipping operation (e.g., Figure 7B the clipped ReLU operation) as per-pixel attention, the output of this activation function is normalized or clipped to [0,1] at the feature level (or pixel level). Then, the normalized / clipped feature map values can be combined with the feature values of the original input feature map 602 (e.g., using a combination engine). For example, the normalized / clipped feature map values can be multiplied by the feature values of the original input feature map 602 to obtain a final output feature map 608 with the same shape C*H*W. Thus, the pipeline of JAM 600 can effectively and efficiently perform an attention mechanism across different feature maps.
[0095] Figure 7A is a graph 705 illustrating an example of the sigmoid operation performed in a squeeze-and-excitation block. As Figure 7A illustrated, the sigmoid is a non-linear operation that normalizes a layer. When the sigmoid operation is performed multiple times, such as when performing object detection on many objects, this operation becomes expensive.
[0096] Figure 7B is a graph 710 illustrating an example of the clipped ReLU operation that can be performed by JAM 600 according to aspects of the present disclosure. As Figure 7B illustrated, the clipped ReLU is a boolean operation that approximates a function of the sigmoid operation and can be implemented in existing hardware without any hardware changes. In some cases, the hardware can be modified to allow the sigmoid operation or the squeeze-and-excitation operation, but adding additional dedicated functional blocks to an integrated circuit (IC) is time-consuming and expensive.
[0097] Figure 8 is a flowchart illustrating an example of a method 800 for processing image data according to certain aspects of the present disclosure. The method 800 can be performed by a computing device (or a component of a computing device) (such as including Figure 3performed by the computing device of system 300). In one illustrative example, the computing device may be a computing system 1000 configured to perform all or part of method 800. In some cases, computing system 1000 may include Figure 3 system 300.
[0098] At block 802, the computing system (e.g., system 300, computing system 1000, etc.) is configured to obtain an image that includes at least a first object. In some aspects, the image may include multiple objects, such as faces, a group of people, mammals, or other moving objects, such as machines or vehicles.
[0099] At block 804, the computing system is configured to generate a feature map based on providing the image to a neural network (e.g., using Figure 3 multi-task detection system 304). At block 806, the computing system (e.g., using Figure 3 FPN 308). For example, the multiple objects may include a first part of the first object. In one aspect, the first part of the first object includes the face of a first person, and the first set of object parts includes at least one body part of the first person.
[0100] In some aspects, the computing system may be configured to generate a first container (e.g., a bounding box or other bounding region) corresponding to the first object based on the first set of object parts. For example, the container may include the face and the at least one body part (e.g., the limbs of the person). In such aspects, the computing system may be configured to generate a corresponding bounding box for each of the multiple objects. In one illustrative aspect, the computing system may generate at least one cost function based on key points and the bounding box, and may map the bounding box and the key points to corresponding parts of the first object. According to this aspect, the first container is generated based on at least one mapped part of the first object. After generating the bounding box, the computing system may determine key points of at least a portion of the first set of object parts and map the key points to the corresponding bounding box.
[0101] In some aspects, when the computing system generates the first container, the computing system may be configured to determine an identification of at least one part of the first set of object parts based on a second object in the image. According to this aspect, the at least one part of the first set of object parts may not be identifiable based on the second object in the image. For example, the second object may occlude the first object, and parts of the first object may appear closer to the second object. In some aspects, to identify the plurality of objects, the computing system may generate corresponding bounding boxes for each of the plurality of objects, determine key points of at least a portion of the first set of object parts, and map the key points to corresponding bounding boxes of other object parts in the same set. The computing system may then determine a corresponding depth associated with each of the plurality of objects, generate at least one cost function based on the key points, the bounding boxes, and the corresponding depth associated with each of the plurality of objects, and map the bounding boxes and the key points to corresponding parts of the first object and the second object based on the at least one cost function. In this case, the first container is generated based on at least one mapped part of the first object.
[0102] In some aspects, the neural network includes at least one joint attention module (e.g., Figure 6 the JAM 600). The joint attention module is configured to generate the feature map and includes at least one 1x1 convolution operation, at least one convolution operation, a cropping function, and a combination engine. The at least one 1x1 convolution operation is configured to obtain an input feature map and generate a first intermediate feature map that includes a reduced number of channels compared to the input feature map. The at least one convolution operation is configured to generate a second intermediate feature map that has the same dimensions as the input feature map, and the cropping function is configured to crop the eigenvalues of the second intermediate feature map. The combination engine is configured to combine the eigenvalues of the input feature map with the cropped eigenvalues to generate the feature map.
[0103] At block 808, the computing system is configured to identify a first set of object parts within the plurality of objects that correspond to the first object. In some aspects, as described above, the first set of objects may be partially occluded by other objects.
[0104] Figure 9 is a flowchart illustrating an example of a method 900 for processing image data according to certain aspects of the present disclosure. Method 900 may be performed by a computing device (or a component of the computing device) having a touch screen, such as a mobile wireless communication device, a smart speaker, a camera, an XR device, a wireless-enabled vehicle, or another computing device. In one illustrative example, computing system 1000 may be configured to perform all or part of method 800.
[0105] At block 902, the computing system (e.g., computing system 1000) is configured to obtain an image that includes at least a first object and a second object. In some aspects, the first object may at least partially occlude the second object, and vice versa. For example, the first object may be in the foreground, and the second object may be in the background.
[0106] At block 904, the computing system is configured to generate a feature map based on providing the image to a neural network. In some aspects, the neural network identifies various features, such as objects of interest within the image. The object may be part of other objects. For example, the object may be a face (as part of a person). The object may also be other types of objects, such as dynamic objects (e.g., vehicles, machines, etc.), other mammals, vegetation, etc.
[0107] At block 906, the computing system is configured to identify a plurality of objects based on the feature map. In some aspects, the plurality of objects includes a first part of the first object and a second part of the second object. The first part of the first object and the second part of the second object at least partially overlap.
[0108] In some aspects, the computing system is configured to determine a respective depth of each of the plurality of objects associated with the feature map.
[0109] At block 908, the computing system is configured to identify a first container corresponding to the first object and a second container corresponding to the second object. In some aspects, the first part of the first object and the second part of the second object are assigned to the first container based on the depth of the first part of the first object and the depth of the second part of the second object. For example, the first object includes a first person, the second object includes a second person, the first part includes a body part or face of the first person, and the second part includes a body part or face of the second person.
[0110] In one aspect, at block 908, to identify the containers, the computing system may at least partially identify the part of the first person and the part of the first person from the plurality of objects based on the respective depth of each object, and then determine a first container for tracking the first person in a subsequent image and a second container for tracking the second person in the subsequent image. In this case, the second part of the second person may overlap or occlude a part of the first person, and the depth may be used to disambiguate the containers of the first person and the second person. The first object is completely contained within the first container, and the second object is completely contained within the second container.
[0111] In some examples, the processes described herein (e.g., methods 800 and 900 and / or other processes described herein) may be performed by a computing device or apparatus. In one example, methods 800 and 900 may be performed by a computing device including a touch screen having the Figure 9 computing architecture of computing system 1000 shown.
[0112] The computing device may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a virtual reality (VR) headset, an augmented reality (AR) headset), AR glasses, a connected watch or smartwatch or other wearable device), a server computer, a computing device of an autonomous vehicle or an autonomous vehicle, a robotic device, a television, and / or any other computing device having the resource capabilities to perform the methods described herein (including methods 500 and 900). In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the methods described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive IP-based data or other types of data.
[0113] The components of the computing device may be implemented in circuitry. For example, the components may include and / or may be implemented using electronic circuits or other electronic hardware (which may include one or more programmable electronic circuits (e.g., a microprocessor, a GPU, a DSP, a CPU, and / or other suitable electronic circuits)), and / or may include and / or may be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein.
[0114] Methods 800 and 900 are illustrated as logic flowcharts, the operations of which represent a sequence of operations that may be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the method.
[0115] Method 800 and 900 and / or other methods or processes described herein may be executed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed jointly on one or more processors, implemented in hardware, or implemented by a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, e.g., in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0116] Figure 10 FIG. is an illustration of an example of a system for implementing certain aspects of the disclosed technology. Specifically, Figure 10 illustrates an example of a computing system 1000, which can be any computing device, such as, for example, any component that constitutes an internal computing system, a remote computing system, a camera, or any combination thereof, where the components of the system communicate with each other using connection 1005. Connection 1005 can be a physical connection using a bus or a direct connection into processor 1010, such as in a chipset architecture. Connection 1005 can also be a virtual connection, a networked connection, or a logical connection.
[0117] In some aspects, computing system 1000 is a distributed system where the functions described in this disclosure can be distributed within one data center, multiple data centers, a peer-to-peer network, etc. In some aspects, one or more of the described system components represent many such components each performing some or all of the functions the component is described for. In some aspects, the components can be physical or virtual devices.
[0118] Example computing system 1000 includes at least one processing unit (CPU or processor) 1010 and connection 1005 that couples various system components including system memory 1015 (such as ROM 1020 and RAM 1025) to processor 1010. Computing system 1000 may include a cache 1012 that is directly connected to, in close proximity to, or integrated as part of processor 1010 of high-speed memory.
[0119] Processor 1010 may include any general-purpose processor and hardware services or software services (such as services 1032, 1034, and 1036 stored in storage device 1030 and configured to control processor 1010), as well as a dedicated processor in which software instructions are incorporated into the actual processor design. Processor 1010 can be substantially a stand-alone computing system that includes multiple cores or processors, buses, memory controllers, caches, etc. The multi-core processor can be symmetric or asymmetric.
[0120] To enable user interaction, computing system 1000 includes an input device 1045 that can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. Computing system 1000 may also include an output device 1035 that can be one or more of a plurality of output mechanisms. In some instances, a multimodal system may enable a user to provide multiple types of input / output to communicate with computing system 1000. Computing system 1000 may include a communication interface 1040 that generally may govern and manage user input and system output. The communication interface may perform or facilitate receiving and / or sending wired or wireless communications using wired and / or wireless transceivers, including using audio jack / plug, microphone jack / plug, universal serial bus (USB) port / plug, port / plug, Ethernet port / plug, fiber optic port / plug, dedicated wired port / plug, wireless signaling, BLE wireless signaling, wireless signaling, RFID wireless signaling, near field communication (NFC) wireless signaling, dedicated short range communication (DSRC) wireless signaling, 802.11WiFi wireless signaling, WLAN signaling, visible light communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), IR communication wireless signaling, public switched telephone network (PSTN) signaling, integrated services digital network (ISDN) signaling, 3G / 4G / 5G / LTE cellular data network wireless signaling, ad hoc network signaling, radio wave signaling, microwave signaling, infrared signaling, visible light signaling, ultraviolet light signaling, wireless signaling along the electromagnetic spectrum, or some combination thereof. The communication interface 1040 may also include one or more global navigation satellite system (GNSS) receivers or transceivers for determining the location of computing system 1000 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States' GPS, Russia's Global Navigation Satellite System (GLONASS), China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There are no restrictions on operating on any particular hardware arrangement, and thus the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.
[0121] The storage device 1030 can be a non-volatile and / or non-transitory and / or computer-readable memory device, and can be a hard disk or other type of computer-readable medium capable of storing data accessible by a computer, such as cassette tapes, flash memory cards, solid state memory devices, digital versatile discs, cartridges, floppy disks, hard disks, magnetic tapes, magnetic stripes, any other magnetic storage medium, flash memory, memristor memory, any other solid state memory, compact disc read-only memory (CD-ROM) optical discs, rewritable compact discs (CD) optical discs, digital video discs (DVD) optical discs, Blu-ray discs (BDD) optical discs, holographic optical discs, another optical medium, secure digital (SD) cards, micro secure digital (microSD) cards, Memory cards, smart card chips, EMV chips, subscriber identity module (SIM) cards, mini / micro / nano / pico SIM cards, another integrated circuit (IC) chip / card, RAM, static RAM (SRAM), dynamic RAM (DRAM), ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASHEPROM), cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase change memory (PCM), spin transfer torque RAM (STT-RAM), another memory chip or cartridge and / or combinations thereof.
[0122] The storage device 1030 may include software services, servers, services, etc. When the code defining such software is executed by the processor 1010, the code causes the system to perform functions. In some aspects, the hardware services for performing specific functions may include software components for performing functions stored in a computer-readable medium connected to necessary hardware components such as the processor 1010, the connection 1005, the output device 1035, etc. The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. The computer-readable medium may include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagated wirelessly or over a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as CDs or DVDs, flash memory, memory, or memory devices. The computer-readable medium may have code and / or machine-executable instructions stored thereon, and the code and / or machine-executable instructions may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. The information, arguments, parameters, data, etc. may be passed, forwarded, or sent via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.
[0123] In some cases, a computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, one or more network interfaces configured to communicate and / or receive data, any combination thereof, and / or other components. One or more network interfaces may be configured to communicate and / or receive wired and / or wireless data, including data according to 3G, 4G, 5G, and / or other cellular standards, data according to the Wi-Fi (802.11x) standard, data according to the Bluetooth TM standard, data according to the IP standard, and / or other types of data.
[0124] Components of a computing device may be implemented in circuitry. For example, each component may include and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein.
[0125] In some aspects, computer-readable storage devices, media, and memories may include cables or wireless signals, etc., that contain bitstreams, etc. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as power consumption, carrier signals, electromagnetic waves, and signals themselves.
[0126] Specific details are provided in the above description to provide an exhaustive understanding of the aspects and examples provided herein. However, one of ordinary skill in the art will understand that these aspects may be practiced without these specific details. For clarity, in some cases, the technology may be presented as including separate functional blocks, including functional blocks that contain devices, device components, steps or routines in a method embodied in software or a combination of hardware and software. Additional components other than those shown in the figures and / or described herein may be used. For example, circuits, systems, networks, processes, and other components may be shown in block diagram form as components to avoid obscuring these aspects in unnecessary detail. In other cases, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the aspects.
[0127] The various aspects may be described above as a process or method, which is depicted as a flowchart, flow diagram, data flow diagram, structure diagram, or block diagram. Although a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. When the operations of a process are completed, the process ends, but the process may have additional steps not included in the figures. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, the termination of the process may correspond to the function returning to the calling function or the main function.
[0128] The processes and methods according to the above examples can be implemented using computer-executable instructions stored or otherwise obtained from a computer-readable medium. Such instructions can include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions. Portions of the computer resources used can be accessed via a network. The computer-executable instructions can be, for example, binary, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, the information used, and / or the information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, etc.
[0129] Devices implementing the processes and methods according to these disclosures can include hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof, and can take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., a computer program product) for performing the necessary tasks can be stored in a computer-readable or machine-readable medium. The processor can execute the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein can also be embodied in a peripheral device or an add-in card. By additional example, such functionality can also be implemented on a circuit board among different chips or different processes executed on a single device.
[0130] Instructions, the media for conveying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.
[0131] In the above description, aspects of the present application are described with reference to their particular aspects, but those skilled in the art will recognize that the present application is not limited thereto. Thus, although the illustrative aspects of the present application have been described in detail herein, it is to be understood that the various inventive concepts can be implemented and adopted in other various ways, and the appended claims are not to be construed as including such variations unless limited by the prior art. The various features and aspects of the above application can be used individually or jointly. Additionally, the aspects can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. Therefore, the specification and drawings should be regarded as illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that in alternative aspects, the methods can be performed in a different order than that described.
[0132] Those of ordinary skill in the art should understand that, without departing from the scope of this specification, the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced with the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively.
[0133] In cases where a component is described as “configured to” perform certain operations, such a configuration can be achieved, for example, by designing an electronic circuit or other hardware to perform the operations, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuit) to perform the operations, or any combination thereof.
[0134] The phrase “coupled to” means that any component is directly or indirectly physically connected to another component, and / or any component directly or indirectly communicates with another component (e.g., connected to another component through a wired or wireless connection and / or other suitable communication interfaces).
[0135] Claim language or other language that recites “at least one of” a set and / or “one or more” in a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language that recites “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, claim language that recites “at least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language “at least one of” a set and / or “one or more” in a set does not limit the set to the items listed in the set. For example, claim language that recites “at least one of A and B” or “at least one of A or B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.
[0136] The various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the aspects disclosed herein may be implemented as electronic hardware, computer software, firmware, or any combination thereof. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and the design constraints imposed on the overall system. A person skilled in the art may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this application.
[0137] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general-purpose computer, a wireless communication device such as a cellular phone, or an integrated circuit device with multiple uses, including applications in wireless communication devices such as cellular phones and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be at least partially implemented by a computer-readable data storage medium including program code, the program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging material. The computer-readable medium may include a memory or data storage medium, such as RAM (such as synchronous dynamic random access memory (SDRAM)), ROM, non-volatile random access memory (NVRAM), EEPROM, flash memory, magnetic or optical data storage media, and the like. Additionally or alternatively, the techniques may be at least partially implemented by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0138] The program code may be executed by a processor, which may include one or more processors, such as one or more DSPs, general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in an alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein.
[0139] Exemplary aspects of the present disclosure include:
[0140] Aspect 1. A method for processing image data, the method comprising: obtaining an image including at least a first object; generating a feature map based on providing the image to a neural network; identifying a plurality of objects based on the feature map, the plurality of objects including a first part of the first object; and identifying a first set of object parts corresponding to the first object within the plurality of objects.
[0141] Aspect 2. The method according to aspect 1, wherein the first part of the first object includes the face of a first person.
[0142] Aspect 3. The method according to aspect 2, wherein the first set of object parts includes at least one body part of the first person.
[0143] Aspect 4. The method according to any one of aspects 1 to 3, the method further comprising: generating a first container corresponding to the first object based on the first set of object parts.
[0144] Aspect 5. The method according to aspect 4, wherein identifying the plurality of objects based on the feature map includes: generating a corresponding bounding box for each of the plurality of objects; determining key points of at least a part of the first set of object parts; and mapping the key points to the corresponding bounding boxes.
[0145] Aspect 6. The method according to aspect 5, the method further comprising: generating at least one cost function based on the key points and the bounding boxes; and mapping the bounding boxes and the key points to corresponding parts of the first object, wherein the first container is generated based on at least one mapped part of the first object.
[0146] Aspect 7. The method according to any one of aspects 5 to 6, wherein identifying the first set of object parts corresponding to the first object within the plurality of objects includes: determining an identification of at least one part of the first set of object parts based on a second object in the image.
[0147] Aspect 8. The method according to aspect 7, wherein the at least one part of the first set of object parts cannot be identified based on the second object in the image.
[0148] Aspect 9. The method according to any one of aspects 7 to 8, wherein identifying the plurality of objects based on the feature map includes: generating a corresponding bounding box for each of the plurality of objects; determining key points of at least a part of the first set of object parts; and mapping the key points to the corresponding bounding boxes.
[0149] Aspect 10. The method according to aspect 9, the method further comprising: determining a respective depth associated with each of the plurality of objects; generating at least one cost function based on the key points, the bounding boxes, and the respective depth associated with each of the plurality of objects; and mapping the bounding boxes and the key points to corresponding parts of the first object and the second object based on the at least one cost function, wherein the first container is generated based on at least one mapped part of the first object.
[0150] Aspect 11. The method according to any one of aspects 1 to 10, wherein the neural network comprises at least one joint attention module.
[0151] Aspect 12. The method according to aspect 11, wherein the at least one joint attention module is configured to generate the feature map, the at least one joint attention module comprising: at least one 1x1 convolution operation configured to obtain an input feature map and generate a first intermediate feature map, the first intermediate feature map comprising a reduced number of channels compared to the input feature map; at least one convolution operation configured to generate a second intermediate feature map having the same dimension as the input feature map; and a cropping function configured to crop the feature values of the second intermediate feature map; and a combination engine configured to combine the feature values of the input feature map with the cropped feature values to generate the feature map.
[0152] Aspect 13. An apparatus for processing an image, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: obtain an image comprising at least a first object; generate a feature map based on providing the image to a neural network; identify a plurality of objects based on the feature map, the plurality of objects comprising a first part of the first object; and identify a first set of object parts within the plurality of objects corresponding to the first object.
[0153] Aspect 14. The apparatus according to aspect 13, wherein the first part of the first object comprises a face of a first person.
[0154] Aspect 15. The apparatus according to aspect 14, wherein the first set of object parts comprises at least one body part of the first person.
[0155] Aspect 16. The apparatus according to any one of aspects 13 to 15, wherein the at least one processor is configured to: generate a first container corresponding to the first object based on the first set of object parts.
[0156] Aspect 17. The apparatus according to aspect 16, wherein the at least one processor is configured to: generate a corresponding bounding box for each of the plurality of objects; determine key points of at least a part of the first set of object parts; and map the key points to the corresponding bounding boxes.
[0157] Aspect 18. The apparatus according to aspect 17, wherein the at least one processor is configured to: generate at least one cost function based on the key points and the bounding boxes; and map the bounding boxes and the key points to the corresponding parts of the first object, wherein the first container is generated based on at least one mapped part of the first object.
[0158] Aspect 19. The apparatus according to any one of aspects 17 to 18, wherein the at least one processor is configured to: determine an identity of at least one part of the first set of object parts based on a second object in the image.
[0159] Aspect 20. The apparatus according to aspect 19, wherein the at least one part of the first set of object parts cannot be identified based on the second object in the image.
[0160] Aspect 21. The apparatus according to any one of aspects 19 to 20, wherein the at least one processor is configured to: generate a corresponding bounding box for each of the plurality of objects; determine key points of at least a part of the first set of object parts; and map the key points to the corresponding bounding boxes.
[0161] Aspect 22. The apparatus according to aspect 21, wherein the at least one processor is configured to: determine a corresponding depth associated with each of the plurality of objects; generate at least one cost function based on the key points, the bounding boxes, and the corresponding depth associated with each of the plurality of objects; and map the bounding boxes and the key points to the corresponding parts of the first object and the second object based on the at least one cost function, wherein the first container is generated based on at least one mapped part of the first object.
[0162] Aspect 23. The apparatus according to any one of aspects 13 to 22, wherein the neural network includes at least one joint attention module.
[0163] Aspect 24. The apparatus according to aspect 23, wherein the at least one joint attention module is configured to generate the feature map, and the at least one joint attention module includes: at least one 1x1 convolution operation configured to obtain an input feature map and generate a first intermediate feature map, the first intermediate feature map including a reduced number of channels compared to the input feature map; at least one convolution operation configured to generate a second intermediate feature map having the same dimension as the input feature map; and a cropping function configured to crop the feature values of the second intermediate feature map; and a combination engine configured to combine the feature values of the input feature map with the cropped feature values to generate the feature map.
[0164] Aspect 25. A method for processing an image, the method comprising: obtaining an image including at least a first object and a second object; generating a feature map based on providing the image to a neural network; identifying a plurality of objects based on the feature map, the plurality of objects including a first part of the first object and a second part of the second object, wherein the first part of the first object and the second part of the second object at least partially overlap; and identifying a first container corresponding to the first object and a second container corresponding to the second object.
[0165] Aspect 26. The method according to aspect 25, the method further comprising: determining a respective depth of each of the plurality of objects associated with the feature map.
[0166] Aspect 27. The method according to any one of aspects 25 to 26, wherein the first part of the first object and the second part of the second object are assigned to the first container based on the depth of the first part of the first object and the depth of the second part of the second object.
[0167] Aspect 28. The method according to any one of aspects 25 to 27, wherein the first object includes a first person, the second object includes a second person, the first part includes a body part or a face of the first person, and the second part includes a body part or a face of the second person.
[0168] Aspect 29. The method according to aspect 28, the method further comprising: identifying the parts of the first person and the parts of the first person from the plurality of objects at least partially based on the respective depths of each object; and determining a first container for tracking the first person in a subsequent image and a second container for tracking the second person in the subsequent image.
[0169] Aspect 30. The method according to any one of aspects 25 to 29, wherein the first object is completely contained within the first container, and wherein the second object is completely contained within the second container.
[0170] Aspect 31. An apparatus for processing an image, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: obtain an image comprising at least a first object and a second object; generate a feature map based on providing the image to a neural network; identify a plurality of objects based on the feature map, the plurality of objects including a first part of the first object and a second part of the second object, wherein the first part of the first object and the second part of the second object at least partially overlap; and identify a first container corresponding to the first object and a second container corresponding to the second object.
[0171] Aspect 32. The apparatus according to aspect 31, wherein the at least one processor is configured to: determine a respective depth of each object of the plurality of objects associated with the feature map.
[0172] Aspect 33. The apparatus according to any one of aspects 31 to 32, wherein the first part of the first object and the second part of the second object are assigned to the first container based on the depth of the first part of the first object and the depth of the second part of the second object.
[0173] Aspect 34. The apparatus according to any one of aspects 31 to 33, wherein the first object includes a first person, the first part includes a body part or a face of the first person, and the second part includes a body part or a face of the second person.
[0174] Aspect 35. The apparatus according to aspect 34, wherein the at least one processor is configured to: identify the part of the first person and the part of the first person from the plurality of objects at least partially based on the respective depth of each object; and determine a first container for tracking the first person in a subsequent image and a second container for tracking the second person in the subsequent image.
[0175] Aspect 36. The apparatus according to any one of aspects 31 to 35, wherein the first object is completely contained within the first container, and wherein the second object is completely contained within the second container.
[0176] Aspect 37: A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the operations according to any one of aspects 1 to 12.
[0177] Aspect 38: A device, the device comprising components for performing the operations according to any one of Aspects 1 to 12.
[0178] Aspect 39: A non-transitory computer-readable medium, the non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the operations according to any one of Aspects 25 to 30.
[0179] Aspect 40: A device, the device comprising components for performing the operations according to any one of Aspects 25 to 30.
Claims
1. A method for processing image data, the method comprising: Obtaining an image including at least a first object; Generating a feature map based on providing the image to a neural network; Identifying a plurality of objects based on the feature map, the plurality of objects including a first part of the first object; And Identifying a first set of object parts corresponding to the first object within the plurality of objects.
2. The method according to claim 1, wherein the first part of the first object includes the face of a first person.
3. The method according to claim 2, wherein the first set of object parts includes at least one body part of the first person.
4. The method according to claim 1, the method further comprising: Generating a first container corresponding to the first object based on the first set of object parts.
5. The method according to claim 4, wherein identifying the plurality of objects based on the feature map includes: Generating a corresponding bounding box for each object in the plurality of objects; Determining key points of at least a part of the first set of object parts; And Mapping the key points to the corresponding bounding boxes.
6. The method according to claim 5, the method further comprising: Generating at least one cost function based on the key points and the bounding boxes; And Mapping the bounding boxes and the key points to the corresponding parts of the first object, wherein the first container is generated based on at least one mapped part of the first object.
7. The method according to claim 5, wherein identifying the first set of object parts corresponding to the first object within the plurality of objects includes: Determining the identification of at least one part of the first set of object parts based on a second object in the image.
8. The method according to claim 7, wherein at least one part of the first set of object parts cannot be identified based on the second object in the image.
9. The method according to claim 7, wherein identifying the plurality of objects based on the feature map includes: Generating a corresponding bounding box for each object in the plurality of objects; Determining key points of at least a part of the first set of object parts; And Mapping the key points to the corresponding bounding boxes.
10. The method according to claim 9, the method further comprising: Determining a corresponding depth associated with each object in the plurality of objects; Generating at least one cost function based on the key points, the bounding boxes and the corresponding depth associated with each object in the plurality of objects; And Mapping the bounding boxes and the key points to the corresponding parts of the first object and the second object based on the at least one cost function, wherein the first container is generated based on at least one mapped part of the first object.
11. The method according to claim 1, wherein the neural network includes at least one joint attention module.
12. The method according to claim 11, wherein the at least one joint attention module is configured to generate the feature map, and the at least one joint attention module comprises: At least one 1x1 convolutional operation, the at least one 1x1 convolutional operation being configured to obtain an input feature map and generate a first intermediate feature map, the first intermediate feature map including a reduced number of channels compared to the input feature map; at least one convolutional operation, the at least one convolutional operation being configured to generate a second intermediate feature map, the second intermediate feature map having the same dimensions as the input feature map; and a cropping function, the cropping function being configured to crop the eigenvalues of the second intermediate feature map; and a combination engine, the combination engine being configured to combine the eigenvalues of the input feature map with the cropped eigenvalues to generate the feature map.
13. An apparatus for processing an image, the apparatus comprising: at least one memory; and at least one processor, the at least one processor being coupled to the at least one memory and being configured to: obtain an image including at least a first object; generate a feature map based on providing the image to a neural network; identify a plurality of objects based on the feature map, the plurality of objects including a first part of the first object; and identify a first set of object parts corresponding to the first object within the plurality of objects.
14. The apparatus according to claim 13, wherein the first part of the first object includes a face of a first person.
15. The apparatus according to claim 14, wherein the first set of object parts includes at least one body part of the first person.
16. The apparatus according to claim 13, wherein the at least one processor is configured to: generate a first container corresponding to the first object based on the first set of object parts.
17. The apparatus according to claim 16, wherein the at least one processor is configured to: generate a corresponding bounding box for each of the plurality of objects; determine key points of at least a part of the first set of object parts; and map the key points to the corresponding bounding boxes.
18. The apparatus according to claim 17, wherein the at least one processor is configured to: generate at least one cost function based on the key points and the bounding boxes; and map the bounding boxes and the key points to corresponding parts of the first object, wherein the first container is generated based on at least one mapped part of the first object.
19. The apparatus according to claim 17, wherein the at least one processor is configured to: determine an identification of at least one part of the first set of object parts based on a second object in the image.
20. The apparatus according to claim 19, wherein the at least one part of the first set of object parts cannot be identified based on the second object in the image.
21. The apparatus according to claim 19, wherein the at least one processor is configured to: generate a corresponding bounding box for each of the plurality of objects; determine key points of at least a part of the first set of object parts; and map the key points to the corresponding bounding boxes.
22. The apparatus according to claim 21, wherein the at least one processor is configured to: Determine a corresponding depth associated with each of the plurality of objects; Generate at least one cost function based on the key points, the bounding boxes, and the corresponding depths associated with each of the plurality of objects; and Map the bounding boxes and the key points to corresponding parts of the first object and the second object based on the at least one cost function, wherein the first container is generated based on at least one mapped part of the first object.
23. The apparatus according to claim 13, wherein the neural network includes at least one joint attention module.
24. The apparatus according to claim 23, wherein the at least one joint attention module is configured to generate the feature map, and the at least one joint attention module comprises: At least one 1x1 convolutional operation configured to obtain an input feature map and generate a first intermediate feature map, the first intermediate feature map including a reduced number of channels compared to the input feature map; at least one convolutional operation configured to generate a second intermediate feature map having the same dimension as the input feature map; And a clipping function configured to clip the eigenvalues of the second intermediate feature map; And a combination engine configured to combine the eigenvalues of the input feature map with the clipped eigenvalues to generate the feature map.
25. A method for processing an image, the method comprising: Obtain an image including at least a first object and a second object; Generate a feature map based on providing the image to a neural network; Identify a plurality of objects based on the feature map, the plurality of objects including a first part of the first object and a second part of the second object, wherein the first part of the first object and the second part of the second object at least partially overlap; And Identify a first container corresponding to the first object and a second container corresponding to the second object.
26. The method according to claim 25, the method further comprising: Determine a corresponding depth of each of the plurality of objects associated with the feature map.
27. The method according to claim 25, wherein the first part of the first object and the second part of the second object are assigned to the first container based on the depth of the first part of the first object and the depth of the second part of the second object.
28. The method according to claim 25, wherein the first object includes a first person, the second object includes a second person, the first part includes a body part or a face of the first person, and the second part includes a body part or a face of the second person.
29. The method according to claim 28, the method further comprising: Identify the parts of the first person and the parts of the first person from the plurality of objects at least partially based on the corresponding depth of each object; And Determine a first container for tracking the first person in a subsequent image and a second container for tracking the second person in the subsequent image.
30. The method according to claim 25, wherein the first object is completely contained within the first container, and wherein the second object is completely contained within the second container.