System and method for rapid object detection

By processing the region of interest in the input image on the mobile electronic device, identifying significant parts of the object and determining its estimated full appearance, the problem of real-time object detection resource limitation on the mobile electronic device is solved, and efficient object detection is achieved.

CN112154476BActive Publication Date: 2025-05-13SAMSUNG ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201980033895.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-05-22
Filing Date
2019-05-22
Publication Date
2025-05-13
Estimated Expiration
2039-05-22

AI Technical Summary

Technical Problem

Real-time object detection on mobile electronic devices is challenging due to resource constraints such as memory and computing power, especially when dealing with objects with different object sizes.

Method used

By processing the region of interest in the input image on an electronic device, a significant portion of the object is identified and the estimated full appearance of the object is determined based on the significant portion and its relationship with the object, thereby achieving rapid object detection.

Benefits of technology

This method improves the efficiency of object detection, can reduce the consumption of computing resources while maintaining high accuracy, and is suitable for real-time object detection on mobile electronic devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112154476B_ABST
    Figure CN112154476B_ABST
Patent Text Reader

Abstract

One embodiment provides a method comprising: identifying a salient portion of an object in an input image based on processing a region of interest (RoI) in the input image at an electronic device. The method further comprises: determining an estimated overall appearance of the object in the input image based on the salient portion and a relationship between the salient portion and the object. Operating the electronic device based on the estimated overall appearance of the object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments relate generally to object detection, and particularly to systems and methods for fast object detection. Background Art

[0002] Object detection generally refers to the process of detecting one or more objects in digital image data. Real-time object detection on mobile electronic devices is challenging due to their resource limitations (e.g., memory and computational limitations). Summary of the invention

[0003] One embodiment provides a method, comprising: identifying a salient portion of an object in an input image based on processing a region of interest (RoI) in the input image at an electronic device. The method also includes determining an estimated overall appearance of the object in the input image based on the salient portion and a relationship between the salient portion and the object. The electronic device is operated based on the estimated overall appearance of the object.

[0004] These and other features, aspects and advantages of one or more embodiments will become understood with reference to the following description, appended claims and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Figure 1 An exemplary computing architecture for implementing an object detection system in one or more embodiments is shown.

[0006] Figure 2 An exemplary object detection system in one or more embodiments is described in detail.

[0007] Figure 3 An exemplary training phase in one or more embodiments is shown.

[0008] Figure 4 An exemplary object detection process in one or more embodiments is shown.

[0009] Figure 5 An exemplary trained multi-label classification network in one or more embodiments is shown.

[0010] Figure 6 A comparison between object detection performed via a conventional cascaded convolutional neural network system and object detection performed via an object detection system in one or more embodiments is shown.

[0011] Figure 7 An exemplary suggestion generation process in one or more embodiments is shown.

[0012] Figure 8 Exemplary applications of an object detection system in one or more embodiments are shown.

[0013] Fig. 9 Another exemplary application of the object detection system in one or more embodiments is shown.

[0014] Fig.10 is a flow diagram of an exemplary process for performing fast object detection in one or more embodiments.

[0015] Fig.11 is a flow chart of an exemplary process for performing fast face detection in one or more embodiments.

[0016] Fig.12 is a high-level block diagram illustrating an information processing system including a computer system for implementing the disclosed embodiments. DETAILED DESCRIPTION

[0017] The following description is made for the purpose of illustrating the general principles of one or an embodiment, and is not meant to limit the inventive concept claimed herein. In addition, the specific features of this article can be used in combination with other features in each of various possible combinations and permutations. Unless otherwise specifically defined in this article, all terms will be given their broadest possible interpretation, including the meaning implied by the specification and the meaning understood by those skilled in the art and / or the meaning defined in dictionaries, papers, etc.

[0018] One or more embodiments generally relate to object detection, and in particular to systems and methods for rapid object detection. One embodiment provides a method that includes: identifying a salient portion of an object in an input image based on processing a region of interest (RoI) in the input image at an electronic device. The method also includes: determining an estimated overall appearance of the object in the input image based on the salient portion and a relationship between the salient portion and the object. Operating the electronic device based on the estimated overall appearance of the object.

[0019] In this specification, the term "input image" generally refers to a digital two-dimensional (2D) image, and the term "input slice" generally refers to a portion of a 2D image segmented from the 2D image. The 2D image may be an image captured via an image sensor (e.g., a camera of a mobile electronic device), a screenshot or screen shot (e.g., a screen shot of a video stream on a mobile electronic device), an image downloaded and stored on a mobile device, etc. The input slice may be segmented from the 2D image using one or more sliding windows or other methods.

[0020] In this specification, the term "region of interest" generally refers to a region of an input image that contains one or more objects (eg, a car, a face, etc.).

[0021] In this specification, the term "face detection" generally refers to the object detection task of detecting one or more faces present in an input image.

[0022] A cascaded convolutional neural network (CNN) is a conventional architecture for object detection. A cascaded CNN consists of multiple stages, where each stage is a CNN-based binary classifier that classifies an input slice as a RoI (e.g., the input slice contains a face) or a non-RoI (i.e., the input slice does not contain a face). From the first stage to the last stage of the cascaded CNN, the CNN used at each stage grows deeper and becomes more discriminative in handling false positives. For example, at the first stage, the input image is scanned to obtain candidate ROIs (e.g., candidate face windows). Since the objects contained in the input image may have different object sizes (i.e., scales), the first stage must address this multi-scale problem by suggesting candidate ROIs that may contain objects of different object sizes.

[0023] A conventional solution to this multi-scale problem is to use a small-sized CNN at the first level. However, due to its limited capacity, a small-sized CNN can only work on a limited range of object sizes. As a result, only objects with object sizes similar to the size of the input slice can be detected by a small-sized CNN. In order to use a small-sized CNN to detect objects with different object sizes, the input image must be resized (i.e., rescaled) to different scales, thereby generating a dense image pyramid with multiple pyramid layers. When scanning each pyramid layer of the dense image pyramid, a fixed-sized input slice is segmented from each pyramid layer, resulting in a slowdown in the overall running time (i.e., low speed).

[0024] Another conventional approach for solving the multi-scale problem is to utilize a large-scale CNN or an ensemble of multiple CNNs that are robust to multi-scale variations. Although this eliminates the need to resize / rescale the input image, this approach still results in a slower overall runtime due to the complex nature of the utilized CNNs.

[0025] The cascaded CNN forwards candidate ROIs from the first stage to the next stage to reduce the number of false positives (i.e., remove candidate ROIs that are not actually ROIs). Before the remaining input slices reach the last stage, most of the input slices are eliminated by the shallower CNNs used in the earlier stages of the cascaded CNN.

[0026] Due to the resource limitations (e.g., memory and computational limitations) of mobile electronic devices, real-time object detection using cascaded CNNs or other conventional methods on mobile electronic devices is challenging (e.g., slow, low accuracy, etc.).

[0027] Objects with different object sizes have different features (i.e., clues). For example, an object with a smaller object size may only have global features (i.e., local features may be missing), while an object with a larger object size may have both global and local features. The two conventional methods described above detect objects as a whole, thereby focusing only on the global features of the object.

[0028] One embodiment provides a system and method for fast object detection, which improves the efficiency of object detection. One embodiment focuses on both global and local features of an object. In order to capture local features of an object, one embodiment extracts salient parts of the object. For example, for facial detection, salient parts of the face include, but are not limited to, eyes, nose, full mouth, left corner of the mouth, right corner of the mouth, and ears, etc. In one embodiment, a deep learning-based method, such as a multi-label classification network, is trained to learn the features of the extracted salient parts (i.e., local features) as well as the features of the entire object (i.e., global features). Unlike conventional methods that classify input slices as objects or non-objects (i.e., binary classifiers), an exemplary multi-label classification network is trained to classify input slices as one of: background, object, or salient part of an object.

[0029] One embodiment increases the speed of deep learning based methods while maintaining high accuracy.

[0030] In one embodiment, if the global features of the object in the input slice are captured, it is determined that the object size is small, and the position of the object is directly obtained from the position of the input slice. Based on this, the position of the object is determined to be the position of the input slice (that is, the position of the input slice is the candidate RoI).

[0031] In one embodiment, if a local feature of the object is captured, it is determined that the object size of the object is large, and the position corresponding to the captured local feature is the position of a significant portion of the entire object. The position of the object is determined based on the position of the input slice and the relationship between the local object (i.e., the significant portion of the entire object) and the entire object.

[0032] In one embodiment, object detection is performed on objects with different object sizes in a single inference, thereby reducing the number of times an input image must be resized / rescaled and thereby reducing the amount / number of pyramid layers included in an image pyramid provided as input to a multi-label classification network, thereby improving efficiency.

[0033] Figure 1An exemplary computing architecture 10 is shown for implementing an object detection system 300 in one or more embodiments. The computing architecture 10 includes an electronic device 100 that includes resources, such as one or more processors 110 and one or more memory units 120. One or more applications can run / operate on the electronic device 100 utilizing the resources of the electronic device 100.

[0034] Examples of electronic device 100 include, but are not limited to, desktop computers, mobile electronic devices (eg, tablet computers, smart phones, laptop computers, etc.), consumer products (such as smart televisions), or any other product that utilizes object detection.

[0035] In one embodiment, the electronic device 100 includes an image sensor 140, such as a camera, integrated in or coupled to the electronic device 100. One or more applications on the electronic device 100 may utilize the image sensor 140 to capture an object presented to the image sensor 140 (e.g., a live video capture of an object, a photograph of an object, etc.).

[0036] In one embodiment, applications on the electronic device 100 include, but are not limited to, an object detection system 300 configured to perform at least one of the following: (1) receiving an input image (e.g., captured via the image sensor 140 and retrieved from the storage device 120); and (2) performing object detection on the input image to detect the presence of one or more objects in the input image. As described in detail later herein, in one embodiment, the object detection system 300 is configured to identify salient portions (e.g., facial parts) of an object (e.g., a face) in the input image based on processing of RoIs in the input image, and determine an estimated overall appearance of the object in the input image based on the salient portions and the relationship between the salient portions and the object. The electronic device 100 can then operate based on the estimated overall appearance of the object.

[0037] In one embodiment, the applications on the electronic device 100 may also include one or more software mobile applications 150 loaded onto or downloaded onto the electronic device 100, such as a camera application, a social media application, etc. The software mobile applications 150 on the electronic device 100 may exchange data with the object detection system 300. For example, the camera application may call the object detection system 300 to perform object detection.

[0038] In one embodiment, the electronic device 100 may also include one or more additional sensors, such as a microphone, a GPS, or a depth sensor. The sensors of the electronic device 100 may be used to capture content and / or capture sensor-based context information. For example, the object detection system 300 and / or the software mobile application 150 may utilize one or more additional sensors of the electronic device 100 to capture content and / or sensor-based context information, such as a microphone for audio data (e.g., voice recordings), a GPS for location data (e.g., location coordinates), or a depth sensor for the shape of an object presented to the image sensor 140.

[0039] In one embodiment, the electronic device 100 includes one or more input / output (I / O) units 130 , such as a keyboard, a keypad, a touch interface, or a display screen, integrated in or coupled to the electronic device 100 .

[0040] In one embodiment, the electronic device 100 is configured to exchange data with one or more remote servers 200 or remote electronic devices via a connection (e.g., a wireless connection such as a WiFi connection or a cellular data connection, a wired connection, or a combination of both). For example, the remote server 200 may be an online platform for hosting one or more online services (e.g., an online social media service) and / or distributing one or more software mobile applications 150. As another example, the object detection system 300 may be loaded onto the electronic device 100 or downloaded to the electronic device 100 from a remote server 200 that maintains and distributes updates for the object detection system 300.

[0041] In one embodiment, computing architecture 10 is a centralized computing architecture. In another embodiment, computing architecture 10 is a distributed computing architecture.

[0042] Figure 2 An exemplary object detection system 300 in one or more embodiments is shown in detail. In one embodiment, the object detection system 300 includes a proposal system 310 configured to determine one or more candidate ROIs in one or more input images according to a deep learning-based approach. As described in detail later herein, in one embodiment, the proposal system 310 utilizes a trained multi-label classification network (MCN) 320 to perform one or more object detection tasks by simultaneously capturing local and global features of an object.

[0043] In one embodiment, the suggestion system 310 includes an optional training phase associated with an optional training system 315. In one embodiment, during the training phase, the training system 315 is configured to receive a set of input images 50 ( Figure 3) (“training image”), and generates a set of input slices 55 by randomly segmenting the input slices 55 from the training image 50 using the image segmentation unit 316. The training system 315 provides the input slices 55 to the initial MCN 317 for use as training data. During the training phase, the MCN 317 is trained to learn local and global features of the object.

[0044] In one embodiment, the training phase may be performed offline (ie, not on the electronic device 100). For example, in one embodiment, the training phase may be performed using a remote server 200 or a remote electronic device.

[0045] In one embodiment, the proposed system 310 includes an operation phase associated with an extraction and detection system 318. In the operation phase, in one embodiment, the extraction and detection system 318 is configured to receive an input image 60 ( Figure 4 ), and by adjusting / re-scaling the input image 60 using the image resizing unit 330 to form a sparse image pyramid 65 ( Figure 4 ). The extraction and detection system 318 provides the sparse image pyramid 65 to the trained MCN 320 (e.g., the trained MCN 320 generated by the training system 315). In response to receiving the sparse image pyramid 65, the MCN 320 generates a set of feature maps 70 ( Figure 4 ), where each feature map 70 is a heat map that indicates one or more regions (i.e., locations) in the input image 60 where features associated with entire objects or significant portions of entire objects (e.g., facial parts) are captured by the MCN 320. As described in detail later herein, in one embodiment, for face detection, the MCN 320 generates a feature map that indicates one or more regions in the input image 60 where features associated with entire faces and / or facial parts are captured.

[0046] The extraction and detection system 318 forwards the feature map 70 to the suggestion generation system 340. The suggestion generation system 340 is configured to generate one or more suggestions for the input image 60 based on the feature map 70 and the predefined object bounding box templates, wherein each suggestion indicates one or more candidate ROIs in the input image 60. As described in detail later herein, in one embodiment, for face detection, the suggestion generation system 340 generates one or more face suggestions based on the feature map 70 of faces and / or face parts and the predefined bounding box templates for different face parts.

[0047] In one embodiment, the object detection system 300 includes a classification and regression unit 350. The classification and regression unit 350 is configured to receive one or more suggestions from the suggestion system 310, classify each suggestion as one of background, a whole object (e.g., a face), or a salient portion of a whole object (e.g., a facial part), and regress each suggestion containing the whole object or the salient portion of the whole object to fit the boundary of the object.

[0048] In one embodiment, the operating phase may occur online (ie, on the electronic device 100 ).

[0049] Figure 3 An exemplary training phase in one or more embodiments is shown. In one embodiment, the training phase includes training an initial MCN 317 for face detection. Specifically, the training system 315 receives a set of training images 50 including faces and facial parts. The image segmentation unit 316 randomly segments a set of input slices 55 including faces and different facial parts (e.g., ears, eyes, full mouth, nose, etc.) from the training images 50. For example, if the training images 50 include an image A showing a first group of runners and a different image B showing a second group of runners, the input slices 55 may include, but are not limited to, one or more of the following: an input slice AA showing an eye segmented from image A, an input slice AB showing a full mouth segmented from image A, an input slice AC showing a nose segmented from image A, an input slice AD ​​showing different eyes segmented from image A, and an input slice BA showing an ear segmented from image B. In one embodiment, the input slice 55 is used to train the initial MCN 317 to classify the input slice 55 into one of the following categories / classifications: (1) background; (2) entire face; (3) eyes; (4) nose; (5) full mouth; (6) left corner of mouth; (7) right corner of mouth; or (8) ear. The MCN 317 is trained to simultaneously capture global features of the entire face and local features of different facial parts (i.e., eyes, nose, full mouth, left corner of mouth, right corner of mouth, or ears). In another embodiment, the initial MCN 317 can be trained to classify the input slice 55 into different or fewer / more categories / classifications than the above eight categories.

[0050] In one embodiment, the trained MCN 320 generated by the training system 315 is fully convolutional. The trained MCN 320 does not require a fixed input size and can receive input images of arbitrary dimensions. The trained MCN 320 eliminates the need to use sliding windows or other methods to segment out input slices. In one embodiment, the input resolution of the trained MCN 320 is 12×12, and the stride width is set to 2.

[0051] Figure 41 shows an exemplary object detection process in one or more embodiments. In one embodiment, the operation phase includes: using the trained MCN 320 to perform face detection. Specifically, the extraction and detection system 318 receives an input image 60 during the operation phase. For example, Figure 4 As shown, the input image 60 may show one or more faces. The image resizing unit 330 resizes the input image 60 to different scales by generating a sparse image pyramid 65 having one or more pyramid levels 66, wherein each pyramid level 66 is encoded with a different scale of the input image 60. The sparse image pyramid 65 has fewer pyramid levels 66 than a dense image pyramid generated using a conventional cascaded CNN. Unlike a dense image pyramid in which each pyramid level of the dense image pyramid corresponds to one specific scale, each pyramid level 66 of the sparse image pyramid 65 corresponds to multiple scales, thereby reducing the amount / number of pyramid levels 66 required and increasing the number of suggestions generated by the suggestion generation system 340 for the input image 60.

[0052] For example, Figure 4 As shown, the sparse image pyramid 65 may include a first pyramid level 66 (level 1) and a second pyramid level 66 (level 2).

[0053] In response to receiving the sparse image pyramid 65 from the image resizing unit 330, the trained MCN 320 generates a set of feature maps 70 for faces and facial parts. Based on the feature maps 70 received from the MCN 320 and the predefined bounding box template for each facial part, the suggestion generation system 340 generates one or more facial suggestions 80. Each facial suggestion 80 indicates one or more candidate face windows 85 (if any), where each candidate face window 85 is a candidate RoI that contains a possible face. For example, Figure 4 As shown, face proposals 80 may include candidate face windows 85 for each face shown in input image 60 .

[0054] If the MCN 320 captures global features of the entire face in the input slice, the suggestion system 310 determines that the location of the entire face is the location of the input slice (i.e., the location of the input slice is the candidate face window 85). If the MCN 320 captures local features of facial parts in the input slice, the suggestion system 310 infers the location of the entire face based on the location of the input slice and the relationship between the facial parts and the entire face.

[0055] Figure 5An exemplary, trained MCN 320 in one or more embodiments is shown. In one embodiment, MCN 320 is fully convolutional. MCN 320 does not require a fixed input size and can receive input images of arbitrary dimensions. MCN 320 includes multiple layers 321 (e.g., one or more convolutional layers, one or more pooling layers), including a final layer 322. Each layer 321 includes a set of receptive fields 323 of a specific size.

[0056] For example, Figure 5 As shown, MCN 320 may include at least the following: (1) a first layer 321 ("layer 1") including a first group of receptive fields 323, wherein each of the first group of receptive fields 323 has a size of 10×10×16; (2) a second layer 321 ("layer 2") including a second group of receptive fields 323, wherein each of the second group of receptive fields 323 has a size of 8×8×16; (3) a third layer 321 ("layer 3") including a third group of receptive fields 323, wherein each of the third group of receptive fields 323 has a size of 8×8×16; (4) a fourth layer 321 ("layer 4") including a fourth group of receptive fields 323, wherein each of the fourth group of receptive fields 323 has a size of 8×8×16. has a size of 6×6×32; (5) a fifth layer 321 (“layer 5”) including a fifth group of receptive fields 323, wherein each of the fifth group of receptive fields 323 has a size of 4×4×32; (6) a sixth layer 321 (“layer 6”) including a sixth group of receptive fields, wherein each of the sixth group of receptive fields 323 has a size of 2×2×32; (7) a seventh layer 321 (“layer 7”) including a seventh group of receptive fields, wherein each of the seventh group of receptive fields 323 has a size of 1×1×64; and (8) a final layer 322 (“layer 8”) including an eighth group of receptive fields, wherein each of the eighth group of receptive fields 323 has a size of 1×1×8.

[0057] In one embodiment, the set of receptive fields 323 of the last layer 322 has a total size of m×n×x, where m×n is the maximum image resolution of the input image 60 that the MCN 320 can receive as input, and x is the number of different categories / classifications that the MCN 320 is trained to classify. For example, if the MCN 320 is trained to classify eight different categories / classifications for face detection (e.g., background, whole face, eyes, nose, whole mouth, left corner of mouth, right corner of mouth, and ears), then each receptive field 323 of the last layer 322 has a size of 1×1×8. If the maximum image resolution is 12×12, then the total size of the last layer 322 is 12×12×8.

[0058] In one embodiment, for each category that the MCN 320 is trained to classify, the last layer 322 is configured to generate a corresponding feature map 70 representing one or more regions in the input image 60 where features associated with the classification are captured by the MCN 320. For example, assume that the MCN 320 is trained to classify at least the following eight categories / classifications for face detection: background, whole face, eyes, nose, full mouth, left corner of mouth, right corner of mouth, and ears. In response to receiving the input image 60, the last layer 322 generates at least the following features: (1) a first feature map 70 (HEAT MAP1) indicating one or more regions in the input image 60 in which features associated with the entire face are captured; (2) a second feature map 70 (HEAT MAP2) indicating one or more regions in the input image 60 in which features associated with the eyes are captured; (3) a third feature map 70 (HEAT MAP3) indicating one or more regions in the input image 60 in which features associated with the nose are captured; (4) a fourth feature map 70 (HEAT MAP4) indicating one or more regions in the input image 60 in which features associated with the entire mouth are captured; (5) a fifth feature map 70 (HEAT MAP5) indicating one or more regions in the input image 60 in which features associated with the left corner of the mouth are captured; (6) a sixth feature map 70 (HEAT MAP6) indicating one or more regions in the input image 60 in which features associated with the right corner of the mouth are captured. MAP6); (7) a seventh feature map 70 (HEAT MAP7) indicating one or more regions in the input image 60 in which features associated with the ear are captured; and (8) an eighth feature map 70 (HEAT MAP8) indicating one or more regions in the input image in which features associated with the background are captured.

[0059] In one embodiment, object detection system 300 is configured to directly identify salient portions of an object when the object size of the object exceeds a processing size (eg, a maximum image resolution of MCN 320).

[0060] Figure 6 4 and the object detection system 300. Assume that the same input image 60 is provided to the conventional cascade CNN system 4 and the object detection system 300. Figure 6 As shown, the input image 60 shows different faces at different areas in the input image 60, such as a first face S at a first area, a second face T at a second area, a third face U at a third area, and a fourth face V at a fourth area.

[0061] In response to receiving the input image 60, the conventional cascaded CNN system 4 generates a dense image pyramid 5 including a plurality of pyramid layers 6, where each pyramid layer 6 corresponds to a specific scale of the input image 60. For example, as Figure 6 shown, the dense image pyramid 5 includes: a first pyramid layer 6 corresponding to the input image 60 at a first scale (Scale 1), a second pyramid layer 6 corresponding to the input image 60 at a second scale (Scale 2), and an Nth pyramid layer 6 corresponding to the input image 60 at an Nth scale (Scale N), where N is a positive integer. The conventional cascaded CNN system 4 provides the dense image pyramid 5 to the cascaded CNN.

[0062] In comparison, in response to receiving the input image 60, the object detection system 300 generates a sparse image pyramid 65 including a plurality of pyramid layers 66, where each pyramid layer 66 corresponds to a different scale of the input image 60. For example, as Figure 6 shown, the sparse image pyramid 65 includes: a first pyramid layer 66 corresponding to the input image 60 at a set of different scales (including the first scale (Scale 1)), and an Mth pyramid layer 66 corresponding to the input image 60 at another set of different scales (including the Mth scale (Scale M)), where M is a positive integer and M < N. When each pyramid layer 66 of the sparse image pyramid 65 is encoded at multiple scales, the sparse image pyramid 65 requires fewer pyramid layers than the dense image pyramid 5. The object detection system 300 provides the sparse image pyramid 65 to the MCN 320.

[0063] For each pyramid layer 6 of the dense image pyramid 5, the cascaded CNN classifies each input slice of the pyramid layer 6 as either a face or just background. For example, as Figure 6 shown, the cascaded CNN classifies as follows: (1) three input slices of the pyramid layer 6A of the dense image pyramid 66 as faces (i.e., faces S, T, and U); and (2) one input slice of the pyramid layer 6B of the dense image pyramid 66 as a face (i.e., face V). Based on this classification, the conventional cascaded CNN system 4 outputs face proposals 8 representing four candidate face windows 85 (i.e., faces S, T, U, and V) in the input image 60.

[0064] In comparison, for each pyramid layer 66 of the sparse image pyramid 65, the MCN 320 classifies each input slice of the pyramid layer 66 as either just background, an entire face, or a specific facial part of the entire face (i.e., eyes, nose, full mouth, left mouth corner, right mouth corner, or ear). For example, as Figure 6As shown, the MCN 320 classifies as follows: (1) one input slice of the pyramid layer 66A of the sparse image pyramid 66 as a mouth (i.e., the mouth of face S); (2) two other input slices of the pyramid layer 66A as eyes (i.e., the eyes of face T and the eyes of face U); and (3) another input slice of the pyramid layer 66A as a face (i.e., face V). The object detection system 300 outputs face proposals 80 representing four candidate face windows 85 (i.e., faces S, T, U, and V) in the input image 60. Therefore, unlike the conventional cascaded CNN system 4, the object detection system 300 is more accurate because it is able to detect the entire face and different face parts.

[0065] Figure 7 1 shows an exemplary suggestion generation process in one or more embodiments. In one embodiment, a set of feature maps 70 generated by MCN 320 in response to receiving input image 60 is forwarded to suggestion generation system 340. For example, if MCN 320 is trained for face detection, then Figure 7 (For ease of illustration, feature maps corresponding to only the background, the entire face, or other facial parts are not shown in Figure 7 ), a set of feature maps 70 may include a first feature map 70A and a second feature map 70B, the first feature map 70A indicating one or more regions in the input image 60 where features associated with the full mouth are captured by the MCN 320, and the second feature map 70B indicating one or more regions in the input image 60 where features associated with the eyes are captured by the MCN 320.

[0066] In one embodiment, the suggestion generation system 340 includes a local maximum unit 341 configured to determine the local maximum of the feature map for each feature map 70 corresponding to a facial part. Let p generally represent a specific facial part, and let τ p Generally, it represents a corresponding predetermined threshold for maintaining strong response points in a local area of ​​the feature map corresponding to the facial part p. In one embodiment, in order to determine the local maximum value of the feature map 70 corresponding to the facial part p, the local maximum value unit 341 applies non-maximum suppression (NMS) to the feature map 70 to obtain one or more strongest response points in one or more local areas of the feature map 70. For example, Figure 7 As shown, the local maximum unit 341 obtains the strongest response point 71A (corresponding to the position of the mouth) for the first feature map 70A and two strongest response points 71BA and 71BB (corresponding to the positions of the left eye and the right eye) for the second feature map 70B.

[0067] In one embodiment, the suggestion generation system 340 includes a bounding box unit 342 configured to: for each feature map 70 corresponding to a facial part, determine one or more bounding boxes for the facial part based on a local maximum of the feature map 70 (e.g., a local maximum determined by the local maximum unit 341) and one or more bounding box templates for the facial part. For each facial part p, the bounding box unit 342 maintains one or more corresponding bounding box templates. The bounding box template corresponding to the facial part is a predefined template area of ​​the facial part. For example, for some facial parts, such as eyes, the bounding box unit 342 may maintain two bounding box templates.

[0068] Assume b i Generally represents the position of the bounding box i, where bi = (x i1 ,y i1 ,x i2 ,y i2 )、(x i1 ,y i1 ) is the coordinate of the upper left vertex of the bounding box i, and (x i2 ,y i2 ) is the coordinate of the lower right vertex of the bounding box i. Let p i Generally, it represents the confidence score of the corresponding bounding box i. In one embodiment, in order to determine the bounding box of the face based on the local maximum value of the feature map 70 corresponding to the facial part p, the bounding box unit 342 sets the corresponding confidence score pi to be equal to its corresponding value in the feature map 70, where the corresponding value is the size of the point on the feature map 70 corresponding to the position of the bounding box. For example, Figure 7 As shown, the bounding box unit 342 determines a bounding box 72A for the first feature map 70A, and four separate bounding boxes 72B for the second feature map 70B (two bounding boxes 70B for the left eye and two bounding boxes 70B for the right eye).

[0069] In one embodiment, the suggestion generation system 340 includes a part box combination (PBC) unit 343, which is configured to infer an area containing a face (i.e., a face region or a face window) from an area containing face parts (i.e., a face part region). In one embodiment, to obtain the face region, face part regions with high overlap are combined by averaging.

[0070] Specifically, given a set of original bounding boxes corresponding to feature maps 70 of different facial parts, the PBC unit 343 initiates the search and merge process by selecting the bounding box with the highest confidence score and identifying the bounding boxes with a confidence score above a threshold τ IoUThe PBC unit 343 merges / combines the selected bounding box and the identified bounding box into a merged bounding box representing the face region by averaging the position coordinates according to the formula (1) provided below:

[0071]

[0072] Among them, C i is a set of highly overlapping bounding boxes, and C i It is defined according to the formula (2) provided below:

[0073] c i = {b i}∪{b j :IoU(b i , b j )>τ IoU} (2).

[0074] The PBC unit 343 determines the corresponding confidence score p for the merged bounding box according to the formula (3) provided below: m,i :

[0075]

[0076] For example, Figure 7 As shown, the PBC unit 343 generates a face proposal 80 including a merged bounding box representing a candidate face window 85, wherein the merged bounding box depends on a set of bounding boxes of the feature maps 70 corresponding to different face parts (e.g., bounding boxes 72A and 72B for feature maps 70A and 70B, respectively). The face proposal 80 is assigned to the merged bounding box, and the bounding box used for merging is removed from the original set. The PBC unit 343 repeats the search and merge process for the remaining bounding boxes in the original set until there are no remaining bounding boxes.

[0077] Figure 8An exemplary application of the object detection system 300 in one or more embodiments is shown. In one embodiment, one or more software mobile applications 150 loaded or downloaded to the electronic device 100 can exchange data with the object detection system 300. In one embodiment, a camera application that controls a camera on the electronic device 100 can call the object detection system 300 to perform object detection. For example, if a user interacts with the shutter of the camera, the camera application may be able to capture a picture (i.e., a photo) only when the object detection system 300 detects the expected features of each object (e.g., a person) within the camera view 400 of the camera. The expected features of each object may include, but are not limited to, objects with open eyes, a smiling mouth, a full face (i.e., a face that is not partially obscured / occluded by an object or shadow), and not partially outside the camera view. In one embodiment, additional learning systems may be used to accurately extract these desired features, such as, but not limited to, open mouth recognition, expression recognition, etc. that can be constructed using supervised labeled data. For example, as Figure 8 As shown, camera view 400 may display four different objects to be captured, G, H, I, and J. Since object G is partially outside of camera view 400, object H has closed eyes, and object I has an open mouth, object detection system 300 only detects the expected features of object J (i.e., objects G, H, and I do not have the expected features).

[0078] If the object detection system 300 detects the expected features of each object to be captured, the camera application enables the picture to be captured; otherwise, the camera application can invoke other actions, such as delaying closing the shutter, providing a warning to the user, etc. Since the size of faces and facial parts can vary with the distance between the object to be captured and the camera, the object detection system 300 can achieve fast face detection with large-scale capabilities.

[0079] Fig. 9 Another exemplary application of the object detection system 300 in one or more embodiments is shown. In one embodiment, the camera application can utilize the object detection system 300 to analyze the current composition 410 of the scene within the camera view of the camera, and provide one or more suggestions to improve the composition of the picture to be taken based on the current composition 410. The object detection system 300 is configured to determine a bounding box of the object so that the boundaries of the bounding box tightly surround the object. Therefore, if the object detection system 300 determines the bounding box of the object, the position and size of the object are also determined. For example, for each object within the camera view, the object detection system 300 can detect the position of the object and the size of the object's face based on the bounding box determined for the object.

[0080] Based on the detected information, the camera application can suggest one or more actions that require a minimal amount of effort from one or more objects within the camera view, such as suggesting that an object move to another location, etc. Fig. 9 As shown, the camera application may provide a first suggestion (Suggestion 1) in which one object is moved further back to create an alternative combination 420 and a second suggestion (Suggestion 2) in which another object is moved further forward to create an alternative combination 430. The suggestions may be presented in various formats using the electronic device 100, including but not limited to visual prompts, voice notifications, etc.

[0081] Fig.10 800 for performing fast object detection in one or more embodiments. Process block 801 includes receiving an input image. Process block 802 includes resizing the input image to different scales to form a sparse image pyramid, wherein the sparse image pyramid is fed to a multi-label classification network (MCN) trained to capture local and global features of objects. Process block 803 includes receiving a set of feature maps generated by the MCN, wherein each feature map corresponds to a specific object classification (e.g., background, entire object, or salient portion). Process block 804 includes determining candidate RoIs in the input image based on the feature maps and a predefined bounding box template of the object. Process block 805 includes generating suggestions indicating candidate RoIs.

[0082] In one embodiment, process blocks 801 - 805 may be performed by one or more components of object detection system 300 , such as MCN 320 , image resizing unit 330 , and suggestion generation system 340 .

[0083] Fig.11 900 is a flow chart of an exemplary process for performing fast face detection in one or more embodiments. Process block 901 includes receiving an input image. Process block 902 resizes the input image to different scales to form a sparse image pyramid, wherein the sparse image pyramid is fed to a multi-label classification network (MCN) trained to capture global features of the entire face and local features of different facial parts. Process block 903 includes receiving a set of feature maps generated by the MCN, wherein each feature map corresponds to a specific object classification (e.g., background, entire face, or facial parts such as eyes, nose, full mouth, left corner of mouth, right corner of mouth, and ears). Process block 904 includes determining candidate facial windows in the input image based on the feature maps of different facial parts and a predefined bounding box template. Process block 905 includes generating a suggestion indicating a candidate facial window.

[0084] In one embodiment, process blocks 901 - 905 may be performed by one or more components of object detection system 300 , such as MCN 320 , image resizing unit 330 , and suggestion generation system 340 .

[0085] Fig.12 6 is a high-level block diagram showing an information processing system including a computer system 600 for implementing the disclosed embodiments. Each system 300, 310, 315, 318, and 350 can be combined into a display device or a server device. The computer system 600 includes one or more processors 601, and may also include an electronic display device 602 (for displaying video, graphics, text, and other data), a main memory 603 (e.g., a random access memory (RAM)), a storage device 604 (e.g., a hard disk drive), a removable storage device 605 (e.g., a removable storage drive, a removable storage module, a tape drive, an optical drive, a computer-readable medium having computer software and / or data stored therein), a viewer interface device 606 (e.g., a keyboard, a touch screen, a keypad, a pointing device), and a communication interface 607 (e.g., a modem, a network interface (e.g., an Ethernet card), a communication port, or a PCMCIA slot and card). The communication interface 607 allows software and data to be transferred between the computer system and external devices. The system 600 also includes a communication infrastructure 608 (eg, a communication bus, a jumper, or a network) to which the above-described devices / modules 601 to 607 are connected.

[0086] The information transmitted via the communication interface 607 may be in the form of signals that can be received by the communication interface 607 via a communication link carrying the signal, such as an electronic, electromagnetic, optical or other signal, and may be implemented using wire or cable, optical fiber, telephone line, cellular telephone link, radio frequency (RF) link and / or other communication channels. The computer program instructions representing the block diagrams and / or flow charts herein may be loaded onto a computer, programmable data processing device or processing device so that a series of operations performed thereon generate a computer-implemented process. In one embodiment, the process 800 ( Fig.10 ) and process 900( Fig.11 ) can be stored as program instructions in the memory 603, the storage device 604 and / or the removable storage device 605 for execution by the processor 601.

[0087] Embodiments have been described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products. Each block or combination of such diagrams / graphs may be implemented by computer program instructions. The computer program instructions produce a machine when provided to a processor so that instructions executed by the processor create a device for implementing the functions / operations specified in the flowchart and / or block diagram. Each block in the flowchart / block diagram may represent a hardware and / or software module or logic. In alternative embodiments, the functions annotated in the blocks may not occur in the order annotated in the accompanying drawings, may occur simultaneously, and the like.

[0088] The terms "computer program medium", "computer usable medium", "computer readable medium" and "computer program product" are generally used to refer to media such as main memory, auxiliary memory, removable storage drive, hard disk installed in a hard drive, and signals. These computer program products are devices for providing software to a computer system. Computer readable media allow a computer system to read data, instructions, messages or message packets, and other computer readable information from the computer readable medium. For example, a computer readable medium may include non-volatile memory such as a floppy disk, ROM, flash memory, disk drive memory, CD-ROM, and other permanent storage. For example, it is useful for transferring information such as data and computer instructions between computer systems. Computer program instructions may be stored in a computer readable medium that can direct a computer, other programmable data processing device, or other device to operate in a specific manner so that the instructions stored in the computer readable medium produce an article of manufacture including instructions that implement the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0089] As will be appreciated by those skilled in the art, aspects of the embodiments may be implemented as systems, methods or computer program products. Therefore, aspects of the embodiments may take the form of complete hardware embodiments, complete software embodiments (including firmware, resident software, microcode, etc.), or embodiments of combined software and hardware aspects, which may generally be referred to herein as "circuits," "modules," or "systems." In addition, aspects of the embodiments may take the form of computer program products implemented in one or more computer-readable media, with one or more computer-readable media having computer-readable program codes implemented thereon.

[0090] Any combination of one or more computer-readable media may be utilized. A computer-readable medium may be a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media will include the following: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that may contain or store a program used by or in conjunction with an instruction execution system, device, or apparatus.

[0091] The computer program code for performing the operation of each aspect of one or more embodiments can be written in any combination of one or more programming languages, including object-oriented programming languages ​​(such as Java, Smalltalk, C++, etc.) and conventional process programming languages ​​(such as C programming language or similar programming languages). The program code can be executed completely on the user's computer, executed as an independent software package part on the user's computer, executed partly on the user's computer and partially on a remote computer, or executed completely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer or to an external computer (for example, the Internet using an Internet service provider) through any type of network (including a local area network (LAN) or a wide area network (WAN)).

[0092] Aspects of one or more embodiments are described above with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products. It should be understood that each block of the flowchart and / or block diagram and the combination of blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a special-purpose computer or other programmable data processing device to produce a machine, so that instructions executed by a processor of the computer or other programmable data processing device create a device for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0093] These computer program instructions may also be stored in a computer-readable medium, which may guide a computer, other programmable data processing device or other device to operate in a specific manner, so that the instructions stored in the computer-readable medium generate the instructions implemented in the flowchart and / or block diagram. Figure 1An artifact of instructions that specifies functions / actions in one or more boxes.

[0094] The computer program instructions may also be loaded onto a computer, other programmable data processing device, or other device to cause a series of operational steps to be performed on the computer, other programmable device, or other device, thereby producing a computer-implemented process, so that the instructions executed on the computer or other programmable device provide a process for implementing the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0095] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments. In this regard, each frame in the flow chart or block diagram may represent a part of a module, segment or instruction, which includes one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the frame may not occur in the order noted in the figure. For example, the two frames shown in succession can actually be performed substantially simultaneously, or these frames can sometimes be performed in reverse order, depending on the functions involved. It will also be noted that each frame illustrated in the block diagram and / or flow chart and the combination of frames in the block diagram and / or flow chart illustration can be implemented by a system based on dedicated hardware that performs a specified function or action or performs a combination of dedicated hardware and computer instructions.

[0096] Unless explicitly stated otherwise, reference to a singular element in a claim does not mean "one and only" but "one or more". All structural and functional equivalents to the elements of the above exemplary embodiments that are now known or later known to a person of ordinary skill in the art are included in the claims presented. Unless the element is explicitly stated using the phrase "means for..." or "step for...", the elements of the claims of this application are not to be interpreted in accordance with the provisions of the sixth paragraph of 35 U.S.C. §112. .

[0097] The terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the present invention. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that when used in this specification, the terms "include" and / or "comprise" specify the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups thereof.

[0098] All means or steps in the appended claims plus corresponding structures, materials, actions and equivalents of functional elements are intended to include any structure, material or action for performing a function in combination with other protection elements specifically claimed. The description of the embodiments has been given for the purpose of illustration and description, but is not intended to be exhaustive or limited to the embodiments of the disclosed form. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the invention.

[0099] Although embodiments have been described with reference to certain forms thereof; however, other forms are possible. Therefore, the spirit and scope of the appended claims should not be limited to the description of the preferred forms contained herein.

Claims

1. A method for object detection, comprising: receiving an input image; identifying one or more salient portions of an object in the input image by classifying a set of input slices of the input image using a multi-label classification network at an electronic device, wherein the multi-label classification network is trained to capture one or more global features of the object and one or more local features of the one or more salient portions of the object based on input slices segmented from a set of training images, the multi-label classification network classifying at least one input slice in the set of input slices as a first object classification representing background, and the multi-label classification network classifying one or more other input slices in the set of input slices as one or more additional object classifications representing the one or more salient portions of the object; determining an estimated overall appearance of the object in the input image based on the one or more salient portions and a relationship between the one or more salient portions and the object as defined by one or more bounding box templates for the one or more salient portions; and The electronic device is operated based on the estimated overall appearance of the object.

2. The method of claim 1, wherein: Identifying one or more salient portions of an object in the input image comprises: resizing the input image by generating a sparse image pyramid comprising one or more pyramid levels, wherein each pyramid level corresponds to the input image at a plurality of different scales; generating a set of feature maps based on the sparse image pyramid; and A region of interest in the input image is determined based on the set of feature maps.

3. The method of claim 2, wherein: Generating the set of feature maps comprises: For each input slice of each pyramid level of the sparse image pyramid, the input slice is classified into one of a plurality of object categories using a multi-label classification network, wherein the multi-label classification network is configured to capture one or more global features of the object and one or more local features of one or more salient parts of the object.

4. The method of claim 1, wherein: The object is a face and the prominent portion is a facial part.

5. The method of claim 3, wherein: The plurality of object categories include at least one of: background, entire face, eyes, nose, entire mouth, left corner of mouth, right corner of mouth, or ear.

6. The method of claim 3, further comprising: In response to the multi-label classification network capturing a global feature of the object in the input slice, determining a position of the object in the input image based on a position of the input slice; as well as In response to the multi-label classification network capturing local features of a salient portion of the object, a position of the object in the input image is inferred based on the position of the input slice and a relationship between the salient portion and the object defined by one or more bounding box templates for the salient portion.

7. The method of claim 1, wherein: Operating the electronic device includes: In response to a request to capture a picture via a camera connected to the electronic device, controlling the capture of the picture based on detecting the presence of one or more expected features in the input image using the multi-label classification network, wherein the input image is a camera view of the camera.

8. The method of claim 1, wherein: Operating the electronic device includes: In response to a request to capture a picture via a camera coupled to the electronic device, determining a current composition of the picture by detecting the position and size of each object in the input image, wherein the input image is a camera view of the camera; and One or more suggestions are provided to change the current composition of the picture based on the current composition.

9. A system for object detection, comprising: at least one processor; as well as a non-transitory processor-readable memory device storing instructions that, when executed by the at least one processor, cause the at least one processor to perform operations comprising: receiving an input image; identifying one or more salient portions of an object in the input image by classifying a set of input slices of the input image using a multi-label classification network at an electronic device, wherein the multi-label classification network is trained to capture one or more global features of the object and one or more local features of the one or more salient portions of the object based on input slices segmented from a set of training images, the multi-label classification network classifying at least one input slice in the set of input slices as a first object classification representing background, and the multi-label classification network classifying one or more other input slices in the set of input slices as one or more additional object classifications representing the one or more salient portions of the object; determining an estimated overall appearance of the object in the input image based on the one or more salient portions and a relationship between the one or more salient portions and the object as defined by one or more bounding box templates for the one or more salient portions; and The electronic device is operated based on the estimated overall appearance of the object.

10. The system of claim 9, wherein: Identifying one or more salient portions of an object in the input image comprises: resizing the input image by generating a sparse image pyramid comprising one or more pyramid levels, wherein each pyramid level corresponds to the input image at a plurality of different scales; generating a set of feature maps based on the sparse image pyramid; and A region of interest in the input image is determined based on the set of feature maps.

11. The system of claim 10, wherein: Generating the set of feature maps comprises: For each input slice of each pyramid level of the sparse image pyramid, the input slice is classified into one of a plurality of object categories using a multi-label classification network, wherein the multi-label classification network is configured to capture one or more global features of the object and one or more local features of one or more salient parts of the object.

12. The system of claim 9, wherein: The object is a face and the prominent portion is a facial part.

13. The system of claim 11, wherein: The plurality of object categories include at least one of: background, entire face, eyes, nose, entire mouth, left corner of mouth, right corner of mouth, or ear.

14. The system of claim 11, wherein: The operations also include: In response to the multi-label classification network capturing a global feature of the object in the input slice, determining a position of the object in the input image based on a position of the input slice; and In response to the multi-label classification network capturing local features of a salient portion of the object, a position of the object in the input image is inferred based on the position of the input slice and a relationship between the salient portion and the object defined by one or more bounding box templates for the salient portion.

Citation Information

Patent Citations

  • Face detecting camera and method

    US20050264658A1