Portrait detection method and apparatus, electronic device, and storage medium

By acquiring images from transportation hubs and performing feature fusion, the distribution of human figures can be automatically detected, solving the problem of high-cost manual statistics and achieving accurate detection of the number and distribution of people in clusters.

CN115731585BActive Publication Date: 2026-05-29SIEMENS (CHINA) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SIEMENS (CHINA) CO LTD
Filing Date
2021-08-26
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Counting the number of people gathered in transportation hubs requires a large amount of manpower, resulting in high costs and an inability to accurately determine the distribution of people.

Method used

By acquiring the image to be detected, generating multiple feature images and fusing features, the distribution of human figures in the image is determined, thereby automatically detecting the number and distribution of people gathered together.

Benefits of technology

There is no need to station staff at each entrance and exit, saving manpower, reducing statistical costs, and improving the accuracy and comprehensiveness of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115731585B_ABST
    Figure CN115731585B_ABST
Patent Text Reader

Abstract

The present application provides a portrait detection method and device, electronic equipment and storage medium, the portrait detection method comprises: obtaining an image to be detected, wherein the image to be detected includes at least one portrait; at least two first feature images of the image to be detected are generated, wherein the first feature image is obtained by extracting features from the image to be detected, and the subsequent first feature image is obtained by extracting features from the previous first feature image; feature fusion is performed on the at least two first feature images to obtain at least two second feature images; and the distribution of the portrait in the image to be detected is determined according to the at least two second feature images. The present scheme reduces the cost of counting the number of gathered personnel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus, electronic device and storage medium for human face detection. Background Technology

[0002] With the rapid development of cities, the passenger flow of urban transportation hubs is increasing. For example, transportation hubs such as subway stations, train stations, and airports have a large passenger flow. In case of accidents or severe weather, a large number of people will gather in a short period of time. The large number of people gathered poses a significant safety hazard. Therefore, it is necessary to determine the number of people gathered in the transportation hub so that flow restriction measures can be taken when the number of people gathered exceeds the carrying capacity of the transportation hub to prevent accidents such as falling onto the platform and stampedes.

[0003] Currently, in order to determine the number of people gathered in a transportation hub, staff members count the number of people entering and leaving the transportation hub at the entrances and exits, and the number of people gathered in the transportation hub is determined based on the count results of each staff member.

[0004] Since transportation hubs typically have multiple entrances and exits, determining the number of people gathered by having staff count the number of people entering and leaving the transportation hub at each entrance and exit requires staff to be stationed at each entrance and exit to count the number of people. Therefore, it requires a lot of manpower, resulting in a high cost for counting the number of people gathered. Summary of the Invention

[0005] In view of this, the facial recognition method, apparatus, electronic device and storage medium provided in this application can reduce the cost of counting the number of people gathered together.

[0006] In a first aspect, embodiments of this application provide a human face detection method, including:

[0007] Obtain an image to be detected, wherein the image to be detected includes at least one human figure;

[0008] At least two first feature images are generated for the image to be detected, wherein the first first feature image is obtained by extracting features from the image to be detected, and the second first feature image is obtained by extracting features from the first first feature image.

[0009] The at least two first feature images are fused to obtain at least two second feature images;

[0010] The distribution of human figures in the image to be detected is determined based on the at least two second feature images.

[0011] Secondly, embodiments of this application also provide a human face detection device, comprising:

[0012] An acquisition module is used to acquire an image to be detected, wherein the image to be detected includes at least one human image;

[0013] The generation module is used to generate at least two first feature images of the image to be detected obtained by the acquisition module, wherein the first first feature image is obtained by extracting features from the image to be detected, and the second first feature image is obtained by extracting features from the first first feature image.

[0014] A fusion module is used to perform feature fusion on the at least two first feature images generated by the generation module to obtain at least two second feature images;

[0015] The detection module is used to determine the distribution of human figures in the image to be detected based on the at least two second feature images obtained by the fusion module.

[0016] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus;

[0017] The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the human face detection method provided in the first aspect above.

[0018] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer instructions, which, when executed by a processor, cause the processor to perform the operations corresponding to the human face detection method provided in the first aspect.

[0019] Fifthly, embodiments of this application also provide a computer program product, which is tangibly stored on a computer-readable medium and includes computer-executable instructions that, when executed, cause at least one processor to perform the human face detection method provided as described in the first aspect or any possible implementation thereof.

[0020] The process involves acquiring an image to be detected that includes human figures, extracting features from the image to obtain multiple first feature images, fusing these first feature images to obtain multiple second feature images, and then determining the distribution of human figures in the image based on these second feature images. Since the image to be detected can be collected from the corresponding location, the human figures in the image can be mapped to the corresponding location. Thus, based on the distribution of human figures in the image, the number and distribution of people gathered in the corresponding location can be determined, achieving automatic detection of the number and distribution of people gathered in the location. This eliminates the need to station staff at each entrance and exit of the location to count the number of people, thereby saving manpower and reducing the cost of counting the number of people gathered in the location.

[0021] Optionally, when obtaining a second feature image by feature fusion of a first feature image, at least two adjacent first feature images are fused according to the generation order of each first feature image to obtain at least two second feature images, wherein different second feature images are obtained by feature fusion of at least two first feature images that are not completely identical.

[0022] Since the first feature images are generated sequentially, each subsequent first feature image is obtained by extracting features from the previous first feature image. Some features in the previous first feature image may be discarded in the subsequent first feature image. The discarded features may be small human figures in the image to be detected. By performing feature fusion on adjacent generated first feature images, it is ensured that the obtained second feature image does not lose the features in the image to be detected. Therefore, when determining the distribution of human figures in the image to be detected based on the second feature image, small human figures in the image to be detected can also be identified, thereby improving the accuracy of detecting the number and distribution of human figures in the image to be detected.

[0023] Optionally, when fusing features of adjacent generated first feature images to obtain a second feature image, the last generated first feature image is first convolved to obtain a second feature image corresponding to the last generated first feature image. Then, the second feature image corresponding to the next generated first feature image is fused with the previous generated first feature image to obtain a second feature image corresponding to the previous generated first feature image.

[0024] Since the second feature image corresponding to the subsequently generated first feature image is fused with the first feature image generated previously to obtain the second feature image corresponding to the first feature image generated previously, for any first feature image, the second feature image corresponding to that first feature image includes all features in the first feature image of a higher order than that first feature image, and no features in the image to be detected are lost. Therefore, when determining the number and distribution of human figures in the image to be detected based on the second feature image, the comprehensiveness of human figure recognition in the image to be detected can be improved, thereby improving the accuracy of detecting the number and distribution of human figures in the image to be detected.

[0025] Optionally, when fusing the second feature image corresponding to the next generated first feature image with the previous generated first feature image to obtain a second feature image corresponding to the previous generated first feature image, convolution processing is performed on the second feature image corresponding to the nth generated first feature image to obtain a third feature image, bilinear interpolation processing is performed on the third feature image to obtain a fourth feature image, feature fusion is performed on the fourth feature image with the (n-1)th generated first feature image to obtain a fifth feature image, and convolution processing is performed on the fifth feature image to obtain a second feature image corresponding to the (n-1)th generated first feature image. Where n is an integer greater than 1 and less than or equal to the total number of first feature images, the size of the second and third feature images corresponding to the nth generated first feature image is C*W*H, where C is the number of channels, W is the width of the image, and H is the height of the image, the size of the fourth feature image is C*2W*2H, the size of the (n-1)th generated first feature image is C*2W*2H, the size of the fifth feature image is 2C*2W*2H, and the size of the second feature image corresponding to the (n-1)th generated first feature image is C*2W*2H.

[0026] Before feature fusion, the third feature image is bilinearly interpolated to obtain a fourth feature image with the same size as the (n-1)th generated first feature image, ensuring smooth feature fusion. After feature fusion, the fifth feature image generated by feature fusion is convolved to obtain a second feature image with the same size as the (n-1)th generated first feature image. This ensures that the input and output feature images have the same size, facilitating the subsequent determination of the distribution of human figures in the image to be detected based on the second feature image, thus enabling smooth human figure detection.

[0027] Optionally, when determining the distribution of human figures in the image to be detected based on the second feature image, the receptive field of each second feature image is first enhanced to obtain a corresponding fifth feature image. Then, feature fusion is performed on each fifth feature image to obtain a sixth feature image. Finally, the distribution of human figures in the image to be detected is determined based on the sixth feature image.

[0028] Since different human figures in the image to be processed have different sizes, the receptive field of the second feature image is enhanced to obtain the fifth feature image, thereby increasing the reference area of ​​the human figures in the fifth feature image in the image to be detected. Because the reference area of ​​the fifth feature image in the image to be detected is increased, the ability to detect human figures of different sizes in the image to be detected can be improved when determining the distribution of human figures in the image to be detected based on the fifth feature image, thereby improving the accuracy of detecting the number and distribution of human figures in the image to be detected.

[0029] Optionally, when performing receptive field enhancement processing on the second feature image to obtain the fifth feature image, for each second feature image, the second feature image is subjected to three convolutional processes to obtain the seventh feature image, two convolutional processes to obtain the eighth feature image, and one convolutional process to obtain the ninth feature image. Then, the seventh, eighth, and ninth feature images are fused to obtain the tenth feature image. Finally, the tenth feature image is convolved to obtain the fifth feature image corresponding to the second feature image. Wherein, for any given second feature image, the size of the second feature image and the corresponding fifth, seventh, eighth, and ninth feature images is C*W*H, and the size of the corresponding tenth feature image is 3C*W*H.

[0030] By performing convolution on the second feature image at different times, the seventh, eighth, and ninth feature images are obtained. Feature fusion is then performed on these three images to obtain the tenth feature image. Finally, convolution is performed on the tenth feature image to obtain the fifth feature image, which corresponds to the second feature image. Since the seventh, eighth, and ninth feature images are obtained by performing convolution on the second feature image at different times, the fifth feature image, based on these three images, has a stronger receptive field than the second feature image. This allows for accurate detection of people of different sizes in the image to be detected, thus ensuring the accuracy of detecting the number and distribution of people in the image.

[0031] Optionally, when determining the distribution of human figures in the image to be detected based on the sixth feature image, the sixth feature image is first normalized. Then, the normalized sixth feature image is input into the pre-trained first classifier, second classifier, and third classifier respectively to obtain the center point information output by the first classifier, the first image frame information output by the second classifier, and the second image frame information output by the third classifier. Then, the distribution of human figures in the image to be detected is determined based on the center point information, the first image frame information, and the second image frame information.

[0032] Center point information indicates the coordinates of the center point of the human head in the image to be detected. First image frame information includes the coordinates of the rectangles used to annotate the human head in the image to be detected. Second image frame information includes the coordinates of the rectangles used to annotate the human body in the image to be detected. Based on the center point information and the first image frame information, the position of the human head in the image to be detected can be determined. Based on the second image frame information, the position of the human body in the image to be detected can be determined. Furthermore, based on the number of rectangles annotating the human head or the number of rectangles annotating the human body, the number of human figures in the image to be detected can be determined. Based on the positions of the rectangles annotating the human head and the human body in the image to be detected, the distribution of human figures in the image to be detected can be determined. By annotating the human head and human body in the image to be processed with rectangles, the number and distribution of human figures in the image to be detected can be determined more accurately. This, in turn, allows for a more accurate determination of the number and distribution of people in a given location, thus improving the user experience.

[0033] Optionally, for any of the above aspects, the normalized sixth feature image can be input into the fourth classifier to obtain the image frame quality information output by the fourth classifier. The image frame quality information is used to indicate the accuracy of the rectangular box used to annotate the human head in the image to be detected. Then, the target center point is selected from the center point information based on the image frame quality information. If the accuracy of the rectangular box used to annotate the human head corresponding to the target center point is less than a preset accuracy threshold, the coordinate value of the target center point is deleted from the center point information.

[0034] Since each bounding box determined by the second classifier corresponds to a center point coordinate in the center point information, when it is determined that a bounding box determined by the second classifier cannot accurately label the human head in the image to be detected, the center point coordinate corresponding to the bounding box is deleted from the center point information, thereby discarding the bounding box that failed to accurately label the human head in the image to be detected, avoiding misidentification of human faces, and thus further improving the accuracy of detecting the number and distribution of human faces in the image to be detected. Attached Figure Description

[0035] Figure 1 This is a flowchart of a human face detection method provided in Embodiment 1 of this application;

[0036] Figure 2 This is a schematic diagram of a feature fusion method provided in Embodiment 2 of this application;

[0037] Figure 3 This is a flowchart of another feature fusion method provided in Embodiment 2 of this application;

[0038] Figure 4 This is a flowchart of a feature fusion method provided in Embodiment 2 of this application;

[0039] Figure 5 This is a schematic diagram of a human face detection method provided in Embodiment 3 of this application;

[0040] Figure 6 This is a schematic diagram of a receptive field enhancement processing method provided in Embodiment 3 of this application;

[0041] Figure 7 This is a flowchart of a method for determining the number and distribution of human images provided in Embodiment 3 of this application;

[0042] Figure 8 This is a schematic diagram of a human face detection device provided in Embodiment 4 of this application;

[0043] Figure 9 This is a schematic diagram of another human face detection device provided in Embodiment 4 of this application;

[0044] Figure 10 This is a schematic diagram of another human face detection device provided in Embodiment 4 of this application;

[0045] Figure 11 This is a schematic diagram of another human face detection device provided in Embodiment 4 of this application;

[0046] Figure 12 This is a schematic diagram of an electronic device provided in Embodiment 5 of this application.

[0047] List of reference numerals in the attached diagram:

[0048] 100: Human face detection methods; 400: Feature fusion methods

[0049] 700: Methods for determining the number and distribution of human faces; 800: Human face detection device.

[0050] 1200: Electronic Equipment A0-A N First feature image

[0051] B0-B N Second feature image C0-C NFifth feature image D: Sixth feature image

[0052] B i Second feature image B i11 B i12 B i21 Feature Image B i13 : Seventh feature image

[0053] B i22 : Eighth feature image B i31 : Ninth Feature Image B i123 : Tenth feature image

[0054] C i Fifth feature image 801: Acquisition module 802: Generation module

[0055] 803: Fusion Module; 804: Detection Module; 8031: Convolution Sub-Module

[0056] 8032: First Fusion Submodule; 8041: Enhancement Submodule; 8042: Second Fusion Submodule

[0057] 8043: Detection Submodule; 805: Calculation Module; 806: Filtering Module

[0058] 807: Module to be removed; 1202: Processor; 1204: Communication interface

[0059] 1206: Memory; 1208: Communication bus; 1210: Program

[0060] 101: Obtain an image to be detected.

[0061] 102: Generate at least two first feature images of the image to be detected.

[0062] 103: Perform feature fusion on each of the first feature images to obtain at least two second feature images.

[0063] 104: Based on each second feature image, determine the distribution of human figures in the image to be detected.

[0064] 401: Input the second feature image corresponding to the nth generated first feature image.

[0065] 402: Perform convolution processing on the second feature image corresponding to the nth first feature image to obtain the third feature image.

[0066] 403: Perform bilinear interpolation on the third feature image to obtain the fourth feature image.

[0067] 404: Perform feature fusion between the fourth feature image and the (n-1)th generated first feature image to obtain the fifth feature image.

[0068] 405: Perform convolution processing on the fifth feature image to obtain the second feature image corresponding to the (n-1)th first feature image.

[0069] 701: Input the normalized sixth feature image into the first classifier to obtain the center point information.

[0070] 702: Input the normalized sixth feature image into the second classifier to obtain the first image box information.

[0071] 703: Input the normalized sixth feature image into the third classifier to obtain the second image box information.

[0072] 704: Determine the distribution of the human image based on the center point information and the information from the first and second image frames. Detailed Implementation

[0073] As mentioned earlier, in high-traffic areas such as subway stations, train stations, and airports, it is necessary to determine the number and distribution of people gathered in these locations to prevent accidents such as falls onto platforms and stampedes. Currently, the number of people entering and leaving these locations is manually counted at the entrances and exits. This method only determines the number of people gathered in the location, not their distribution. Determining the distribution requires on-site manual inspection. Furthermore, high-traffic areas typically have multiple entrances and exits; for example, subway stations usually have four entrances and exits, and train stations have multiple entrances and exits. Manually counting people requires personnel at each entrance and exit, resulting in significant manpower costs and high overall costs associated with determining the number of people gathered in these locations.

[0074] In this embodiment, for a location where the number and distribution of gathered people need to be determined, an image including human figures is collected from the location. Features are extracted from the image to obtain multiple first feature images. Then, feature fusion is performed on each first feature image to obtain multiple second feature images. The distribution of human figures in the image to be detected is then determined based on each second feature image. Since the image to be detected is collected from the location where the number and distribution of gathered people need to be determined, the number and distribution of gathered people in the location can be determined based on the distribution of human figures in the image to be detected. Therefore, by collecting an image including human figures from the location where the number and distribution of gathered people need to be counted, and by processing the image to determine the distribution of human figures in the image, the number and distribution of gathered people in the location can be determined. This eliminates the need to have staff at each entrance and exit of the location to count people, thus saving manpower and reducing the cost of counting the number of gathered people in the location.

[0075] It should be noted that in this embodiment, feature images are extracted from the image to be detected. By performing various types of processing on the feature images, such as feature extraction, feature fusion, and receptive field enhancement, the number and distribution of people in the image to be detected are determined. Each feature image involved (including the first feature image, the second feature image, ... the Nth feature image, etc.) refers to the feature map in the convolutional layer.

[0076] The facial recognition method, apparatus, and electronic device provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0077] Example 1

[0078] Figure 1 This is a flowchart of a human face detection method 100 provided in Embodiment 1 of this application, as follows: Figure 1 As shown, the human face detection method 100 includes the following steps:

[0079] Step 101: Obtain an image to be detected.

[0080] The image to be detected is the image that needs to be used for facial recognition, and it contains at least one human image. When determining the number and distribution of people gathered in a high-traffic area, the image to be detected is an image from within that high-traffic area, such as one captured by a camera positioned high up within the area.

[0081] Step 102: Generate at least two first feature images of the image to be detected.

[0082] After obtaining the image to be detected, features are first extracted from the image to obtain a first feature image. Then, features are extracted from the obtained first feature image to obtain a new first feature image. That is, the first first feature image is obtained after extracting features from the image to be detected, and the next first feature image is obtained after extracting features from the previous first feature image.

[0083] For example, features are extracted from the image to be detected to obtain a first feature image 1, features are extracted from the first feature image 1 to obtain a first feature image 2, features are extracted from the first feature image 2 to obtain a first feature image 3, and features are extracted from the first feature image 3 to obtain a first feature image 4. That is, the first feature image 1 is obtained after extracting features from the image to be detected, the first feature image 2 is obtained after extracting features from the first feature image 1, the first feature image 3 is obtained after extracting features from the first feature image 2, and the first feature image 4 is obtained after extracting features from the first feature image 3.

[0084] Step 103: Perform feature fusion on each first feature image to obtain at least two second feature images.

[0085] After acquiring multiple first feature images, feature fusion is performed on two or more first feature images to obtain at least two second feature images, wherein different second feature images are obtained by fusing at least two first feature images that are not completely identical.

[0086] The purpose of feature fusion is to combine the features extracted from an image into a single feature that is more discriminative than the input. Specifically, it involves fusing features from at least two first feature images to obtain a second feature image that is more discriminative than each of the first feature images used individually. Feature fusion can employ either a sequential feature fusion strategy or a parallel feature fusion strategy. A sequential feature fusion strategy directly concatenates two features; if the dimensions of the two input features x and y are p and q, the dimension of the output feature z is p+q. A parallel feature fusion strategy combines two feature vectors into a negative vector; for input features x and y, the output feature z = x + iy, where i is an imaginary unit.

[0087] It should be noted that, in addition to the series feature fusion strategy or parallel feature fusion strategy mentioned above, other types of feature fusion methods can also be used to obtain the second feature image by performing feature fusion on the first feature image. The specific method of feature fusion is not limited in the embodiments of this application.

[0088] Step 104: Determine the distribution of human figures in the image to be detected based on each second feature image.

[0089] Since the second feature image is obtained by feature fusion of the first feature image, and the first feature image is extracted directly or indirectly from the image to be detected, the second feature image includes information reflecting the position, outline, and size of the human figure in the image to be detected. Therefore, the distribution of the human figure in the image to be detected can be determined based on each second feature image.

[0090] In this embodiment, after acquiring the image to be detected, which includes human images, features are extracted from the image to be detected to obtain multiple first feature images. Then, the first feature images are fused to obtain multiple second feature images. The distribution of human images in the image to be detected is then determined based on the second feature images. Since the image to be detected can be collected from the corresponding location, the human images in the image to be detected can be mapped to the corresponding location. Thus, based on the distribution of human images in the image to be detected, the number of people gathered in the corresponding location and the distribution of people are determined. This achieves automatic detection of the number of people gathered and the distribution of people, eliminating the need to equip each entrance and exit of the location with staff to count the number of people, thereby saving manpower and reducing the cost of counting the number of people gathered in the location.

[0091] It should be noted that since the first feature image is obtained by extracting features from the image to be detected, and each subsequent feature image is obtained by extracting features from the previous one, the later feature images are obtained, the higher their order. Higher-order feature images have stronger semantic information but lower resolution and poorer ability to perceive details, causing small objects to be lost in higher-order feature images. By fusing features from different feature images to obtain the second feature image, it is ensured that the second feature image includes higher-order semantic information without losing small objects. This ensures that smaller human figures in the image to be detected can be identified, thereby ensuring the accuracy of detecting the number and distribution of human figures in the image to be detected.

[0092] It should also be noted that, in the embodiments of this application and subsequent embodiments, the distribution of human figures in the image to be detected may include the positional distribution of human figures in the image to be detected, and may include the number of human figures in the image to be detected.

[0093] Example 2

[0094] Based on the human image detection method 100 provided in Example 1, when performing feature fusion on the first feature image to obtain the second image, feature fusion can be performed on at least two adjacent first feature images according to the generation order of each first feature image to obtain at least two second feature images. Different second feature images are obtained by feature fusion of at least two first feature images that are not completely identical.

[0095] In this embodiment, since the subsequent first feature image is obtained by extracting features from the previous first feature image, small objects in the previous first feature image may be lost in the subsequent first feature image. According to the generation order of the first feature images, at least two adjacent first feature images are fused to obtain a second feature image, ensuring that the second feature image includes small objects. Thus, when determining the number and distribution of human figures in the image to be detected based on the second feature image, smaller human figures in the image to be detected can be identified, thereby ensuring the accuracy of detecting the number and distribution of human figures in the image to be detected.

[0096] In one example, following the generation order of the first feature images from front to back, the first feature images are first feature image 1, first feature image 2, first feature image 3, and first feature image 4. When generating the second feature image by feature fusion of the first feature images, feature fusion can be performed on first feature image 1 and first feature image 2, feature fusion can be performed on first feature image 2 and first feature image 3, feature fusion can be performed on first feature image 3 and first feature image 4, feature fusion can be performed on first feature image 1, first feature image 2 and first feature image 3, feature fusion can be performed on first feature image 2, first feature image 3 and first feature image 4, and feature fusion can be performed on first feature image 1, first feature image 2, first feature image 3 and first feature image 4. Each feature fusion can obtain a second feature image.

[0097] It should be understood that when performing feature fusion on first feature image 1, first feature image 2, and first feature image 3, feature fusion can be performed on first feature image 1 and first feature image 2 first, and then the feature fusion result can be fused with first feature image 3 to obtain the second feature image. Similarly, when performing feature fusion on first feature image 2, first feature image 3, and first feature image 4, feature fusion can be performed on first feature image 2 and first feature image 3 first, and then the feature fusion result can be fused with first feature image 4 to obtain the second feature image. Finally, when performing feature fusion on first feature image 1, first feature image 2, first feature image 3, and first feature image 4, feature fusion can be performed on first feature image 1 and first feature image 2 first to obtain feature fusion result 1, then feature fusion result 1 can be fused with first feature image 3 to obtain feature fusion result 2, and finally feature fusion result 2 can be fused with first feature image 4 to obtain the second feature image.

[0098] In one possible implementation, when performing feature fusion on the first feature image to generate the second feature image, the second feature image corresponding to the subsequent first feature image can be fused with the previous first feature image to obtain the second feature image corresponding to the previous first feature image. Figure 2 This is a schematic diagram of a feature fusion method provided in Embodiment 2 of this application, as shown below. Figure 2 As shown, there are a total of N first feature images. According to the generation order of each first feature image, the first first feature image A1 is extracted from the image to be detected A0, and the nth first feature image A1 is extracted from the image to be detected A0. n From the first feature image A n-1 Extracted from, where n is an integer greater than 1 and less than or equal to N. For the Nth generated first feature image A... N Perform convolution processing to obtain the first feature image A generated by the Nth generation. N The corresponding second feature image B N The first feature image A generated by the nth generation will be compared with this image. n The corresponding second feature image B n , and the first feature image A generated at the (n-1)th time n-1 Perform feature fusion to obtain the first feature image A generated at the (n-1)th time. n-1 The corresponding second feature image B n-1 .

[0099] Figure 3 This is a schematic diagram of another feature fusion method provided in Embodiment 2 of this application, as shown below. Figure 3 As shown, there are a total of four first feature images. Following the generation order of these first feature images, the first feature image A1 is extracted from the image to be detected A0, the second feature image A2 is extracted from the first feature image A1, the third feature image A3 is extracted from the first feature image A2, and the fourth feature image A4 is extracted from the first feature image A3. Convolution processing is performed on the first feature image A4 to obtain the second feature image B4 corresponding to it. Feature fusion is performed on the first feature image A3 and the second feature image B4 to obtain the second feature image B3 corresponding to the first feature image A3. Feature fusion is performed on the first feature image A2 and the second feature image B3 to obtain the second feature image B2 corresponding to the first feature image A2. Feature fusion is performed on the first feature image A1 and the second feature image B2 to obtain the second feature image B1 corresponding to the first feature image A1.

[0100] In this embodiment, the Nth generated first feature image is convolved to obtain a second feature image corresponding to the Nth first feature image. The second feature image corresponding to the nth generated first feature image is then fused with the (n-1)th generated first feature image to obtain a second feature image corresponding to the (n-1)th generated first feature image. This ensures that the sum of all the second feature images includes all the feature information in the image to be detected, thereby improving the comprehensiveness of human image recognition in the image to be detected and ensuring the accuracy of detecting the number and distribution of human images in the image to be detected.

[0101] In this embodiment, for the later-acquired first feature image, the corresponding second feature image has a lower resolution, includes fewer features, and is smaller in size. This allows for the rapid identification of larger human figures in the image to be detected. Conversely, for the earlier-acquired first feature image, the corresponding second feature image has a higher resolution, includes more features, and is larger in size. This allows for the identification of smaller human figures in the image to be detected. Therefore, the obtained second feature images have different resolutions. Lower-resolution second feature images include higher-order features and can be used to quickly identify larger human figures in the image to be detected, while higher-resolution second feature images include more image information and can be used to identify smaller human figures in the image to be detected. Thus, determining the distribution of human figures in the image to be detected using these second feature images not only improves the efficiency of human figure identification but also ensures the accuracy of human figure identification.

[0102] In one possible implementation, when fusing the second feature image corresponding to the nth generated first feature image with the (n-1)th generated first feature image to obtain the second feature image corresponding to the (n-1)th generated first feature image, bilinear interpolation can be used to make the first and second feature images of the feature fusion have the same size, so as to ensure that the feature fusion of the first and second feature images can be performed smoothly. Figure 4 This is a flowchart of a feature fusion method 400 provided in Embodiment 2 of this application, as follows: Figure 4 As shown, the feature fusion method 400 includes the following steps:

[0103] Step 401: Input the second feature image corresponding to the nth generated first feature image.

[0104] The size of the second feature image corresponding to the nth generated first feature image is C*W*H, where C is the number of channels, W is the width of the image, and H is the height of the image.

[0105] It should be noted that defining the size of the second feature image corresponding to the nth generated first feature image as C*W*H is only to illustrate the changes in the size and number of channels of each feature image during the feature fusion process, and is not to specifically limit the size and number of channels of the third feature image, because different first feature images have different sizes, and the second feature images corresponding to different first feature images also have different sizes.

[0106] Step 402: Perform convolution processing on the second feature image corresponding to the nth generated first feature image to obtain the third feature image.

[0107] When obtaining the second feature image corresponding to the n-1 generated first feature images, the second feature image corresponding to the nth generated first feature image is first convolved to obtain the third feature image. See also Figure 3 For example, when obtaining the second feature image B3 corresponding to the first feature image A3, the second feature image B4 is first convolved to obtain the third feature image.

[0108] When performing convolution processing on the second feature image corresponding to the nth generated first feature image, the size of the resulting third feature image is also C*W*H. Furthermore, when performing convolution processing on the second feature image corresponding to the nth generated first feature image, the size of the convolution kernel can be C*3*3.

[0109] Step 403: Perform bilinear interpolation on the third feature image to obtain the fourth feature image.

[0110] When the size of the second feature image corresponding to the nth generated first feature image is C*W*H, the size of the (n-1)th generated first feature image is C*2W*2H. In order to perform feature fusion with the (n-1)th generated first feature image, bilinear interpolation is performed on the third feature image to obtain a fourth feature image with a size of C*2W*2H.

[0111] When performing bilinear interpolation on a third feature image of size C*W*H, an upsampling layer bilinear interpolation can be performed on the third feature image to obtain a fourth feature image of size C*2W*2H.

[0112] Step 404: Perform feature fusion between the fourth feature image and the (n-1)th generated first feature image to obtain the fifth feature image.

[0113] When the size of the second feature image corresponding to the nth generated first feature image is C*W*H, the size of the (n-1)th generated first feature image is C*2W*2H, and the size of the fourth feature image is also C*2W*2H. By performing feature fusion on the fourth feature image and the (n-1)th generated first feature image, a fifth feature image with a size of 2C*2W*2H is obtained.

[0114] Step 405: Perform convolution processing on the fifth feature image to obtain the second feature image corresponding to the (n-1)th generated first feature image.

[0115] Since the size of the fifth feature image is 2C*2W*2H, and the size of the (n-1)th generated first feature image is C*2W*2H, the second feature image corresponding to the (n-1)th generated first feature image should have the same size as the (n-1)th generated first feature image. Therefore, the fifth feature image is convolved to obtain the second feature image corresponding to the (n-1)th generated first feature image with a size of C*2W*2H.

[0116] When performing convolution processing on the fifth feature image, the size of the convolution kernel can be C*3*3.

[0117] In this embodiment, before feature fusion, the third feature image is bilinearly interpolated to obtain a fourth feature image, ensuring that the fourth feature image has the same size as the (n-1)th generated first feature image, thus enabling successful feature fusion. After feature fusion, the fifth feature image is convolved to obtain a second feature image corresponding to the (n-1)th generated first feature image, ensuring that the second feature image corresponding to the (n-1)th generated first feature image has the same size as the (n-1)th generated first feature image. This guarantees that the input and output feature images have the same size, facilitating the subsequent determination of the distribution of human figures in the image to be detected based on the second feature image, thus enabling successful human figure detection.

[0118] Example 3

[0119] When an image to be detected contains multiple human figures, the size of the human figures in the image is uncertain due to the distance between the people and the image acquisition device. People closer to the image acquisition device have larger human figures in the image to be detected, while people farther away from the image acquisition device have smaller human figures in the image to be detected. In order to identify human figures of different sizes from the image to be detected, the receptive field of the second feature image can be enhanced. Then, the distribution of human figures in the image to be detected can be determined based on the second feature image after the receptive field enhancement.

[0120] Figure 5This is a schematic diagram of a human face detection method provided in Embodiment 3 of this application, as shown below. Figure 5 As shown, receptive field enhancement processing is performed on each second feature image to obtain the corresponding fifth feature image. Specifically, receptive field enhancement processing is performed on second feature image B1 to obtain fifth feature image C1, and receptive field enhancement processing is performed on second feature image B2 to obtain fifth feature image C2. n-1 The fifth feature image C is obtained by performing receptive field enhancement processing. n-1 For the second feature image B n The fifth feature image C is obtained by performing receptive field enhancement processing. n For the second feature image B N The fifth feature image C is obtained by performing receptive field enhancement processing. N After obtaining the fifth feature image corresponding to each second feature image, feature fusion is performed on each fifth feature image to obtain a sixth feature image D. Then, the distribution of human figures in the image to be detected is determined based on the sixth feature image D.

[0121] In this embodiment, since the distance between people and the image acquisition device varies, the size of the human figures in the image to be detected varies. By performing receptive field enhancement processing on the second feature image to obtain the fifth feature image, the reference area of ​​the human figures in the fifth feature image in the image to be detected can be increased. Therefore, when determining the distribution of human figures in the image to be detected based on the fifth feature image, the ability to detect human figures of different sizes in the image to be detected can be improved, thereby improving the accuracy of detecting the number and distribution of human figures in the image to be detected.

[0122] In one possible implementation, when performing receptive field enhancement processing on the second feature image to obtain the fifth feature image, the second feature image can be subjected to convolution processing of different numbers of times, and then the multiple feature images obtained through different numbers of convolution processing can be fused to obtain the fifth feature image.

[0123] Figure 6 This is a schematic diagram of a receptive field enhancement processing method provided in Embodiment 3 of this application, as shown below. Figure 6 As shown, for each second feature image B i The second feature image B is processed through three parallel convolutional processes. i Convolution processing is performed, and then the feature images obtained from the three parallel convolution processing flows are fused to obtain the fifth feature image C. i Here, the second feature image B is defined. i The dimensions are C*W*H, where C is the number of channels, W is the width of the image, and H is the height of the image.

[0124] In the first convolution processing step, the second feature image B is first processed by a convolution kernel of size C*3*3. i Perform convolution processing to obtain a feature image B of size C*W*H. i11 Then, the feature image B is processed by a convolution kernel of size C*3*3. i11 Perform convolution processing to obtain a feature image B of size C*W*H. i12 Then, the feature image B is processed by a convolution kernel of size C*3*3. i12 Perform convolution processing to obtain the seventh feature image B with size C*W*H. i13 .

[0125] In the second convolution processing step, the second feature image B is first processed by a convolution kernel of size C*3*3. i Perform convolution processing to obtain a feature image B of size C*W*H. i21 Then, the feature image B is processed by a convolution kernel of size C*3*3. i21 Perform convolution processing to obtain the eighth feature image B with size C*W*H. i22 .

[0126] In the third convolution processing step, the second feature image B is processed by a convolution kernel of size C*3*3. i Perform convolution processing to obtain the ninth feature image B with size C*W*H. i31 .

[0127] It should be noted that in the above three convolution processing flows, the size of the convolution kernel used in a total of 6 convolution processing flows is C*3*3. The 6 convolution processing flows can use the same or different convolution kernels, or some of the convolution processing flows can use the same convolution kernel. This application embodiment does not limit this.

[0128] In obtaining the seventh feature image B i13 Eighth feature image B i22 and the ninth feature image B i31 Then, for the seventh feature image B i13 Eighth feature image B i22 and the ninth feature image B i31 Feature fusion is performed to obtain the tenth feature image B with a size of 3C*W*H. i123 Then, a convolutional kernel of size C*1*1 is used to convolve the tenth feature image to obtain the result that is convolved with the second feature image B. i The corresponding fifth feature image C i Fifth feature image C i Size and second feature image B i The same, also C*W*H.

[0129] In this embodiment, by performing convolution processing on the second feature image at different times, a seventh feature image, an eighth feature image, and a ninth feature image are obtained. After feature fusion of the seventh, eighth, and ninth feature images to obtain a tenth feature image, convolution processing is then performed on the tenth feature image to obtain a fifth feature image with the same size as the second feature image. This results in the fifth feature image having a stronger receptive field than the second feature image, thereby enabling accurate detection of people of different sizes in the image to be detected based on the fifth feature image. This ensures the accuracy of detecting the number and distribution of people in the image to be detected.

[0130] In one possible implementation, when determining the distribution of human figures in the image to be detected based on the sixth feature image, the sixth feature image can be input into multiple pre-trained classifiers. Each classifier determines the coordinates of the center point of the human figure in the image to be detected and the bounding box of the human figure. Then, based on the coordinates of the center point of the human figure and the bounding box of the human figure, the distribution of human figures in the image to be detected can be determined.

[0131] Figure 7 This is a flowchart of a method 700 for determining the number and distribution of human images according to Embodiment 3 of this application, as shown below. Figure 7 As shown, the method 700 for determining the number and distribution of human figures includes the following steps:

[0132] Step 701: Input the normalized sixth feature image into the first classifier to obtain the center point information output by the first classifier.

[0133] After obtaining the sixth feature image, it is first normalized to facilitate its input into a pre-trained classifier. The classifier uses the normalized sixth feature image to identify human figures in the target image. Specifically, the sixth feature image can be normalized using a group normalization process.

[0134] A first classifier is pre-trained using image samples. This classifier determines the coordinates of the center point of the human head in the corresponding original image based on the input feature image. After normalizing the sixth feature image, this normalized image is input into the first classifier to obtain the center point information output by the classifier. This center point information indicates the coordinates of the center point of the human head in the image to be detected. Based on the center point coordinates output by the first classifier, the center point of the human head can be marked on the image to be detected.

[0135] Step 702: Input the normalized sixth feature image into the second classifier to obtain the first image box information output by the second classifier.

[0136] A second classifier is pre-trained using image samples. This second classifier determines the bounding box used to annotate the human head in the corresponding original image based on the input feature image. After normalizing the sixth feature image, the normalized sixth feature image is input into the second classifier to obtain the first image bounding box information output by the second classifier. This first image bounding box information includes the coordinate values ​​of the bounding box used to annotate the human head in the image to be detected.

[0137] In one example, the first image frame information includes the coordinates of the top left corner and the bottom right corner of the image frame. These coordinates are offset values ​​relative to the center point of the human head. Since the image frame defined by the first image frame information is used to label the human head in the image to be detected, by combining the center point information output by the first classifier and the first image frame information, the human head can be labeled on the image to be detected by a rectangular box.

[0138] Step 703: Input the normalized sixth feature image into the third classifier to obtain the second image frame information output by the third classifier.

[0139] A third classifier is pre-trained using sample images. This classifier determines the bounding boxes used to label the human body in the corresponding original image based on the input feature image. After normalizing the sixth feature image, the normalized sixth feature image is input into the third classifier to obtain the second image bounding box information output by the third classifier. The second image bounding box information includes the coordinate values ​​of the bounding boxes used to label the human body in the image to be detected.

[0140] In one example, the second image frame information includes the coordinates of the top left corner and the bottom right corner of the image frame. Since the image frame defined by the second image frame information is used to annotate the human body in the image to be detected, each human body can be annotated on the image to be detected by rectangles based on the second image frame information.

[0141] It should be noted that due to occlusion or other reasons, the image to be detected may not include a complete human body. For example, the image to be detected may only include the head or only the head and upper body. The third classifier trained by the image samples can predict the position of the entire human body in the image to be detected based on the head, and then output the coordinate values ​​of the image box used to label the human body in the image to be detected.

[0142] Step 704: Determine the distribution of human figures in the image to be detected based on the center point information, the first image frame information, and the second image frame information.

[0143] Since the center point information is used to indicate the coordinates of the center point of the human head in the image to be detected, the first image frame information is used to indicate the rectangular frame marking the human head in the image to be detected, and the second image frame information is used to indicate the rectangular frame marking the human body in the image to be detected, the position of the human head in the image to be detected can be determined based on the center point information and the first image frame information, and the position of the human body in the image to be detected can be determined based on the second image frame information. Furthermore, the number of human heads or human bodies in the image to be detected can be determined based on the number of human images in the image to be detected, and the distribution of human images in the image to be detected can be determined based on the position of the human heads and the position of the human bodies in the image to be detected.

[0144] In this embodiment, multiple classifiers are pre-trained. The normalized sixth feature image is input into each classifier to obtain the center point information, first image frame information, and second image frame information output by each classifier. Based on the center point information and the first image frame information, the position of the human head in the image to be detected can be determined. Based on the second image frame information, the position of the human body in the image to be detected can be determined. Furthermore, based on the number of rectangles annotating the human head or the number of rectangles annotating the human body, the number of human figures in the image to be detected can be determined. Based on the positions of the rectangles annotating the human head and the rectangles annotating the human body in the image to be detected, the distribution of human figures in the image to be detected can be determined. Determining the number and distribution of human figures in the image to be detected based on the center point information, and using coordinate deviation values ​​to determine the rectangles annotating the human head, can improve the computational speed of the second classifier. Based on the bounding boxes labeled with human heads and human bodies, the number and distribution of human figures in the image to be detected can be determined. This avoids conflicts between human head features and human body features, thus enabling a more accurate identification of the human figures in the image to be detected. Furthermore, by mapping the bounding boxes labeled with human heads and standard human bodies in the image to the corresponding locations, the number and distribution of people gathered in the corresponding locations can be accurately determined.

[0145] Optionally, in Figure 7Based on the method 700 for determining the number and distribution of human figures, a fourth classifier can be pre-trained using image samples. This fourth classifier determines the accuracy of the bounding boxes used to annotate the human head in the corresponding original image, based on the input feature image. After normalizing the sixth feature image, the normalized image is input into the fourth classifier to obtain the image frame quality information output by the classifier. This image frame quality information indicates the accuracy of the bounding boxes used to annotate the human head in the image to be detected. After obtaining the image frame quality information, a target center point can be determined from the center point information. If the accuracy of the bounding box corresponding to the target center point is less than a preset accuracy threshold, the coordinates of the target center point are removed from the center point information.

[0146] In this embodiment, a pre-trained fourth classifier is used to detect whether the bounding boxes determined by the second classifier can accurately label the human head. Since each bounding box determined by the second classifier corresponds to a center point coordinate in the center point information, when it is determined that a bounding box determined by the second classifier cannot accurately label the human head in the image to be detected, the center point coordinate corresponding to the bounding box is deleted from the center point information, thereby discarding the bounding box that failed to accurately label the human head in the image to be detected, avoiding misidentification of human faces, and thus further improving the accuracy of detecting the number and distribution of human faces in the image to be detected.

[0147] Example 4

[0148] Figure 8 This is a schematic diagram of a human face detection device 800 provided in Embodiment 4 of this application, as shown below. Figure 8 As shown, the human face detection device 800 includes:

[0149] The acquisition module 801 is used to acquire an image to be detected, wherein the image to be detected includes at least one human image;

[0150] The generation module 802 is used to generate at least two first feature images of the image to be detected obtained by the acquisition module 801, wherein the first first feature image is obtained by extracting features from the image to be detected, and the second first feature image is obtained by extracting features from the previous first feature image.

[0151] The fusion module 803 is used to perform feature fusion on at least two first feature images generated by the generation module 802 to obtain at least two second feature images;

[0152] The detection module 804 is used to determine the distribution of human figures in the image to be detected based on at least two second feature images obtained by the fusion module 803.

[0153] In this embodiment, the acquisition module 801 can be used to execute step 101 in the first embodiment above, the generation module 802 can be used to execute step 102 in the first embodiment above, the fusion module 803 can be used to execute step 103 in the first embodiment above, and the detection module 804 can be used to execute step 104 in the first embodiment above.

[0154] In one possible implementation, such as Figure 8 As shown, the fusion module 803 is used to perform feature fusion on at least two adjacent first feature images according to the generation order of each first feature image to obtain at least two second feature images, wherein different second feature images are obtained by feature fusion of at least two first feature images that are not completely the same.

[0155] Figure 9 This is a schematic diagram of another human face detection device 800 provided in Embodiment 4 of this application, as shown below. Figure 9 As shown, the fusion module 803 includes:

[0156] The convolution submodule 8031 ​​is used to perform convolution processing on the Nth generated first feature image according to the generation order of each first feature image to obtain a second feature image corresponding to the Nth generated first feature image, where N is the number of first feature images;

[0157] The first fusion submodule 8032 is used to fuse the second feature image corresponding to the nth generated first feature image obtained by the convolution submodule 8031 ​​with the (n-1)th generated first feature image to obtain the second feature image corresponding to the (n-1)th generated first feature image, where n is an integer greater than 1 and less than or equal to N.

[0158] In one possible implementation, such as Figure 9 As shown, the first fusion submodule 8032 is used to perform the following operations:

[0159] The second feature image corresponding to the nth generated first feature image is convolved to obtain the third feature image. The dimensions of the second feature image and the third feature image corresponding to the nth generated first feature image are both C*W*H, where C is the number of channels, W is the width of the image, and H is the height of the image.

[0160] The third feature image is subjected to bilinear interpolation to obtain the fourth feature image, wherein the size of the fourth feature image is C*2W*2H;

[0161] The fourth feature image is fused with the (n-1)th generated first feature image to obtain the fifth feature image, wherein the size of the (n-1)th generated first feature image is C*2W*2H and the size of the fifth feature image is 2C*2W*2H.

[0162] The fifth feature image is convolved to obtain the second feature image corresponding to the (n-1)th generated first feature image, wherein the size of the second feature image corresponding to the (n-1)th generated first feature image is C*2W*2H.

[0163] Figure 10 This is a schematic diagram of another human face detection device 800 provided in Embodiment 4 of this application, as shown below. Figure 10 As shown, the detection module 804 includes:

[0164] The enhancement submodule 8041 is used to perform receptive field enhancement processing on each second feature image to obtain the corresponding fifth feature image;

[0165] The second fusion submodule 8042 is used to perform feature fusion on the fifth feature images obtained by the enhancement submodule 8041 to obtain a sixth feature image;

[0166] The detection submodule 8043 is used to determine the distribution of human figures in the image to be detected based on the sixth feature image obtained by the second fusion submodule 8042.

[0167] In one possible implementation, such as Figure 10 As shown, the enhancement submodule 8041 performs the following processing for each second feature image:

[0168] The second feature image is subjected to three convolution processes to obtain the sixth feature image. The dimensions of the second and sixth feature images are both C*W*H, where C is the number of channels, W is the width of the image, and H is the height of the image.

[0169] The second feature image is subjected to two convolution processes to obtain the seventh feature image, wherein the size of the seventh feature image is C*W*H;

[0170] The second feature image is convolved once to obtain the eighth feature image, where the size of the eighth feature image is C*W*H;

[0171] The sixth, seventh, and eighth feature images are fused to obtain the ninth feature image, which has a size of 3C*W*H.

[0172] The ninth feature image is convolved to obtain the fifth feature image corresponding to the second feature image, wherein the size of the fifth feature image is C*W*H.

[0173] In one possible implementation, such as Figure 10 As shown, the detection submodule 8043 is used to perform the following processing:

[0174] The normalized sixth feature image is input into the first classifier to obtain the center point information output by the first classifier. The first classifier is used to determine the center point coordinates of the human head in the original image corresponding to the input feature image. The center point information is used to indicate the center point coordinates of the human head in the image to be detected.

[0175] The normalized sixth feature image is input into the second classifier to obtain the first image box information output by the second classifier. The second classifier is used to determine the rectangular box used to mark the human head in the original image corresponding to the feature image based on the input feature image. The first image box information includes the coordinate values ​​of the rectangular box used to mark the human head in the image to be detected.

[0176] The normalized sixth feature image is input into the third classifier to obtain the second image box information output by the third classifier. The third classifier is used to determine the rectangular box used to mark the human body in the original image corresponding to the feature image based on the input feature image. The second image box information includes the coordinate values ​​of the rectangular box used to mark the human body in the image to be detected.

[0177] Based on the center point information, the first image frame information, and the second image frame information, the distribution of human figures in the image to be detected is determined.

[0178] Figure 11 This is a schematic diagram of another human face detection device 800 provided in Embodiment 4 of this application, as shown below. Figure 11 As shown, the human face detection device 800 also includes:

[0179] The calculation module 805 is used to input the normalized sixth feature image into the fourth classifier to obtain the image frame quality information output by the fourth classifier. The fourth classifier is used to determine the accuracy of the rectangular box used to annotate the human head in the original image corresponding to the feature image based on the input feature image. The image frame quality information is used to indicate the accuracy of the rectangular box used to annotate the human head in the image to be detected.

[0180] The filtering module 806 is used to determine the target center point from the center point information based on the image frame quality information obtained by the calculation module 805, wherein the accuracy of the rectangle used to annotate the human head corresponding to the target center point is less than a preset accuracy threshold.

[0181] The deletion module 807 is used to delete the coordinate values ​​of the target center point determined by the filtering module 806 from the center point information.

[0182] It should be noted that the information interaction and execution process between the modules and sub-modules in the above-mentioned face detection device are based on the same concept as the aforementioned face detection method embodiment. For details, please refer to the description in the aforementioned face detection method embodiment, and it will not be repeated here.

[0183] Example 5

[0184] Figure 12 This is a schematic diagram of an electronic device provided in Embodiment 5 of this application. The specific embodiments of this application do not limit the specific implementation of the electronic device. See also... Figure 12 The electronic device 1200 provided in this application embodiment includes: a processor 1202, a communications interface 1204, a memory 1206, and a communication bus 1208. Wherein:

[0185] The processor 1202, communication interface 1204, and memory 1206 communicate with each other via communication bus 1208.

[0186] Communication interface 1204 is used for communication with other electronic devices or servers.

[0187] The processor 1202 is used to execute program 1210, which can specifically execute the relevant steps in any of the aforementioned portrait detection method embodiments.

[0188] Specifically, program 1210 may include program code that includes computer operation instructions.

[0189] The processor 1202 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The smart device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.

[0190] Memory 1206 is used to store program 1210. Memory 1206 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0191] Specifically, program 1210 can be used to cause processor 1202 to execute the human face detection method in any of the foregoing embodiments.

[0192] The specific implementation of each step in program 1210 can be found in the corresponding steps and units described in any of the foregoing embodiments of the human face detection method, and will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.

[0193] The electronic device of this application acquires an image to be detected, including human images, extracts features from the image to be detected to obtain multiple first feature images, and then fuses the first feature images to obtain multiple second feature images. The distribution of human images in the image to be detected is then determined based on the second feature images. Since the image to be detected can be collected from the corresponding location, the human images in the image to be detected can be mapped to the corresponding location. Thus, the number and distribution of people gathered in the corresponding location can be determined based on the distribution of human images in the image to be detected. This achieves automatic detection of the number and distribution of people gathered in the location, eliminating the need to equip each entrance and exit of the location with staff to count the number of people, thereby saving manpower and reducing the cost of counting the number of people gathered in the location.

[0194] This application also provides a computer-readable storage medium storing instructions for causing a machine to perform the image detection method as described herein. Specifically, a system or apparatus equipped with a storage medium storing software program code that implements the functions of any of the embodiments described above, and enabling the computer (or CPU or MPU) of the system or apparatus to read and execute the program code stored in the storage medium.

[0195] In this case, the program code read from the storage medium can itself implement the function of any of the above embodiments, and therefore the program code and the storage medium storing the program code constitute part of this application.

[0196] Examples of storage media used to provide program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, program code can be downloaded from a server computer via a communication network.

[0197] Furthermore, it should be clear that not only can the program code read by the computer be executed, but also the operating system or other components operating on the computer can be instructed based on the program code to perform some or all of the actual operations, thereby realizing the function of any of the embodiments described above.

[0198] Furthermore, it is understood that the program code read from the storage medium is written to the memory set in the expansion board inserted into the computer or to the memory set in the expansion module connected to the computer. Then, based on the instructions of the program code, the CPU or other components installed on the expansion board or expansion module execute some and all of the actual operations, thereby realizing the function of any of the above embodiments.

[0199] This application also provides a computer program product, which is tangibly stored on a computer-readable medium and includes computer-executable instructions. When executed, the computer-executable instructions cause at least one processor to perform the human face detection methods provided in the above embodiments. It should be understood that the solutions in this embodiment have the corresponding technical effects in the above method embodiments, which will not be repeated here.

[0200] It should be noted that not all steps and modules in the above processes and system structure diagrams are mandatory; some steps or modules can be omitted as needed. The execution order of each step is not fixed and can be adjusted as required. The system structure described in the above embodiments can be a physical structure or a logical structure. That is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or they may be jointly implemented by certain components in multiple independent devices.

[0201] In the above embodiments, the hardware modules can be implemented mechanically or electrically. For example, a hardware module may include permanent dedicated circuitry or logic (such as a dedicated processor, FPGA, or ASIC) to perform the corresponding operations. The hardware module may also include programmable logic or circuitry (such as a general-purpose processor or other programmable processor), which can be temporarily configured by software to perform the corresponding operations. The specific implementation method (mechanical, dedicated permanent circuitry, or temporarily configured circuitry) can be determined based on cost and time considerations.

[0202] The present application has been shown and described in detail above with reference to the accompanying drawings and preferred embodiments. However, the present application is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art will know that more embodiments of the present application can be obtained by combining the code review methods in the different embodiments above. These embodiments are also within the protection scope of the present application.

Claims

1. A human face detection method (100), characterized in that, include: Obtain (101) an image to be detected, wherein the image to be detected includes at least one human image; Generate at least two first feature images of the image to be detected (102), wherein the first first feature image is obtained by extracting features from the image to be detected, and the second first feature image is obtained by extracting features from the first first feature image; Feature fusion (103) is performed on the at least two first feature images to obtain at least two second feature images; Based on the at least two second feature images, determine (104) the distribution of human figures in the image to be detected; The feature fusion (103) of the at least two first feature images to obtain at least two second feature images includes: performing feature fusion on at least two adjacent first feature images according to the generation order of each first feature image to obtain at least two second feature images, wherein different second feature images are obtained by feature fusion of at least two first feature images that are not completely identical; The step of fusing features of at least two adjacent first feature images according to the generation order of each first feature image to obtain at least two second feature images includes: performing convolution processing on the Nth generated first feature image according to the generation order of each first feature image to obtain a second feature image corresponding to the Nth generated first feature image, where N is the number of first feature images; fusing the second feature image corresponding to the nth generated first feature image with the (n-1)th generated first feature image to obtain a second feature image corresponding to the (n-1)th generated first feature image, where n is an integer greater than 1 and less than or equal to N; The process of fusing the second feature image corresponding to the nth generated first feature image with the (n-1)th generated first feature image to obtain the second feature image corresponding to the (n-1)th generated first feature image includes: performing convolution processing (402) on the second feature image corresponding to the nth generated first feature image to obtain a third feature image, wherein the dimensions of the second feature image corresponding to the nth generated first feature image and the third feature image are both C*W*H, where C is the number of channels, W is the width of the image, and H is the height of the image; and performing bilinear interpolation processing (403) on the third feature image. The process involves: obtaining a fourth feature image, wherein the size of the fourth feature image is C*2W*2H; fusing the fourth feature image with the (n-1)th generated first feature image (404) to obtain a fifth feature image, wherein the size of the (n-1)th generated first feature image is C*2W*2H and the size of the fifth feature image is 2C*2W*2H; and performing convolution processing (405) on the fifth feature image to obtain a second feature image corresponding to the (n-1)th generated first feature image, wherein the size of the second feature image corresponding to the (n-1)th generated first feature image is C*2W*2H.

2. The method according to claim 1, characterized in that, Determining the distribution of human figures in the image to be detected based on the at least two second feature images includes: Each second feature image is subjected to receptive field enhancement processing to obtain the corresponding fifth feature image; The fifth feature images are fused to obtain a sixth feature image; Based on the sixth feature image, the distribution of human figures in the image to be detected is determined.

3. The method according to claim 2, characterized in that, The step of performing receptive field enhancement processing on each of the second feature images to obtain the corresponding fifth feature image includes: For each of the second feature images, perform the following: The second feature image is subjected to three convolution processes to obtain the seventh feature image. The dimensions of the second feature image and the seventh feature image are both C*W*H, where C is the number of channels, W is the width of the image, and H is the height of the image. The second feature image is subjected to two convolution processes to obtain the eighth feature image, wherein the size of the eighth feature image is C*W*H; The second feature image is convolved once to obtain the ninth feature image, wherein the size of the ninth feature image is C*W*H; The seventh feature image, the eighth feature image, and the ninth feature image are fused to obtain a tenth feature image, wherein the size of the tenth feature image is 3C*W*H; The tenth feature image is convolved to obtain the fifth feature image corresponding to the second feature image, wherein the size of the fifth feature image is C*W*H.

4. The method according to claim 2, characterized in that, The step of determining the distribution of human figures in the image to be detected based on the sixth feature image includes: The normalized sixth feature image is input into the first classifier (701) to obtain the center point information output by the first classifier. The first classifier is used to determine the center point coordinates of the human head in the original image corresponding to the feature image based on the input feature image. The center point information is used to indicate the center point coordinates of the human head in the image to be detected. The normalized sixth feature image is input into the second classifier (702) to obtain the first image frame information output by the second classifier. The second classifier is used to determine the rectangular frame used to mark the human head in the original image corresponding to the feature image based on the input feature image. The first image frame information includes the coordinate values ​​of the rectangular frame used to mark the human head in the image to be detected. The normalized sixth feature image is input into the third classifier (703) to obtain the second image box information output by the third classifier. The third classifier is used to determine the rectangular box used to mark the human body in the original image corresponding to the feature image based on the input feature image. The second image box information includes the coordinate values ​​of the rectangular box used to mark the human body in the image to be detected. Based on the center point information, the first image frame information and the second image frame information, the distribution of the human figure in the image to be detected is determined (704).

5. The method according to claim 4, characterized in that, The method further includes: The normalized sixth feature image is input into the fourth classifier to obtain the image frame quality information output by the fourth classifier. The fourth classifier is used to determine, based on the input feature image, the accuracy of the rectangular box used to annotate the human head in the original image corresponding to the feature image. The image frame quality information is used to indicate the accuracy of the rectangular box used to annotate the human head in the image to be detected. Based on the image frame quality information, a target center point is determined from the center point information, wherein the accuracy of the rectangular frame used to annotate the human head corresponding to the target center point is less than a preset accuracy threshold. Delete the coordinate value of the target center point from the center point information.

6. A human face detection device (800), characterized in that, Includes modules for performing the operations in the method as described in any one of claims 1-5.

7. An electronic device (1200), characterized in that, include: The system includes a processor (1202), a communication interface (1204), a memory (1206), and a communication bus (1208). The processor (1202), the memory (1206), and the communication interface (1204) communicate with each other through the communication bus (1208). The memory (1206) is used to store at least one executable instruction, which causes the processor (1202) to perform the operation corresponding to the human face detection method as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, cause the processor to perform the method of any one of claims 1-5.

9. A computer program product, characterized in that, The computer program product is tangibly stored on a computer-readable medium and includes computer-executable instructions that, when executed, cause at least one processor to perform the method according to any one of claims 1-5.