Face occlusion image detection method and device, equipment and storage medium

By introducing the Atrous Spatial Pyramid Pooling (ASPP) layer into the CenterNet network model, the feature extraction and detection capabilities are enhanced, the problem of insufficient precision in face occluded image detection is solved, and higher detection accuracy and robustness are achieved.

CN120599677APending Publication Date: 2025-09-05AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510672310.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The CenterNet algorithm has insufficient detection accuracy when dealing with face occlusion and overlapping objects, and is particularly prone to missing objects in occlusion situations.

Method used

The Atrous Spatial Pyramid Pooling (ASPP) layer is introduced into the CenterNet network model to enhance feature extraction capabilities. Semantic features are extracted through the backbone network, and multi-scale context information is obtained through the neck module and ASPP layer. The predictor is combined to predict the center point and size, thereby improving detection accuracy.

Benefits of technology

The accuracy and robustness of face occlusion image detection are improved, and it can better adapt to detection tasks under different scales and occlusion conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599677A_ABST
    Figure CN120599677A_ABST
Patent Text Reader

Abstract

The invention discloses a face occlusion image detection method and device, equipment and a storage medium, and the method comprises the steps: obtaining a to-be-detected face occlusion image, and inputting the face occlusion image to a target Center Net network model; wherein a check module in the target Center Net network model is integrated with an ASPP (Application Specific Private Protocol) layer; and through the target Center Net network model, extracting feature information corresponding to the face occlusion image, and according to a feature extraction result, outputting face detection information in the face occlusion image. According to the technical scheme provided by the embodiment of the invention, the detection performance of the target Center Net network model on the face occlusion image can be improved, and the accuracy of the detection result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method, device, equipment and storage medium for detecting face occlusion images. Background Art

[0002] In actual applications, faces are often blocked by various objects (such as masks, hats, glasses, scarves, etc.), resulting in the loss of some facial feature information, making it difficult for the face recognition system to make accurate judgments.

[0003] This occlusion situation greatly increases the difficulty of face recognition. Although the CenterNet algorithm is widely used in face recognition, it has limited processing of occlusion and overlapping objects. When there is occlusion or overlap between targets, the CenterNet algorithm may encounter some difficulties.

[0004] Because the CenterNet algorithm detects objects by predicting their center points, if the center points of two objects completely overlap or nearly overlap, the algorithm may not accurately detect the two objects. Since it does not provide a feature map for each category of objects, but uses a shared feature map for all objects, it may miss some objects in some cases. Summary of the Invention

[0005] The present invention provides a method, apparatus, device and storage medium for detecting face occluded images, which can improve the detection performance of a target CenterNet network model for face occluded images and improve the accuracy of the detection results.

[0006] According to one aspect of the present invention, a method for detecting face occlusion images is provided, the method comprising:

[0007] Obtain a face occlusion image to be detected, and input the face occlusion image into the target CenterNet network model;

[0008] The neck module in the target CenterNet network model integrates the atrous spatial pyramid pooling layer ASPP;

[0009] The target CenterNet network model is used to extract feature information corresponding to the face-occluded image, and face detection information in the face-occluded image is output based on the feature extraction result.

[0010] Optionally, extracting feature information corresponding to the face occluded image through the target CenterNet network model includes:

[0011] Extracting semantic features from face occluded images through the backbone network in the target CenterNet network model;

[0012] Inputting the semantic features corresponding to the face occlusion image into the neck module;

[0013] Through the neck module and the ASPP layer, multi-scale context information corresponding to the face occlusion image is obtained according to the semantic features, and a multi-scale fusion feature map corresponding to the face occlusion image is output according to the multi-scale context information.

[0014] Optionally, outputting face detection information in the face-occluded image based on the feature extraction result includes:

[0015] Input the multi-scale fusion feature map corresponding to the face occlusion image into the predictor in the target CenterNet network model;

[0016] According to the output result of the predictor, face detection information in the face occlusion image is determined.

[0017] Optionally, inputting the multi-scale fusion feature map corresponding to the face occlusion image into the predictor in the target CenterNet network model includes:

[0018] The multi-scale fusion feature map corresponding to the face occlusion image is input into the center point predictor, center point offset predictor and width and height predictor in the target CenterNet network model to obtain the center point position information, center point offset and target face width and height corresponding to the face occlusion image.

[0019] Optionally, determining face detection information in the face occlusion image according to an output result of the predictor includes:

[0020] Adjusting the center point position information corresponding to the face occlusion image according to the center point offset;

[0021] The frame vertex position information is determined based on the adjusted center point position information and the width and height of the target face, and a detection frame is generated in the face occlusion image based on the frame vertex position information.

[0022] Optionally, the ASPP layer is composed of a normal convolutional layer and a plurality of dilated convolutional layers with different expansion rates;

[0023] Obtaining multi-scale context information corresponding to the face occlusion image according to the semantic features through the neck module and the ASPP layer, including:

[0024] The semantic features are convolved through the neck module, the ordinary convolution layer in the ASPP layer, and multiple dilation-rate-varying dilation layers, and then the convolution results are added on the channels to obtain multi-scale contextual information corresponding to the face occluded image.

[0025] According to another aspect of the present invention, a facial occlusion image detection device is provided, the device comprising:

[0026] An image acquisition module is used to acquire an occluded face image to be detected and input the occluded face image into a target CenterNet network model;

[0027] The neck module in the target CenterNet network model integrates the atrous spatial pyramid pooling layer ASPP;

[0028] The image processing module is used to extract feature information corresponding to the face occlusion image through the target CenterNet network model, and output face detection information in the face occlusion image based on the feature extraction result.

[0029] According to another aspect of the present invention, an electronic device is provided, comprising:

[0030] at least one processor; and

[0031] a memory communicatively connected to the at least one processor; wherein,

[0032] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the face occlusion image detection method described in any embodiment of the present invention.

[0033] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the face occlusion image detection method described in any embodiment of the present invention when executed.

[0034] According to another aspect of the present invention, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the computer program implements the face occlusion image detection method according to any embodiment of the present invention.

[0035] The technical solution provided by the embodiment of the present invention obtains a face occlusion image to be detected, inputs the face occlusion image into a target CenterNet network model, and integrates an ASPP layer in the neck module of the target CenterNet network model. The feature information corresponding to the face occlusion image is extracted through the target CenterNet network model, and the face detection information in the face occlusion image is output according to the feature extraction result. This technical means can improve the detection performance of the target CenterNet network model for the face occlusion image and improve the accuracy of the detection results.

[0036] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0038] Figure 1 is a flowchart of a method for detecting face occlusion images according to an embodiment of the present invention;

[0039] Figure 2a is a flowchart of another method for detecting face occlusion images provided according to an embodiment of the present invention;

[0040] Figure 2b 1 is a schematic diagram of the structure of a target CenterNet network model provided according to an embodiment of the present invention;

[0041] Figure 2c 1 is a schematic structural diagram of a neck module in a target CenterNet network model provided according to an embodiment of the present invention;

[0042] Figure 2d 1 is a schematic diagram of the structure of a predictor in a target CenterNet network model provided according to an embodiment of the present invention;

[0043] Figure 3 2 is a schematic structural diagram of a face occlusion image detection device provided by an embodiment of the present invention;

[0044] Figure 4 It is a structural diagram of an electronic device for implementing the face occlusion image detection method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0045] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0046] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0047] Figure 1 This is a flowchart of a face occlusion image detection method provided by an embodiment of the present invention. This embodiment is applicable to the case of detecting occluded face images. The method can be performed by a face occlusion image detection device. The face occlusion image detection device can be implemented in the form of hardware and / or software. The device can be configured in an electronic device. Figure 1 As shown, the method includes:

[0048] Step 110: Obtain a face occlusion image to be detected, and input the face occlusion image into a target CenterNet network model.

[0049] In practical applications, the CenterNet network model is an anchor-free object detection architecture. Compared to traditional anchor-based detection methods, it avoids the extensive anchor generation and classification process, thereby improving detection speed and accuracy. The CenterNet network model directly locates the target by detecting its center point and regressing its size, making the entire detection process simpler and more efficient. The CenterNet network model has a wide range of applications, including facial and vehicle recognition, helping to implement smart access control and vehicle management, and can also be used in the medical field.

[0050] At the same time, the CenterNet network model also has some shortcomings. First, while the CenterNet network model has improved in speed and accuracy, its performance may still need further improvement compared to some more complex detection networks. Second, because the CenterNet network model relies on deep learning and convolutional neural networks, it requires a large amount of labeled data for training, which can be challenging in some cases. Furthermore, for some special or complex scenarios, the CenterNet network model may require further optimization and improvement to adapt to different needs.

[0051] Therefore, in order to improve the accuracy of face occlusion image detection results, this embodiment improves the CenterNet network model to obtain a target CenterNet network model. The neck module in the target CenterNet network model integrates the Atrous Spatial Pyramid Pooling (ASPP) layer. The neck module is located between the backbone network and the output layer in the target CenterNet network model. Its main function is to further process and integrate the features extracted by the backbone network to extract higher-level feature information and pass it to the subsequent classifier or detector.

[0052] Step 120: Extract feature information corresponding to the face occlusion image through the target CenterNet network model, and output face detection information in the face occlusion image based on the feature extraction result.

[0053] In this embodiment, after the face occlusion image is input into the target CenterNet network model, the backbone network in the target CenterNet network model can extract features of the face occlusion image and then input the extracted features into the neck module.

[0054] The ASPP layer in the neck module can capture multi-scale contextual information. By performing convolution operations at different void rates, ASPP can effectively extract features of different scales and pass them to the subsequent classifier or detector in the model. The classifier or detector outputs face detection information based on the feature extraction results.

[0055] The benefit of this setup is that by introducing the ASPP layer into the neck module, the target CenterNet network model can better cope with objects with large scale variations. Therefore, in complex scenes, the target CenterNet network model incorporating the ASPP layer can more accurately locate and identify objects.

[0056] Secondly, the ASPP layer enhances the feature representation capabilities of the target CenterNet network model. The multiple branches in the ASPP layer can extract features at different scales. These features complement each other in the subsequent fusion process, improving the feature representation capabilities of the entire network. This helps the target CenterNet network model more accurately identify the category and location of the object.

[0057] Finally, the Target CenterNet network model incorporating the ASPP layer also exhibits improved robustness. Because the ASPP layer can capture multi-scale contextual information, the Target CenterNet network model can still accurately identify and locate the target even when the target is occluded or deformed. This makes the Target CenterNet network model more reliable and stable in practical applications.

[0058] The technical solution provided by the embodiment of the present invention obtains a face occlusion image to be detected, inputs the face occlusion image into a target CenterNet network model, and integrates an ASPP layer in the neck module of the target CenterNet network model. The feature information corresponding to the face occlusion image is extracted through the target CenterNet network model, and the face detection information in the face occlusion image is output according to the feature extraction result. This technical means can improve the detection performance of the target CenterNet network model for the face occlusion image and improve the accuracy of the detection results.

[0059] Figure 2a A flowchart of another method for detecting face occlusion images provided by an embodiment of the present invention is shown in FIG. Figure 2a As shown, the method includes:

[0060] Step 210: Obtain a face occlusion image to be detected, and input the face occlusion image into a target CenterNet network model.

[0061] Step 220: Extract semantic features from the face occluded image through the backbone network in the target CenterNet network model.

[0062] In this embodiment, Figure 2b It can be a structural diagram of a target CenterNet network model, such as Figure 2bAs shown, after the face occlusion image is input into the target CenterNet network model, the backbone network of the model is used to extract high-dimensional semantic features from the input image, providing key information for subsequent target detection. Specifically, the backbone network can be a classic convolutional neural network, such as the Residual Neural Network (ResNet) or the Visual Geometry Group Network (VGG), etc., which is not limited in this embodiment.

[0063] Step 230: Input the semantic features corresponding to the face occlusion image into the neck module, obtain the multi-scale context information corresponding to the face occlusion image according to the semantic features through the neck module and the ASPP layer, and output the multi-scale fusion feature map corresponding to the face occlusion image according to the multi-scale context information.

[0064] In practical applications, the existing CenterNet performs three consecutive deconvolution operations with a stride of 2 on the features extracted by the backbone network. Each operation doubles the length and width of the feature map. The resulting feature map is eight times the size of the backbone network output and one-quarter the size of the input image. However, in applications such as face recognition, collected facial images come in a variety of sizes. Directly upsampling the backbone network output may lose information at other scales, which can affect the detection of occluded objects.

[0065] Compared with the prior art, this embodiment introduces the ASPP layer in the neck module of the model to process the output of the backbone network, which can retain more global information. In one implementation of this embodiment, Figure 2c This is a structural diagram of the neck module in a target CenterNet network model, as shown in Figure 2c As shown, the ASPP layer consists of a normal convolutional layer and multiple dilated convolutional layers with different expansion rates. Through the neck module and the ASPP layer, multi-scale contextual information corresponding to the face-occluded image is obtained based on the semantic features. This includes: convolving the semantic features through the neck module, the normal convolutional layers in the ASPP layer, and multiple dilated convolutional layers with different expansion rates, and then performing an addition operation on the convolution results on the channels to obtain the multi-scale contextual information corresponding to the face-occluded image.

[0066] This contextual information at different scales helps the predictor more accurately identify faces of varying sizes and improves detection of occluded objects. The improved neck module upsamples the fused multi-scale feature maps to generate high-resolution feature maps for final object detection. This improvement enables the target CenterNet network model to maintain its ability to localize small objects while also improving detection performance for multi-scale faces. By introducing the ASPP layer structure, the target CenterNet model is better suited to face detection tasks under varying sizes and occlusion conditions.

[0067] In a specific embodiment, the dilated convolution controls the change of the receptive field through the dilation rate dr. The elements in the convolution kernel will skip a fixed distance of pixels on the input feature map during the convolution process, and the size of this distance is determined by the dilation rate. Therefore, ordinary convolution can also be regarded as a dilated convolution with a dilation rate of 1. For two-dimensional dilated convolution, the calculation formula of the receptive field RF (width and height are the same) can be simplified to: RF = 1 + (k-1) × dr. The dilated convolution can increase the receptive field of the model without compressing the size of the feature map by setting the padding value. The corresponding receptive field size between the two layers (the receptive field of the previous layer is 1) is shown in the following formula, which is determined by two parameters: dilation rate dr and convolution kernel size k:

[0068] l=k+(k-1)*(dr-1)

[0069] Dilated convolutions with different dilation rates can achieve completely different receptive fields while maintaining the same output dimension. For example, when dr is 6, the receptive field size is 13x13; when dr is 12, the receptive field increases to 25x25; and when dr is 18, the receptive field further expands to 37x37. This expansion of the receptive field is crucial for capturing a wider range of contextual information, especially when processing large-scale objects in images. However, dilated convolutions are sparse in their information acquisition, selectively processing only a subset of pixels in the input feature map. When the dilation rate dr is close to the size of the input feature map, the convolution kernel may only cover a portion of the input image due to edge effects, resulting in information loss. To address this issue, in addition to dilated convolutions with different dilation rates, the ASPP layer also incorporates a global average pooling operation in parallel. Global average pooling captures global features from the entire input feature map, compensating for the sparsity of information acquisition by dilated convolutions and improving the model's ability to perceive global information.

[0070] In this embodiment, if Figure 2cAs shown in the figure, ASPP sets up four dilated convolution branches, and selects dilation rates of 3, 6, 12, and 18 respectively. Specifically, they may include 3×3 convolutions with a dilation rate of 3, 3×3 convolutions with a dilation rate of 6, 3×3 convolutions with a dilation rate of 12, and 3×3 convolutions with a dilation rate of 18. This setting can enhance the localization of large targets while maintaining good classification performance for small targets. Through the improved neck module, the model can better adapt to the characteristics of face occlusion image datasets and improve overall object detection performance. These parts are stacked together, and then a 1×1 convolution is used to adjust the number of channels, and finally the features output by ASPP are obtained. The main purpose of the ASPP layer is to capture contextual information of different scales by performing convolution operations at different dilation rates, thereby extracting as many features as possible.

[0071] Step 240: Input the multi-scale fusion feature map corresponding to the face occlusion image into the predictor in the target CenterNet network model.

[0072] In one implementation of this embodiment, Figure 2b As shown, the multi-scale fusion feature map corresponding to the face occlusion image is input into the predictor in the target CenterNet network model, including: inputting the multi-scale fusion feature map corresponding to the face occlusion image into the center point predictor, center point offset predictor and width and height predictor in the target CenterNet network model to obtain the center point position information, center point offset and target face width and height corresponding to the face occlusion image.

[0073] In a specific embodiment, Figure 2d This is a schematic diagram of the predictor structure in a target CenterNet network model in this embodiment. The multi-scale fusion and high-resolution feature map obtained by the neck module are input into three scale-invariant branches, which are used to predict the center point and category of the target face, the offset of the center point, and the width and height of the target, respectively.

[0074] The number of channels C output by the first branch is the number of categories, indicating the presence of an object in the grid, its type, and its confidence level. The feature points in the output feature map represent the scores of the corresponding categories. Max pooling is used to compare each feature point with its surrounding feature points to find the feature point with the highest score within a certain area. The point with the highest score is the center point, so this branch is also called the heat map branch.

[0075] The size of the original face occlusion image will be compressed after passing through the backbone network, and precision loss will occur in the process of rounding down the feature map. Therefore, although the upsampling operation in the neck module attempts to restore the original size, it cannot accurately restore all position information. Even if the convolution operation in the predictor can keep the scale unchanged, the detected center point position may still deviate from the actual position. In order to overcome this challenge, the predictor of the CenterNet network model provided in this embodiment is designed with a second branch, which is specifically used to calculate the offset of the center point on the x-axis and y-axis. The number of channels output by its convolution layer is 2, corresponding to the offsets on the two coordinate axes respectively. Using these offsets, the coordinates of the center point detected by the first branch (heat map branch) can be fine-tuned to enhance the accuracy of positioning.

[0076] In addition, the third branch is used to predict the width and height of the object to which the center point belongs. Similar to the second branch, the number of output channels of the convolutional layer of this branch is also 2, which directly corresponds to the two predicted values ​​of width and height. It is worth noting that the target CenterNet network model does not adopt the design of allocating 2 channels to each category (that is, the number of channels of 2xC) when predicting the offset and width and height of the center point. Instead, it directly uses two channels and predicts the offset and width and height of the center point of the target category determined by the first branch by default. Since the center point position of the first branch on the feature map already has strong category characteristics, and the three predictors maintain scale consistency and invariance during the two convolution processes, the outputs of the other two branches are closely related to the features of these positions.

[0077] The advantage of this setting is that by simultaneously predicting the target's center point and classification, center point offset, and width and height, the target CenterNet network model can more accurately locate the target's position and shape, which helps reduce positioning errors and improve detection accuracy. Secondly, since the CenterNet network model mainly relies on center point detection, even if there is occlusion or overlap between targets, as long as the center point can be accurately identified, the complete shape of the target can be restored through offset and width and height predictors, which helps reduce missed detections or false detections caused by occlusion or overlap.

[0078] Step 250: Determine face detection information in the face occlusion image according to the output result of the predictor.

[0079] In one implementation of this embodiment, face detection information in the face occlusion image is determined based on the output result of the predictor, including: adjusting the center point position information corresponding to the face occlusion image based on the center point offset; determining the border vertex position information based on the adjusted center point position information and the width and height of the target face, and generating a detection frame in the face occlusion image based on the border vertex position information.

[0080] The technical solution provided by the embodiment of the present invention obtains a face occlusion image to be detected, inputs the face occlusion image into a target CenterNet network model, extracts semantic features from the face occlusion image through the backbone network in the target CenterNet network model, inputs the semantic features corresponding to the face occlusion image into the neck module, obtains multi-scale context information corresponding to the face occlusion image according to the semantic features through the neck module and the ASPP layer, and outputs a multi-scale fusion feature map corresponding to the face occlusion image according to the multi-scale context information, inputs the multi-scale fusion feature map corresponding to the face occlusion image into the predictor in the target CenterNet network model, and determines the face detection information in the face occlusion image according to the output result of the predictor. This technical solution can improve the detection performance of the target CenterNet network model for face occlusion images and improve the accuracy of the detection results.

[0081] Figure 3 This is a schematic diagram of the structure of a face occlusion image detection device provided by an embodiment of the present invention, which is applied to electronic devices such as Figure 3 As shown, the device includes: an image acquisition module 310 and an image processing module 320.

[0082] An image acquisition module 310 is configured to acquire an occluded face image to be detected and input the occluded face image into a target CenterNet network model;

[0083] The neck module in the target CenterNet network model integrates the atrous spatial pyramid pooling layer ASPP;

[0084] The image processing module 320 is used to extract feature information corresponding to the face occlusion image through the target CenterNet network model, and output face detection information in the face occlusion image based on the feature extraction result.

[0085] The technical solution provided by the embodiment of the present invention obtains a face occlusion image to be detected, inputs the face occlusion image into a target CenterNet network model, and integrates an ASPP layer in the neck module of the target CenterNet network model. The feature information corresponding to the face occlusion image is extracted through the target CenterNet network model, and the face detection information in the face occlusion image is output according to the feature extraction result. This technical means can improve the detection performance of the target CenterNet network model for the face occlusion image and improve the accuracy of the detection results.

[0086] Based on the above embodiment, the ASPP layer consists of a common convolutional layer and a plurality of dilation convolutional layers with different expansion rates.

[0087] The image processing module 320 includes:

[0088] A semantic feature extraction unit, configured to extract semantic features from the face occluded image through the backbone network in the target CenterNet network model;

[0089] A feature map output unit is configured to input the semantic features corresponding to the face occlusion image into a neck module; obtain multi-scale context information corresponding to the face occlusion image based on the semantic features through the neck module and the ASPP layer, and output a multi-scale fused feature map corresponding to the face occlusion image based on the multi-scale context information;

[0090] A convolution processing unit, configured to perform convolution processing on the semantic features through the neck module, a common convolution layer in the ASPP layer, and a plurality of dilated convolution layers with different expansion rates, and then perform an addition operation on the convolution processing results on the channels to obtain multi-scale context information corresponding to the face occluded image;

[0091] A feature map input unit, configured to input the multi-scale fusion feature map corresponding to the face occlusion image into a predictor in a target CenterNet network model;

[0092] an information determining unit, configured to determine face detection information in the face occlusion image based on an output result of the predictor;

[0093] A predictor input unit is used to input the multi-scale fusion feature map corresponding to the face occlusion image into the center point predictor, center point offset predictor, and width and height predictor in the target CenterNet network model to obtain the center point position information, center point offset, and target face width and height corresponding to the face occlusion image;

[0094] A detection frame generation unit is used to adjust the center point position information corresponding to the face occlusion image according to the center point offset; determine the border vertex position information according to the adjusted center point position information and the width and height of the target face; and generate a detection frame in the face occlusion image according to the border vertex position information.

[0095] The above device can execute the methods provided by all the above embodiments of the present invention, and has the corresponding functional modules and beneficial effects of executing the above methods. For technical details not fully described in the embodiments of the present invention, please refer to the methods provided by all the above embodiments of the present invention.

[0096] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0097] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0098] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0099] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the face occlusion image detection method.

[0100] In some embodiments, the face occlusion image detection method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the face occlusion image detection method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the face occlusion image detection method in any other appropriate manner (for example, by means of firmware).

[0101] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0102] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0103] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0104] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0105] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0106] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0107] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0108] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A method for detecting face occlusion images, characterized in that: The method comprises: Obtain a face occlusion image to be detected, and input the face occlusion image into the target CenterNet network model; The neck module in the target CenterNet network model integrates the atrous spatial pyramid pooling layer ASPP; The target CenterNet network model is used to extract feature information corresponding to the face-occluded image, and face detection information in the face-occluded image is output based on the feature extraction result.

2. The method according to claim 1, characterized in that Extracting feature information corresponding to the face occluded image through the target CenterNet network model includes: Extracting semantic features from face occluded images through the backbone network in the target CenterNet network model; Inputting the semantic features corresponding to the face occlusion image into the neck module; Through the neck module and the ASPP layer, multi-scale context information corresponding to the face occlusion image is obtained according to the semantic features, and a multi-scale fusion feature map corresponding to the face occlusion image is output according to the multi-scale context information.

3. The method according to claim 2, characterized in that Outputting face detection information in the face occluded image based on the feature extraction result includes: Input the multi-scale fusion feature map corresponding to the face occlusion image into the predictor in the target CenterNet network model; According to the output result of the predictor, face detection information in the face occlusion image is determined.

4. The method according to claim 3, characterized in that Inputting the multi-scale fusion feature map corresponding to the face occlusion image into the predictor in the target CenterNet network model includes: The multi-scale fusion feature map corresponding to the face occlusion image is input into the center point predictor, center point offset predictor and width and height predictor in the target CenterNet network model to obtain the center point position information, center point offset and target face width and height corresponding to the face occlusion image.

5. The method according to claim 4, characterized in that Determining face detection information in the face occlusion image according to an output result of the predictor, including: Adjusting the center point position information corresponding to the face occlusion image according to the center point offset; The frame vertex position information is determined based on the adjusted center point position information and the width and height of the target face, and a detection frame is generated in the face occlusion image based on the frame vertex position information.

6. The method according to claim 2, characterized in that The ASPP layer consists of a normal convolutional layer and multiple dilation convolutional layers with different expansion rates; Obtaining multi-scale context information corresponding to the face occlusion image according to the semantic features through the neck module and the ASPP layer, including: The semantic features are convolved through the neck module, the ordinary convolution layer in the ASPP layer, and multiple dilation-rate-varying dilation layers, and then the convolution results are added on the channels to obtain multi-scale contextual information corresponding to the face occluded image.

7. A face occlusion image detection device, characterized in that: The device comprises: An image acquisition module is used to acquire an occluded face image to be detected and input the occluded face image into a target CenterNet network model; The neck module in the target CenterNet network model integrates the atrous spatial pyramid pooling layer ASPP; The image processing module is used to extract feature information corresponding to the face occlusion image through the target CenterNet network model, and output face detection information in the face occlusion image based on the feature extraction result.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the face occlusion image detection method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the face occlusion image detection method according to any one of claims 1 to 6 when executed.

10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the face occlusion image detection method according to any one of claims 1 to 6.