A face image detection method and device, electronic equipment and storage medium

By applying decorrelation regularization to the convolutional kernels of the face detection model, the problem of overfitting was solved, and the detection accuracy and generalization ability of the model in different scenarios were improved.

CN116152874BActive Publication Date: 2026-02-27DOUYIN VISION CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111371285.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-18
Publication Date
2026-02-27
Estimated Expiration
2041-11-18

AI Technical Summary

Technical Problem

Existing face liveness detection models suffer from poor feature generalization and low generalization due to overfitting to data, which affects the detection accuracy in different sampling scenarios.

Method used

During the training of the face detection model, the convolution kernels in the preset convolutional layers are subjected to decorrelation regularization, and the target loss function is determined based on the regularization result in order to reduce the correlation between features and improve the generalization ability of the model.

Benefits of technology

By using decorrelation regularization, the detection accuracy and generalization ability of the face detection model under different sampling scenarios are improved, resulting in higher image monitoring accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116152874B_ABST
    Figure CN116152874B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a face image detection method and device, electronic equipment and storage medium, wherein the method comprises: obtaining a face image to be detected; inputting the face image to be detected into a pre-trained face detection model to obtain a face image detection result; wherein in the training process of the face detection model, the correlation of the features corresponding to the convolution kernels in the preset convolution layer in the face detection model is regularized, and the target loss function of the face detection model is determined according to the regularized result. The technical scheme disclosed in the embodiments of the present disclosure solves the problem that the existing face image detection model has overfitting and low model generalization ability, and can obtain a face image detection result with higher monitoring accuracy in the face image detection process, thereby improving the generalization ability of the target training model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The embodiment of the present disclosure relates to the technical field of artificial intelligence, in particular to a face image detection method and device, electronic equipment and storage medium. BACKGROUND

[0002] At present, in the face living body detection technology, the features of the living body face image are usually learned in depth through a convolutional neural network model, and a face living body detection model is obtained by pre-training. Then, the model obtained by training is used to identify the living body face image. In the process of model training, the strategy of dense pixel supervision, pseudo-depth map, reflection map and texture map and other auxiliary supervision can make the model effectively learn more representations.

[0003] However, the face living body detection model obtained by the auxiliary supervision strategy usually tends to overfit the data. Due to the overfitting of the data, the features identified by the model have poor generality, so the model has low generalization and the accuracy of the face living body monitoring result for different data domains needs to be improved. SUMMARY

[0004] The embodiment of the present disclosure provides a face image detection method, device, electronic equipment and storage medium, which uses a face detection model that has been subjected to feature decorrelation regularization processing in the training process, so as to improve the detection accuracy of face images in different sampling scenarios.

[0005] In a first aspect, the embodiment of the present disclosure provides a face image detection method, comprising:

[0006] obtaining a face image to be detected;

[0007] inputting the face image to be detected into a pre-trained face detection model to obtain a face image detection result;

[0008] In the training process of the face detection model, the features corresponding to the convolution kernels in the preset convolution layer in the face detection model are subjected to decorrelation regularization processing, and the target loss function of the face detection model is determined according to the regularization processing result.

[0009] In a second aspect, the embodiment of the present disclosure further provides a face image detection device, comprising:

[0010] an image obtaining module, configured to obtain a face image to be detected;

[0011] an image detection module, configured to input the face image to be detected into a pre-trained face detection model to obtain a face image detection result;

[0012] In the training process of the face detection model, the features corresponding to the convolution kernels in the preset convolution layer in the face detection model are subjected to decorrelation regularization processing, and a target loss function of the face detection model is determined according to the regularization processing result.

[0013] In a third aspect, the present disclosure also provides an electronic device, which comprises:

[0014] one or more processors;

[0015] a storage device configured to store one or more programs,

[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the face image detection method according to any of the embodiments of the present disclosure.

[0017] In a fourth aspect, the present disclosure also provides a storage medium containing computer executable instructions, which, when executed by a computer processor, are used to perform the face image detection method according to any of the embodiments of the present disclosure.

[0018] The technical solution of the embodiments of the present disclosure obtains a face image to be detected, inputs the face image to be detected into a pre-trained face detection model, and obtains a face image detection result. In the training process of the face detection model, the features corresponding to the convolution kernels in the preset convolution layer in the face detection model are subjected to decorrelation regularization processing, and a target loss function of the face detection model is determined according to the regularization processing result. Therefore, the face detection model has higher generalization ability and can obtain more accurate image detection results. The technical solution disclosed in the embodiments of the present disclosure solves the problem of overfitting and low model generalization ability of the existing face image detection model, can obtain image detection results with higher monitoring accuracy in the process of face image detection, and improves the generalization ability of the target training model. BRIEF DESCRIPTION OF DRAWINGS

[0019] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn according to the scale.

[0020] Figure 1 A flowchart of a face image detection method provided by the first embodiment of the present disclosure;

[0021] Figure 2 A flowchart of a face image detection method provided by the second embodiment of the present disclosure;

[0022] Figure 3 A structure schematic diagram of a face image detection device provided by Embodiment Three of the present disclosure;

[0023] Figure 4 A structure schematic diagram of an electronic device provided by Embodiment Four of the present disclosure. DETAILED DESCRIPTION

[0024] Embodiments of the present disclosure will be described in more detail with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments set forth herein, but rather the embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.

[0025] It should be understood that each step described in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0026] The term "comprising" and variations thereof as used in the present disclosure are open-ended, that is, "including but not limited to". The term "based on" is "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related definitions will be given in the description below.

[0027] It should be noted that the "first", "second", and the like concepts mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.

[0028] It should be noted that the modification of "one", "multiple" mentioned in the present disclosure is illustrative and not limiting, and those skilled in the art should understand that unless the context clearly indicates otherwise, it should be understood as "one or more".

[0029] Embodiment One

[0030] Figure 1 A flowchart of a face image detection method provided by Embodiment One of the present disclosure, the embodiments of the present disclosure are suitable for the process of recognizing and classifying face images, and are particularly suitable for recognizing live face images. The method can be performed by a face image detection device, which can be implemented in the form of software and / or hardware, and the device can be configured in an electronic device, such as a mobile terminal or a server device.

[0031] As Figure 1 shown, the face image detection method provided by the embodiment includes:

[0032] S110, acquiring a face image to be detected.

[0033] The face image to be detected can be a face living body image, which is an image collected in a scene requiring identity recognition and authentication. For example, in a scene of account login or transaction information determination, face recognition verification is required, and the image of the user's face is collected in real time through the camera of the terminal where the application client is located. Therefore, it is necessary to identify whether the collected image is a living face image. Or, in other detection scenes of non-living face images requiring image recognition, the technical solution of the embodiment can also be applicable.

[0034] Specifically, when a user collects a face image to be detected through a terminal device, the user can be in various image collection environments. For example, indoor environment, outdoor environment. Further, in the outdoor environment, there can be strong sunlight or overcast conditions. Or, in some scenes, there are obvious landmark features near the face when collecting images, which can be identified as image features and affect the result of face image detection. Therefore, when performing face image detection, the adaptability and generalization ability of the face image detection model have certain requirements.

[0035] S120, inputting the face image to be detected into a pre-trained face detection model to obtain a face image detection result.

[0036] In the training process of the face detection model, the corresponding features of the convolution kernel in the pre-set convolution layer in the face detection model are subjected to decorrelation regularization processing, and the target loss function of the face detection model is determined according to the regularization processing result.

[0037] In the embodiment, the face detection model used for image recognition is subjected to decorrelation processing of the corresponding features of the convolution kernel in the training process, which reduces the correlation degree between image features, and the convolution kernel can learn more features, thereby improving the generalization ability of the face detection model. Therefore, the image detection result obtained by the face detection model in the embodiment has higher accuracy.

[0038] Further, in the training process of the face detection model, each convolutional layer in the preset convolutional neural network model (initial face detection model not trained) extracts features of the model training samples, and different convolutional kernels can extract different feature maps; then, supervised deep learning is performed according to the extracted features and labels of the model training samples. Generally, the features extracted from the same model training sample set have commonalities, such as sample images collected under the same lighting conditions, or the collection objects of the sample images maintain consistent gestures or actions. Then, the target convolutional neural network model trained based on the samples in the model training sample set has a relatively accurate recognition result for test images collected under the same conditions. However, for face live images collected under different conditions, the accuracy of the image recognition result is low, that is, in the face live monitoring scenario, the model trained based on the samples in the single model training sample set has poor feature generality, and the performance of detecting face attacks under different image collection conditions (images collected under different conditions) is not ideal, and the model generalization is low.

[0039] Considering that the cost of collecting live face images under different collection scenarios is high, and the data size and type of many data sets are limited, the model is prone to overfitting to the content of the data set, thereby reducing the generalization ability of the face detection model. In this step, the generalization ability of the model is improved by performing decorrelation regularization processing on the convolutional layers in the preset convolutional neural network. It should be noted that the feature decorrelation processing refers to representation decorrelation processing. The feature refers to the feature value in the feature map, and the representation refers to the form of the feature value in the image. The decorrelation refers to reducing the correlation between the representations.

[0040] Specifically, in the preset convolutional neural network model, some convolutional layers can be selectively regularized, that is, some convolutional layers participate in the regularization process, and some convolutional layers do not participate in the regularization process; or all convolutional layers can be regularized. The convolutional layers that need to be regularized can be determined and set in advance. Then, for each convolutional layer to be regularized, first, the feature correlation matrix is determined according to the average value of the sample features extracted by each convolutional kernel in the convolutional layer; then, the regularization term of the preset convolutional neural network model is determined according to the L1 regularization of the feature correlation matrix and the L1 regularization of the diagonal matrix of the feature correlation matrix. In this embodiment, the type of the preset convolutional neural network model is not limited, that is, the number of convolutional layers of the convolutional neural network and the number of convolutional kernels in each convolutional layer are not limited, which can be adjusted according to the face image detection effect, or the model parameters can be initialized and set according to the relevant experience.

[0041] Further, in determining the feature correlation matrix, for each layer of the convolutional layers that need to be regularized, the product value of the mean of the sample features extracted by each two convolution kernels in the same convolutional layer can be calculated; then, the calculated product value forms the feature correlation matrix. The feature correlation matrix of each layer can be expressed by the formula: The correlation between the features of each convolutional layer can be expressed as wherein L represents the number of convolutional layers to which the regularization is added, and l represents the serial number of the convolutional layer to which the regularization is added, and takes a value of 1-L. N represents the number of convolution kernels of the lth convolutional layer, and h i and h j represent the average values of the feature maps taken by the ith and jth convolution kernels in the lth convolutional layer, and the values of i and j range from 1 to N. In some optional embodiments, the feature correlation can also be expressed by other numerical relationships between h i and h j , such as a ratio relationship, using h i / h j .

[0042] Therefore, in the model training process, the goal is to reduce the correlation between the feature maps. In the back propagation process, the derivative function related to the feature maps is as follows, The formula shows that each feature map will be punished by the average values of other feature maps. When the convolution kernel outputs a feature map with a high average value, the output of other convolution kernels will be inhibited, so as to promote the convolution kernel to obtain a de-correlated representation. In addition, considering the collaborative representation of the convolution kernel, we constrain the correlation degree to be lower than the self-correlation degree of other features. That is, when a feature map has a correlation degree with other feature maps lower than its own correlation degree, the purpose of feature de-correlation is achieved.

[0043] A(l) is a symmetric square and also a semi-definite matrix. We use the L1 regularization of A(l) to obtain the most relevant group of correlation relationships with other features. In summary, the final de-correlation regularization term can be expressed as

[0044] wherein diagA(l) represents a diagonal matrix, and ‖diagA(h_l)‖1 represents the maximum value of the self-correlation. R dr is a feature de-correlation regularization. In the training process, by minimizing R dr , the correlation between the features can be reduced. When R dr takes a value of zero in the training process, it means that the correlation degree of the feature corresponding to the convolution kernel with other features is lower than its self-correlation degree, and the purpose of de-correlation has been achieved. Finally, based on the above regularization term, the target loss function can be determined as argmin θL(y, f(x, 0)) + λ l R dr wherein x and y respectively represent the model training sample and the label, L(y, f(x, 0)) represents the original loss function, 0 represents the weight of the model, and λ l represents the coefficient of the convolution layer selected to join the decorrelation regularization, which is a hyperparameter.

[0045] In the process of model training, the preset convolutional neural network model compares the output result with the label of the model training sample. If the comparison result does not satisfy the condition set by the target loss function, the parameter iteration model training process of the preset convolutional neural network model is adjusted. Until the output result of the preset convolutional neural network model satisfies the target loss function, the model training process is completed, and the target convolutional neural network model, i.e., the face detection model, is obtained. The target convolutional neural network model has undergone feature decorrelation regularization processing in the training process, and therefore its generalization ability is improved.

[0046] The technical scheme of the embodiment of the present disclosure can obtain a face image to be detected, input the face image to be detected into a pre-trained face detection model, and obtain a face image detection result. In the training process of the face detection model, the sample features extracted by each convolution kernel are subjected to decorrelation regularization processing for the convolution layer to be regularized in the convolution layer structure of the face detection model, and a target loss function is determined according to the regularization processing result. When the output result of the preset convolutional neural network model satisfies the target loss function, the model training process is completed, and the final neural network model for face detection is obtained. The technical scheme disclosed in the embodiment of the present disclosure solves the problem of overfitting of the existing face image detection model and low generalization ability of the model, can obtain a face image detection result with higher monitoring accuracy in the process of face image detection, and improves the generalization ability of the target training model.

[0047] Embodiment Two

[0048] The embodiment of the present disclosure can be combined with each optional scheme of the face image detection method provided in the above embodiments. The face image detection method provided in the embodiment is optimized on the basis of the above embodiments, and further describes the process of performing redundancy reduction processing and face image detection after sample feature decorrelation processing.

[0049] Figure 2 A flowchart of a face image detection method provided in Embodiment Two of the present disclosure is shown in FIG. 2. As shown in FIG. 2, the face image detection method provided in the embodiment includes the following steps. Figure 2

[0050] ​S210. Obtain model training samples and input the model training samples into a preset convolutional neural network model for model training.

[0051] The training samples for the model can be live face images. The preset convolutional neural network model is then used to recognize these live faces, suitable for identity recognition and authentication scenarios, such as account login and transaction information verification, or other scenarios requiring image recognition. This embodiment does not limit the type of the preset convolutional neural network model, i.e., it does not limit the number of convolutional layers or the number of kernels in each layer. These can be adjusted based on the face image detection results, or the model parameters can be initialized based on relevant experience.

[0052] S220. For the convolutional layers in the preset convolutional neural network model that need to be regularized, the sample features extracted by each convolutional kernel are subjected to decorrelation regularization.

[0053] Considering the high cost of collecting live face images from different acquisition scenarios, and the limited data size and type of many datasets, models are prone to overfitting within these datasets, thus reducing the generalization ability of face detection models. In this step, decorrelation regularization is applied to the convolutional layers in the pre-defined convolutional neural network to improve the model's generalization ability. The specific processing procedure can be found in the details of the above embodiments.

[0054] S230. The feature correlation matrix determined during the regularization process is subjected to sparse processing, and the regularization term is updated based on the sparsely processed feature correlation matrix. The target loss function is determined based on the updated regularization term.

[0055] Decorrelated features prevent co-adaptation between convolutional kernels, but R dr Only update the weights of convolutional kernels with a fixed pattern. Here, "fixed pattern" refers to the weights of kernels with a fixed pattern for h. i To be honest, h j Since it's fixed, it follows a fixed update pattern. Therefore, non-fixed-pattern convolutional kernels become redundant because they fail to learn features. Therefore, we can apply R to a subset of convolutional kernels in the regularized convolutional layer by randomly selecting kernels. dr To break the fixed update pattern and reduce overfitting, the regularization process involves sparse processing of the feature correlation matrix determined during the regularization process. This process can be called the regularization process of Decorrelated Sparse Representation (DSR).

[0056] Specifically, we relax R by applying sparse weights to the correlation matrix A(l). dr It can be written as follows: S ij is a sparse symmetric matrix with the same dimension as the feature correlation matrix, which can be expressed as S ij = d*r, d = {0, 1}, r ~ U(0, 1). Wherein, r is the first value, indicating the correlation coefficient, which is sampled from a uniform distribution; d represents the second value, indicating whether to select the convolution kernel, which depends on the sparsity of S ij . The setting of matrix sparsity is a hyperparameter. When the preset sparsity hyperparameter of S ij is 0.5, it means that half of the convolution kernels in the l-th convolution layer are randomly selected. That is, the sparse convolution kernel is randomly selected, multiplied by the random coefficient S ij to change the update mode of the convolution kernel. Therefore, the DSR regularization term can be expressed as:

[0057] During the training process, the target loss function can be determined as argmin θ L(y, f(x, θ)) + λ l R dsr . Wherein, x, y respectively represent the model training sample and label, L(y, f(x, θ)) represents the original loss function, θ represents the weight of the model, and λ l represents the coefficient of the convolution layer selected to join the decorrelation regularization, which is a hyperparameter.

[0058] S240, when the output result of the preset convolutional neural network model satisfies the target loss function, the model training process is completed, and a target convolutional neural network model is obtained.

[0059] During the model training process, after several iterations, each convolution kernel has decorrelated and non-redundant features. When the output result of the preset convolutional neural network model satisfies the target loss function, the model training process is completed, and a target convolutional neural network model with improved generalization ability and without overfitting can be obtained.

[0060] It can be understood that the structure of each preset convolutional neural network model is fixed, and after reducing the correlation and redundancy of several features, the convolution kernel can have the opportunity to learn other features, and the learned features become more extensive. Therefore, when using the model trained by the target convolutional neural network model obtained in the process across the data domain, the accuracy of image recognition can be improved.

[0061] S250, obtaining a face image to be detected.

[0062] When the training sample of the target convolutional neural network is a face live image, the face image to be detected can also be a face live image, including face images collected in different image collection environments.

[0063] S260, input the face image to be detected into the target convolutional neural network model, and obtain a face image detection result.

[0064] The target convolutional neural network is a model that has undergone feature decorrelation and redundancy removal processing in a training process, and the image detection effect of the model is better and the model has good generalization ability.

[0065] The technical scheme of the embodiment of the present disclosure obtains model training samples, inputs the model training samples into a preset convolutional neural network model for model training, and in the process of model training, for a convolution layer to be regularized in the preset convolutional neural network model, performs decorrelation regularization processing on the sample features extracted by each convolution kernel, and further performs redundancy removal processing on the basis of the regularization processing by using a random sparse matrix, and determines a target loss function according to the regularization processing result of the decorrelation sparse representation. When the result output by the preset convolutional neural network model satisfies the target loss function, the model training process is completed, and a target neural network model is obtained. The technical scheme disclosed in the embodiment of the present disclosure solves the problem that the convolutional neural network model tends to overfit data in the model training process and has low model generalization ability, can reduce the correlation between the extracted features in the process of model training, and thus improves the generalization ability of the target training model.

[0066] Embodiment three

[0067] Figure 3 A structure schematic diagram of a face image detection device provided in the third embodiment of the present disclosure. The face image detection device provided in the embodiment is suitable for detecting and classifying face image categories, and is particularly suitable for live face detection.

[0068] As shown in Figure 3 The face image detection device includes an image acquisition module 310 and an image detection module 320.

[0069] The image acquisition module 310 is configured to acquire a face image to be detected. The image detection module 320 is configured to input the face image to be detected into a pre-trained face detection model, and obtain a face image detection result. In the training process of the face detection model, the features corresponding to the convolution kernels in the preset convolutional layer of the face detection model are subjected to decorrelation regularization processing, and a target loss function of the face detection model is determined according to the regularization processing result.

[0070] The technical scheme of the embodiment of the present disclosure is that a face image to be detected is acquired, the face image to be detected is input into a pre-trained face detection model, and a face image detection result is acquired; in the training process of the face detection model, the features corresponding to the convolution kernels in the preset convolution layer in the face detection model are subjected to decorrelation regularization processing, and a target loss function of the face detection model is determined according to the regularization processing result, so that the face detection model has higher generalization ability and can obtain more accurate image detection results. The technical scheme disclosed in the embodiment of the present disclosure solves the problem that the existing face image detection model has overfitting and low model generalization ability, can obtain an image detection result with higher monitoring accuracy in the process of face image detection, and improves the generalization ability of the target training model.

[0071] In some optional implementations, the face image detection apparatus further includes a model training module configured to perform decorrelation regularization processing on features corresponding to convolution kernels in a preset convolution layer in the face detection model, and determine a target loss function of the face detection model according to a regularization processing result.

[0072] The model training module includes a regularization processing submodule configured to determine, for each convolution layer in the preset convolution layer, a feature correlation matrix according to average values of sample features extracted by convolution kernels in the convolution layer.

[0073] The regularization term of the preset convolution neural network model is determined according to L1 regularization of the feature correlation matrix and L1 regularization of a diagonal matrix of the feature correlation matrix.

[0074] In some optional implementations, the regularization processing submodule is specifically configured to:

[0075] Calculate a product value of the average values of the sample features extracted by each two convolution kernels in the same convolution layer.

[0076] The product value calculated is used to form the feature correlation matrix.

[0077] In some optional implementations, the regularization processing submodule is further configured to:

[0078] The original loss function of the face detection model and the regularization term are superimposed as the target loss function.

[0079] In some optional implementations, the model training module includes a redundancy removal processing submodule configured to:

[0080] After the de-correlation regularization processing is performed on the convolution kernel corresponding features in the preset convolution layer in the face detection model, the convolution kernel corresponding features after the de-correlation regularization processing are subjected to de-redundancy sparse processing, and a target loss function of the face detection model is updated according to a de-redundancy sparse processing result.

[0081] In some optional implementations, the de-redundancy processing submodule is specifically configured to:

[0082] randomly generate a sparse symmetric matrix with the same dimension as the feature correlation matrix in the de-correlation regularization processing result;

[0083] multiply the sparse symmetric matrix and the feature correlation matrix.

[0084] In some optional implementations, the de-redundancy processing submodule is further configured to:

[0085] for each element in the sparse symmetric matrix, randomly sample a first value in a uniform distribution of 0-1;

[0086] determine a second value corresponding to the element as 1 or 0 randomly according to a preset sparsity hyperparameter, where the preset sparsity hyperparameter represents the number of elements with the second value of 1;

[0087] multiply the first value and the second value to obtain a product result as an element value, so as to generate the sparse symmetric matrix.

[0088] The face image detection apparatus provided by the embodiments of the present disclosure can perform the face image detection method provided by any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of the execution method.

[0089] It should be noted that each unit and module included in the apparatus is only divided according to the function logic, but is not limited to the above division, as long as the corresponding function can be implemented; in addition, the specific names of each functional unit are only for convenient distinction, and do not limit the protection scope of the embodiments of the present disclosure.

[0090] Embodiment Four

[0091] Reference will now be made to the following description Figure 4 which shows an electronic device (e.g., a mobile phone) suitable for implementing the embodiments of the present disclosure. Figure 4The diagram below shows the structure of the terminal device or server 400. The terminal device in this embodiment may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and vehicle terminals (e.g., vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 4 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0092] like Figure 4 As shown, electronic device 400 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 406 into random access memory (RAM) 403. RAM 403 also stores various programs and data required for the operation of electronic device 400. The processing unit 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.

[0093] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, magnetic tapes, hard disks, etc.; and communication devices 409. Communication device 409 allows electronic device 400 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 4 An electronic device 400 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0094] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 409, or installed from storage device 406, or installed from ROM 402. When the computer program is executed by processing device 401, it performs the functions defined in the face image detection method of embodiments of this disclosure.

[0095] The electronic device provided by the embodiments of the present disclosure and the face image detection method provided by the above embodiments belong to the same disclosure concept, and the technical details not described in detail in the present embodiment can be referred to the above embodiments, and the present embodiment has the same beneficial effects as the above embodiments.

[0096] Embodiment Five

[0097] The present disclosure provides a computer storage medium, which stores a computer program, and the program is executed by a processor to implement the face image detection method provided by the above embodiments.

[0098] It should be noted that the computer readable medium of the present disclosure can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of computer readable storage media can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or a flash memory (FLASH), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as a part of a carrier wave, which carries computer readable program code. Such a propagated data signal can take many forms, including but not limited to an electromagnetic signal, an optical signal or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium other than the computer readable storage medium, which can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or apparatus. The program code contained in the computer readable medium can be transmitted by any suitable medium, including but not limited to a wire, a cable, an RF (radio frequency) or the like, or any suitable combination of the above.

[0099] In some embodiments, the client, server, or both can communicate using any known or later developed network protocols, such as the HyperText Transfer Protocol (HTTP), and can be interconnected with any form or medium of digital data communication (for example, a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), the Internet, and peer-to-peer networks (for example, ad hoc peer-to-peer networks), as well as any then currently known or later developed networks.

[0100] The computer-readable medium described above can be included in the electronic device described above; or can exist independently of the electronic device and be not assembled into the electronic device.

[0101] The computer-readable medium described above carries one or more programs, which, when executed by the electronic device, cause the electronic device to:

[0102] Obtain a face image to be detected;

[0103] Input the face image to be detected into a pre-trained face detection model to obtain a face image detection result;

[0104] In the training process of the face detection model, the features corresponding to the convolution kernels in the preset convolution layer in the face detection model are subjected to decorrelation regularization processing, and a target loss function of the face detection model is determined according to the regularization processing result.

[0105] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network ("LAN") or a wide area network ("WAN"), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0106] The computer program product of the first aspect can include a computer readable storage medium. The computer readable storage medium can include instructions. The instructions can include one or both of: instructions that configure the one or more processors to perform operations described herein; and instructions that configure the one or more processors to cause performance of the operations described herein.

[0107] The units described in the embodiments of the present disclosure can be implemented by software or by hardware. In some cases, the names of the units and modules do not constitute a limitation on the units and modules themselves. For example, a data generation module can also be described as a "video data generation module".

[0108] The functions described in the above description can be performed at least in part by one or more hardware logic components. For example, non-limiting examples of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0109] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0110] According to one or more embodiments of the present disclosure, Example One provides a face image detection method, which comprises:

[0111] obtaining a face image to be detected;

[0112] inputting the face image to be detected into a pre-trained face detection model to obtain a face image detection result;

[0113] In the training process of the face detection model, the features corresponding to the convolution kernels in the preset convolution layers in the face detection model are subjected to decorrelation regularization processing, and a target loss function of the face detection model is determined according to the regularization processing result.

[0114] According to one or more embodiments of the present disclosure, Example Two provides a face image detection method, which further comprises:

[0115] In some optional implementation manners, the decorrelation regularization processing of the features corresponding to the convolution kernels in the preset convolution layers in the face detection model comprises:

[0116] for each convolution layer in the preset convolution layers, determining a feature correlation matrix according to the average values of the sample features extracted by the convolution kernels in the convolution layer;

[0117] determining a regularization term of the preset convolution neural network model according to the L1 regularization of the feature correlation matrix and the L1 regularization of the diagonal matrix of the feature correlation matrix.

[0118] According to one or more embodiments of the present disclosure, Example Three provides a face image detection method, which further comprises:

[0119] In some optional implementation manners, the determining the feature correlation matrix according to the average values of the sample features extracted by the respective convolution kernels in the convolution layer comprises:

[0120] calculating a product value of the average values of the sample features extracted by each two convolution kernels in the same convolution layer;

[0121] composing the product value calculated into the feature correlation matrix.

[0122] According to one or more embodiments of the present disclosure, Example Four provides a face image detection method, further comprising:

[0123] In some optional implementation manners, the determining the target loss function of the face detection model according to the regularization processing result comprises:

[0124] superimposing the original loss function of the face detection model and the regularization term as the target loss function.

[0125] According to one or more embodiments of the present disclosure, Example Five provides a face image detection method, further comprising:

[0126] In some optional implementation manners, after the de-correlation regularization processing is performed on the convolution kernel corresponding features in the preset convolution layer in the face detection model, the method further comprises:

[0127] performing de-redundancy sparse processing on the convolution kernel corresponding features after the de-correlation regularization processing, and updating the target loss function of the face detection model according to the de-redundancy sparse processing result.

[0128] According to one or more embodiments of the present disclosure, Example Six provides a face image detection method, further comprising:

[0129] In some optional implementation manners, the performing de-redundancy sparse processing on the convolution kernel corresponding features after the de-correlation regularization processing comprises:

[0130] randomly generating a sparse symmetric matrix with the same dimension as the feature correlation matrix in the de-correlation regularization processing result;

[0131] multiplying the sparse symmetric matrix and the feature correlation matrix.

[0132] According to one or more embodiments of the present disclosure, Example Seven provides a face image detection method, further comprising:

[0133] In some optional implementation manners, the randomly generating a sparse symmetric matrix with the same dimension as the feature correlation matrix in the de-correlation regularization processing result comprises:

[0134] For each element in the sparse symmetric matrix, a first numerical value is determined by random sampling in a uniform distribution of 0-1;

[0135] According to a preset sparsity hyperparameter, a second numerical value corresponding to the element is randomly determined to be 1 or 0, wherein the preset sparsity hyperparameter represents the number of elements with the second numerical value of 1;

[0136] The product result of the first numerical value and the second numerical value is taken as an element value to generate the sparse symmetric matrix.

[0137] According to one or more embodiments of the present disclosure,

Example Eight

[0138] An image acquisition module is configured to acquire a face image to be detected;

[0139] An image detection module is configured to input the face image to be detected into a pre-trained face detection model to obtain a face image detection result.

[0140] In the training process of the face detection model, the features corresponding to the convolution kernels in the preset convolution layer in the face detection model are subjected to decorrelation regularization processing, and a target loss function of the face detection model is determined according to the regularization processing result.

[0141] According to one or more embodiments of the present disclosure,

Example Nine

[0142] In some optional implementations, the face image detection device further comprises a model training module configured to perform decorrelation regularization processing on features corresponding to convolution kernels in a preset convolution layer in the face detection model, and determine a target loss function of the face detection model according to a regularization processing result.

[0143] The model training module comprises a regularization processing submodule configured to, for each convolution layer in the preset convolution layer, determine a feature correlation matrix according to average values of sample features extracted by convolution kernels in the convolution layer.

[0144] The regularization term of the preset convolutional neural network model is determined according to L1 regularization of the feature correlation matrix and L1 regularization of a diagonal matrix of the feature correlation matrix.

[0145] According to one or more embodiments of the present disclosure,

Example Ten

[0146] In some optional implementations, the regularization processing submodule is specifically configured to:

[0147] The product value of the mean values of the sample features extracted by each two convolution kernels in the same convolution layer is calculated;

[0148] The calculated product value is used to form the feature correlation matrix.

[0149] According to one or more embodiments of the present disclosure, example eleven provides a face image detection device, further comprising:

[0150] In some optional implementations, the regularization processing submodule is further configured to:

[0151] The original loss function of the face detection model is superimposed with the regularization term as the target loss function.

[0152] According to one or more embodiments of the present disclosure, example twelve provides a face image detection device, further comprising:

[0153] In some optional implementations, the model training module comprises a redundancy removal processing submodule configured to:

[0154] After the de-correlation regularization processing of the convolution kernel corresponding features in the preset convolution layer of the face detection model, the de-correlation regularization processed convolution kernel corresponding features are subjected to redundancy removal sparse processing, and the target loss function of the face detection model is updated according to the redundancy removal sparse processing result.

[0155] According to one or more embodiments of the present disclosure, example thirteen provides a face image detection device, further comprising:

[0156] In some optional implementations, the redundancy removal processing submodule is specifically configured to:

[0157] A sparse symmetric matrix with the same dimension as the feature correlation matrix in the de-correlation regularization processing result is randomly generated;

[0158] The sparse symmetric matrix is multiplied with the feature correlation matrix.

[0159] According to one or more embodiments of the present disclosure, example fourteen provides a face image detection device, further comprising:

[0160] In some optional implementations, the redundancy removal processing submodule is further configured to:

[0161] For each element in the sparse symmetric matrix, a first value is randomly sampled in a uniform distribution of 0-1;

[0162] According to a preset sparsity hyper-parameter, the second value corresponding to the element is randomly determined as 1 or 0, wherein the preset sparsity hyper-parameter represents the number of elements with the second value of 1;

[0163] The product result of the first value and the second value is taken as an element value to generate the sparse symmetric matrix.

[0164] The above description is merely preferred embodiments of the present disclosure and a description of the principles of the technology applied. It should be understood by those skilled in the art that the disclosed range of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.

[0165] In addition, although each operation is described in a particular order, this should not be understood as requiring the operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be implemented separately or in any suitable subcombination.

[0166] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. A method for detecting human faces, characterized in that, include: Acquire the image of the face to be detected; The face image to be detected is input into a pre-trained face detection model to obtain the face image detection result; In the training process of the face detection model, the features corresponding to the convolution kernels in the preset convolutional layers of the face detection model are subjected to decorrelation regularization, and the target loss function of the face detection model is determined based on the regularization result. The process of performing decorrelation regularization on the features corresponding to the convolution kernels in the preset convolutional layer includes: determining the feature correlation matrix based on the average value of the sample features extracted by each convolution kernel in the convolutional layer; and determining the regularization term of the preset convolutional neural network model based on the L1 regularization of the feature correlation matrix and the L1 regularization of the diagonal matrix of the feature correlation matrix.

2. The method according to claim 1, characterized in that, The step of determining the feature correlation matrix based on the average value of the sample features extracted by each convolutional kernel in the convolutional layer includes: Calculate the product of the mean values ​​of the sample features extracted by every two convolutional kernels in the same convolutional layer; The calculated product values ​​are used to form the feature correlation matrix.

3. The method according to claim 1, characterized in that, The target loss function of the face detection model was determined based on the regularization result, including: The original loss function of the face detection model is superimposed with the regularization term to form the target loss function.

4. The method according to any one of claims 1-3, characterized in that, After performing decorrelation regularization on the features corresponding to the convolution kernels in the preset convolutional layers of the face detection model, the method further includes: The features corresponding to the convolution kernels after decorrelation regularization are subjected to redundancy removal and sparsity processing, and the target loss function of the face detection model is updated based on the redundancy removal and sparsity processing results.

5. The method according to claim 4, characterized in that, The process of performing redundancy removal and sparsity processing on the features corresponding to the convolution kernels after decorrelation regularization includes: Randomly generate a sparse symmetric matrix with the same dimension as the feature correlation matrix in the decorrelation regularization result; Multiply the sparse symmetric matrix with the feature correlation matrix.

6. The method according to claim 5, characterized in that, Randomly generate a sparse symmetric matrix with the same dimension as the feature correlation matrix in the decorrelation regularization result, including: For each element in the sparse symmetric matrix, a first value is determined by random sampling from a uniform distribution of 0-1. According to a preset sparsity hyperparameter, the second value corresponding to an element is randomly determined to be 1 or 0, wherein the preset sparsity hyperparameter represents the number of elements whose second value is 1; The product of the first value and the second value is used as the element value to generate the sparse symmetric matrix.

7. A face image detection device, characterized in that, include: The image acquisition module is used to acquire images of the face to be detected. The image detection module is used to input the face image to be detected into a pre-trained face detection model to obtain the face image detection result; In the training process of the face detection model, the features corresponding to the convolution kernels in the preset convolutional layers of the face detection model are subjected to decorrelation regularization, and the target loss function of the face detection model is determined based on the regularization result. The decorrelation regularization process for the features corresponding to the convolution kernels in the preset convolutional layers includes: determining the feature correlation matrix based on the average value of the sample features extracted by each convolution kernel in the preset convolutional layer; and determining the regularization term of the preset convolutional neural network model based on the L1 regularization of the feature correlation matrix and the L1 regularization of the diagonal matrix of the feature correlation matrix.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the face image detection method as described in any one of claims 1-6.

9. A storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the face image detection method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Classification model training method, image processing method and device

    CN111242222A