Face Recognition Method, Device, Electronic Device and Computer Readable Storage Medium

By extracting and transforming features of the image to be processed, non-illuminated features are separated, and the problem of face recognition errors under different lighting conditions is solved, and the accuracy of face recognition is improved.

CN114332993BActive Publication Date: 2025-06-10SHENZHEN XUMI YUNTU SPACE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111552557.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-17
Publication Date
2025-06-10
Estimated Expiration
2041-12-17

AI Technical Summary

Technical Problem

In the prior art, different lighting conditions have a great impact on the face recognition results, resulting in recognition errors.

Method used

By extracting and transforming the image to be processed, a spatial illumination semantic map and a channel illumination semantic map are obtained, and illumination separation is performed based on these two semantic maps to obtain non-illumination features, and finally facial recognition is performed based on non-illumination features.

Benefits of technology

This method reduces the interference of light on face recognition by separating light features and improves the accuracy of face recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332993B_ABST
    Figure CN114332993B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of computer vision technology, and provides a face recognition method, apparatus, electronic device, and computer-readable storage medium. The method includes: obtaining an image to be processed; extracting features from the image to be processed to obtain a feature image of the image to be processed; performing feature transformation on the feature image to obtain a spatial illumination semantic map corresponding to the image to be processed; performing feature transformation on the feature image based on the spatial illumination semantic map to obtain a channel illumination semantic map corresponding to the image to be processed; separating illumination from the image to be processed according to the spatial illumination semantic map and the channel illumination semantic map to obtain non-illumination features of the image to be processed; and performing face recognition based on the non-illumination features to obtain a face recognition result. This method separates non-illumination features that do not contain illumination features from the image to be processed for face recognition, which can minimize the interference of illumination features in the image on face recognition as much as possible, obtain better face recognition results, and improve the accuracy of face recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision technology, and particularly to a face recognition method, apparatus, electronic device, and computer-readable storage medium. Background Art

[0002] Face recognition is a biometric identification technology that identifies a person's identity based on their facial feature information. A series of related technologies that use a camera or webcam to collect images or video streams containing human faces, automatically detect and track faces in the images, and then perform face recognition on the detected faces are usually also called portrait recognition or facial recognition. With the development of computer vision technology, face recognition technology has become increasingly mature.

[0003] Face recognition is relatively sensitive to the lighting conditions of the images or video streams containing human faces that are collected. It often requires good lighting conditions to obtain better face recognition results. Once adverse situations such as overexposure or backlighting occur, it is easy to cause problems with face recognition errors. Summary of the Invention

[0004] In view of this, embodiments of the present disclosure provide a face recognition method, apparatus, electronic device, and computer-readable storage medium to solve the problem in the prior art that different lighting conditions have a greater impact on face recognition results.

[0005] In a first aspect of the embodiments of the present disclosure, a face recognition method is provided, including:

[0006] Obtain an image to be processed;

[0007] Extract features from the image to be processed to obtain a feature image of the image to be processed;

[0008] Perform feature transformation on the feature image to obtain a spatial lighting semantic map corresponding to the image to be processed;

[0009] Based on the spatial lighting semantic map, perform feature transformation on the feature image to obtain a channel lighting semantic map corresponding to the image to be processed;

[0010] According to the spatial lighting semantic map and the channel lighting semantic map, perform lighting separation on the image to be processed to obtain non-lighting features of the image to be processed;

[0011] Perform face recognition based on the non-lighting features to obtain a face recognition result.

[0012] In a second aspect of the embodiments of the present disclosure, a face recognition apparatus is provided, including:

[0013] An obtaining module configured to obtain an image to be processed;

[0014] A feature extraction module, configured to perform feature extraction on an image to be processed to obtain a feature image of the image to be processed;

[0015] A first feature transformation module, configured to perform feature transformation on the feature image to obtain a spatial illumination semantic map corresponding to the image to be processed;

[0016] A second feature transformation module, configured to perform feature transformation on the feature image based on the spatial illumination semantic map to obtain a channel illumination semantic map corresponding to the image to be processed;

[0017] An illumination separation module, configured to perform illumination separation on the image to be processed according to the spatial illumination semantic map and the channel illumination semantic map to obtain a non-illumination feature of the image to be processed;

[0018] A face recognition module, configured to perform face recognition based on the non-illumination feature to obtain a face recognition result.

[0019] In a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0020] In a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0021] The beneficial effects of the embodiments of the present disclosure compared with the prior art are as follows: By performing feature extraction and feature transformation on the image to be processed, a spatial illumination semantic map and a channel illumination semantic map of the image to be processed are obtained, and illumination separation is performed according to the spatial illumination semantic map and the channel illumination semantic map to obtain a non-illumination feature in the image to be processed. Finally, face recognition is performed on the non-illumination feature to obtain a face recognition result. This method calculates the spatial illumination semantic map and the channel illumination semantic map, uses these two illumination semantic maps for illumination separation, separates non-illumination features that do not include illumination features from the image to be processed for face recognition, and can minimize the interference of illumination features in the image on face recognition, thereby obtaining a better face recognition result and improving the accuracy of face recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 It is a schematic diagram of the application scenario of the embodiment of the present disclosure;

[0024] Figure 2 It is a schematic flowchart of a face recognition method provided by the embodiment of the present disclosure;

[0025] Figure 3 It is a schematic flowchart of a method for separating illumination from an image to be processed according to a spatial illumination semantic map and a channel illumination semantic map provided by the embodiment of the present disclosure;

[0026] Figure 4 It is another schematic flowchart of a method for separating illumination from an image to be processed according to a spatial illumination semantic map and a channel illumination semantic map provided by the embodiment of the present disclosure;

[0027] Figure 5 It is a schematic flowchart of a training process of a network provided by the embodiment of the present disclosure;

[0028] Figure 6 It is a schematic structural diagram of a face recognition device provided by the embodiment of the present disclosure;

[0029] Figure 7 It is a schematic structural diagram of an electronic device provided by the embodiment of the present disclosure. Detailed Embodiments

[0030] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0031] A face recognition method and device according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0032] Figure 1 It is a schematic diagram of the application scenario of the embodiment of the present disclosure. This application scenario may include terminal devices 1, 2, and 3, server 4, and network 5.

[0033] The terminal devices 1, 2, and 3 can be hardware or software. When the terminal devices 1, 2, and 3 are hardware, they can be various electronic devices with a display screen and supporting communication with the server 4, including but not limited to smartphones, tablets, laptop computers, desktop computers, etc.; when the terminal devices 1, 2, and 3 are software, they can be installed in the above-mentioned electronic devices. The terminal devices 1, 2, and 3 can be implemented as multiple software or software modules, or can also be implemented as a single software or software module, and the embodiments of the present disclosure do not limit this. Further, various applications can be installed on the terminal devices 1, 2, and 3, such as data processing applications, instant messaging tools, social platform software, search applications, shopping applications, etc.

[0034] The server 4 can be a server that provides various services. For example, it can be a background server that receives requests sent by terminal devices with which it establishes a communication connection. This background server can receive and analyze requests sent by terminal devices and generate processing results. The server 4 can be a single server, or can also be a server cluster composed of several servers, or can also be a cloud computing service center, and the embodiments of the present disclosure do not limit this.

[0035] It should be noted that the server 4 can be hardware or software. When the server 4 is hardware, it can be various electronic devices that provide various services for the terminal devices 1, 2, and 3. When the server 4 is software, it can be multiple software or software modules that provide various services for the terminal devices 1, 2, and 3, or can also be a single software or software module that provides various services for the terminal devices 1, 2, and 3, and the embodiments of the present disclosure do not limit this.

[0036] The network 5 can be a wired network connected by coaxial cables, twisted pairs, and optical fibers, or can also be a wireless network that can interconnect various communication devices without wiring, such as Bluetooth, Near Field Communication (NFC), Infrared, etc., and the embodiments of the present disclosure do not limit this.

[0037] The user can establish a communication connection with the server 4 via the network 5 through the terminal devices 1, 2, and 3 to receive or send information, etc. Specifically, after the user imports the image to be processed into the server 4, the server 4 extracts features from the image to be processed to obtain a feature image, performs feature transformation on the feature image, respectively obtains the spatial illumination semantic map and the spatial illumination semantic map of the image to be processed, combines the spatial illumination semantic map and the spatial illumination semantic map for illumination separation, obtains the non-illumination features of the image to be processed, and performs face recognition on the non-illumination features to obtain a face recognition result.

[0038] It should be noted that the specific types, quantities, and combinations of the terminal devices 1, 2, and 3, the server 4, and the network 5 can be adjusted according to the actual requirements of the application scenario, and the embodiments of the present disclosure do not limit this.

[0039] Figure 2 It is a schematic flowchart of a face recognition method provided by an embodiment of the present disclosure. Figure 2 The face recognition method can be executed by Figure 1 the server. As Figure 2 shown, the face recognition method includes:

[0040] S201, obtain an image to be processed.

[0041] Among them, the image to be processed is an image that needs to perform face recognition. Understandably, the image to be processed contains a face. Among them, the image to be processed can be obtained by any method.

[0042] S202, perform feature extraction on the image to be processed to obtain a feature image of the image to be processed.

[0043] Feature extraction is a method of extracting required features through image analysis and transformation. A feature is the corresponding (essential) characteristic or property that distinguishes a certain type of object from other types of objects, or a set of these characteristics and properties. A feature is data that can be extracted through measurement or processing. For an image, each image has its own features that can distinguish it from other types of images. Some are natural features that can be intuitively felt, such as brightness, edges, textures, and colors; some are features that need to be obtained through transformation or processing, such as moments, histograms, and principal components. In this embodiment, the feature obtained by performing feature extraction on the image to be processed is recorded as a feature image.

[0044] Among them, performing feature extraction on the image to be processed can be achieved by any method. In some embodiments, a neural network can be used to perform feature extraction on the input image to be processed to obtain a feature image, such as a residual module; in a specific embodiment, the residual network includes two downsampling layers; further, two residual modules are used to perform feature extraction on the image to be processed. The first residual module includes a downsampling layer with a stride of 2 and a channel number of 64, and the second residual module includes 2 downsampling layers with a stride of 1 and a channel of 64. In some embodiments, the size of the image to be processed is 112*112, then the dimension of the feature image obtained by performing feature extraction through the above feature extraction module is (56, 56, 64).

[0045] Among them, the downsampling principle: For an image I with a size of M*N, performing s-fold downsampling on it results in a low-resolution image with a size of (M / s)*(N / s). s is the greatest common divisor of M and N. For an image in matrix form, it means changing the image within an s*s window of the original image into a single pixel, and the value of this pixel is the average of all pixels within the window.

[0046] In other embodiments, other methods can be used to extract features from the image to be processed to obtain a feature image.

[0047] S203. Perform feature transformation on the feature image to obtain the spatial illumination semantic map corresponding to the image to be processed.

[0048] Feature transformation is a process of further processing and transforming the extracted feature image to obtain new image features. The process of performing feature transformation on the feature image to obtain the spatial illumination semantic map of the image to be processed will be described in detail in subsequent embodiments.

[0049] S204. Based on the spatial illumination semantic map, perform feature transformation on the feature image to obtain the channel illumination semantic map corresponding to the image to be processed.

[0050] After obtaining the spatial illumination semantic map, perform feature transformation on the feature image in combination with the spatial illumination semantic map to obtain the channel illumination semantic map of the image to be processed. This process will be described in detail in subsequent embodiments.

[0051] In some embodiments, the spatial illumination semantic map and the channel illumination semantic map respectively represent the illumination features in the image to be processed extracted by different methods.

[0052] S205. According to the spatial illumination semantic map and the channel illumination semantic map, perform illumination separation on the image to be processed to obtain the non-illumination features of the image to be processed.

[0053] Illumination separation means distinguishing the illumination features and non-illumination features in the feature image. It can be understood that in this embodiment, the illumination features represent the features related to illumination in the image, and the non-illumination features represent the features unrelated to illumination in the image. The non-illumination features do not change with illumination.

[0054] In this embodiment, the input image to be processed is subjected to illumination separation to obtain the non-illumination features therein. In some embodiments, based on the spatial illumination semantic map and the channel illumination semantic map, the input image to be processed is subjected to illumination separation to obtain the non-illumination features of the input image to be processed, including: performing illumination separation on the input image to be processed according to the spatial illumination semantic map and the channel illumination semantic map to obtain the illumination features of the input image to be processed, and then obtaining the non-illumination features based on the illumination features of the input image to be processed. Further, in some embodiments, obtaining the non-illumination features based on the illumination features of the input image to be processed includes: determining the difference between 1 and the illumination features as the non-illumination features of the input image to be processed. Among them, performing illumination separation on the input image to be processed according to the spatial illumination semantic map and the channel illumination semantic map to obtain the illumination features of the input image to be processed will be described in detail in the subsequent embodiments.

[0055] S206. Perform face recognition based on the non-illumination features to obtain a face recognition result.

[0056] Face recognition is a biometric identification technology that performs identity recognition based on the facial feature information of a person. For an image or video stream containing a human face, the face is automatically detected and tracked in the image, and then face recognition is performed on the detected face. Among them, face recognition can be implemented in any way. For example, a face recognition result can be obtained by performing face recognition on the non-illumination features through a pre-trained face recognition model. In this embodiment, since the input image to be processed is subjected to illumination separation to separate the non-illumination features that do not include illumination features therein, and face recognition is performed on the non-illumination features, the illumination-related features in the image are removed, avoiding the influence of illumination on face recognition in the input image to be processed, thereby improving the accuracy of face recognition.

[0057] According to the technical solution provided by the embodiment of the present disclosure, by performing feature extraction and feature transformation on the input image to be processed, a spatial illumination semantic map and a channel illumination semantic map of the input image to be processed are obtained, and illumination separation is performed according to the spatial illumination semantic map and the channel illumination semantic map to obtain the non-illumination features in the input image to be processed. Finally, face recognition is performed on the non-illumination features to obtain a face recognition result. This method calculates the spatial illumination semantic map and the channel illumination semantic map, uses these two illumination semantic maps for illumination separation, separates the non-illumination features that do not include illumination features from the input image to be processed for face recognition, and can minimize the interference of illumination features in the image on face recognition, thereby obtaining a better face recognition result and improving the accuracy of face recognition.

[0058] In some embodiments, according to the spatial illumination semantic map and the channel illumination semantic map, the illumination of the image to be processed is separated to obtain the non-illumination features of the image to be processed, and at the same time, the illumination features of the image to be processed are obtained. In this embodiment, the face recognition method further includes: determining the illumination category of the image to be processed according to the illumination features; when the illumination category meets the preset illumination conditions, entering the step of performing face recognition based on the non-illumination features to obtain the face recognition result.

[0059] In this embodiment, the illumination features and non-illumination features of the image to be processed are obtained by separating the illumination of the feature image. At the same time, the corresponding illumination category is determined by using the illumination features, and the illumination category is judged to determine whether face recognition needs to be performed.

[0060] Among them, the illumination category means that the illumination is divided into different categories according to the illumination conditions. In some embodiments, the illumination conditions may include the light source. In a specific embodiment, the illumination can be divided into different categories such as natural light on a sunny day, natural light on a cloudy day, white light, yellow light, etc. according to the light source. In other embodiments, the illumination conditions include the illumination angle. In a specific embodiment, the image can be divided into different categories such as front light, side light, back light, etc. according to the illumination angle. In other embodiments, the illumination conditions can also include the illumination intensity. In a specific embodiment, the illumination of the image can be divided into strong light, normal light, weak light, etc. according to the illumination intensity. In other embodiments, the illumination category also includes the category of exposure (which refers to the amount of light entering the lens and irradiating on the photosensitive element during the photography process, and is controlled by the combination of the aperture, shutter, and sensitivity).

[0061] Determining the illumination category of the image to be processed according to the illumination features can be achieved by any method. The preset illumination conditions can be set according to the actual situation. In some embodiments, the preset illumination conditions can be set to that the illumination intensity reaches the preset illumination intensity threshold, the illumination category is front light, etc. Only when the illumination category determined by the illumination features of the image to be processed meets the above conditions, is face recognition allowed to run. In other embodiments, the preset illumination conditions can also be set to other conditions.

[0062] In other embodiments, when the illumination category does not meet the preset illumination conditions, the face recognition of the image to be processed is stopped.

[0063] The face recognition task is usually applied in scenarios for authentication, such as logging in to an account via face recognition or making a payment via face recognition. In such scenarios, successful authentication is achieved through face recognition, and then the account can be logged in successfully, the payment can be completed, etc. These scenarios may all involve user privacy and payment security, etc. Therefore, a higher accuracy rate is required for the face recognition task to ensure user privacy and security. The technical solution provided by the embodiments of the present disclosure, while performing face recognition only using the non-lighting features of the image to be processed, extracts the lighting features from the image to be processed to identify the lighting category of the image to be processed, and restricts the face recognition task based on the lighting category. Only when the lighting category meets the preset lighting conditions is the face recognition task allowed, which can reduce the impact of lighting conditions on face recognition by others and prevent the situation of forging or counterfeiting face images to conduct face recognition and infringing on the rights of others.

[0064] In some embodiments, performing feature transformation on the feature image to obtain the spatial lighting semantic map corresponding to the image to be processed includes: performing parallel multi-scale convolution calculations on the feature image to obtain spatial lighting feature images with different fields of view; and fusing the spatial lighting feature images to obtain the spatial lighting semantic map.

[0065] Among them, convolution is a feature extraction method. From a functional perspective, the convolution process is a process of linearly transforming and mapping to new values at each position of the image. Feature fusion is to optimize and combine different feature vectors extracted from the same pattern, and there are serial and parallel methods, etc.

[0066] In some embodiments, performing parallel multi-scale convolution operations on the feature image includes: respectively performing convolution operations on the feature image of the image to be processed through convolution modules with multiple different convolution kernels simultaneously to obtain spatial lighting feature images with different fields of view. In a specific embodiment, parallel convolution operations are performed on the feature image at 3 scales to obtain spatial lighting feature images with three fields of view.

[0067] Furthermore, in some embodiments, each scale of performing convolution operations on the feature image is completed through a computing module respectively. In a specific embodiment, the structure of each computing module includes, in sequence: a convolution kernel k, a convolution layer with 32 channels, batch normalization, a relu (Rectified Linear Unit) activation layer, and a convolution layer with a convolution kernel of 1×1 and 1 channel, a batch normalization layer, and a softmax layer (normalized exponential calculation). Among them, the convolution kernel k corresponding to each scale can be set differently. Taking three scales corresponding to three computing modules as an example, the convolution kernel k can be set to 3, 5, and 7. In other embodiments, the convolution kernel size of each computing module can be set to other values according to the actual situation.

[0068] In a specific embodiment, the above calculation module can be expressed as:

[0069] f(x; k) = softmax(bn(conv(relu(bn(conv(x; k))); 1)))

[0070] Wherein, x represents the input feature image, k represents the convolution kernel of the calculation module where it is located, conv represents the convolution operation, bn represents the batch normalization operation, and relu represents the activation operation.

[0071] The spatial illumination semantic map can be expressed as:

[0072] A1 = softmax(conv([f(p0; 3), f(p0; 5), f(p0; 7)]; 1))

[0073] Wherein, A1 represents the spatial illumination semantic map, and p0 represents the feature image of the image to be processed.

[0074] Furthermore, in some embodiments, the spatial illumination features corresponding to each field of view are superimposed to obtain the spatial illumination semantic map.

[0075] In the technical solution provided by the embodiments of the present disclosure, by performing parallel multi-scale convolution calculations on the feature image, the spatial illumination features of different fields of view of the image to be processed are obtained, and the spatial illumination features of each field of view are fused to obtain the spatial illumination semantic map, making the calculation more accurate and better extracting the spatial illumination features in the image.

[0076] In some embodiments, based on the spatial illumination semantic map, feature transformation is performed on the feature image to obtain the channel illumination semantic map corresponding to the image to be processed, including: performing convolution on the feature image to obtain an intermediate feature image; using the spatial illumination semantic map as a weight, performing parallel pooling and dimensional transformation on the intermediate feature image to obtain the channel illumination feature and the local channel illumination feature; performing matrix operations on the channel illumination feature and the local channel illumination feature to obtain the channel illumination semantic map.

[0077] In a specific embodiment, performing convolution on the feature image to obtain the intermediate feature image includes: the feature image undergoes convolution calculation with a 1*1 convolution kernel and 64 channels to obtain the intermediate feature image.

[0078] Wherein, pooling is also subsampling, which reduces the size of the data; common pooling methods include maximum pooling, average pooling, etc. In some embodiments, parallel pooling is performed on the intermediate feature image using the spatial illumination semantic map as a weight, that is, multiplying the spatial illumination semantic map by the intermediate feature image and then performing parallel pooling on the product.

[0079] In some embodiments, performing parallel pooling on the weighted feature image includes performing two parallel poolings; in a specific embodiment, taking the dimension of the image features of the image to be processed as (56, 56, 64) as an example, the dimension of the features obtained by one pooling is (7, 7, 64), and the dimension of the features obtained by the other pooling is (1, 1, 64). Further, feature transformation is performed on the features obtained by pooling to obtain channel illumination features with a dimension of (49, 64), and local channel illumination features with a dimension of (1, 64).

[0080] In a specific embodiment, performing matrix operations on the channel illumination features and the local channel illumination features to obtain a channel illumination semantic map includes: performing matrix multiplication on the channel illumination features and the local channel illumination features, performing softmax, then performing matrix multiplication with the local channel illumination features, and then performing softmax to obtain the channel illumination semantic map.

[0081] Let the pooling and dimension transformation operations be P, and the above process of determining the channel illumination semantic map can be expressed as:

[0082] B 0 = P(conv(p 0 )A 1 ; 49)

[0083] B 1 = P(conv(p 0 )A 1 ; 1)

[0084] where p0 represents the feature image to be processed, conv represents the convolution operation, B0 represents the local channel illumination features, and B1 represents the channel illumination features.

[0085]

[0086] where, represents transposing B 0 .

[0087] The technical solution provided by the embodiments of the present disclosure, by multiplying the feature image after convolution with the spatial illumination semantic map, performing parallel pooling operations and dimension transformation operations, obtains channel illumination features and local channel illumination features, and performs matrix operations on the two channel illumination features, calculates the query and encoding of the global channel illumination features for the local channel illumination features, and can obtain the channel illumination semantic map for subsequent illumination separation.

[0088] In some embodiments, such as Figure 3As shown, according to the spatial illumination semantic map and the channel illumination semantic map, the illumination of the image to be processed is separated, and while obtaining the non-illumination features of the image to be processed, the illumination features of the image to be processed are also obtained, including steps S301 to S303, where:

[0089] S301, according to the spatial illumination semantic map and the channel illumination semantic map, respectively obtain a spatial non-illumination semantic map and a channel non-illumination semantic map.

[0090] In some embodiments, according to the spatial illumination semantic map and the channel illumination semantic map, respectively obtaining a spatial non-illumination semantic map and a channel non-illumination semantic map includes: determining the difference between 1 and the spatial illumination semantic map as the spatial non-illumination semantic map, and determining the difference between 1 and the channel illumination semantic map as the channel non-illumination semantic map.

[0091] In a specific embodiment, taking A 1 to represent the spatial illumination semantic map and B 2 to represent the channel illumination semantic map as an example, calculating the spatial non-illumination semantic map and the channel non-illumination semantic map can be expressed as:

[0092] A - = 1 - A 1

[0093] B - = 1 - B 2

[0094] A - represents the spatial non-illumination semantic map, and B - represents the channel non-illumination semantic map.

[0095] S302, multiply the feature image by the spatial non-illumination semantic map and the channel non-illumination semantic map to obtain non-illumination features.

[0096] S303, multiply the feature image by the spatial illumination semantic map and the channel illumination semantic map to obtain illumination features.

[0097] In this embodiment, the feature obtained by multiplying the feature image by the spatial illumination semantic map and the channel illumination semantic map is denoted as the initial illumination feature, and the feature obtained by multiplying the feature image by the spatial non-illumination semantic map and the channel non-illumination semantic map is denoted as the initial non-illumination feature. At the same time, in this embodiment, the process of multiplying the feature image by the spatial and channel semantic maps is denoted as the first stage of illumination separation.

[0098] The determination of the initial illumination feature and the initial non-illumination feature in the above first stage can be expressed as:

[0099] p′ 0 = p 0 ·A 1 ·B2

[0100] p″ 0 = p 0 ·A - ·B -

[0101] p′ 0 represents the initial illumination feature, p″ 0 represents the initial non-illumination feature, p 0 represents the feature image.

[0102] The technical solution provided by the embodiments of the present disclosure, after obtaining the spatial illumination semantic map and the channel illumination semantic map, further determines the spatial non-illumination semantic map and the channel non-illumination semantic map, and performs feature transformation on the feature image by combining each semantic map, separating the illumination feature and the non-illumination feature from the feature image of the image to be processed, dividing the features in the image to be processed into illumination features related to illumination and non-illumination features unrelated to illumination, facilitating face recognition using the illumination-independent illumination features, reducing the influence of illumination on face recognition, and improving the accuracy of face recognition.

[0103] Furthermore, in some embodiments, the non-illumination feature is the stage non-illumination feature of the first stage, and the illumination feature is the stage illumination feature of the first stage; as Figure 4 shown, according to the spatial illumination semantic map and the channel illumination semantic map, performing illumination separation on the image to be processed further includes steps S401 to S405, where:

[0104] S401, taking the stage illumination feature and the stage non-illumination feature of the first stage as the initial illumination feature and the initial non-illumination feature, and inputting them into the second stage.

[0105] Wherein, the first stage and the second stage represent different stages of performing illumination separation. In this embodiment, illumination separation is performed multiple times in each stage to make the effect of illumination separation better as much as possible.

[0106] S402, performing residual processing and illumination separation on the initial illumination feature to obtain the first intermediate illumination feature and the first intermediate non-illumination feature of the second stage, and performing residual processing and illumination separation on the initial non-illumination feature to obtain the second intermediate illumination feature and the second intermediate non-illumination feature of the second stage.

[0107] For the initial illumination feature of the input second stage (the stage illumination feature of the first stage), residual processing and illumination separation are performed in sequence. The initial illumination feature is further split into an illumination feature and a non-illumination feature, which are denoted as the first intermediate illumination feature and the first intermediate non-illumination feature in this embodiment. Similarly, after performing residual processing and illumination separation on the initial non-illumination feature of the input second stage in sequence, a second intermediate illumination feature and a second intermediate non-illumination feature are obtained.

[0108] In a specific embodiment, performing residual processing on the initial non-illumination feature and the initial illumination feature includes: passing through a residual module with a downsampling factor of 2 and an output channel of 128, and then passing through 3 residual modules with a downsampling factor of 1 and an output channel of 128 to obtain the non-illumination feature after residual processing and the illumination feature after residual processing.

[0109] Further, in a specific embodiment, performing illumination separation on the non-illumination feature after residual processing and the illumination feature after residual processing is similar to the process of steps S202 to S204. First, a spatial semantic feature map and a channel semantic feature map are separated from the feature, and then the illumination feature and the non-illumination feature are determined using the semantic map, which will not be elaborated here.

[0110] S403. Obtain the stage illumination feature of the second stage based on the first intermediate illumination feature and the second intermediate illumination feature, and obtain the stage non-illumination feature of the second stage based on the first intermediate non-illumination feature and the second intermediate non-illumination feature.

[0111] In some embodiments, obtaining the stage illumination feature of the second stage based on the first intermediate illumination feature and the second intermediate illumination feature includes: performing weighted summation on the first intermediate illumination feature and the second intermediate illumination feature based on a weight parameter to obtain the stage illumination feature of the second stage.

[0112] In a specific embodiment, performing weighted summation on the first intermediate illumination feature and the second intermediate illumination feature based on a weight parameter to obtain the stage illumination feature of the second stage includes: multiplying the second intermediate illumination feature (the intermediate illumination feature split from the stage non-illumination feature of the first stage) by the weight parameter, and determining the sum value of the product and the first intermediate illumination feature as the stage illumination feature of the second stage.

[0113] Determining the stage illumination feature of the second stage can be expressed as:

[0114] p′ 1 =g′ 1 +σ·h′ 1

[0115] where σ is the weight parameter, g′ 1 represents the first intermediate illumination feature, h′ 1Represents the second intermediate illumination feature, p′ 1 Represents the stage illumination feature of the second stage.

[0116] In some other embodiments, obtaining the stage non-illumination feature of the second stage based on the first intermediate non-illumination feature and the second intermediate non-illumination feature includes: performing weighted summation on the first intermediate non-illumination feature and the second intermediate non-illumination feature based on a weight parameter to obtain the stage non-illumination feature of the second stage.

[0117] In a specific embodiment, performing weighted summation on the first intermediate non-illumination feature and the second intermediate non-illumination feature based on a weight parameter to obtain the stage non-illumination feature of the second stage includes: multiplying the first intermediate non-illumination feature (the intermediate non-illumination feature obtained by splitting the stage illumination feature of the first stage) by the weight parameter, and determining the sum value of the product and the second non-intermediate illumination feature as the stage non-illumination feature of the second stage.

[0118] Determining the stage non-illumination feature of the second stage can be expressed as:

[0119] p″ 1 =σ·g″ 1 +h″ 1

[0120] where σ is the weight parameter, g″ 1 represents the first intermediate non-illumination feature, h″ 1 represents the second intermediate non-illumination feature, p′ 1 represents the stage non-illumination feature of the second stage.

[0121] Among them, the weight parameter can be set to a fixed value according to the actual situation, or can be determined through training; in some embodiments, the weight parameters of different stages can be set to be the same or can be set to be different.

[0122] S404, using the stage non-illumination feature and stage illumination feature of the second stage as the initial non-illumination feature and initial illumination feature, enter the next stage, and repeat the same operations as the second stage.

[0123] After performing the illumination separation of the second stage, it is possible to continue according to the operations of the second stage and perform the illumination separation of the third stage and the fourth stage again, so as to further improve the accuracy of the illumination separation and make the finally obtained illumination feature and non-illumination feature better distinguished. Therefore, in this embodiment, the stage illumination feature and stage non-illumination feature calculated in the second stage are used as the initial illumination feature and initial non-illumination feature of the next stage, and the operations similar to those of the second stage are repeatedly executed, that is, the third stage of the illumination separation.

[0124] After performing the stages corresponding to the preset number, the obtained stage illumination features and stage non-illumination features are the illumination features and non-illumination features respectively.

[0125] Among them, the preset number can be set according to the actual situation. For example, in a specific embodiment, the preset number is set to 4. That is, the stage illumination feature of the fourth stage obtained after the end of the fourth stage is the final illumination feature, and the stage non-illumination feature of the fourth stage is the final non-illumination feature. In other embodiments, the preset number can also be set to other values.

[0126] In some embodiments, the operation processes of the third stage and the fourth stage are similar to that of the second stage, and the parameters therein can be set differently. For example, the number of residual modules and the number of channels of the residual modules are set differently according to the actual situation.

[0127] In the technical solution provided by the embodiments of the present disclosure, after performing illumination semantic separation on the feature image in the first stage and using the semantic map to perform illumination separation on the feature image, the illumination features and non-illumination features obtained in the first stage are used for illumination separation in the second stage, the third stage, and so on. The illumination features and non-illumination features calculated in the first stage are respectively subjected to multiple illumination separations, so that the effect of illumination separation is better. As the number of stages increases, the number of times of illumination separation increases, and the illumination features in the obtained non-illumination features become fewer and fewer. When performing face recognition, a more accurate recognition result can be obtained.

[0128] In some embodiments, the method is implemented by a neural network, and the neural network is obtained by training a preset neural network; in some embodiments, the preset neural network includes at least a first stage, a second stage, a third stage, and a fourth stage.

[0129] Further, as Figure 5 shown, the training process of the neural network includes steps S501 to S505:

[0130] S501, obtain a sample image.

[0131] Among them, the sample image is used to train the preset neural network. In some embodiments, the sample image is an image including a human face.

[0132] Further, in some embodiments, the sample image carries an illumination category annotation and a face category annotation. In a specific embodiment, the illumination category annotation includes: classified by light source as sunny natural light, cloudy natural light, white light, yellow light, and classified by illumination angle as front light, side light, back light. The illumination category annotation includes a total of 4*3 + 1 (exposure) = 13 types. The face category annotation is the user ID (identity identification number) corresponding to the sample image.

[0133] S502. Use the initial weight as the current weight parameter, input the sample image into the preset neural network, and obtain the first prediction probability of face recognition and the first prediction probability of illumination category of the sample image output by the preset neural network.

[0134] Among them, the initial weight can be set according to the actual situation. In a specific embodiment, the initial weight is set to 0.5. The preset neural network extracts features from the input sample image to obtain a sample feature image, and performs at least the first stage, the second stage, the third stage, and the fourth stage of illumination semantic separation and illumination separation on the sample feature image to obtain a sample illumination feature and a sample non-illumination feature. Determine the category of the sample illumination feature, that is, the illumination prediction category of the sample image, and determine the category of the sample non-illumination feature, that is, the face prediction category of the sample image. According to the illumination category annotation and illumination prediction category of the sample image, the first prediction probability of the illumination category of the sample image by the preset neural network can be obtained. According to the face category annotation and face prediction category of the sample image, the first prediction probability of face recognition of the sample image by the preset neural network can be obtained.

[0135] Among them, the processing process of the preset neural network for the sample image, which is the same as the type of the processing process of the trained neural network for the image to be processed, will not be elaborated here.

[0136] S503. Gradually reduce the current weight parameter according to the preset step size to obtain the second prediction probability of face recognition and the second prediction probability of illumination category of the sample image by the preset neural network.

[0137] Adjust the current weight parameter of the preset neural network, and predict the face recognition result and illumination category recognition result of the sample image again, and correspondingly obtain the second prediction probability of face recognition and the second prediction probability of illumination category.

[0138] In some embodiments, gradually reducing the current weight parameter according to the preset step size includes: starting from the fourth stage, each time reducing the current weight parameter of one stage based on the preset step size, and the stages include the second stage, the third stage, or the fourth stage.

[0139] Among them, the preset step size can be set according to the actual situation. In a specific embodiment, the preset step size is set to 0.1. In a specific embodiment, starting from the fourth stage, the current weight parameter of each stage is lowered by one each time. For example, in a specific embodiment, the initial weight parameters of the second, third, and fourth stages are all set to 0.5. Calculate the first prediction probability of face recognition and the first prediction probability of illumination category. Lower the current weight parameter of the fourth stage to 0.4, and the current weight parameters of the second and third stages remain unchanged at 0.5. Calculate the second prediction probability of face recognition and the second prediction probability of illumination category, and so on. Further, in a specific embodiment, after the current weight parameter of the second stage is adjusted to 0.4 and the training converges, return to lower the current weight parameter of the fourth stage again.

[0140] S504. On the premise of ensuring that the second prediction probability of face recognition is the same as the first prediction probability of face recognition, and the second prediction probability of illumination category and the first prediction probability of illumination category are the same, train the preset neural network to convergence, and return to the step of gradually lowering the current weight parameter according to the preset step size.

[0141] In some embodiments, a loss function is set to supervise that the second prediction probability of face recognition is the same as the first prediction probability of face recognition, and the second prediction probability of illumination category and the first prediction probability of illumination category are the same:

[0142] L = max(p 1 -p 3 , 0) + max(p 2 -p 4 , 0)

[0143] Among them, p 1 represents the first prediction probability of face recognition, p 3 represents the second prediction probability of face recognition, p 2 represents the first prediction probability of illumination category, p 4 represents the second prediction probability of illumination category. The value of L is the sum of the larger value of p 1 -p 3 and 0 and the larger value of p 2 -p 4 and 0. In order to ensure that p 1 = p 3 , p 2 = p 4 then p 1 -p 3 = 0, p 2 -p 4 = 0, that is, L needs to be 0 to ensure that the second prediction probability of face recognition is the same as the first prediction probability of face recognition, and the second prediction probability of illumination category and the first prediction probability of illumination category are the same.

[0144] When L meets the corresponding conditions, train the neural network until it converges, and then it can return to adjust the current weight parameters of the next stage.

[0145] Among them, the convergence of the neural network can be supervised by the loss of the predicted value of the illumination feature and the true category of the illumination by the illumination feature extraction network, and the non-illumination feature extraction network is supervised by the loss of the predicted value of the non-illumination feature and the true category of the face ID. When the value of the loss function meets the set conditions, it is determined that the neural network converges. Among them, the loss function can adopt any loss function such as cross-entropy loss.

[0146] S505. Until the preset neural network no longer converges, stop reducing the current weight parameters to obtain the neural network.

[0147] After reducing the current weight parameters, if the neural network can be trained to converge, then it can return to adjust the current weight parameters of the next stage. However, if the neural network cannot be trained to converge, then the current weight parameters will no longer be adjusted, and the current weight parameters are determined as the weight parameters in the neural network.

[0148] The technical solution provided by the embodiments of the present disclosure proposes a method of dynamically associating weights with losses, dynamically reducing the current weight parameters of each stage during the training process, and gradually training the preset neural network corresponding to each weight parameter value until it converges. If the current weight parameters are reduced while the accuracy remains unchanged, then the neural network separates the illumination features and non-illumination features in the image better, thereby training a model with better feature separation effect, which can reduce the influence of illumination features in face recognition more and improve the accuracy of face recognition.

[0149] In other embodiments, the weight parameters of each stage can also be set to fixed values. For example, in a specific embodiment, the weight parameters of the second stage, the third stage, and the fourth stage are set to 0.3, 0.2, and 0.1 respectively, and the preset neural network corresponding to the weight parameters is trained until it converges to obtain the neural network.

[0150] All the above optional technical solutions can be combined arbitrarily to form the optional embodiments of the present application, which will not be elaborated here one by one.

[0151] In a specific embodiment, the steps of the above face recognition method are described with a complete embodiment.

[0152] Train the preset neural network to determine the illumination separation neural network, and the training process is as follows:

[0153] The input of the preset neural network is a face image with a size of 112x112. Among them, the face image carries illumination category annotations. The illumination categories are divided into 4 types from the light source, including natural light on sunny days, natural light on cloudy days, white light, and yellow light; and 3 types from the illumination angle, including frontal light, side light, and back light. There are a total of 4x3 = 12 types, plus exposure, a total of 13 types. The face image training set is labeled with illumination categories according to these 13 types. At the same time, each face image in the face image training set also carries a user ID category annotation, that is, who this face is.

[0154] The face image passes through a residual module (BottleNeck in ResNet) with a downsampling factor of 2 and an output channel of 64, and then passes through 2 residual modules with a downsampling factor of 1 and an output channel of 64 to obtain a feature image p0 of (56, 56, 64).

[0155] Then, p0 enters the "illumination semantic separation" module M0 as the input. p0 first needs to perform "spatial illumination semantics" calculation in M0. The "spatial illumination semantics" calculation is as follows: p0 enters three calculation modules in parallel. The process of each calculation module is: convolution calculation with a convolution kernel k and a channel number of 32, batch normalization, relu activation calculation, and then convolution calculation with a convolution kernel of 1x1 and a channel of 1, batch normalization, and softmax calculation. There is a parameter k above. The k values of the three calculation modules are 3, 5, and 7 respectively. The three calculation modules have three outputs, and the dimensions of the three outputs are all (56, 56, 1). The three outputs are superimposed to obtain A0 with a dimension of (56, 56, 3). A0 passes through convolution calculation with a convolution kernel of 1x1 and a channel number of 1, and then performs softmax calculation to obtain A1 with a dimension of (56, 56, 1). A1 is the spatial illumination semantic map.

[0156] The spatial illumination semantic map can calculate illumination feature images with different fields of view through parallel multi-scale convolution, and then perform fusion calculation on the feature images with different fields of view, making such calculations more accurate.

[0157] Secondly, p0 needs to perform "channel illumination semantics" calculation in M0. The "channel illumination semantics" calculation is as follows: p0 passes through convolution calculation with a convolution kernel of 1x1 and a channel number of 64, then multiplies by A1, and then performs two parallel pooling operations. One output B0 has a dimension changed to (7, 7, 64), and one output B1 has a dimension changed to (1, 1, 64). Here, p0 performs convolution and multiplies by A1 and then pools, which is equivalent to using the spatial illumination semantic map as the weight of local pooling. Then, perform dimension transformation, change the dimension of B0 to (49, 64), change the dimension of B1 to (1, 64), perform matrix multiplication between B1 and B0, then perform softmax, and then perform matrix multiplication with B0 to obtain B2(1, 64), and then perform softmax, and the dimension is further changed to (1, 1, 64).

[0158] In the calculation of the channel illumination semantic map, B1 can be regarded as the channel illumination feature, B0 as the local channel illumination feature, and subsequent matrix operations calculate the query and encoding of the global illumination feature for the local illumination feature. B2 is the channel illumination semantic map.

[0159] Furthermore, the spatial non-illumination semantics A - and the channel non-illumination semantics B - can be obtained by subtracting 1 from them.

[0160] Perform illumination semantic separation on the feature image p0: p0 is multiplied by the spatial illumination semantic map and the channel illumination semantic map to obtain the illumination feature p′ 0 ; p0 is multiplied by the spatial non-illumination semantic map and the channel non-illumination semantic map to obtain the non-illumination feature p″ 0 .

[0161] The above overall calculation process is the first stage of the "illumination separation neural network". The input of the first stage is an image, and the output is the illumination feature and non-illumination feature corresponding to the image.

[0162] In the second stage, a branch calculation scheme is designed. The illumination feature and non-illumination feature obtained from the first stage will respectively pass through two branches (including several residual modules to extract features and an illumination semantic separation module to separate features). Specifically as follows:

[0163] p′ 0 Passes through a residual module with a downsampling factor of 2 and an output channel of 128, and then passes through 3 residual modules with a downsampling factor of 1 and an output channel of 128 to obtain a feature map g of (28, 28, 128) 1 , then, g 1 Enters the illumination separation module to obtain the illumination feature map g′ 1 and the non-illumination feature map g″ 1 ;

[0164] p″ 0 Passes through a residual module with a downsampling factor of 2 and an output channel of 128, and then passes through 3 residual modules with a downsampling factor of 1 and an output channel of 128 to obtain a feature map h of (28, 28, 128) 1 , then, h 1 Enters the illumination separation module to obtain the illumination feature map h′ 1 and the non-illumination feature map h″ 1 ;

[0165] Add g′ 1 and h′ 1 to obtain the stage illumination feature map p′ of the second stage 1 ;

[0166] Add g″ 1 and h″ 1 to obtain the stage non-illumination feature map p″ 1 ;

[0167] When adding, a weight parameter σ is introduced and it can be expressed as:

[0168] p′ 1 = g′ 1 + σ·h′ 1

[0169] p″ 1 = σ·g″ 1 + h″ 1

[0170] In this embodiment, the calculations of the third and fourth stages of the illumination separation neural network are similar to those of the second stage, except for some parameters. The numbers of residual modules are 6 and 4 respectively, and the numbers of channels of the residual modules are 256 and 512 respectively.

[0171] In the fourth stage, the outputs p′ 3 and p″ 3 of the neural network are obtained. Pooling operations are respectively performed on these two feature maps, and then the operations of connecting two fully connected layers are carried out for classification prediction. The illumination feature map p′ 3 is pooled and fully connected to predict the illumination category of the face image, and the non-illumination feature map p″ 3 is pooled and fully connected to predict the ID category of the face image. The extraction network of the illumination feature is supervised by the cross-entropy loss between the predicted illumination feature value and the true illumination category; the extraction network of the non-illumination feature is supervised by the cross-entropy loss between the predicted non-illumination feature value and the true face ID category.

[0172] In the above technical solution, in the second, third, and fourth stages of the neural network, there are two branches. One branch calculates the illumination feature map, and the other branch calculates the non-illumination feature map.

[0173] In this embodiment, in order to achieve the purpose of separating the illumination feature and ensure the accuracy of model prediction, a dynamic weight correlation loss scheme is proposed:

[0174] First, set the weight parameter σ of the second, third, and fourth stages to 0.5, and let the preset neural network be trained until convergence at this time.

[0175] Second, for the input pictures, first calculate the probabilities p 1 and p 2 of the two branches with σ all being 0.5 for the output of this category, and then adjust the σ of the fourth stage to 0.4 to obtain the probabilities p 3 and p4 . In this embodiment, after reducing σ, the probability of the output corresponding category cannot decrease, which indicates that the model can well separate the illumination feature and the non-illumination feature. The following loss function can be used to constrain that the probability of the output corresponding category cannot decrease after reducing σ:

[0176] L = max(p 1 - p 3 , 0) + max(p 2 - p 4 , 0)

[0177] If the accuracy remains unchanged after reducing σ, then the neural network separates features better. Further, after the model converges, the third stage can be reduced to 0.4, and then the second stage can be reduced. And so on. If the model cannot converge to the same accuracy after reducing σ in a certain stage, then this stage will no longer be reduced, and the final illumination separation neural network is obtained.

[0178] Secondly, after the model training is completed, the trained neural network is used to perform illumination separation and face recognition on the image to be processed. Input a face image (image to be processed) into the neural network, and the corresponding illumination feature and non-illumination feature can be obtained. Among them, the illumination type of the image to be processed can be determined from the illumination feature; secondly, due to the existence of the illumination separation technology, the non-illumination feature can avoid the interference of illumination on other significant features of the face. Therefore, the features excluding the illumination interference can be well extracted, and the non-illumination feature can be used for face comparison or face recognition to obtain better face recognition results. Further, according to the illumination category, if the preset illumination condition is not satisfied, the face recognition can be rejected. At this time, it may be that the network is attacked, or the illumination is too extreme and the recognition is inaccurate, which can prevent unknown risks.

[0179] The following is an embodiment of the device of the present disclosure, which can be used to execute the method embodiment of the present disclosure. For the details not disclosed in the device embodiment of the present disclosure, please refer to the method embodiment of the present disclosure.

[0180] Figure 6 is a schematic diagram of a face recognition device provided by an embodiment of the present disclosure. As Figure 6 shown, the face recognition device includes:

[0181] An acquisition module 601, configured to acquire an image to be processed.

[0182] A feature extraction module 602, configured to perform feature extraction on the image to be processed to obtain a feature image of the image to be processed.

[0183] A first feature transformation module 603, configured to perform feature transformation on the feature image to obtain a spatial illumination semantic map corresponding to the image to be processed.

[0184] The second feature transformation module 604 is configured to perform feature transformation on the feature image based on the spatial illumination semantic map to obtain the channel illumination semantic map corresponding to the image to be processed.

[0185] The illumination separation module 605 is configured to perform illumination separation on the image to be processed according to the spatial illumination semantic map and the channel illumination semantic map to obtain the non-illumination features of the image to be processed.

[0186] The face recognition module 606 is configured to perform face recognition based on the non-illumination features to obtain the face recognition result.

[0187] According to the technical solution provided by the embodiments of the present disclosure, by performing feature extraction and feature transformation on the image to be processed, the spatial illumination semantic map and the channel illumination semantic map of the image to be processed are obtained, and illumination separation is performed according to the spatial illumination semantic map and the channel illumination semantic map to obtain the non-illumination features in the image to be processed. Finally, face recognition is performed on the non-illumination features to obtain the face recognition result. The device calculates the spatial illumination semantic map and the channel illumination semantic map, uses these two illumination semantic maps for illumination separation, separates the non-illumination features that do not include illumination features from the image to be processed for face recognition, and can minimize the interference of the illumination features in the image on face recognition, thereby obtaining a better face recognition result and improving the accuracy of face recognition.

[0188] In some embodiments, the illumination separation module 605 of the above device is further configured to obtain the illumination features of the image to be processed; in this embodiment, the above device further includes: an illumination category prediction module 607, configured to determine the illumination category of the image to be processed according to the illumination features; when the illumination category meets the preset illumination condition, the face recognition module 606 enters the step of performing face recognition based on the non-illumination features to obtain the face recognition result.

[0189] In some embodiments, the first feature transformation module 603 of the above device includes: a convolution sub-module 608, configured to perform parallel multi-scale convolution calculations on the feature image to obtain spatial illumination feature images with different fields of view; a fusion sub-module 609, configured to fuse the spatial illumination feature images to obtain the spatial illumination semantic map.

[0190] In some embodiments, the second feature transformation module 604 of the above device includes: a convolution sub-module 610, configured to perform convolution on the feature image to obtain an intermediate feature image; a feature processing sub-module 611, configured to use the spatial illumination semantic map as a weight to perform parallel pooling and dimensional transformation on the intermediate feature image to obtain channel illumination features and local channel illumination features; a matrix operation sub-module 612, configured to perform matrix operations on the channel illumination features and the local channel illumination features to obtain the channel illumination semantic map.

[0191] In some embodiments, the light separation module 605 of the above device includes: a non-light semantic map determination sub-module 613, configured to obtain a spatial non-light semantic map and a channel non-light semantic map according to the spatial light semantic map and the channel light semantic map respectively; a non-light feature determination sub-module 614, configured to multiply the feature image by the spatial non-light semantic map and the channel non-light semantic map to obtain a non-light feature; a light feature determination sub-module 615, configured to multiply the feature image by the spatial light semantic map and the channel light semantic map to obtain a light feature.

[0192] In some embodiments, the non-light feature is the stage non-light feature of the first stage, and the light feature is the stage light feature of the first stage. In this embodiment, the light separation module 605 of the above device further includes: inputting the stage light feature and the stage non-light feature of the first stage as the initial light feature and the initial non-light feature into the second stage; a light separation sub-module 616, configured to perform residual processing and light separation on the initial light feature to obtain the first intermediate light feature and the first intermediate non-light feature of the second stage, and perform residual processing and light separation on the initial non-light feature to obtain the second intermediate light feature and the second intermediate non-light feature of the second stage; a stage feature determination sub-module 617, configured to obtain the stage light feature of the second stage according to the first intermediate light feature and the second intermediate light feature, and obtain the stage non-light feature of the second stage according to the first intermediate non-light feature and the second intermediate non-light feature; a loop sub-module 618, configured to use the stage non-light feature and the stage light feature of the second stage as the initial non-light feature and the initial light feature, enter the next stage, and repeat the same operations as in the second stage; after performing the stages corresponding to the preset number, the obtained stage light feature and stage non-light feature are the light feature and the non-light feature respectively.

[0193] In some embodiments, the stage feature determination sub-module of the above device is further configured to: perform weighted summation on the first intermediate light feature and the second intermediate light feature based on the weight parameter to obtain the stage light feature of the second stage; perform weighted summation on the first intermediate non-light feature and the second intermediate non-light feature based on the weight parameter to obtain the stage non-light feature of the second stage.

[0194] In some embodiments, the above device further includes a model training module 619, including:

[0195] an image acquisition sub-module, configured to acquire a sample image;

[0196] A prediction probability determination sub-module, configured to use the initial weight as the current weight parameter for each stage, input a sample image into a preset neural network, and obtain a first face recognition prediction probability and a first illumination category prediction probability of the sample image output by the preset neural network;

[0197] A weight adjustment sub-module, configured to gradually reduce the current weight parameter step by step according to a preset step size, and obtain a second face recognition prediction probability and a second illumination category prediction probability of the sample image by the preset neural network;

[0198] A training sub-module, configured to train the preset neural network to convergence on the premise that the second face recognition prediction probability is the same as the first face recognition prediction probability, and the second illumination category prediction probability is the same as the first illumination category prediction probability, and return the step of gradually reducing the current weight parameter step by step according to the preset step size; until the preset neural network no longer converges, stop reducing the current weight parameter to obtain a neural network.

[0199] In some embodiments, the weight adjustment sub-module of the above device is further configured to: starting from the fourth stage, each time reduce the current weight parameter of one stage based on the preset step size, and the stages include the second stage, the third stage or the fourth stage.

[0200] It should be understood that the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.

[0201] Figure 7 is a schematic diagram of the electronic device 7 provided by the embodiments of the present disclosure. As Figure 7 shown, the electronic device 7 of this embodiment includes: a processor 701, a memory 702, and a computer program 703 stored in the memory 702 and operable on the processor 701. When the processor 701 executes the computer program 703, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 701 executes the computer program 703, the functions of each module / unit in the above-mentioned device embodiments are implemented.

[0202] Exemplarily, the computer program 703 can be divided into one or more modules / units. One or more modules / units are stored in the memory 702 and executed by the processor 701 to complete the present disclosure. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 703 in the electronic device 7.

[0203] The electronic device 7 can be a desktop computer, a notebook, a palm computer, a cloud server, or other electronic devices. The electronic device 7 may include, but is not limited to, a processor 701 and a memory 702. Those skilled in the art can understand that Figure 7 These are merely examples of the electronic device 7 and do not constitute a limitation on the electronic device 7. It may include more or fewer components than those shown in the figure, or combine certain components, or have different components. For example, the electronic device may also include input / output devices, network access devices, a bus, etc.

[0204] The processor 701 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0205] The memory 702 can be an internal storage unit of the electronic device 7. For example, the hard disk or memory of the electronic device 7. The memory 702 can also be an external storage device of the electronic device 7. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the electronic device 7. Further, the memory 702 can also include both the internal storage unit and the external storage device of the electronic device 7. The memory 702 is used to store computer programs and other programs and data required by the electronic device. The memory 702 can also be used to temporarily store data that has been output or is to be output.

[0206] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0207] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0208] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this disclosure.

[0209] In the embodiments provided by this disclosure, it should be understood that the disclosed device / electronic device and method can be implemented in other ways. For example, the device / electronic device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there can be other division methods. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0210] The unit described as a separated component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0211] In addition, in each embodiment of the present disclosure, each functional unit may be integrated into one processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0212] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above-mentioned embodiment methods of the present disclosure may also be completed by instructing relevant hardware through a computer program. The computer program may be stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments may be implemented. The computer program may include computer program code, and the computer program code may be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0213] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present disclosure, and should all be included in the protection scope of the present disclosure.

Claims

1. A face recognition method, characterized in that, comprising: obtaining an image to be processed; performing feature extraction on the image to be processed to obtain a feature image of the image to be processed; performing feature transformation on the feature image to obtain a spatial illumination semantic map corresponding to the image to be processed; performing feature transformation on the feature image based on the spatial illumination semantic map to obtain a channel illumination semantic map corresponding to the image to be processed; performing illumination separation on the image to be processed according to the spatial illumination semantic map and the channel illumination semantic map to obtain a non-illumination feature of the image to be processed; performing face recognition based on the non-illumination feature to obtain a face recognition result; The performing feature transformation on the feature image to obtain a spatial illumination semantic map corresponding to the image to be processed includes: performing parallel multi-scale convolution calculation on the feature image to obtain spatial illumination feature images with different fields of view; fusing the spatial illumination feature images to obtain the spatial illumination semantic map; The performing feature transformation on the feature image based on the spatial illumination semantic map to obtain a channel illumination semantic map corresponding to the image to be processed includes: performing convolution on the feature image to obtain an intermediate feature image; using the spatial illumination semantic map as a weight to perform parallel pooling and dimensional transformation on the intermediate feature image to obtain a channel illumination feature and a local channel illumination feature; performing matrix operations on the channel illumination feature and the local channel illumination feature to obtain the channel illumination semantic map; The performing illumination separation on the image to be processed according to the spatial illumination semantic map and the channel illumination semantic map includes: respectively obtaining a spatial non-illumination semantic map and a channel non-illumination semantic map according to the spatial illumination semantic map and the channel illumination semantic map; multiplying the feature image by the spatial non-illumination semantic map and the channel non-illumination semantic map to obtain the non-illumination feature; multiplying the feature image by the spatial illumination semantic map and the channel illumination semantic map to obtain the illumination feature.

2. The method according to claim 1, characterized in that, when performing illumination separation on the image to be processed according to the spatial illumination semantic map and the channel illumination semantic map to obtain a non-illumination feature of the image to be processed, an illumination feature of the image to be processed is also obtained; The method further includes: determining an illumination category of the image to be processed according to the illumination feature; when the illumination category meets a preset illumination condition, entering the step of performing face recognition based on the non-illumination feature to obtain a face recognition result.

3. The method according to claim 1, characterized in that, the non-illumination feature is a stage non-illumination feature of the first stage, and the illumination feature is a stage illumination feature of the first stage; The performing illumination separation on the image to be processed according to the spatial illumination semantic map and the channel illumination semantic map further includes: inputting the stage illumination feature and the stage non-illumination feature of the first stage as an initial illumination feature and an initial non-illumination feature into a second stage; Perform residual processing and light separation on the initial light feature to obtain the first intermediate light feature and the first intermediate non-light feature in the second stage. Perform residual processing and light separation on the initial non-light feature to obtain the second intermediate light feature and the second intermediate non-light feature in the second stage; Obtain the stage light feature in the second stage according to the first intermediate light feature and the second intermediate light feature, and obtain the stage non-light feature in the second stage according to the first intermediate non-light feature and the second intermediate non-light feature; Use the stage non-light feature and the stage light feature in the second stage as the initial non-light feature and the initial light feature, enter the next stage, and repeat the same operations as in the second stage; After performing the stages corresponding to the preset number, the obtained stage light feature and stage non-light feature are the light feature and the non-light feature respectively.

4. The method according to claim 3, wherein: The obtaining the stage light feature in the second stage according to the first intermediate light feature and the second intermediate light feature includes: performing weighted summation on the first intermediate light feature and the second intermediate light feature based on the weight parameter to obtain the stage light feature in the second stage; The obtaining the stage non-light feature in the second stage according to the first intermediate non-light feature and the second intermediate non-light feature includes: performing weighted summation on the first intermediate non-light feature and the second intermediate non-light feature based on the weight parameter to obtain the stage non-light feature in the second stage.

5. The method according to claim 4, wherein, The method is implemented by a neural network; the neural network is obtained by training a preset neural network; the training process of the neural network includes: Obtain a sample image; Using the initial weight as the current weight parameter for each stage, input the sample image into the preset neural network to obtain the first prediction probability of face recognition and the first prediction probability of light category of the sample image output by the preset neural network; Gradually reduce the current weight parameter according to the preset step size to obtain the second prediction probability of face recognition and the second prediction probability of light category of the sample image by the preset neural network; On the premise that the second prediction probability of face recognition is the same as the first prediction probability of face recognition, and the second prediction probability of light category is the same as the first prediction probability of light category, train the preset neural network until it converges, and return to the step of gradually reducing the current weight parameter according to the preset step size; Until the preset neural network no longer converges, stop reducing the current weight parameter to obtain the neural network.

6. The method according to claim 5, wherein, The gradually reducing the current weight parameter according to the preset step size includes: Starting from the fourth stage, each time reduce the current weight parameter of one stage based on the preset step size, and the stage includes the second stage, the third stage or the fourth stage.

7. A face recognition device, wherein, comprises: An acquisition module configured to acquire an image to be processed; A feature extraction module, configured to extract features from the to-be-processed image to obtain a feature image of the to-be-processed image; A first feature transformation module, configured to perform feature transformation on the feature image to obtain a spatial illumination semantic map corresponding to the to-be-processed image, including: performing parallel multi-scale convolution calculations on the feature image to obtain spatial illumination feature images with different fields of view; fusing the spatial illumination feature images to obtain the spatial illumination semantic map; A second feature transformation module, configured to perform feature transformation on the feature image based on the spatial illumination semantic map to obtain a channel illumination semantic map corresponding to the to-be-processed image, including: performing convolution on the feature image to obtain an intermediate feature image; using the spatial illumination semantic map as a weight to perform parallel pooling and dimensional transformation on the intermediate feature image to obtain a channel illumination feature and a local channel illumination feature; performing matrix operations on the channel illumination feature and the local channel illumination feature to obtain the channel illumination semantic map; An illumination separation module, configured to perform illumination separation on the to-be-processed image according to the spatial illumination semantic map and the channel illumination semantic map to obtain a non-illumination feature of the to-be-processed image, including: obtaining a spatial non-illumination semantic map and a channel non-illumination semantic map according to the spatial illumination semantic map and the channel illumination semantic map respectively; multiplying the feature image by the spatial non-illumination semantic map and the channel non-illumination semantic map to obtain the non-illumination feature; multiplying the feature image by the spatial illumination semantic map and the channel illumination semantic map to obtain the illumination feature; A face recognition module, configured to perform face recognition based on the non-illumination feature to obtain a face recognition result.

8. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium storing a computer program, wherein, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Face recognition method and system based on deep learning

    CN112800872A

  • Image processing method, model training method, device and equipment

    CN113516592A