Face Recognition Method, Device, Equipment and Storage Medium

By introducing a lightweight multi-scale feature fusion structure into the face recognition network, the problem of slow running speed caused by the large number of parameters of the deep neural network face recognition model is solved, and the demand for high-performance face recognition on mobile devices is realized.

CN115359521BActive Publication Date: 2025-06-24PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210866085.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-22
Publication Date
2025-06-24
Estimated Expiration
2042-07-22

AI Technical Summary

Technical Problem

In the prior art, when using neural networks with high accuracy for face recognition, the model parameters are large, which reduces the model's running speed and makes the model inconvenient to deploy on mobile devices with high real-time requirements.

Method used

By inputting the face image dataset to a face recognition network based on lightweight multi-scale feature fusion for model training, the image features are extracted using the lightweight network, and the adaptive feature fusion structure is introduced to fuse features of different scales to reduce the model size while maintaining recognition accuracy.

Benefits of technology

While significantly reducing the size of the model, the accuracy of the model is maintained, the memory and computing amount required for the face recognition network is reduced, and it is suitable for high-performance face recognition on mobile and embedded devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359521B_ABST
    Figure CN115359521B_ABST
Patent Text Reader

Abstract

The present invention relates to artificial intelligence technology and discloses a face recognition method, including: training a face recognition model through a face image dataset; constructing a loss function model by using the true bounding box of a face target, the width, height, and center point coordinates of the true bounding box of the face target, and the identity category of the true bounding box of the face target; performing optimization training on the loss function model by using the gradient descent method; inputting a face image to be recognized into a face recognition network model based on lightweight multi-scale feature fusion for prediction, to obtain a predicted bounding box of the face target in the face image to be recognized, an identity category corresponding to the predicted bounding box, and a recognition accuracy corresponding to the identity category. The present invention also relates to blockchain technology, and the face image dataset is stored in a blockchain. The present invention can solve the problems in the prior art that the number of model parameters is large, the running speed of the model is reduced, and the model is not convenient to be deployed on mobile devices with high real-time requirements, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and particularly to a face recognition method, device, equipment and storage medium. Background Art

[0002] Face recognition is a biometric-based recognition technology for identity authentication. Since its development in the 1970s, it has been one of the most popular research topics in the field of computer vision. With the proposal of concepts such as smart cities and human-computer interaction, the research on face recognition technology is of great significance and has been widely applied in fields such as security monitoring, mobile payment, and driverless.

[0003] The development of face recognition technology mainly goes through three stages: The first stage uses methods such as subspace, geometric features, and template matching. However, these methods require manual feature extraction and cannot achieve fully automatic face recognition; the second stage uses a combination of a face feature extractor and a feature classifier. However, the features extracted by these methods are too single to meet complex face recognition scenarios; the third stage, that is, the current stage, with the wide application of deep learning technology in the field of images, face recognition algorithms based on deep neural networks have obvious advantages in detection accuracy compared with traditional algorithms.

[0004] In order to pursue better performance, researchers usually use deeper neural networks for face recognition. Although this method can improve the accuracy of face recognition algorithms, it also increases the number of model parameters, reduces the model running speed, and makes the model not easy to be deployed on mobile devices with high real-time requirements. Therefore, based on the above problems, a method needs to be designed to make the face recognition model occupy less memory and maintain a high recognition accuracy for mobile embedded devices with limited storage and computing power. Summary of the Invention

[0005] The present invention provides a face recognition method, device, equipment and storage medium, and its main purpose is to solve the problems in the prior art that when using a neural network with high accuracy for face recognition, the number of model parameters is large, the model running speed is reduced, and the model is not easy to be deployed on mobile devices with high real-time requirements.

[0006] In a first aspect, to achieve the above object, a face recognition method provided by the present invention includes:

[0007] Input the face image dataset into the face recognition network based on lightweight multi-scale feature fusion for training the face recognition model; wherein, the face image dataset includes face pictures, the true bounding boxes of face targets annotated on the face pictures, the widths, heights and center point coordinates of the true bounding boxes of face targets, and the identity categories marked on the true bounding boxes of face targets;

[0008] Within the face recognition network after training the face recognition model, construct a loss function model through the true bounding boxes of face targets, the widths, heights and center point coordinates of the true bounding boxes of face targets, and the identity categories of the true bounding boxes of face targets;

[0009] Use the gradient descent method to optimize and train the loss function model. When the loss function of the loss function model reaches a preset threshold, obtain the face recognition network model based on lightweight multi-scale feature fusion; wherein,

[0010] The face recognition network model based on lightweight multi-scale feature fusion includes: a backbone network layer for extracting features of three different image dimensions from a face image, a pooling layer for performing pooling processing on the third output feature obtained by the backbone network layer, a feature fusion layer for performing feature fusion processing on the first output feature, the second output feature obtained by the backbone network layer, and the pooled third output feature obtained by the pooling layer respectively, and a detection head layer for generating a prediction result according to the three fused features obtained by the feature fusion layer;

[0011] Input the face image to be recognized into the face recognition network model based on lightweight multi-scale feature fusion for face recognition, and obtain the predicted bounding box of the face target in the face image to be recognized, the identity category corresponding to the predicted bounding box, and the recognition accuracy corresponding to the identity category.

[0012] In a second aspect, to solve the above problems, the present invention also provides a face recognition device, and the device includes:

[0013] A training module, configured to input the face image dataset into the face recognition network based on lightweight multi-scale feature fusion for training the face recognition model; wherein, the face image dataset includes face pictures, the true bounding boxes of face targets annotated on the face pictures, the widths, heights and center point coordinates of the true bounding boxes of face targets, and the identity categories marked on the true bounding boxes of face targets;

[0014] A loss function construction module, which is used to construct a loss function model in the face recognition network after the face recognition model is trained, by using the true bounding box of the face target, the width, height, and center point coordinates of the true bounding box of the face target, and the identity category of the true bounding box of the face target;

[0015] An optimization module, which is used to optimize and train the loss function model by using the gradient descent method. When the loss function of the loss function model reaches a preset threshold, a face recognition network model based on lightweight multi-scale feature fusion is obtained; wherein,

[0016] The face recognition network model based on lightweight multi-scale feature fusion includes: a backbone network layer for extracting features of three different image dimensions of a face image, a pooling layer for performing pooling processing on the third output feature obtained by the backbone network layer, a feature fusion layer for respectively performing feature fusion processing on the first output feature, the second output feature obtained by the backbone network layer, and the pooled third output feature obtained by the pooling layer, and a detection head layer for generating a prediction result according to the three fusion features obtained by the feature fusion layer;

[0017] A prediction module, which inputs the face image to be recognized into the face recognition network model based on lightweight multi-scale feature fusion for face recognition, and obtains the predicted bounding box of the face target in the face image to be recognized, the identity category corresponding to the predicted bounding box, and the recognition accuracy corresponding to the identity category.

[0018] In a third aspect, to solve the above problems, the present invention further provides an electronic device, which includes:

[0019] At least one processor; and,

[0020] A memory communicatively connected to the at least one processor; wherein,

[0021] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the face recognition method as described above.

[0022] In a fourth aspect, to solve the above problems, the present invention further provides a computer-readable storage medium, storing a computer program, and when the computer program is executed by a processor, it implements the face recognition method as described above.

[0023] The face recognition method, device, equipment and storage medium proposed by the present invention input a face image data set into a face recognition network based on lightweight multi-scale feature fusion for face recognition training, and obtain a face recognition network model based on lightweight multi-scale feature fusion. Because it uses a lightweight network to extract image features and introduces an adaptive feature fusion structure to fuse features of different scales (different dimensions), it can maintain the accuracy of the model while significantly reducing the model size, realizing high-performance face recognition on mobile terminals and embedded devices. The face recognition network model trained by the present invention has fewer parameters, reduces the memory size required for the face recognition network, reduces the computational amount of the network, strengthens the feature fusion effect, shortens the model training time, improves the accuracy of the model, and has universality, and can be extended to similar tasks such as pedestrian detection and expression recognition. Description of the Drawings

[0024] Figure 1 It is a schematic flowchart of the face recognition method provided by an embodiment of the present invention;

[0025] Figure 1.1 It is an overall structure diagram of the face recognition network based on lightweight multi-scale feature fusion provided by an embodiment of the present invention;

[0026] Figure 1.2 It is an overall structure diagram of the structure of the original YOLOv4 network;

[0027] Figure 2 It is a module schematic diagram of the face recognition device provided by an embodiment of the present invention;

[0028] Figure 3 It is an internal structure schematic diagram of an electronic device for implementing the face recognition method provided by an embodiment of the present invention;

[0029] The realization, functional characteristics and advantages of the purpose of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments

[0030] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0031] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.

[0032] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0033] The present invention provides a face recognition method. Refer to Figure 1 As shown, it is a schematic flowchart of the face recognition method provided by an embodiment of the present invention. This method can be executed by a device, and the device can be implemented by software and / or hardware.

[0034] In this embodiment, the face recognition method includes:

[0035] Step S110: Input the face image dataset into a face recognition network based on lightweight multi-scale feature fusion for face recognition model training; wherein, the face image dataset includes face pictures, the true bounding boxes of face targets marked on the face pictures, the widths, heights, and center point coordinates of the true bounding boxes of face targets, and the identity categories marked on the true bounding boxes of face targets.

[0036] Specifically, the face image dataset includes multiple face pictures, and each face picture includes at least one face image. Its sources can be face pictures collected through social networks or face pictures specially taken for actual purposes. Each face on each face picture is marked with the true bounding box of the face target, and the width, height, and center point coordinates of each true bounding box are calculated. Moreover, the identity category of each true bounding box needs to be marked. Among them, the identity category can be the identity information of the face, such as the ID number and name, or other identity information required for face recognition, such as the job number code of the corresponding employee in the company's employee attendance device, etc.

[0037] As an optional embodiment of the present invention, the face image dataset is stored in the blockchain. Before inputting the face image dataset into a face recognition network based on lightweight multi-scale feature fusion for face recognition model training, it further includes:

[0038] Collect face image samples to obtain a face image sample set;

[0039] Mark the true bounding boxes of face targets in the face image sample set, calculate the widths, heights, and center point coordinates of the marked true bounding boxes of face targets, and mark the identity categories of the true bounding boxes of face targets to obtain the face image dataset.

[0040] Specifically, according to the actual application scenarios of face recognition, corresponding face image samples are collected. For example, for places that require identity verification such as stations and airports, face images and corresponding information can be collected through the system. At this time, each picture can include only one person. For employee attendance, employee face pictures can be directly collected, and each picture can have one face or multiple faces. The multiple face image samples collected constitute a face image sample set; for each face image on each face picture in the face image sample set, true border annotation is performed, which can be completed using picture files in different formats. For example, for picture files in TXT format, the face coordinates are recorded in the picture file, and the true border of the face target is marked through the face coordinates. Then, the width, height, and center point coordinates of the true border are calculated, and the identity category of the true border of each marked face target is marked.

[0041] As an optional embodiment of the present invention, for the true border of the face target in the face image sample set, annotation is performed, the width, height, and center point coordinates of the true border of the annotated face target are calculated, and the identity category of the true border of the face target is marked, obtaining a face image data set including:

[0042] Define the face image sample set as: {Data k (x,y), k∈[1,K], x∈[1,X], y∈[1,Y]}; where Data k (x,y) represents the pixel information of the x-th row and y-th column of the k-th picture in the face image sample set, K represents the number of pictures in the face image sample set, X represents the number of rows of pixels in the pictures in the face image sample set, and Y represents the number of columns of pixels in the pictures in the face image sample set;

[0043] For each face target's true border in each picture in the defined face image sample set, annotation is performed; where the true border of the face target is defined as:

[0044]

[0045]

[0046] Where represents the upper left corner coordinates of the true border of the n-th face target in the k-th picture in the face image sample set, represents the abscissa of the upper left corner coordinate point of the true border of the n-th face target in the k-th picture in the face image sample set, represents the ordinate of the upper left corner coordinate point of the true border of the n-th face target in the k-th picture in the face image sample set; Denote the right-bottom coordinate of the ground truth bounding box of the n-th face object in the k-th picture in the face image sample set. Denote the abscissa of the right-bottom coordinate point of the ground truth bounding box of the n-th face object in the k-th picture in the face image sample set. Denote the ordinate of the right-bottom coordinate point of the ground truth bounding box of the n-th face object in the k-th picture in the face image sample set; K denotes the number of pictures in the face image sample set; N k Denote the number of ground truth bounding boxes of the face objects in the k-th picture in the face image training dataset selected for training the preset face recognition network model in the face image dataset.

[0047] According to the preset ground truth bounding box width calculation formula, the preset ground truth bounding box height calculation formula, and the preset ground truth bounding box center point coordinate calculation formula, calculate the width, height, and center point coordinates of the ground truth bounding box of the face object respectively, and mark the identity category of the ground truth bounding box of the face object to obtain the face image dataset; where

[0048] The preset ground truth bounding box width calculation formula is:

[0049] The preset ground truth bounding box height calculation formula is:

[0050] The preset ground truth bounding box center point coordinate calculation formula is:

[0051]

[0052] The identity category of the ground truth bounding box of the face object is defined as:

[0053] Where Denote the width of the ground truth bounding box of the n-th face object in the k-th image in the face image sample set. Denote the height of the ground truth bounding box of the n-th face object in the k-th image in the face image sample set, and t represents ground truth.

[0054] Specifically, by first defining the face image sample set, each picture in the face image sample set and each face image on each picture can be represented by their corresponding positions, defining the ground truth bounding box of the face object according to the position of the face image, then calculating its height, width, and center coordinates based on the ground truth bounding box, and defining the identity category of the ground truth bounding box, the picture information is transformed into digital information that can be operated on, thus facilitating subsequent calculations.

[0055] Step S120: Within the face recognition network after training the face recognition model, construct a loss function model based on the true bounding box of the face target, the width and height of the true bounding box of the face target, the center point coordinates, and the identity category of the true bounding box of the face target.

[0056] Specifically, the loss function model is used to estimate the degree of inconsistency between the predicted value f(x) of the face recognition network obtained through training and the true value Y. It is a non - negative real - valued function, denoted by L(Y, f(x)). The smaller the loss function, the better the robustness of the trained network. The loss function is the core part of the empirical risk function and an important part of the structural risk function. The structural risk function of the model includes an empirical risk term and a regularization term.

[0057] To ensure the accuracy of the face recognition network model obtained after training the face recognition network, it is necessary to construct a loss function model within the face recognition network obtained after training. During the process of training the model, the face image dataset can be divided into two parts, one as the face image training dataset and the other as the face image verification and optimization dataset. First, train the face recognition model of the face recognition network through the face image training dataset, and then optimize it through the face image verification and optimization dataset.

[0058] As an optional embodiment of the present invention, the loss function model includes: a target bounding box loss model;

[0059] The calculation formula of the target bounding box loss model is:

[0060]

[0061] Where, represents the Euclidean distance between the center points of the predicted bounding box and the true bounding box of the face target, c represents the diagonal length of the smallest rectangle that can cover the predicted bounding box and the true bounding box of the face target, and the calculation method of c is:

[0062]

[0063] IOU represents the intersection - over - union ratio of the predicted bounding box and the true box of the face target, C h and C W respectively represent the height and width of the smallest rectangle that can cover the predicted bounding box and the true bounding box of the face target, h and h gt respectively represent the height of the predicted bounding box of the face target and the height of the true bounding box of the face target, w and w gtrespectively represent the height of the predicted bounding box of the face target and the width of the ground truth bounding box of the face target. The calculation methods of \(v\) and \(\alpha\) in the formula are as follows:

[0064]

[0065] The width and height of the predicted bounding box of each face target in each face image in the face image dataset are respectively defined as:

[0066] and

[0067] The center point coordinates of the predicted bounding box of each face target in each face image in the face image dataset are defined as:

[0068]

[0069] Specifically, through the above-created target bounding box loss model, the accuracy of the labeled position of the predicted bounding box of the face target in the face image to be recognized can be optimized.

[0070] As an optional embodiment of the present invention, the loss function model further includes: a target confidence loss model;

[0071] The calculation formula of the target confidence loss model is:

[0072]

[0073] where represents the identity category within the ground truth bounding box of the \(n\)th face target in the \(k\)th image in the face image dataset, represents the identity category within the predicted bounding box of the \(n\)th face target in the \(k\)th image in the face image dataset, and \(\lambda\) noobject represents the confidence penalty weight when there is no target in the predicted bounding box of the face target.

[0074] Specifically, through the above-created target confidence loss model, the accuracy of the predicted target of the predicted bounding box of the face image to be recognized can be optimized.

[0075] As an optional embodiment of the present invention, the loss function model further includes: a target category loss model;

[0076] The calculation formula of the target category loss model is:

[0077]

[0078] where It represents the confidence of the identity category within the true bounding box of the nth face target in the kth image of the face image dataset. It represents the confidence of the identity category within the predicted bounding box of the nth face target in the kth image of the face image dataset;

[0079] The predicted bounding box of each face target in each face image of the face image dataset is defined as:

[0080]

[0081]

[0082] Among them, It represents the upper - left coordinate of the predicted bounding box of the nth face target in the kth image of the face image dataset, It represents the abscissa of the upper - left coordinate point of the predicted bounding box of the nth face target in the kth image of the face image dataset, It represents the ordinate of the upper - left coordinate point of the predicted bounding box of the nth face target in the kth image of the face image dataset; It represents the lower - right coordinate of the predicted bounding box of the nth face target in the kth image of the face image dataset, It represents the abscissa of the lower - right coordinate point of the predicted bounding box of the nth face target in the kth image of the face image dataset, It represents the ordinate of the lower - right coordinate point of the predicted bounding box of the nth face target in the kth image of the face image dataset, N k ′ represents the number of predicted bounding boxes of the face targets in the face image training dataset selected for training the preset face recognition network model in the face image dataset.

[0083] Specifically, the accuracy of the identity category prediction in the predicted bounding box of the face image to be recognized can be optimized through the above - created target category loss model.

[0084] Step S130: Optimize and train the loss function model using the gradient descent method. When the loss function of the loss function model reaches the preset threshold, obtain the face recognition network model based on lightweight multi - scale feature fusion; among them,

[0085] The face recognition network model based on lightweight multi-scale feature fusion includes: a backbone network layer for extracting features of a face image in three different image dimensions, a pooling layer for performing pooling processing on the third output feature obtained by the backbone network layer, a feature fusion layer for performing feature fusion processing on the first output feature, the second output feature obtained by the backbone network layer, and the pooled third output feature obtained by the pooling layer respectively, and a detection head layer for generating a prediction result according to the three fused features obtained by the feature fusion layer.

[0086] Specifically, the loss function of the loss function model of the face recognition network based on lightweight multi-scale feature fusion may only include one of the object confidence loss model, the object bounding box loss model, and the object category loss model, or may include two of them, or may even include the above three loss function models at the same time. When including the above three models at the same time, the loss function of the face recognition network based on lightweight multi-scale feature fusion is: loss(object) = loss(box) + loss(confidence) + loss(type); when loss(object) reaches a preset threshold, complete optimization training is performed to obtain the face recognition network model based on lightweight multi-scale feature fusion.

[0087] Since the loss function model of the face recognition network based on lightweight multi-scale feature fusion in the present invention is trained by using the face recognition network based on lightweight multi-scale feature fusion, it has the advantages of few calculation parameters, smaller memory occupation, high accuracy, etc. As Figure 1.1 and Figure 1.2 shown, it can be clearly seen that the present invention improves the overall structure of the original YOLOv4 network structure, thereby obtaining a face recognition network structure based on lightweight multi-scale feature fusion. Using a lightweight network to extract image features and introducing an adaptive feature fusion structure to fuse features of different scales can maintain the accuracy of the model while significantly reducing the model size, realizing high-performance face recognition on mobile and embedded devices.

[0088] Step S140: Input the face image to be recognized into the face recognition network model based on lightweight multi-scale feature fusion for face recognition, and obtain the prediction bounding box of the face target in the face image to be recognized, the identity category corresponding to the prediction bounding box, and the recognition accuracy corresponding to the identity category.

[0089] Specifically, by inputting the face image to be recognized into the face recognition network model based on lightweight multi-scale feature fusion, through the processing of the face image by each functional structure layer in the face recognition network model based on lightweight multi-scale feature fusion, finally, the prediction bounding box of the face target in the face image to be recognized, the identity category corresponding to the prediction bounding box, and the recognition accuracy corresponding to the identity category are output.

[0090] As an optional embodiment of the present invention, inputting the face image to be recognized into the face recognition network model based on lightweight multi-scale feature fusion for face recognition, and obtaining the prediction bounding box of the face target in the face image to be recognized, the identity category corresponding to the prediction bounding box, and the recognition accuracy corresponding to the identity category includes:

[0091] Through the backbone network layer, feature extraction of the face image to be recognized is performed in three different image dimensions to obtain three different-dimensional features, namely feature X1, feature X2, and feature X3, and convolution operations are respectively performed on feature X1, feature X2, and feature X3 to obtain the first output feature Level1, the second output feature Level2, and the third output feature respectively;

[0092] Through the pooling layer, pooling processing is performed on the third output feature to obtain the pooled third output feature Level3;

[0093] Through the feature fusion layer, feature fusion processing is performed on the first output feature Level1, the second output feature Level2, and the pooled third output feature Level3 to obtain the first fusion feature ASFF1, the second fusion feature ASFF2, and the third fusion feature ASFF3 respectively; where

[0094] The first fusion feature ASFF1 is obtained by multiplying Level1, Level2, Level3 by the parameters α1, β1, γ1 respectively and then adding them; ASFF2 is obtained by multiplying Level1, Level2, Level3 by the parameters α2, β2, γ2 respectively and then adding them; ASFF3 is obtained by multiplying Level1, Level2, Level3 by the parameters α3, β3, γ3 respectively and then adding them;

[0095] Through the detection head layer, according to the first fusion feature ASFF1, the second fusion feature ASFF2, and the third fusion feature ASFF3, the prediction bounding box of the face target in the face image to be recognized, the identity category corresponding to the prediction bounding box, and the recognition accuracy corresponding to the identity category are generated.

[0096] Specifically, when the face image to be recognized is input into the face recognition network model based on lightweight multi-scale feature fusion, the face image is feature-extracted in the backbone network layer according to three different size dimensions. At this time, the single features of the face are extracted, such as eyes, nose, mouth, etc. Then, the single features are synthesized through convolution calculation to form three output features, namely the first output feature Level1, the second output feature Level2, and the third output feature. Since the feature layer of the third output feature is connected to the pooling layer, the third output feature of the pooling layer is pooled, and after dimensionality reduction, the calculation of the third output layer is simplified to obtain the pooled third output feature Level3. Then, the first output feature Level1, the second output feature Level2, and the pooled third output feature Level3 are subjected to feature fusion processing through the feature fusion layer, where α1, β1, γ1, α2, β2, γ2, α3, β3, γ3 are all known parameters. Finally, through the detection head layer, the prediction result of the face image to be recognized is obtained according to the fused features.

[0097] As Figure 2 shown, it is the functional module diagram of the face recognition device according to an embodiment of the present invention.

[0098] The face recognition device 200 according to the present invention can be installed in an electronic device. According to the functions achieved, the face recognition device may include a training module 210, a loss function construction module 220, an optimization module 230, and a prediction module 240. The modules according to the present invention may also be referred to as units, which refer to a series of computer program segments that can be executed by the processor of the electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0099] In this embodiment, the functions of each module / unit are as follows:

[0100] The training module 210 is used to input the face image dataset into the face recognition network based on lightweight multi-scale feature fusion for face recognition model training.

[0101] Among them, the face image dataset includes face pictures, the true bounding boxes of face targets marked on the face pictures, the widths, heights, and center point coordinates of the true bounding boxes of face targets, and the identity categories marked on the true bounding boxes of face targets.

[0102] Specifically, the face image dataset includes multiple face pictures, and each face picture includes at least one face image. The source can be face pictures collected through social networks or face pictures specially taken for actual purposes. For each face on each face picture, the true border of the face target is marked, and the width, height, and center point coordinates of each true border are calculated. Moreover, it is also necessary to mark the identity category of each true border, where the identity category can be the identity information of the face, such as the ID number and name, or other identity information required for face recognition authentication, such as the job number code of the corresponding employee in the company's employee attendance device, etc.

[0103] As an optional embodiment of the present invention, the face image dataset is stored in the blockchain, and the face recognition device 200 further includes: a face image sample collection module and a marking module (not shown in the figure). Among them,

[0104] The face image sample collection module is used to collect face image samples to obtain a face image sample set;

[0105] The marking module is used to mark the true border of the face target in the face image sample set, calculate the width, height, and center point coordinates of the marked true border of the face target, and mark the identity category of the true border of the face target to obtain the face image dataset.

[0106] Specifically, through the face image sample collection module, according to the actual face recognition application scenario, corresponding face image samples are collected. For example, for places such as stations and airports that require identity verification, face images and corresponding information can be collected through the public security system. At this time, each picture can include only one portrait; for employee attendance, employee face pictures can be directly collected, and each picture can have one face or multiple faces. The multiple collected face image samples constitute the face image sample set; through the marking module, the true border of each face image on each face picture in the face image sample set is marked. Different formats of picture files can be used to complete this, such as picture files in TXT format. In this format of picture file, the face coordinates are recorded. Through the face coordinates, the true border of the face target is marked, and then the width, height, and center point coordinates of the true border are calculated, and the identity category of the true border of each marked face target is marked.

[0107] As an optional embodiment of the present invention, the marking module further includes: a sample set definition unit, a true border marking unit, and a calculation unit (not shown in the figure). Among them,

[0108] The sample set definition unit is used to define the face image sample set as: {Datak (x, y), k ∈

[0109] [1, K], x ∈ [1, X], y ∈ [1, Y]}; where, Data k (x, y) represents the pixel information of the x-th row and y-th column of the k-th picture in the face image sample set, K represents the number of pictures in the face image sample set, X represents the number of rows of pixels of the pictures in the face image sample set, and Y represents the number of columns of pixels of the pictures in the face image sample set;

[0110] The ground truth bounding box annotation unit is used to annotate the ground truth bounding box of each face target in each picture in the defined face image sample set; where, the ground truth bounding box of the face target is defined as:

[0111]

[0112]

[0113] where, represents the upper left corner coordinates of the ground truth bounding box of the n-th face target in the k-th picture in the face image sample set, represents the abscissa of the upper left corner coordinate point of the ground truth bounding box of the n-th face target in the k-th picture in the face image sample set, represents the ordinate of the upper left corner coordinate point of the ground truth bounding box of the n-th face target in the k-th picture in the face image sample set; represents the lower right corner coordinates of the ground truth bounding box of the n-th face target in the k-th picture in the face image sample set, represents the abscissa of the lower right corner coordinate point of the ground truth bounding box of the n-th face target in the k-th picture in the face image sample set, represents the ordinate of the lower right corner coordinate point of the ground truth bounding box of the n-th face target in the k-th picture in the face image sample set; K represents the number of pictures in the face image sample set; N k represents the number of ground truth bounding boxes of the face targets in the k-th picture in the face image training dataset selected for training the preset face recognition network model;

[0114] The calculation unit is used to calculate the width, height and center point coordinates of the ground truth bounding box of the face target respectively according to the preset ground truth bounding box width calculation formula, preset ground truth bounding box height calculation formula and preset ground truth bounding box center point coordinate calculation formula, and mark the identity category of the ground truth bounding box of the face target to obtain the face image dataset; where,

[0115] The preset ground truth bounding box width calculation formula is:

[0116] The formula for calculating the preset true border height is as follows:

[0117] The formula for calculating the coordinates of the center point of the preset true border is as follows:

[0118]

[0119] The identity category of the true border of the face target is defined as:

[0120] where represents the width of the true border of the nth face target in the kth image in the face image sample set, represents the height of the true border of the nth face target in the kth image in the face image sample set, and t represents true.

[0121] Specifically, first, the face image sample set is defined by the sample set definition unit, so that each picture in the face image sample set and each face image on each picture can be represented by its corresponding position. Then, through the true border annotation unit, the true border of the face target is defined according to the position of the face image. Next, through the calculation unit, its height, width, and center coordinates are calculated based on the true border, and the identity category of the true border is defined, so that the picture information is transformed into digital information that can be operated on, thus facilitating subsequent calculations.

[0122] The loss function construction module 220 is used to construct a loss function model within the face recognition network after the face recognition model is trained, based on the true border of the face target, the width, height, and center point coordinates of the true border of the face target, and the identity category of the true border of the face target.

[0123] Specifically, the loss function model is used to estimate the degree of inconsistency between the predicted value f(x) of the trained face recognition network and the true value Y. It is a non - negative real - valued function, denoted by L(Y, f(x)). The smaller the loss function, the better the robustness of the trained network. The loss function is the core part of the empirical risk function and an important part of the structural risk function. The structural risk function of the model includes an empirical risk term and a regularization term.

[0124] To ensure the accuracy of the face recognition network model obtained after training the face recognition network, it is necessary to construct a loss function model within the face recognition network obtained after training. During the process of training the model, the face image dataset can be divided into two parts, one as the face image training dataset and the other as the face image verification and optimization dataset. First, use the face image training dataset to train the face recognition model of the face recognition network, and then perform optimization through the face image verification and optimization dataset.

[0125] As an optional embodiment of the present invention, the loss function model includes: a target bounding box loss model;

[0126] The calculation formula of the target bounding box loss model is:

[0127]

[0128] where represents the Euclidean distance between the center points of the predicted bounding box and the true bounding box of the face target, c represents the diagonal length of the smallest rectangle that can cover the predicted bounding box and the true bounding box of the face target, and the calculation method of c is:

[0129]

[0130] IOU represents the intersection over union of the predicted bounding box and the true box of the face target, C h and C w respectively represent the height and width of the smallest rectangle that can cover the predicted bounding box and the true bounding box of the face target, h and h gt respectively represent the height of the predicted bounding box of the face target and the height of the true bounding box of the face target, w and w gt respectively represent the height of the predicted bounding box of the face target and the width of the true bounding box of the face target, and the calculation methods of v and α in the formula are:

[0131]

[0132] The width and height of the predicted bounding box of each face target in each face image in the face image dataset are respectively defined as:

[0133] and

[0134] The center point coordinates of the predicted bounding box of each face target in each face image in the face image dataset are defined as:

[0135]

[0136] Specifically, through the target bounding box loss model created above, the accuracy of the labeled position of the predicted bounding box of the face target in the face image to be recognized can be optimized.

[0137] As an alternative embodiment of the present invention, the loss function model further includes: a target confidence loss model;

[0138] The calculation formula of the target confidence loss model is:

[0139]

[0140] Wherein, represents the identity category within the true bounding box of the nth face target in the kth picture in the face image dataset, represents the identity category within the predicted bounding box of the nth face target in the kth picture in the face image dataset, and λ noobject represents the confidence penalty weight when there is no target in the predicted bounding box of the face target.

[0141] Specifically, through the target confidence loss model created above, the accuracy of the predicted target of the predicted bounding box of the face image to be recognized can be optimized.

[0142] As an alternative embodiment of the present invention, the loss function model further includes: a target category loss model;

[0143] The calculation formula of the target category loss model is:

[0144]

[0145] Wherein, represents the confidence of the identity category within the true bounding box of the nth face target in the kth picture in the face image dataset, represents the confidence of the identity category within the predicted bounding box of the nth face target in the kth picture in the face image dataset;

[0146] The predicted bounding box of each face target in each face image in the face image dataset is defined as:

[0147]

[0148]

[0149] Wherein, represents the upper left coordinate of the predicted bounding box of the nth face target in the kth picture in the face image dataset, represents the abscissa of the upper left coordinate point of the predicted bounding box of the nth face target in the kth picture in the face image dataset, Denote the ordinate of the upper left coordinate point of the predicted bounding box of the n-th face target in the k-th picture in the face image dataset; Denote the lower right corner coordinates of the predicted bounding box of the n-th face target in the k-th picture in the face image dataset, Denote the abscissa of the lower right coordinate point of the predicted bounding box of the n-th face target in the k-th picture in the face image dataset, Denote the ordinate of the lower right coordinate point of the predicted bounding box of the n-th face target in the k-th picture in the face image dataset, N k ′ Denote the number of predicted bounding boxes of face targets in the face image training dataset selected for training the preset face recognition network model in the face image dataset.

[0150] Specifically, the target category loss model created above can optimize the accuracy of the identity category prediction in the predicted bounding box of the face image to be recognized.

[0151] The optimization module 230 is used to optimize and train the loss function model by using the gradient descent method. When the loss function of the loss function model reaches the preset threshold, a face recognition network model based on lightweight multi-scale feature fusion is obtained; where,

[0152] The face recognition network model based on lightweight multi-scale feature fusion includes: a backbone network layer for extracting features of three different image dimensions of a face image, a pooling layer for performing pooling processing on the third output feature obtained by the backbone network layer, and a feature fusion layer for performing feature fusion processing on the first output feature, the second output feature obtained by the backbone network layer, and the pooled third output feature obtained by the pooling layer respectively, and a detection head layer for generating a prediction result according to the three fusion features obtained by the feature fusion layer.

[0153] Specifically, the loss function of the loss function model of the face recognition network based on lightweight multi-scale feature fusion can include only one of the target confidence loss model, the target bounding box loss model, and the target category loss model, or can include two of them, and can even include the above three loss function models at the same time. When including the above three models at the same time, the loss function of the face recognition network based on lightweight multi-scale feature fusion is: loss(object) = loss(box) + loss(confidence) + loss(type); when loss(object) reaches the preset threshold, complete optimization training is performed to obtain a face recognition network model based on lightweight multi-scale feature fusion.

[0154] The loss function model of the face recognition network based on lightweight multi-scale feature fusion in the present invention is trained using the face recognition network based on lightweight multi-scale feature fusion. Therefore, it has the advantages of few computing parameters, smaller memory occupation, and high accuracy. As Figure 1.1 and Figure 1.2 shown, it can be clearly seen that the present invention improves the overall structure of the original YOLOv4 network structure to obtain a structure of a face recognition network based on lightweight multi-scale feature fusion. The lightweight network is used to extract image features, and an adaptive feature fusion structure is introduced to fuse features of different scales, which can maintain the accuracy of the model while significantly reducing the model size, and achieve high-performance face recognition on mobile and embedded devices.

[0155] The prediction module 240 is configured to input the face image to be recognized into the face recognition network model based on lightweight multi-scale feature fusion for face recognition, and obtain the prediction bounding box of the face target in the face image to be recognized, the identity category corresponding to the prediction bounding box, and the recognition accuracy corresponding to the identity category.

[0156] Specifically, by inputting the face image to be recognized into the face recognition network model based on lightweight multi-scale feature fusion, through the processing of the face image by each functional structure layer in the face recognition network model based on lightweight multi-scale feature fusion, finally, the prediction bounding box of the face target in the face image to be recognized, the identity category corresponding to the prediction bounding box, and the recognition accuracy corresponding to the identity category are output.

[0157] As an optional embodiment of the present invention, the prediction module 240 further includes: a feature extraction unit, a pooling unit, a feature fusion unit, and a prediction unit (not shown in the figure). Among them,

[0158] The feature extraction unit is configured to extract features of three different image dimensions from the face image to be recognized through the backbone network layer, obtain three different-dimensional features, namely feature X1, feature X2, and feature X3, and perform convolution operation processing on feature X1, feature X2, and feature X3 respectively to obtain the first output feature Level1, the second output feature Level2, and the third output feature respectively;

[0159] The pooling unit is configured to perform pooling processing on the third output feature through the pooling layer to obtain the pooled third output feature Level3;

[0160] A feature fusion unit is used to perform feature fusion processing on the first output feature Level1, the second output feature Level2, and the pooled third output feature Level3 through a feature fusion layer, respectively obtaining a first fusion feature ASFF1, a second fusion feature ASFF2, and a third fusion feature ASFF3; where

[0161] The first fusion feature ASFF1 is obtained by multiplying Level1, Level2, Level3 by parameters α1, β1, γ1 respectively and then adding them together; ASFF2 is obtained by multiplying Level1, Level2, Level3 by parameters α2, β2, γ2 respectively and then adding them together; ASFF3 is obtained by multiplying Level1, Level2, Level3 by parameters α3, β3, γ3 respectively and then adding them together;

[0162] Through a detection head layer, based on the first fusion feature ASFF1, the second fusion feature ASFF2, and the third fusion feature ASFF3, a predicted bounding box of the face target in the face image to be recognized, the identity category corresponding to the predicted bounding box, and the recognition accuracy corresponding to the identity category are generated.

[0163] Specifically, when the face image to be recognized is input into the face recognition network model based on lightweight multi-scale feature fusion, the feature extraction unit extracts features of the face image according to three different size dimensions in the backbone network layer. At this time, the extracted features are single features of the face, such as eyes, nose, mouth, etc. Then, through convolutional calculation, the single features are integrated to form three output features, namely the first output feature Level1, the second output feature Level2, and the third output feature. Since the feature layer of the third output feature is connected to the pooling layer, the third output feature is pooled by the pooling unit. After dimensionality reduction, the calculation of the third output layer is simplified to obtain the pooled third output feature Level3. Then, through the feature fusion unit, feature fusion processing is performed on the first output feature Level1, the second output feature Level2, and the pooled third output feature Level3. Among them, α1, β1, γ1, α2, β2, γ2, α3, β3, γ3 are all known parameters. Finally, through the detection unit, in the detection head layer, based on the fused features, the prediction result of the face image to be recognized is obtained.

[0164] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing the face recognition method according to an embodiment of the present invention.

[0165] The electronic device 1 may include a processor 10, a memory 11, and a bus. It may also include a computer program stored in the memory 11 and executable on the processor 10, such as a face recognition program 12.

[0166] Among them, the memory 11 includes at least one type of readable storage medium, which includes flash memory, mobile hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), magnetic memories, magnetic disks, optical disks, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device 1, such as the mobile hard disk of the electronic device 1. In some other embodiments, the memory 11 may also be an external storage device of the electronic device 1, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 may also include both an internal storage unit and an external storage device of the electronic device 1. The memory 11 can be used not only to store application software installed on the electronic device 1 and various types of data, such as the code of the face recognition program, but also to temporarily store data that has been output or will be output.

[0167] In some embodiments, the processor 10 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including the combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (Control Unit) of the electronic device, connecting all components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules (such as the face recognition program, etc.) stored in the memory 11, and calling data stored in the memory 11, to perform various functions of the electronic device 1 and process data.

[0168] The bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to enable connection and communication between the memory 11 and at least one processor 10, etc.

[0169] Figure 3 Only an electronic device with components is shown. Those skilled in the art can understand that Figure 3 the shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0170] For example, although not shown, the electronic device 1 may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to the at least one processor 10 through a power management device, so as to implement functions such as charging management, discharging management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device 1 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0171] Furthermore, the electronic device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.

[0172] Optionally, the electronic device 1 may further include a user interface. The user interface may be a display (Display), an input unit (such as a keyboard (Keyboard)). Optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.

[0173] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0174] The face recognition program 12 stored in the memory 11 in the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can implement:

[0175] Input the face image dataset into the face recognition network based on lightweight multi-scale feature fusion for training the face recognition model. Among them, the face image dataset includes face pictures, the true bounding boxes of face targets marked on the face pictures, the widths and heights of the true bounding boxes of face targets, the center point coordinates, and the identity categories marked on the true bounding boxes of face targets.

[0176] Within the face recognition network after training the face recognition model, construct a loss function model through the true bounding box of the face target, the width and height of the true bounding box of the face target, the center point coordinates, and the identity category of the true bounding box of the face target.

[0177] Use the gradient descent method to optimize and train the loss function model. When the loss function of the loss function model reaches the preset threshold, obtain the face recognition network model based on lightweight multi-scale feature fusion. Among them,

[0178] The face recognition network model based on lightweight multi-scale feature fusion includes: a backbone network layer for extracting features of three different image dimensions for face images, a pooling layer for performing pooling processing on the third output feature obtained by the backbone network layer, a feature fusion layer for performing feature fusion processing on the first output feature, the second output feature obtained by the backbone network layer, and the pooled third output feature obtained by the pooling layer respectively, and a detection head layer for generating a prediction result according to the three fusion features obtained by the feature fusion layer.

[0179] Input the face image to be recognized into the face recognition network model based on lightweight multi-scale feature fusion for face recognition, and obtain the predicted bounding box of the face target in the face image to be recognized, the identity category corresponding to the predicted bounding box, and the recognition accuracy corresponding to the identity category.

[0180] Specifically, the specific implementation method of the processor 10 for the above instructions can refer to Figure 1 The description of the relevant steps in the corresponding embodiments will not be repeated here. It should be emphasized that to further ensure the privacy and security of the above-mentioned face image dataset, the above-mentioned face image dataset can also be stored in a node of a blockchain.

[0181] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).

[0182] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation.

[0183] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0184] In addition, in each embodiment of the present invention, the functional modules can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.

[0185] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-mentioned exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.

[0186] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.

[0187] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.

[0188] In addition, obviously, the term "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or apparatuses stated in the system claims can also be implemented by one unit or apparatus through software or hardware. Words such as "second" are used to denote names and do not denote any particular order.

[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A face recognition method, applied to an electronic device, characterized in that The method includes: Inputting a face image dataset into a face recognition network based on lightweight multi-scale feature fusion for training a face recognition model; wherein, the face image dataset includes face pictures, the true bounding boxes of face targets marked on the face pictures, the width, height, and center point coordinates of the true bounding boxes of face targets, and the identity categories marked on the true bounding boxes of face targets; Within the face recognition network after training the face recognition model, constructing a loss function model through the true bounding boxes of face targets, the width, height, and center point coordinates of the true bounding boxes of face targets, and the identity categories of the true bounding boxes of face targets; Using the gradient descent method to optimize and train the loss function model, and when the loss function of the loss function model reaches a preset threshold, obtaining a face recognition network model based on lightweight multi-scale feature fusion; wherein, The face recognition network model based on lightweight multi-scale feature fusion includes: a backbone network layer for extracting features of three different image dimensions for a face image, a pooling layer for performing pooling processing on the third output feature obtained by the backbone network layer, a feature fusion layer for respectively performing feature fusion processing on the first output feature, the second output feature obtained by the backbone network layer, and the pooled third output feature obtained by the pooling layer, and a detection head layer for generating a prediction result according to the three fusion features obtained by the feature fusion layer; Inputting the face image to be recognized into the face recognition network model based on lightweight multi-scale feature fusion for face recognition, and obtaining the predicted bounding box of the face target in the face image to be recognized, the identity category corresponding to the predicted bounding box, and the recognition accuracy corresponding to the identity category; Wherein, the loss function model includes: a target bounding box loss model; The calculation formula of the target bounding box loss model is: Among them, represents the Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box of the face target, c represents the diagonal length of the smallest rectangle that can cover the predicted bounding box and the ground truth bounding box of the face target, and the calculation method of c is as follows: IOU represents the intersection over union of the predicted bounding box of the face target and the ground truth box, C h and C w respectively represent the height and width of the smallest rectangle that can cover the predicted bounding box and the ground truth bounding box of the face target. h and h gt respectively represent the height of the predicted bounding box of the face target and the height of the ground truth bounding box of the face target, and w and w gt respectively represent the height of the predicted bounding box of the face target and the width of the ground truth bounding box of the face target. The calculation methods of v and α in the formula are as follows:

2. The face recognition method according to claim 1, wherein The face image dataset is stored in a blockchain, and before inputting the face image dataset into the face recognition network based on lightweight multi-scale feature fusion for training the face recognition model, it further includes: Collecting face image samples to obtain a face image sample set; Annotating the true bounding boxes of face targets in the face image sample set, calculating the width, height, and center point coordinates of the annotated true bounding boxes of face targets, and marking the identity categories of the true bounding boxes of face targets to obtain a face image dataset.

3. The face recognition method according to claim 2, wherein The annotating the true bounding boxes of face targets in the face image sample set, calculating the width, height, and center point coordinates of the annotated true bounding boxes of face targets, and marking the identity categories of the true bounding boxes of face targets to obtain a face image dataset includes: Define the face image sample set as: {Data k (x,y), k ∈ [1, K], x ∈ [1, X], y ∈ [1, Y]}; where Data k (x, y) represents the pixel information of the x-th row and y-th column of the k-th picture in the face image sample set, K represents the number of pictures in the face image sample set, X represents the number of rows of pixels of the pictures in the face image sample set, and Y represents the number of columns of pixels of the pictures in the face image sample set; Annotating the true bounding boxes of each face target in each picture in the defined face image sample set; wherein, the true bounding box of the face target is defined as: Among them, represents the abscissa of the upper left corner coordinate of the ground truth bounding box of the nth face target in the kth picture of the face image sample set, represents the abscissa of the upper left corner coordinate point of the ground truth bounding box of the nth face target in the kth picture of the face image sample set, represents the ordinate of the upper left corner coordinate point of the ground truth bounding box of the nth face target in the kth picture of the face image sample set; represents the abscissa of the lower right corner coordinate of the ground truth bounding box of the nth face target in the kth picture of the face image sample set, represents the abscissa of the lower right corner coordinate point of the ground truth bounding box of the nth face target in the kth picture of the face image sample set, represents the ordinate of the lower right corner coordinate point of the ground truth bounding box of the nth face target in the kth picture of the face image sample set; K represents the number of pictures in the face image sample set; N k represents the number of ground truth bounding boxes of the face targets in the kth picture of the face image training dataset selected for training the preset face recognition network model in the face image dataset; According to the preset true border width calculation formula, the preset true border height calculation formula, and the preset true border center point coordinate calculation formula, calculate the width, height, and center point coordinates of the true border of the face target respectively, and mark the identity category of the true border of the face target to obtain a face image dataset; wherein, The formula for calculating the preset true border width is as follows: The formula for calculating the preset true border height is as follows: The preset true border center point coordinate calculation formula is: The identity category of the true bounding box of the face target is defined as: wherein, represents the width of the true bounding box of the nth face target in the kth image in the face image sample set, represents the height of the true bounding box of the nth face target in the kth image in the face image sample set, and t represents true.

4. The face recognition method according to claim 3, wherein The width and height of the predicted border of each face target in each face picture in the face image dataset are respectively defined as: and The center point coordinate of the predicted border of each face target in each face picture in the face image dataset is defined as:

5. The face recognition method according to claim 4, wherein The loss function model further includes: an object confidence loss model; The calculation formula of the object confidence loss model is: wherein, represents the identity category within the true bounding box of the n-th face target in the k-th picture of the face image dataset, represents the identity category within the predicted bounding box of the n-th face target in the k-th picture of the face image dataset, λ noobject represents the confidence penalty weight when there is no target in the predicted bounding box of the face target.

6. The face recognition method according to claim 5, wherein The loss function model further includes: an object category loss model; The calculation formula of the object category loss model is: Among them, represents the identity class confidence within the true bounding box of the nth face target in the kth picture of the face image dataset, represents the identity class confidence within the predicted bounding box of the nth face target in the kth picture of the face image dataset; The predicted border of each face target in each face image in the face image dataset is defined as: Among them, represents the abscissa of the upper left corner coordinate of the predicted bounding box of the nth face target in the kth picture of the face image dataset, represents the abscissa of the upper left corner coordinate point of the predicted bounding box of the nth face target in the kth picture of the face image dataset, represents the ordinate of the upper left corner coordinate point of the predicted bounding box of the nth face target in the kth picture of the face image dataset; represents the abscissa of the lower right corner coordinate of the predicted bounding box of the nth face target in the kth picture of the face image dataset, represents the abscissa of the lower right corner coordinate point of the predicted bounding box of the nth face target in the kth picture of the face image dataset, represents the ordinate of the lower right corner coordinate point of the predicted bounding box of the nth face target in the kth picture of the face image dataset, N k ' represents the number of predicted bounding boxes of the face targets in the face image training dataset selected for training the preset face recognition network model.

7. The face recognition method according to claim 1, wherein Inputting the face image to be recognized into the face recognition network model based on lightweight multi-scale feature fusion for face recognition, and obtaining the predicted border of the face target in the face image to be recognized, the identity category corresponding to the predicted border, and the recognition accuracy corresponding to the identity category includes: Performing feature extraction on the face image to be recognized through the backbone network layer in three different image dimensions to obtain three different-dimensional features, namely feature X1, feature X2, and feature X3, and performing convolution operation processing on the feature X1, feature X2, and feature X3 respectively to obtain the first output feature Level1, the second output feature Level2, and the third output feature respectively; Performing pooling processing on the third output feature through the pooling layer to obtain the pooled third output feature Level3; Performing feature fusion processing on the first output feature Level1, the second output feature Level2, and the pooled third output feature Level3 through the feature fusion layer to obtain the first fusion feature ASFF1, the second fusion feature ASFF2, and the third fusion feature ASFF3 respectively; wherein, The first fusion feature ASFF1 is obtained by multiplying Level1, Level2, Level3 by parameters α1, β1, γ1 respectively and then adding them; ASFF2 is obtained by multiplying Level1, Level2, Level3 by parameters α2, β2, γ2 respectively and then adding them; ASFF3 is obtained by multiplying Level1, Level2, Level3 by parameters α3, β3, γ3 respectively and then adding them; Through the detection head layer, according to the first fusion feature ASFF1, the second fusion feature ASFF2, and the third fusion feature ASFF3, generate the predicted border of the face target in the face image to be recognized, the identity category corresponding to the predicted border, and the recognition accuracy corresponding to the identity category.

8. A face recognition device, characterized in that, The device includes: A training module for inputting a face image dataset into a face recognition network based on lightweight multi-scale feature fusion for training a face recognition model; wherein, the face image dataset includes face pictures, true bounding boxes of face targets annotated on the face pictures, the widths, heights, and center point coordinates of the true bounding boxes of the face targets, and the identity categories marked on the true bounding boxes of the face targets; A loss function construction module for constructing a loss function model within the face recognition network after training the face recognition model, based on the true bounding boxes of the face targets, the widths, heights, and center point coordinates of the true bounding boxes of the face targets, and the identity categories of the true bounding boxes of the face targets; An optimization module for optimizing and training the loss function model using the gradient descent method, and obtaining a face recognition network model based on lightweight multi-scale feature fusion when the loss function of the loss function model reaches a preset threshold; wherein, The face recognition network model based on lightweight multi-scale feature fusion includes: a backbone network layer for extracting features of three different image dimensions from a face image, a pooling layer for performing pooling processing on the third output feature obtained by the backbone network layer, a feature fusion layer for performing feature fusion processing on the first output feature, the second output feature obtained by the backbone network layer, and the pooled third output feature obtained by the pooling layer respectively, and a detection head layer for generating a prediction result based on the three fusion features obtained by the feature fusion layer; A prediction module for inputting a face image to be recognized into the face recognition network model based on lightweight multi-scale feature fusion for face recognition, and obtaining a predicted bounding box of the face target in the face image to be recognized, the identity category corresponding to the predicted bounding box, and the recognition accuracy corresponding to the identity category; Wherein, the loss function model includes: a target bounding box loss model; the calculation formula of the target bounding box loss model is: Among them, represents the Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box of the face target. c represents the diagonal length of the smallest rectangle that can cover the predicted bounding box and the ground truth bounding box of the face target. The calculation method of c is as follows: The IOU represents the intersection over union of the predicted bounding box of the face target and the ground truth box, C h and C w respectively represent the height and width of the smallest rectangle that can cover the predicted bounding box and the ground truth bounding box of the face target, h and h gt respectively represent the height of the predicted bounding box of the face target and the height of the ground truth bounding box of the face target, and w and w gt respectively represent the height of the predicted bounding box of the face target and the width of the ground truth bounding box of the face target. The calculation methods of v and α in the formula are as follows:

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the face recognition method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, The computer program, when executed by a processor, implements the face recognition method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • In-plane rotation-invariant face detection method and device, and storage medium

    CN111695522A

  • Face recognition method for simultaneous detection and feature extraction

    CN113378675A