Facial recognition method and apparatus, and computer-readable storage medium and robot

By using margin-based loss function in the training process of the face recognition model and adjusting the value of margins according to the image quality of the sample image, the problem of low robustness of the face recognition model in the prior art is solved, and higher robustness and accuracy are achieved.

WO2025112130A1PCT designated stage expired Publication Date: 2025-06-05UBTECH ROBOTICS CORP LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/140716
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2023-12-21
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

The existing face recognition model uses fixed margins for all samples during training, resulting in low robustness of face recognition.

Method used

The margin-based loss function is used during model training. The value of the margin is positively correlated with the image quality of the sample image used for training, so that the value of the margin is flexibly adjusted according to the image quality.

Benefits of technology

By adjusting the value of margins according to image quality, the robustness and accuracy of the face recognition model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023140716_05062025_PF_FP_ABST
    Figure CN2023140716_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of image processing, and particularly relates to a facial recognition method and apparatus, and a computer-readable storage medium and a robot. The method comprises: acquiring a target image to be recognized; and using a preset facial recognition model to perform facial recognition on the target image, so as to obtain a facial recognition result, wherein the facial recognition model uses a margin-based loss function in a training process, and the value of a margin is positively correlated with the image quality of a sample image for training. By means of the method, the value of a margin can be determined on the basis of the image quality of a sample image in a training process, and the value of the margin is positively correlated with the image quality of the sample image, such that the value of the margin can be flexibly adjusted on the basis of the image quality, thereby facilitating an improvement in the robustness of facial recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Face recognition method, device, computer-readable storage medium, and robot

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 28, 2023, with application number 202311611147.X and invention name “A face recognition method, device, computer-readable storage medium and robot”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application belongs to the field of image processing technology, and in particular relates to a face recognition method, device, computer-readable storage medium, and robot. Background Art

[0003] Facial recognition technology is widely used in modern life. In the field of robotics, facial recognition can be used to achieve enhanced security monitoring, emotion recognition, and human-computer interaction. However, currently used facial recognition methods (such as the CosFace algorithm) typically employ a margin-based angular space strategy during model training to improve facial recognition performance. However, this model uses a fixed margin for all samples, resulting in low robustness in facial recognition. Technical issues

[0004] In view of this, the embodiments of the present application provide a face recognition method, device, computer-readable storage medium and robot to solve the problem that the existing face recognition model uses a fixed margin for all samples during the training process, resulting in low robustness of face recognition. Technical Solutions

[0005] A first aspect of an embodiment of the present application provides a face recognition method, which may include:

[0006] Acquire the target image to be identified;

[0007] Performing face recognition on the target image using a preset face recognition model to obtain a face recognition result;

[0008] The face recognition model uses a margin-based loss function during training, and the value of the margin is positively correlated with the image quality of the sample images used for training.

[0009] In a specific implementation of the first aspect, the training process of the face recognition model may include:

[0010] Determining a face recognition training sample set; wherein the face recognition training sample set includes a plurality of sample images;

[0011] Calculate the image quality of each sample image separately;

[0012] determining a margin of each sample image according to an image quality of each sample image;

[0013] The initial model is trained using the face recognition training sample set and the margin-based loss function to obtain the face recognition model.

[0014] In a specific implementation of the first aspect, respectively calculating the image quality of each sample image may include:

[0015] Determine the feature vector of each sample image respectively;

[0016] The image quality of each sample image is calculated according to the feature vector of each sample image.

[0017] In a specific implementation of the first aspect, determining the margin of each sample image according to the image quality of each sample image may include:

[0018] Determine the upper limit value of image quality, the lower limit value of image quality, the upper limit value of margin and the lower limit value of margin respectively;

[0019] The margin of each sample image is determined according to the image quality upper limit value, the image quality lower limit value, the margin upper limit value, the margin lower limit value, and the image quality of each sample image.

[0020] In a specific implementation of the first aspect, determining the margin of each sample image based on the image quality upper limit, the image quality lower limit, the margin upper limit, the margin lower limit, and the image quality of each sample image may include:

[0021] Limiting the image quality of each sample image according to the image quality upper limit and the image quality lower limit to obtain the image quality of each sample image after the range is limited;

[0022] Calculating an image quality span according to the image quality upper limit value and the image quality lower limit value;

[0023] Calculate the margin span according to the margin upper limit value and the margin lower limit value;

[0024] Calculating a span ratio coefficient according to the margin span and the image quality span;

[0025] The margin of each sample image is determined according to the span ratio coefficient, the image quality lower limit, the margin lower limit, and the image quality of each sample image after range restriction.

[0026] In a specific implementation of the first aspect, determining the face recognition training sample set may include:

[0027] Get the preset original sample set;

[0028] The original sample set is subjected to data enhancement processing according to a preset data enhancement method to obtain the face recognition training sample set.

[0029] In a specific implementation of the first aspect, performing data enhancement processing on the original sample set according to a preset data enhancement method to obtain the face recognition training sample set may include:

[0030] The original sample set is subjected to color space enhancement processing, and / or motion blur enhancement processing, and / or out-of-focus blur enhancement processing to obtain the face recognition training sample set.

[0031] A second aspect of the embodiments of the present application provides a face recognition device, which may include:

[0032] An image acquisition module is used to acquire a target image to be identified;

[0033] A face recognition module is used to perform face recognition on the target image using a preset face recognition model to obtain a face recognition result; wherein, the face recognition model uses a margin-based loss function during training, and the value of the margin is positively correlated with the image quality of the sample image used for training.

[0034] In a specific implementation of the second aspect, the face recognition module may include:

[0035] A training sample set determination submodule is used to determine a face recognition training sample set; wherein the face recognition training sample set includes a plurality of sample images;

[0036] An image quality calculation submodule, used to calculate the image quality of each sample image respectively;

[0037] An image margin determination submodule, configured to determine the margin of each sample image according to the image quality of each sample image;

[0038] The initial model training submodule is used to train the initial model using the face recognition training sample set and the margin-based loss function to obtain the face recognition model.

[0039] In a specific implementation of the second aspect, the image quality calculation submodule may include:

[0040] a feature vector determination unit, configured to determine a feature vector for each sample image;

[0041] The image quality calculation unit is used to calculate the image quality of each sample image according to the feature vector of each sample image.

[0042] In a specific implementation of the second aspect, the image margin determination submodule may include:

[0043] a range determination unit, configured to respectively determine an upper limit value of image quality, a lower limit value of image quality, an upper limit value of margin, and a lower limit value of margin;

[0044] The image margin determining unit is configured to determine the margin of each sample image according to the image quality upper limit value, the image quality lower limit value, the margin upper limit value, the margin lower limit value, and the image quality of each sample image.

[0045] In a specific implementation of the second aspect, the image margin determining unit may include:

[0046] a quality range limiting subunit, configured to limit the image quality of each sample image according to the image quality upper limit value and the image quality lower limit value, to obtain the image quality of each sample image after the range is limited;

[0047] a quality span calculation subunit, configured to calculate an image quality span according to the image quality upper limit value and the image quality lower limit value;

[0048] a margin span calculation subunit, configured to calculate the margin span according to the margin upper limit value and the margin lower limit value;

[0049] a scale factor calculation subunit, configured to calculate a span scale factor according to the margin span and the image quality span;

[0050] The image margin determination subunit is used to determine the margin of each sample image according to the span ratio coefficient, the image quality lower limit, the margin lower limit and the image quality of each sample image after range restriction.

[0051] In a specific implementation of the second aspect, the training sample set determination submodule may include:

[0052] A sample set acquisition unit, used to acquire a preset original sample set;

[0053] The enhancement processing unit is used to perform data enhancement processing on the original sample set according to a preset data enhancement method to obtain the face recognition training sample set.

[0054] In a specific implementation of the second aspect, the enhancement processing unit can be used to perform color space enhancement processing on the original sample set, and / or, perform motion blur enhancement processing on the original sample set, and / or, perform out-of-focus blur enhancement processing on the original sample set, to obtain the face recognition training sample set.

[0055] A third aspect of an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned face recognition methods are implemented.

[0056] A fourth aspect of an embodiment of the present application provides a robot comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the above-mentioned face recognition methods when executing the computer program.

[0057] A fifth aspect of the embodiments of the present application provides a computer program product, which, when run on a robot, enables the robot to perform the steps of any one of the above-mentioned face recognition methods. Beneficial effects

[0058] Compared with the prior art, the embodiments of the present application have the following advantages: the embodiments of the present application obtain a target image to be identified; perform face recognition on the target image using a preset face recognition model to obtain a face recognition result; wherein, the face recognition model uses a margin-based loss function during training, and the value of the margin is positively correlated with the image quality of the sample image used for training. Through the embodiments of the present application, the value of the margin can be determined based on the image quality of the sample image during training; wherein, the value of the margin is positively correlated with the image quality of the sample image, so that the value of the margin can be flexibly adjusted according to the image quality, which helps to improve the robustness of face recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0060] Figure 1 is a schematic diagram of an image in a real application scenario;

[0061] FIG2 is a schematic flow chart of the training process of the face recognition model according to an embodiment of the present application;

[0062] FIG3 is a schematic diagram of the circular area mean filtering effect;

[0063] FIG4 is a schematic diagram of a face recognition training sample set;

[0064] FIG5 is a schematic diagram of image quality measurement values;

[0065] FIG6 is a flow chart of an embodiment of a face recognition method in an embodiment of the present application;

[0066] FIG7 is a structural diagram of an embodiment of a face recognition device according to an embodiment of the present application;

[0067] FIG8 is a schematic block diagram of a robot in an embodiment of the present application. Modes for Carrying Out the Invention

[0068] In order to make the purpose, features, and advantages of the invention of this application more obvious and easy to understand, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described below are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0069] It will be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0070] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0071] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0072] As used in this specification and the appended claims, the term "if" can be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0073] In addition, in the description of the present application, the terms "first", "second", "third", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0074] Facial recognition technology is widely used in modern life. In the field of robotics, facial recognition can be used to achieve enhanced security monitoring, emotion recognition, and human-computer interaction. However, currently used facial recognition methods (such as the CosFace algorithm) typically employ a margin-based angular space strategy during model training to improve facial recognition performance. However, this model uses a fixed margin for all samples, resulting in low robustness in facial recognition.

[0075] In view of this, embodiments of the present application provide a face recognition method, device, computer-readable storage medium, and robot to solve the problem that existing face recognition methods use a fixed margin for all samples, resulting in low robustness of face recognition.

[0076] It should be noted that the executor of the method of this application is a robot, which may include but is not limited to any common robot in the prior art, such as a guiding robot, an inspection robot, a food delivery robot, and an educational robot.

[0077] In existing face recognition methods represented by the CosFace algorithm, a fixed-margin angular space strategy is usually used to calculate the loss function when training face recognition models. However, in real application scenarios, images may have problems such as light disturbance, motion blur, and out-of-focus blur, as shown in Figure 1. This leads to inconsistent image quality in real application scenarios, and the feature spaces corresponding to images of different qualities may also be different. Therefore, the face recognition model obtained by using this training strategy does not perform well in real application scenarios.

[0078] In the embodiment of the present application, a margin-based loss function can be used during the model training process, and the value of the margin is positively correlated with the image quality of the sample images used for training. Therefore, the value of the margin can be flexibly adjusted according to the image quality to obtain a face recognition model with higher robustness and accuracy. Specifically, referring to Figure 2, the training process of the face recognition model of the embodiment of the present application may include the following steps:

[0079] Step S201: Determine a face recognition training sample set.

[0080] The face recognition training sample set may include multiple sample images.

[0081] Face recognition is a metric learning task. A face recognition model with good recognition effect usually requires millions of data for training. Due to the large scale of training data, if offline stored data is used for training, it may take up too much memory space and increase training costs. Therefore, in the embodiment of the present application, model training can be performed through online training.

[0082] Specifically, a preset raw sample set can be obtained; wherein the raw sample set can include a preset number of raw sample images. Considering that real images in real application scenarios may have problems such as light disturbance, motion blur, and out-of-focus blur, as shown in Figure 1, and that the sample images in the acquired raw sample set are usually high-definition images with relatively ideal image data, which may differ significantly from images in real application scenarios, after obtaining the raw sample set, data enhancement processing can be performed on the raw sample set according to a preset data enhancement method to obtain a face recognition training sample set.

[0083] Here, the data enhancement processing of the original sample set may include performing color space enhancement processing on the original sample set, and / or performing motion blur enhancement processing on the original sample set, and / or performing out-of-focus blur enhancement processing on the original sample set to obtain a face recognition training sample set.

[0084] Optionally, a linear contrast enhancement method can be used to perform color space enhancement on the original sample images in the original sample set. If the width of the input image I is W, the height is H, and the channel is C, a pixel in I can be represented as I(h,w,c), where the height of the pixel is h∈[0,h) and the width of the pixel is w∈[0,W). Then the enhancement strategy can be: O(h,w,c)=αI(h,w,c)+β

[0085] Wherein, O(h,w,c) is the pixel point output by I(h,w,c) after color space enhancement processing, and α and β are preset linear coefficients and constants, respectively. Here, α can be preferably set to any value between (0.35,1.65), and β can be set to any value between (0,20).

[0086] Considering that the value of O(h,w,c) may be out of bounds, the maximum value of the pixel of the output image O can be recorded as V max , and normalize the output image O:

[0087] Among them, O'(h,w,c) is the normalized pixel point of O(h,w,c).

[0088] Based on this, the pixel values ​​of the output image can be restricted to a reasonable value range.

[0089] Optionally, motion blur enhancement can be performed on the original sample set. The key to motion blur enhancement lies in simulating the blur of motion direction and intensity, which is mainly determined by the intensity (degree) and angle (angle) of the convolution kernel. Therefore, the effect of motion blur enhancement can be controlled by setting the intensity and angle of the convolution kernel. Here, a diagonal convolution kernel (kernel) can be constructed. For example, a convolution kernel with an intensity of 3 [1,0,0; 0,1,0; 0,0,1] can be constructed. Then, an affine transformation matrix can be constructed according to a preset angle, and the convolution kernel can be affine transformed according to the affine transformation matrix to obtain a motion blur convolution kernel (motion_kernel). After that, the original sample image can be convolved with the motion blur convolution kernel to obtain an image after motion blur enhancement. In addition, in order to simulate actual motion as richly as possible, the intensity range of the convolution kernel can be set between (3, 9) and the angle range of the convolution kernel can be set between (0, 90°).

[0090] Optionally, the original sample set can be subjected to out-of-focus blur enhancement processing. When a person or a camera acquisition device moves, if the focal length of the camera acquisition device is not appropriate, out-of-focus blur is likely to occur. Since the imaging of an ideal point after passing through the camera can be described by a point spread function (PSF), a defocus convolution kernel (blur_kernel) can be used here to simulate the effect of out-of-focus blur; specifically, please refer to Figure 3. Compared with ordinary mean filtering, the mean filtering of the circular area can better simulate the effect of out-of-focus blur, so a circular mean filter kernel function can be used to construct a defocus convolution kernel. After that, the defocus convolution kernel can be used to perform a convolution operation on the original sample image to obtain an image after out-of-focus blur enhancement processing. In addition, in order to simulate the actual motion situation as richly as possible, the intensity range of the defocus convolution kernel can be set between (3, 5).

[0091] In addition, in order to construct sample images of varying quality, the data of the original sample set can be enhanced according to a preset enhancement ratio; wherein the enhancement ratio can be specific and situational according to actual needs, and this application does not limit this. For example, the original sample set can be enhanced in color space at an enhancement ratio of 15%, and / or the original sample set can be enhanced in motion blur at an enhancement ratio of 5%, and / or the original sample set can be enhanced in out-of-focus blur at an enhancement ratio of 5%, as shown in Figure 4, to obtain a richer face recognition training sample set that is closer to real application scenarios.

[0092] Step S202: Calculate the image quality of each sample image.

[0093] In the embodiment of the present application, the feature vector of each sample image may be determined separately.

[0094] Here, a preset feature extraction algorithm can be used to extract features from each sample image to obtain a feature vector of each sample image; wherein the preset feature extraction algorithm can be any common feature extraction algorithm in the prior art, including but not limited to color histogram, local binary pattern (LBP), histogram of oriented gradient (HOG) and other feature extraction algorithms.

[0095] It should be understood that the feature space of a high-quality image is highly aggregated, while the feature space of a low-quality image is relatively divergent. Therefore, in the embodiments of the present application, the image quality of a sample image can be calculated based on its feature vector. Referring to FIG. 5 , generally, if the image quality is good, the image quality measure calculated using the image's feature vector will be larger; if the image quality is poor, the image quality measure calculated using the image's feature vector will be smaller. Based on the calculated image quality measure, the image quality can be easily judged.

[0096] Specifically, the image quality of each sample image can be calculated based on the feature vector of each sample image. The method for calculating image quality can be set according to actual circumstances and is not specifically limited in this application. For example, the L1 norm of the feature vector of the sample image can be calculated, and this L1 norm can be used as a measure of the image quality of the sample image. For another example, the L2 norm of the feature vector of the sample image can be calculated, and this L2 norm can be used as a measure of the image quality of the sample image.

[0097] Step S203: Determine the margin of each sample image according to the image quality of each sample image.

[0098] In an embodiment of the present application, the margin of the sample image can be designed as a convex function that is positively correlated with the quality of the sample image. The setting method of this function can be specific and situational according to actual conditions, and this application does not limit this.

[0099] As just an example, in an embodiment of the present application, the upper limit value of image quality, the lower limit value of image quality, the upper limit value of margin and the lower limit value of margin can be determined respectively, and the margin of each sample image can be determined respectively based on the upper limit value of image quality, the lower limit value of image quality, the upper limit value of margin, the lower limit value of margin and the image quality of each sample image.

[0100] Specifically, the upper limit of image quality u can be α and the lower limit of image quality l α Determine them as 110 and 10 respectively; and the upper limit value of margin u m and the lower margin value l m Determined to be 0.8 and 0.45 respectively.

[0101] Thereafter, the margin of each sample image may be determined according to the image quality upper limit value, the image quality lower limit value, the margin upper limit value, the margin lower limit value, and the image quality of each sample image.

[0102] It should be noted that since the feature space of images with good quality has a high degree of aggregation, while the feature space of images with poor quality is relatively divergent, the margin of sample images with good quality can be set to a larger value to increase the cosine similarity between the same categories and further improve the category separability, while the margin of sample images with poor quality can be set to a smaller value to reduce the cosine similarity between the same categories and help better distinguish different faces.

[0103] Specifically, the image quality upper limit value u α , image quality lower limit l α Limit the image quality of the sample images to obtain the image quality of each sample image after the range is limited:

[0104] Among them, α i is the image quality of the i-th sample image, α i ' is the image quality of the i-th sample image after range limitation.

[0105] Afterwards, the image quality upper limit u α and the lower limit of image quality l α Calculate the image quality span u α -l α =100; you can also set the upper limit of the margin u m and the lower margin value l m Calculate the margin span u m -l m =0.35; then, the span ratio coefficient can be calculated based on the image quality span and margin span

[0106] The margin of the sample image can be determined based on the span ratio coefficient, the image quality lower limit, the margin lower limit, and the image quality after the sample image range is restricted:

[0107] Among them, m(α i') is the margin of the i-th sample image. Based on this, we can ensure that the margin of the worst quality sample image is 0.45, and the margin of the best quality sample image is 0.8. As the image quality improves, the image margin can also be adaptively increased, so that difficult samples (sample images with poor quality) can be better learned to adapt to more complex scenes.

[0108] In actual applications, the upper limit of image quality, the lower limit of image quality, the upper limit of margin and the lower limit of margin can also be specified and contextualized according to actual needs, including but not limited to the above-mentioned settings.

[0109] Step S204: Use the face recognition training sample set and the margin-based loss function to train the initial model to obtain a face recognition model.

[0110] In the existing CosFace algorithm, the margin-based loss function L is:

[0111] Among them, y i is the true category label of the i-th sample image, is the angle between the feature vector of the i-th sample and its true category label, m is the preset margin, s is the preset scaling parameter used to adjust the range of cosine similarity, j is the category label other than the category to which the i-th sample belongs, θ j is the angle between the feature vector of the i-th sample and its true category label.

[0112] In the embodiment of the present application, the margin-based loss function can be adjusted to obtain the margin-based loss function L' of the embodiment of the present application:

[0113] Among them, the margin m(a i The value of ') is positively correlated with the image quality of the sample image used for training, and can be obtained according to the method discussed in the above steps, which will not be repeated here.

[0114] In an embodiment of the present application, a face recognition training sample set and a margin-based loss function can be used to train an initial model to obtain a face recognition model. The initial model can be any artificial intelligence model in the prior art, including but not limited to convolutional neural network (CNN), recurrent neural network (RNN), long short-term memory network (LSTM), generative adversarial network (GAN), attention mechanism and other artificial intelligence models.

[0115] During the training process of the initial model, the above-mentioned margin-based loss function can be used to calculate the training loss value, and the parameters of the initial model can be adjusted based on the training loss value to obtain the face recognition model of the present application.

[0116] After the face recognition model is trained, it can be applied to face recognition tasks in real-world scenarios. Specifically, referring to Figure 6, the process of using the face recognition model for face recognition may include:

[0117] Step S601: Acquire a target image to be identified.

[0118] In a specific implementation of an embodiment of the present application, a preset camera acquisition device can be used to acquire images, and the acquired images can be stored in a preset storage module; when face recognition extraction is required, the target image to be identified can be obtained from the preset storage module.

[0119] In another specific implementation of the embodiment of the present application, a preset camera acquisition device can be used to perform real-time image acquisition to obtain the target image to be identified.

[0120] Step S602: Use a preset face recognition model to perform face recognition on the target image to obtain a face recognition result.

[0121] Here, the target image can be used as the input of the face recognition model, so that the face recognition result output by the face recognition model can be obtained.

[0122] In summary, the embodiment of the present application obtains a target image to be identified; performs face recognition on the target image using a preset face recognition model to obtain a face recognition result; wherein, the face recognition model uses a margin-based loss function during the training process, and the value of the margin is positively correlated with the image quality of the sample image used for training. Through the embodiment of the present application, the value of the margin can be determined based on the image quality of the sample image during the training process; wherein, the value of the margin is positively correlated with the image quality of the sample image, so that the value of the margin can be flexibly adjusted according to the image quality, which helps to improve the robustness of face recognition.

[0123] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0124] Corresponding to the face recognition method described in the above embodiment, FIG7 shows a structural diagram of an embodiment of a face recognition device provided in an embodiment of the present application.

[0125] In an embodiment of the present application, a face recognition device may include:

[0126] Image acquisition module 701, used to acquire the target image to be identified;

[0127] The face recognition module 702 is used to perform face recognition on the target image using a preset face recognition model to obtain a face recognition result; wherein, the face recognition model uses a margin-based loss function during the training process, and the value of the margin is positively correlated with the image quality of the sample image used for training.

[0128] In a specific implementation of the embodiment of the present application, the face recognition module may include:

[0129] A training sample set determination submodule is used to determine a face recognition training sample set; wherein the face recognition training sample set includes a plurality of sample images;

[0130] An image quality calculation submodule, used to calculate the image quality of each sample image respectively;

[0131] An image margin determination submodule, configured to determine the margin of each sample image according to the image quality of each sample image;

[0132] The initial model training submodule is used to train the initial model using the face recognition training sample set and the margin-based loss function to obtain the face recognition model.

[0133] In a specific implementation of the embodiment of the present application, the image quality calculation submodule may include:

[0134] a feature vector determination unit, configured to determine a feature vector for each sample image;

[0135] The image quality calculation unit is used to calculate the image quality of each sample image according to the feature vector of each sample image.

[0136] In a specific implementation of the embodiment of the present application, the image margin determination submodule may include:

[0137] a range determination unit, configured to respectively determine an upper limit value of image quality, a lower limit value of image quality, an upper limit value of margin, and a lower limit value of margin;

[0138] The image margin determining unit is configured to determine the margin of each sample image according to the image quality upper limit value, the image quality lower limit value, the margin upper limit value, the margin lower limit value, and the image quality of each sample image.

[0139] In a specific implementation of the embodiment of the present application, the image margin determination unit may include:

[0140] a quality range limiting subunit, configured to limit the image quality of each sample image according to the image quality upper limit value and the image quality lower limit value, to obtain the image quality of each sample image after the range is limited;

[0141] a quality span calculation subunit, configured to calculate an image quality span according to the image quality upper limit value and the image quality lower limit value;

[0142] a margin span calculation subunit, configured to calculate the margin span according to the margin upper limit value and the margin lower limit value;

[0143] a scale factor calculation subunit, configured to calculate a span scale factor according to the margin span and the image quality span;

[0144] The image margin determination subunit is used to determine the margin of each sample image according to the span ratio coefficient, the image quality lower limit, the margin lower limit and the image quality of each sample image after range restriction.

[0145] In a specific implementation of the embodiment of the present application, the training sample set determination submodule may include:

[0146] A sample set acquisition unit, used to acquire a preset original sample set;

[0147] The enhancement processing unit is used to perform data enhancement processing on the original sample set according to a preset data enhancement method to obtain the face recognition training sample set.

[0148] In a specific implementation of an embodiment of the present application, the enhancement processing unit can be used to perform color space enhancement processing on the original sample set, and / or, perform motion blur enhancement processing on the original sample set, and / or, perform out-of-focus blur enhancement processing on the original sample set to obtain the face recognition training sample set.

[0149] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices, modules and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0150] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0151] Figure 8 shows a schematic block diagram of a robot provided in an embodiment of the present application. For ease of explanation, only the parts related to the embodiment of the present application are shown.

[0152] As shown in FIG8 , the robot 8 of this embodiment includes: a processor 80, a memory 81, and a computer program 82 stored in the memory 81 and executable on the processor 80. When the processor 80 executes the computer program 82, it implements the steps in the aforementioned face recognition method embodiments, such as steps S701 to S702 shown in FIG7 . Alternatively, when the processor 80 executes the computer program 82, it implements the functions of the modules / units in the aforementioned device embodiments, such as the functions of modules 701 to 702 shown in FIG7 .

[0153] For example, the computer program 82 may be divided into one or more modules / units, which are stored in the memory 81 and executed by the processor 80 to implement the present application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 82 in the robot 8.

[0154] Those skilled in the art will understand that Figure 8 is merely an example of the robot 8 and does not constitute a limitation on the robot 8. The robot 8 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the robot 8 may also include input and output devices, network access devices, buses, etc.

[0155] The processor 80 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0156] The memory 81 can be an internal storage unit of the robot 8, such as a hard drive or memory of the robot 8. The memory 81 can also be an external storage device of the robot 8, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the robot 8. Furthermore, the memory 81 can include both an internal storage unit of the robot 8 and an external storage device. The memory 81 is used to store the computer program and other programs and data required by the robot 8. The memory 81 can also be used to temporarily store data that has been output or is about to be output.

[0157] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0158] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0159] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0160] In the embodiments provided in this application, it should be understood that the disclosed devices / robots and methods can be implemented in other ways. For example, the device / robot embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0161] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0162] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0163] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable storage media do not include electric carrier signals and telecommunication signals.

[0164] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A face recognition method, characterized in that, comprising: Obtaining a target image to be recognized; Performing face recognition on the target image using a preset face recognition model to obtain a face recognition result; Wherein, in the training process of the face recognition model, a margin-based loss function is used, and the value of the margin is positively correlated with the image quality of the sample images used for training.

2. The face recognition method according to claim 1, characterized in that, The training process of the face recognition model includes: Determining a face recognition training sample set; wherein, the face recognition training sample set includes multiple sample images; Calculating the image quality of each sample image respectively; Determining the margin of each sample image respectively according to the image quality of each sample image; Training an initial model using the face recognition training sample set and the margin-based loss function to obtain the face recognition model.

3. The face recognition method according to claim 2, characterized in that, The calculating the image quality of each sample image respectively includes: Determining the feature vector of each sample image respectively; Calculating the image quality of each sample image respectively according to the feature vector of each sample image.

4. The face recognition method according to claim 2, characterized in that, The determining the margin of each sample image respectively according to the image quality of each sample image includes: Determining an image quality upper limit value, an image quality lower limit value, a margin upper limit value and a margin lower limit value respectively; Determining the margin of each sample image respectively according to the image quality upper limit value, the image quality lower limit value, the margin upper limit value, the margin lower limit value and the image quality of each sample image.

5. The face recognition method according to claim 4, characterized in that, The determining the margin of each sample image respectively according to the image quality upper limit value, the image quality lower limit value, the margin upper limit value, the margin lower limit value and the image quality of each sample image includes: Performing range limitation on the image quality of each sample image according to the image quality upper limit value and the image quality lower limit value to obtain the range-limited image quality of each sample image; Calculating an image quality span according to the image quality upper limit value and the image quality lower limit value; Calculating a margin span according to the margin upper limit value and the margin lower limit value; Calculating a span ratio coefficient according to the margin span and the image quality span; Determining the margin of each sample image respectively according to the span ratio coefficient, the image quality lower limit value, the margin lower limit value and the range-limited image quality of each sample image .

6. The face recognition method according to any one of claims 2 to 5, characterized in that, The determining the face recognition training sample set includes: Obtaining a preset original sample set; Performing data augmentation processing on the original sample set according to a preset data augmentation method to obtain the face recognition training sample set.

7. The face recognition method according to claim 6, characterized in that, The performing data augmentation processing on the original sample set according to a preset data augmentation method to obtain the face recognition training sample set includes: Perform color space enhancement processing on the original sample set, and / or perform motion blur enhancement processing on the original sample set, and / or perform defocus blur enhancement processing on the original sample set to obtain the face recognition training sample set.

8. A face recognition device Characterized in that It includes: An image acquisition module for acquiring a target image to be recognized; A face recognition module for performing face recognition on the target image using a preset face recognition model to obtain a face recognition result; wherein, in the training process of the face recognition model, a margin-based loss function is used, and the value of the margin is positively correlated with the image quality of the sample images used for training.

9. A computer-readable storage medium storing a computer program Characterized in that When the computer program is executed by a processor, the steps of the face recognition method according to any one of claims 1 to 7 are implemented.

10. A robot, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor Characterized in that When the processor executes the computer program, the steps of the face recognition method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Face recognition method and device, electronic equipment and storage medium

    CN112052789A

  • Face recognition model training method and device, face recognition method and device and related equipment

    CN114255354A

  • Face recognition method and device, electronic equipment and computer readable storage medium

    CN116152870A