Face recognition method and system based on multi-model cooperation

Through a multi-model collaborative approach, feature extraction and classification models are trained by selecting training data sets processed with multiple different mask strategies based on scene features, which solves the problem of low recognition accuracy of a single model and achieves highly adaptable and accurate face recognition in diverse environments.

CN120656216AActive Publication Date: 2025-09-16GUANGZHOU VIDEO STAR INTELLIGENT CO LTD

Patent Information

Application Number
CN202510439815.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-09-16
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

In existing face recognition technology, the single model recognition method based on a single training data set results in low recognition accuracy and efficiency, insufficient intelligence and automation, and cannot adapt to diverse environments.

Method used

A multi-model collaborative method is adopted to train feature extraction and classification models based on a training data set processed by multiple different mask strategies through scene feature selection. Multiple feature extraction models and classification models are combined to determine the classification and recognition results of face images.

Benefits of technology

It improves the adaptability and accuracy of face recognition, makes the recognition results more robust, suitable for diverse environments, and improves the intelligence and automation of face recognition services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656216A_ABST
    Figure CN120656216A_ABST
Patent Text Reader

Abstract

The invention discloses a face recognition method and system based on multi-model cooperation. The method comprises the following steps: acquiring face image data to be recognized and corresponding scene features; determining a plurality of corresponding feature extraction models and classification models according to the scene features; the plurality of feature extraction models and the classification model are obtained through training of a training data set processed by a plurality of different mask strategies corresponding to the scene features; inputting the face image data into the plurality of feature extraction models to obtain a plurality of image features corresponding to the face image; and determining a classification recognition result corresponding to the face image data according to the plurality of image features and the classification model. Therefore, the method can improve the adaptability and accuracy of face recognition in different scenes, enables the recognition result to be more robust, is suitable for diversified environments, and improves the intelligent degree and automation degree of face recognition service.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular to a face recognition method and system based on multi-model collaboration. Background Art

[0002] In recent years, with the widespread use of facial recognition devices, "face scanning" has become an essential part of daily life. With the rapid development of algorithmic technology, improving facial recognition accuracy through this technology has become a major technical challenge. However, most existing facial recognition technologies rely solely on a single model trained with a single training dataset, failing to fully leverage the algorithmic advantages of multiple models trained with a multi-mask strategy. Consequently, these technologies exhibit low accuracy and efficiency, and lack both intelligence and automation. Clearly, existing technologies have shortcomings that urgently need to be addressed. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a face recognition method and system based on multi-model collaboration, which can improve the adaptability and accuracy of face recognition in different scenarios, make the recognition results more robust and applicable to diverse environments, and improve the intelligence and automation level of face recognition services.

[0004] In order to solve the above technical problems, the first aspect of the present invention discloses a face recognition method based on multi-model collaboration, the method comprising: Obtaining facial image data to be recognized and corresponding scene features; Determining a plurality of corresponding feature extraction models and classification models according to the scene features; the plurality of feature extraction models and classification models are trained by training data sets processed by a plurality of different masking strategies corresponding to the scene features; Inputting the facial image data into the multiple feature extraction models to obtain multiple image features corresponding to the facial image; A classification recognition result corresponding to the facial image data is determined based on the multiple image features and the classification model.

[0005] As an optional embodiment, in the first aspect of the present invention, the scene characteristics include at least one of identification location, ambient light, scene type, user information and identification purpose.

[0006] As an optional embodiment, in the first aspect of the present invention, determining the corresponding multiple feature extraction models and classification models based on the scene features includes: Determining a mask prediction parameter corresponding to the scene feature; For each candidate feature extraction model, determining the mask strategy parameters corresponding to the training data set corresponding to the candidate feature extraction model; Calculating a first similarity between the mask strategy parameter and the mask prediction parameter to obtain a first model priority corresponding to the candidate feature extraction model; Screening out all the candidate feature models whose first model priority is greater than a preset first priority threshold, and obtaining corresponding multiple feature extraction models; At least one classification model is determined according to the mask strategy parameters corresponding to the training data sets corresponding to the multiple feature extraction models.

[0007] As an optional implementation manner, in the first aspect of the present invention, determining the occlusion prediction parameter corresponding to the scene feature includes: The scene features are input into a trained mask parameter prediction neural network to obtain mask prediction parameters corresponding to the scene features; the mask parameter prediction neural network is trained by a training data set including multiple training scene images and corresponding scene feature annotations and mask situation annotations.

[0008] As an optional embodiment, in the first aspect of the present invention, the mask strategy parameter and the mask prediction parameter both include at least one of a mask image position, a mask portion size, a mask portion area ratio, and three-dimensional information of the mask portion; and the first similarity is calculated by the following steps: Inputting the mask strategy parameters into a trained mask vector feature extraction neural network to obtain a first effect feature vector corresponding to the mask strategy parameters; Inputting the mask prediction parameter into the mask vector feature extraction neural network to obtain a second effect feature vector corresponding to the mask prediction parameter; the mask vector feature extraction neural network is trained using a training data set including a plurality of training mask parameters and corresponding pre-mask image annotations and post-mask image annotations; The inverse of the vector distance between the first effect feature vector and the second effect feature vector is calculated to obtain the first similarity.

[0009] As an optional embodiment, in the first aspect of the present invention, determining at least one classification model based on the mask strategy parameters corresponding to the training data sets corresponding to the multiple feature extraction models includes: For each candidate classification model, calculating, based on the mask vector feature extraction neural network and the vector distance algorithm, a second similarity between the mask strategy parameters corresponding to the training data set of the candidate classification model and the mask strategy parameters corresponding to the training data set corresponding to each of the feature extraction models; Calculating an average of the second similarities corresponding to the candidate classification model and each of the feature extraction models to obtain a second model priority corresponding to the candidate classification model; All the candidate classification models whose second model priority is greater than a preset second priority threshold are screened out to obtain at least one corresponding classification model.

[0010] As an optional embodiment, in the first aspect of the present invention, the candidate feature extraction model and the candidate classification model are trained by the following steps: Determining a plurality of training image datasets after masking using different mask strategy parameters; A preset prediction basic model is trained based on each of the training image data sets to obtain a plurality of trained preliminary prediction models; the prediction basic model and the preliminary prediction model include a feature extraction model and a classification model; determining all of the training image data sets as a common data set; Based on the common data set, all the preliminary prediction models are jointly trained, and model parameters are optimized based on a common loss function related to the mask strategy parameters to obtain the trained candidate feature extraction model and the candidate classification model; the common loss function is a weighted sum of the loss function values ​​of each of the preliminary prediction models; wherein the weighted weight corresponding to the loss function value corresponding to each of the preliminary prediction models is proportional to the masking degree of the mask strategy parameters corresponding to the training image data set corresponding to the preliminary prediction model; wherein the masking degree is calculated by the following steps: Inputting the mask strategy parameters corresponding to the training image data set corresponding to the preliminary prediction model into the mask vector feature extraction neural network to obtain a third effect feature vector corresponding to the mask strategy parameters; Calculating a first vector distance between the third effect feature vector and a first reference feature vector corresponding to a standard unobstructed image; Calculating a second vector distance between the third effect feature vector and a second reference feature vector corresponding to the standard fully occluded image; the first reference feature vector and the second reference feature vector are both obtained by pre-analyzing vector extraction results of a preset data set based on the occlusion vector feature extraction neural network; The ratio of the first vector distance to the second vector distance is calculated to obtain the masking degree corresponding to the preliminary prediction model.

[0011] As an optional embodiment, in the first aspect of the present invention, determining the classification recognition result corresponding to the facial image data based on the multiple image features and the classification model includes: Based on a random combination algorithm, the plurality of image features are randomly sampled and combined to obtain a plurality of combined feature sets; Inputting each of the combined feature sets into the classification model to obtain an output set recognition result; The intersection of all the set recognition results is calculated to obtain the classification recognition result corresponding to the face image data.

[0012] A second aspect of an embodiment of the present invention discloses a face recognition system based on multi-model collaboration, the system comprising: An acquisition module is used to obtain the face image data to be recognized and the corresponding scene features; A determination module, configured to determine a plurality of corresponding feature extraction models and classification models based on the scene features; the plurality of feature extraction models and classification models are trained using a training data set processed using a plurality of different masking strategies corresponding to the scene features; an extraction module, configured to input the facial image data into the plurality of feature extraction models to obtain a plurality of image features corresponding to the facial image; The recognition module is used to determine the classification recognition result corresponding to the facial image data based on the multiple image features and the classification model.

[0013] As an optional implementation, in the second aspect of the present invention, the scene feature includes at least one of an identification location, ambient light, scene type, user information, and identification purpose.

[0014] As an optional embodiment, in the second aspect of the present invention, the specific manner in which the determination module determines the corresponding multiple feature extraction models and classification models based on the scene features includes: Determining a mask prediction parameter corresponding to the scene feature; For each candidate feature extraction model, determining the mask strategy parameters corresponding to the training data set corresponding to the candidate feature extraction model; Calculating a first similarity between the mask strategy parameter and the mask prediction parameter to obtain a first model priority corresponding to the candidate feature extraction model; Screening out all the candidate feature models whose first model priority is greater than a preset first priority threshold, and obtaining corresponding multiple feature extraction models; At least one classification model is determined according to the mask strategy parameters corresponding to the training data sets corresponding to the multiple feature extraction models.

[0015] As an optional implementation, in the second aspect of the present invention, the specific manner in which the determination module determines the occlusion prediction parameter corresponding to the scene feature includes: The scene features are input into a trained mask parameter prediction neural network to obtain mask prediction parameters corresponding to the scene features; the mask parameter prediction neural network is trained by a training data set including multiple training scene images and corresponding scene feature annotations and mask situation annotations.

[0016] As an optional embodiment, in the second aspect of the present invention, the mask strategy parameter and the mask prediction parameter both include at least one of a masked image position, a masked portion size, a masked portion area ratio, and three-dimensional information of the masked portion; and the first similarity is calculated by the following steps: Inputting the mask strategy parameters into a trained mask vector feature extraction neural network to obtain a first effect feature vector corresponding to the mask strategy parameters; Inputting the mask prediction parameter into the mask vector feature extraction neural network to obtain a second effect feature vector corresponding to the mask prediction parameter; the mask vector feature extraction neural network is trained using a training data set including a plurality of training mask parameters and corresponding pre-mask image annotations and post-mask image annotations; The inverse of the vector distance between the first effect feature vector and the second effect feature vector is calculated to obtain the first similarity.

[0017] As an optional embodiment, in the second aspect of the present invention, the determination module determines a specific manner of at least one classification model based on the mask strategy parameters corresponding to the training data sets corresponding to the multiple feature extraction models, including: For each candidate classification model, calculating, based on the mask vector feature extraction neural network and the vector distance algorithm, a second similarity between the mask strategy parameters corresponding to the training data set of the candidate classification model and the mask strategy parameters corresponding to the training data set corresponding to each of the feature extraction models; Calculating an average of the second similarities corresponding to the candidate classification model and each of the feature extraction models to obtain a second model priority corresponding to the candidate classification model; All the candidate classification models whose second model priority is greater than a preset second priority threshold are screened out to obtain at least one corresponding classification model.

[0018] As an optional embodiment, in the second aspect of the present invention, the candidate feature extraction model and the candidate classification model are trained by the following steps: Determining a plurality of training image datasets after masking using different mask strategy parameters; A preset prediction basic model is trained based on each of the training image data sets to obtain a plurality of trained preliminary prediction models; the prediction basic model and the preliminary prediction model include a feature extraction model and a classification model; determining all of the training image data sets as a common data set; Based on the common data set, all the preliminary prediction models are jointly trained, and model parameters are optimized based on a common loss function related to the mask strategy parameters to obtain the trained candidate feature extraction model and the candidate classification model; the common loss function is a weighted sum of the loss function values ​​of each of the preliminary prediction models; wherein the weighted weight corresponding to the loss function value corresponding to each of the preliminary prediction models is proportional to the masking degree of the mask strategy parameters corresponding to the training image data set corresponding to the preliminary prediction model; wherein the masking degree is calculated by the following steps: Inputting the mask strategy parameters corresponding to the training image data set corresponding to the preliminary prediction model into the mask vector feature extraction neural network to obtain a third effect feature vector corresponding to the mask strategy parameters; Calculating a first vector distance between the third effect feature vector and a first reference feature vector corresponding to a standard unobstructed image; Calculating a second vector distance between the third effect feature vector and a second reference feature vector corresponding to the standard fully occluded image; the first reference feature vector and the second reference feature vector are both obtained by pre-analyzing vector extraction results of a preset data set based on the occlusion vector feature extraction neural network; The ratio of the first vector distance to the second vector distance is calculated to obtain the masking degree corresponding to the preliminary prediction model.

[0019] As an optional embodiment, in the second aspect of the present invention, the specific manner in which the recognition module determines the classification recognition result corresponding to the facial image data based on the multiple image features and the classification model includes: Based on a random combination algorithm, the plurality of image features are randomly sampled and combined to obtain a plurality of combined feature sets; Inputting each of the combined feature sets into the classification model to obtain an output set recognition result; The intersection of all the set recognition results is calculated to obtain the classification recognition result corresponding to the face image data.

[0020] A third aspect of the present invention discloses another face recognition system based on multi-model collaboration, the system comprising: a memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute part or all of the steps in the face recognition method based on multi-model collaboration disclosed in the first aspect of the present invention.

[0021] The fourth aspect of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute part or all of the steps in the face recognition method based on multi-model collaboration disclosed in the first aspect of the present invention.

[0022] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: The present invention can select the corresponding feature extraction model and classification model obtained by training the training data set processed by multiple different mask strategies based on scene features, and use multiple feature extraction models to extract multiple image features of the face image, and combine the classification model to determine the classification recognition result, thereby improving the adaptability and accuracy of face recognition in different scenarios, making the recognition results more robust and applicable to diversified environments, and improving the intelligence and automation of face recognition services. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0024] Figure 1 This is a flow chart of a face recognition method based on multi-model collaboration disclosed in an embodiment of the present invention.

[0025] Figure 2 This is a structural diagram of a face recognition system based on multi-model collaboration disclosed in an embodiment of the present invention.

[0026] Figure 3 This is a structural diagram of another face recognition system based on multi-model collaboration disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0028] The terms "first," "second," and so on, in the description and claims of the present invention and the accompanying drawings are used to distinguish between different objects, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or device.

[0029] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0030] The present invention discloses a face recognition method and system based on multi-model collaboration. This method selects, based on scene features, corresponding feature extraction models and classification models trained using training data sets processed using multiple different masking strategies. The system then uses the multiple feature extraction models to extract multiple image features from facial images, and combines these models with the classification model to determine classification and recognition results. This method improves the adaptability and accuracy of face recognition in different scenarios, making the recognition results more robust and applicable to diverse environments, and enhancing the intelligence and automation of face recognition services. These are described in detail below.

[0031] Example 1 See also Figure 1 , Figure 1 This is a flow chart of a face recognition method based on multi-model collaboration disclosed in an embodiment of the present invention. Figure 1 The face recognition method based on multi-model collaboration described above can be applied to a data processing system / data processing device / data processing server (wherein the server includes a local processing server or a cloud processing server). Figure 1 As shown, the face recognition method based on multi-model collaboration may include the following operations: 101. Obtain facial image data to be recognized and corresponding scene features.

[0032] 102. Determine multiple corresponding feature extraction models and classification models based on scene features. Optionally, multiple feature extraction models and classification models are trained using training data sets processed using multiple different masking strategies corresponding to scene features.

[0033] 103. Input the facial image data into multiple feature extraction models to obtain multiple image features corresponding to the facial image. 104. Determine a classification recognition result corresponding to the facial image data based on multiple image features and a classification model.

[0034] It can be seen that the above-mentioned embodiments of the invention can select the corresponding feature extraction model and classification model obtained by training the training data set processed by multiple different mask strategies based on scene features, and use multiple feature extraction models to extract multiple image features of the face image, and combine the classification model to determine the classification recognition results, thereby improving the adaptability and accuracy of face recognition in different scenarios, making the recognition results more robust and applicable to diverse environments, and improving the intelligence and automation of face recognition services.

[0035] As an optional embodiment, in the above steps, the scene features include at least one of identification location, ambient light, scene type, user information and identification purpose.

[0036] It can be seen that through the above optional embodiments, the content of the scene features is limited to comprehensively characterize the features related to the face recognition scene, so as to facilitate subsequent accurate face recognition, assist in improving the adaptability and accuracy of face recognition in different scenarios, make the recognition results more robust and applicable to diverse environments, and improve the intelligence and automation level of face recognition services.

[0037] As an optional embodiment, in the above steps, determining corresponding multiple feature extraction models and classification models according to scene features includes: Determining occlusion prediction parameters corresponding to scene features; For each candidate feature extraction model, determining the mask strategy parameters corresponding to the training data set corresponding to the candidate feature extraction model; Calculating a first similarity between the mask strategy parameter and the mask prediction parameter to obtain a first model priority corresponding to the candidate feature extraction model; Screening out all candidate feature models whose first model priority is greater than a preset first priority threshold, and obtaining corresponding multiple feature extraction models; At least one classification model is determined according to mask strategy parameters corresponding to training data sets corresponding to the plurality of feature extraction models.

[0038] It can be seen that through the above optional embodiments, the mask prediction parameters are determined based on the scene features, and the similarity between the mask strategy parameters of the candidate feature extraction model and the mask prediction parameters is calculated, and the feature extraction models with a priority higher than the threshold are screened. At the same time, the classification model is determined according to the mask strategy parameters of the selected feature extraction model to determine a more suitable and accurate feature extraction model and classification model, so as to assist in improving the adaptability and accuracy of face recognition in different scenarios, make the recognition results more robust and applicable to diverse environments, and improve the intelligence and automation level of face recognition services.

[0039] As an optional embodiment, in the above step, determining the occlusion prediction parameter corresponding to the scene feature includes: The scene features are input into a trained mask parameter prediction neural network to obtain mask prediction parameters corresponding to the scene features; the mask parameter prediction neural network is trained by a training data set including multiple training scene images and corresponding scene feature annotations and mask situation annotations.

[0040] It can be seen that through the above-mentioned optional embodiments, the mask parameter prediction neural network can be input based on the scene features to obtain the corresponding mask prediction parameters, and the correspondence between the scene features and the mask conditions in the training data set can be learned to improve the accuracy and adaptability of the mask prediction parameters, thereby assisting in improving the adaptability and accuracy of face recognition in different scenarios, making the recognition results more robust and applicable to diverse environments, and improving the intelligence and automation of face recognition services.

[0041] As an optional embodiment, in the above steps, the mask strategy parameter and the mask prediction parameter both include at least one of the mask image position, the mask portion size, the mask portion area ratio, and the mask portion three-dimensional information; and the first similarity is calculated by the following steps: Inputting the mask strategy parameters into the trained mask vector feature extraction neural network to obtain the first effect feature vector corresponding to the mask strategy parameters; Inputting the mask prediction parameter into a mask vector feature extraction neural network to obtain a second effect feature vector corresponding to the mask prediction parameter; optionally, the mask vector feature extraction neural network is trained using a training data set including a plurality of training mask parameters and corresponding pre-mask image annotations and post-mask image annotations; The inverse of the vector distance between the first effect feature vector and the second effect feature vector is calculated to obtain a first similarity.

[0042] It can be seen that through the above optional embodiments, based on the mask vector feature extraction neural network, the mask strategy parameters and the mask prediction parameters are respectively converted into effect feature vectors, and the similarity is calculated through vector distance to quantify the degree of matching between different masking strategies and predicted masking situations, thereby improving the adaptability and accuracy of feature extraction model selection, assisting in improving the adaptability and accuracy of face recognition in different scenarios, making the recognition results more robust and applicable to diverse environments, and improving the intelligence and automation of face recognition services.

[0043] As an optional embodiment, in the above step, determining at least one classification model based on mask strategy parameters corresponding to training data sets corresponding to multiple feature extraction models includes: For each candidate classification model, calculating, based on the mask vector feature extraction neural network and the vector distance algorithm, a second similarity between the mask strategy parameters corresponding to the training data set of the candidate classification model and the mask strategy parameters corresponding to the training data set corresponding to each feature extraction model; Calculating an average of the second similarities corresponding to the candidate classification model and each feature extraction model to obtain the second model priority corresponding to the candidate classification model; All candidate classification models whose second model priority is greater than a preset second priority threshold are screened out to obtain at least one corresponding classification model.

[0044] It can be seen that through the above optional embodiments, based on the mask vector feature extraction neural network and vector distance algorithm, the similarity between the mask strategy parameters of the candidate classification model and the mask strategy parameters of the feature extraction model is calculated, and the priority of the classification model is determined by the average similarity to screen out a classification model that better matches the feature extraction model, thereby improving the overall recognition accuracy, assisting in improving the adaptability and accuracy of face recognition in different scenarios, making the recognition results more robust and applicable to diverse environments, and improving the intelligence and automation of face recognition services.

[0045] As an optional embodiment, in the above steps, the candidate feature extraction model and the candidate classification model are trained by the following steps: Determining a plurality of training image datasets after masking using different mask strategy parameters; A preset prediction basic model is trained based on each training image data set to obtain a plurality of trained preliminary prediction models; optionally, the prediction basic model and the preliminary prediction model include a feature extraction model and a classification model; All training image datasets are determined as a common dataset; Based on a common data set, all preliminary prediction models are jointly trained, and model parameters are optimized based on a common loss function related to the mask strategy parameters to obtain trained candidate feature extraction models and candidate classification models; optionally, the common loss function is a weighted sum of loss function values ​​of each preliminary prediction model; wherein the weighted weight corresponding to the loss function value corresponding to each preliminary prediction model is proportional to the masking degree of the mask strategy parameters corresponding to the training image data set corresponding to the preliminary prediction model; wherein the masking degree is calculated by the following steps: Inputting the mask strategy parameters corresponding to the training image data set corresponding to the preliminary prediction model into the mask vector feature extraction neural network to obtain a third effect feature vector corresponding to the mask strategy parameters; Calculating a first vector distance between the third effect feature vector and a first reference feature vector corresponding to the standard unobstructed image; Calculating a second vector distance between the third effect feature vector and a second reference feature vector corresponding to the standard fully occluded image; optionally, the first reference feature vector and the second reference feature vector are both obtained by pre-analyzing vector extraction results of a preset data set based on a mask vector feature extraction neural network; The ratio of the first vector distance to the second vector distance is calculated to obtain the masking degree corresponding to the preliminary prediction model.

[0046] It can be seen that through the above optional embodiments, multiple training image data sets are generated based on different mask strategy parameters, and all preliminary prediction models are jointly trained through a common data set, and the model parameters are optimized using a common loss function related to the mask strategy parameters to train feature extraction models and classification models that are more suitable for masking situations, thereby assisting in improving the adaptability and accuracy of face recognition in different scenarios, making the recognition results more robust and applicable to diverse environments, and improving the intelligence and automation of face recognition services.

[0047] As an optional embodiment, in the above step, determining the classification recognition result corresponding to the facial image data based on multiple image features and a classification model includes: Based on the random combination algorithm, multiple image features are randomly sampled and combined to obtain multiple combined feature sets; Input each combined feature set into the classification model to obtain the output set recognition result; Calculate the intersection of all set recognition results to obtain the classification recognition results corresponding to the face image data.

[0048] It can be seen that through the above optional embodiments, multiple image features are sampled and combined based on a random combination algorithm, and the recognition results of each combined feature set are calculated through a classification model, and then the intersection of all recognition results is taken to reduce feature interference and misclassification, improve the adaptability and accuracy of face recognition in different scenarios, make the recognition results more robust and applicable to diverse environments, and improve the intelligence and automation level of face recognition services.

[0049] Example 2 See also Figure 2 , Figure 2 This is a structural diagram of a face recognition system based on multi-model collaboration disclosed in an embodiment of the present invention. Figure 2 The face recognition system based on multi-model collaboration described above can be applied to a data processing system / data processing device / data processing server (wherein the server includes a local processing server or a cloud processing server). Figure 2 As shown, the face recognition system based on multi-model collaboration may include: The acquisition module 201 is used to acquire the facial image data to be recognized and the corresponding scene features.

[0050] The determination module 202 is used to determine corresponding multiple feature extraction models and classification models according to scene features. Optionally, multiple feature extraction models and classification models are trained using training data sets processed using multiple different masking strategies corresponding to scene features.

[0051] The extraction module 203 is used to input the facial image data into multiple feature extraction models to obtain multiple image features corresponding to the facial image. The recognition module 204 is used to determine the classification recognition result corresponding to the face image data based on multiple image features and a classification model.

[0052] It can be seen that the above-mentioned embodiments of the invention can select the corresponding feature extraction model and classification model obtained by training the training data set processed by multiple different mask strategies based on scene features, and use multiple feature extraction models to extract multiple image features of the face image, and combine the classification model to determine the classification recognition results, thereby improving the adaptability and accuracy of face recognition in different scenarios, making the recognition results more robust and applicable to diverse environments, and improving the intelligence and automation of face recognition services.

[0053] As an optional embodiment, the scene feature includes at least one of an identification location, ambient light, scene type, user information, and identification purpose.

[0054] It can be seen that through the above optional embodiments, the content of the scene features is limited to comprehensively characterize the features related to the face recognition scene, so as to facilitate subsequent accurate face recognition, assist in improving the adaptability and accuracy of face recognition in different scenarios, make the recognition results more robust and applicable to diverse environments, and improve the intelligence and automation level of face recognition services.

[0055] As an optional embodiment, the specific manner in which the determination module determines the corresponding multiple feature extraction models and classification models according to the scene features includes: Determining occlusion prediction parameters corresponding to scene features; For each candidate feature extraction model, determining the mask strategy parameters corresponding to the training data set corresponding to the candidate feature extraction model; Calculating a first similarity between the mask strategy parameter and the mask prediction parameter to obtain a first model priority corresponding to the candidate feature extraction model; Screening out all candidate feature models whose first model priority is greater than a preset first priority threshold, and obtaining corresponding multiple feature extraction models; At least one classification model is determined according to mask strategy parameters corresponding to training data sets corresponding to the plurality of feature extraction models.

[0056] It can be seen that through the above optional embodiments, the mask prediction parameters are determined based on the scene features, and the similarity between the mask strategy parameters of the candidate feature extraction model and the mask prediction parameters is calculated, and the feature extraction models with a priority higher than the threshold are screened. At the same time, the classification model is determined according to the mask strategy parameters of the selected feature extraction model to determine a more suitable and accurate feature extraction model and classification model, so as to assist in improving the adaptability and accuracy of face recognition in different scenarios, make the recognition results more robust and applicable to diverse environments, and improve the intelligence and automation level of face recognition services.

[0057] As an optional embodiment, the specific manner in which the determination module determines the occlusion prediction parameter corresponding to the scene feature includes: The scene features are input into a trained mask parameter prediction neural network to obtain mask prediction parameters corresponding to the scene features; the mask parameter prediction neural network is trained by a training data set including multiple training scene images and corresponding scene feature annotations and mask situation annotations.

[0058] It can be seen that through the above-mentioned optional embodiments, the mask parameter prediction neural network can be input based on the scene features to obtain the corresponding mask prediction parameters, and the correspondence between the scene features and the mask conditions in the training data set can be learned to improve the accuracy and adaptability of the mask prediction parameters, thereby assisting in improving the adaptability and accuracy of face recognition in different scenarios, making the recognition results more robust and applicable to diverse environments, and improving the intelligence and automation of face recognition services.

[0059] As an optional embodiment, the mask strategy parameter and the mask prediction parameter both include at least one of a mask image position, a mask portion size, a mask portion area ratio, and three-dimensional information of the mask portion; and the first similarity is calculated by the following steps: Inputting the mask strategy parameters into the trained mask vector feature extraction neural network to obtain the first effect feature vector corresponding to the mask strategy parameters; Inputting the mask prediction parameter into a mask vector feature extraction neural network to obtain a second effect feature vector corresponding to the mask prediction parameter; optionally, the mask vector feature extraction neural network is trained using a training data set including a plurality of training mask parameters and corresponding pre-mask image annotations and post-mask image annotations; The inverse of the vector distance between the first effect feature vector and the second effect feature vector is calculated to obtain a first similarity.

[0060] It can be seen that through the above optional embodiments, based on the mask vector feature extraction neural network, the mask strategy parameters and the mask prediction parameters are respectively converted into effect feature vectors, and the similarity is calculated through vector distance to quantify the degree of matching between different masking strategies and predicted masking situations, thereby improving the adaptability and accuracy of feature extraction model selection, assisting in improving the adaptability and accuracy of face recognition in different scenarios, making the recognition results more robust and applicable to diverse environments, and improving the intelligence and automation of face recognition services.

[0061] As an optional embodiment, the determination module determines a specific manner of determining at least one classification model based on mask strategy parameters corresponding to training data sets corresponding to multiple feature extraction models, including: For each candidate classification model, calculating, based on the mask vector feature extraction neural network and the vector distance algorithm, a second similarity between the mask strategy parameters corresponding to the training data set of the candidate classification model and the mask strategy parameters corresponding to the training data set corresponding to each feature extraction model; Calculating an average of the second similarities corresponding to the candidate classification model and each feature extraction model to obtain the second model priority corresponding to the candidate classification model; All candidate classification models whose second model priority is greater than a preset second priority threshold are screened out to obtain at least one corresponding classification model.

[0062] It can be seen that through the above optional embodiments, based on the mask vector feature extraction neural network and vector distance algorithm, the similarity between the mask strategy parameters of the candidate classification model and the mask strategy parameters of the feature extraction model is calculated, and the priority of the classification model is determined by the average similarity to screen out a classification model that better matches the feature extraction model, thereby improving the overall recognition accuracy, assisting in improving the adaptability and accuracy of face recognition in different scenarios, making the recognition results more robust and applicable to diverse environments, and improving the intelligence and automation of face recognition services.

[0063] As an optional embodiment, the candidate feature extraction model and the candidate classification model are trained through the following steps: Determining a plurality of training image datasets after masking using different mask strategy parameters; A preset prediction basic model is trained based on each training image data set to obtain a plurality of trained preliminary prediction models; optionally, the prediction basic model and the preliminary prediction model include a feature extraction model and a classification model; All training image datasets are determined as a common dataset; Based on a common data set, all preliminary prediction models are jointly trained, and model parameters are optimized based on a common loss function related to the mask strategy parameters to obtain trained candidate feature extraction models and candidate classification models; optionally, the common loss function is a weighted sum of loss function values ​​of each preliminary prediction model; wherein the weighted weight corresponding to the loss function value corresponding to each preliminary prediction model is proportional to the masking degree of the mask strategy parameters corresponding to the training image data set corresponding to the preliminary prediction model; wherein the masking degree is calculated by the following steps: Inputting the mask strategy parameters corresponding to the training image data set corresponding to the preliminary prediction model into the mask vector feature extraction neural network to obtain a third effect feature vector corresponding to the mask strategy parameters; Calculating a first vector distance between the third effect feature vector and a first reference feature vector corresponding to the standard unobstructed image; Calculating a second vector distance between the third effect feature vector and a second reference feature vector corresponding to the standard fully occluded image; optionally, the first reference feature vector and the second reference feature vector are both obtained by pre-analyzing vector extraction results of a preset data set based on a mask vector feature extraction neural network; The ratio of the first vector distance to the second vector distance is calculated to obtain the masking degree corresponding to the preliminary prediction model.

[0064] It can be seen that through the above optional embodiments, multiple training image data sets are generated based on different mask strategy parameters, and all preliminary prediction models are jointly trained through a common data set, and the model parameters are optimized using a common loss function related to the mask strategy parameters to train feature extraction models and classification models that are more suitable for masking situations, thereby assisting in improving the adaptability and accuracy of face recognition in different scenarios, making the recognition results more robust and applicable to diverse environments, and improving the intelligence and automation of face recognition services.

[0065] As an optional embodiment, the recognition module determines the classification recognition result corresponding to the facial image data based on multiple image features and a classification model in a specific manner including: Based on the random combination algorithm, multiple image features are randomly sampled and combined to obtain multiple combined feature sets; Input each combined feature set into the classification model to obtain the output set recognition result; Calculate the intersection of all set recognition results to obtain the classification recognition results corresponding to the face image data.

[0066] It can be seen that through the above optional embodiments, multiple image features are sampled and combined based on a random combination algorithm, and the recognition results of each combined feature set are calculated through a classification model, and then the intersection of all recognition results is taken to reduce feature interference and misclassification, improve the adaptability and accuracy of face recognition in different scenarios, make the recognition results more robust and applicable to diverse environments, and improve the intelligence and automation level of face recognition services.

[0067] Example 3 See also Figure 3 , Figure 3 This is another face recognition system based on multi-model collaboration disclosed in an embodiment of the present invention. Figure 3 The face recognition system based on multi-model collaboration is applied to a data processing system / data processing device / data processing server (wherein the server includes a local processing server or a cloud processing server). Figure 3 As shown, the face recognition system based on multi-model collaboration may include: A memory 301 storing executable program code; a processor 302 coupled to the memory 301; The processor 302 calls the executable program code stored in the memory 301 to execute the steps of the face recognition method based on multi-model collaboration described in the first embodiment.

[0068] Example 4 An embodiment of the present invention discloses a computer-readable storage medium storing a computer program for electronic data exchange, wherein the computer program enables a computer to execute the steps of the face recognition method based on multi-model collaboration described in the first embodiment.

[0069] Example 5 An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute the steps of the face recognition method based on multi-model collaboration described in Example 1.

[0070] The foregoing description of specific embodiments of the present disclosure is intended to illustrate a method for performing a multi-tasking process. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0071] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0072] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0073] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0074] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0075] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0076] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0077] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0078] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0079] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0080] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0081] This specification may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.

[0082] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

[0083] Finally, it should be noted that the face recognition method and system based on multi-model collaboration disclosed in the embodiments of the present invention are only preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments can still be modified, or some of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A face recognition method based on multi-model collaboration, characterized in that: The method comprises: Obtaining facial image data to be recognized and corresponding scene features; Determining a plurality of corresponding feature extraction models and classification models according to the scene features; the plurality of feature extraction models and classification models are trained by training data sets processed by a plurality of different masking strategies corresponding to the scene features; Inputting the facial image data into the multiple feature extraction models to obtain multiple image features corresponding to the facial image; A classification recognition result corresponding to the facial image data is determined based on the multiple image features and the classification model.

2. The face recognition method based on multi-model collaboration according to claim 1, characterized in that: The scene features include at least one of an identification location, ambient light, scene type, user information, and identification purpose.

3. The face recognition method based on multi-model collaboration according to claim 1, characterized in that: Determining a plurality of corresponding feature extraction models and classification models based on the scene features includes: Determining a mask prediction parameter corresponding to the scene feature; For each candidate feature extraction model, determining the mask strategy parameters corresponding to the training data set corresponding to the candidate feature extraction model; Calculating a first similarity between the mask strategy parameter and the mask prediction parameter to obtain a first model priority corresponding to the candidate feature extraction model; Screening out all the candidate feature models whose first model priority is greater than a preset first priority threshold, and obtaining corresponding multiple feature extraction models; At least one classification model is determined according to the mask strategy parameters corresponding to the training data sets corresponding to the multiple feature extraction models.

4. The face recognition method based on multi-model collaboration according to claim 3, characterized in that: The determining of the mask prediction parameter corresponding to the scene feature includes: The scene features are input into a trained mask parameter prediction neural network to obtain mask prediction parameters corresponding to the scene features; the mask parameter prediction neural network is trained by a training data set including multiple training scene images and corresponding scene feature annotations and mask situation annotations.

5. The face recognition method based on multi-model collaboration according to claim 3, characterized in that: The mask strategy parameter and the mask prediction parameter both include at least one of a masked image position, a masked portion size, a masked portion area ratio, and three-dimensional information of the masked portion; and the first similarity is calculated by the following steps: Inputting the mask strategy parameters into a trained mask vector feature extraction neural network to obtain a first effect feature vector corresponding to the mask strategy parameters; Inputting the mask prediction parameter into the mask vector feature extraction neural network to obtain a second effect feature vector corresponding to the mask prediction parameter; the mask vector feature extraction neural network is trained using a training data set including a plurality of training mask parameters and corresponding pre-mask image annotations and post-mask image annotations; The inverse of the vector distance between the first effect feature vector and the second effect feature vector is calculated to obtain the first similarity.

6. The face recognition method based on multi-model collaboration according to claim 5, characterized in that: The determining of at least one classification model based on mask strategy parameters corresponding to the training data sets corresponding to the multiple feature extraction models includes: For each candidate classification model, calculating, based on the mask vector feature extraction neural network and the vector distance algorithm, a second similarity between the mask strategy parameters corresponding to the training data set of the candidate classification model and the mask strategy parameters corresponding to the training data set corresponding to each of the feature extraction models; Calculating an average of the second similarities corresponding to the candidate classification model and each of the feature extraction models to obtain a second model priority corresponding to the candidate classification model; All the candidate classification models whose second model priority is greater than a preset second priority threshold are screened out to obtain at least one corresponding classification model.

7. The face recognition method based on multi-model collaboration according to claim 5, characterized in that: The candidate feature extraction model and the candidate classification model are trained by the following steps: Determining a plurality of training image datasets after masking using different mask strategy parameters; A preset prediction basic model is trained based on each of the training image data sets to obtain a plurality of trained preliminary prediction models; the prediction basic model and the preliminary prediction model include a feature extraction model and a classification model; determining all of the training image data sets as a common data set; Based on the common data set, all the preliminary prediction models are jointly trained, and model parameters are optimized based on a common loss function related to the mask strategy parameters to obtain the trained candidate feature extraction model and the candidate classification model; the common loss function is a weighted sum of the loss function values ​​of each of the preliminary prediction models; wherein the weighted weight corresponding to the loss function value corresponding to each of the preliminary prediction models is proportional to the masking degree of the mask strategy parameters corresponding to the training image data set corresponding to the preliminary prediction model; wherein the masking degree is calculated by the following steps: Inputting the mask strategy parameters corresponding to the training image data set corresponding to the preliminary prediction model into the mask vector feature extraction neural network to obtain a third effect feature vector corresponding to the mask strategy parameters; Calculating a first vector distance between the third effect feature vector and a first reference feature vector corresponding to a standard unobstructed image; Calculating a second vector distance between the third effect feature vector and a second reference feature vector corresponding to the standard fully occluded image; wherein the first reference feature vector and the second reference feature vector are both obtained by pre-analyzing vector extraction results of a preset data set based on the occlusion vector feature extraction neural network; The ratio of the first vector distance to the second vector distance is calculated to obtain the masking degree corresponding to the preliminary prediction model.

8. The face recognition method based on multi-model collaboration according to claim 1, characterized in that: Determining the classification recognition result corresponding to the facial image data based on the multiple image features and the classification model includes: Based on a random combination algorithm, the plurality of image features are randomly sampled and combined to obtain a plurality of combined feature sets; Inputting each of the combined feature sets into the classification model to obtain an output set recognition result; The intersection of all the set recognition results is calculated to obtain the classification recognition result corresponding to the face image data.

9. A face recognition system based on multi-model collaboration, characterized in that: The system comprises: An acquisition module is used to obtain the face image data to be recognized and the corresponding scene features; A determination module, configured to determine a plurality of corresponding feature extraction models and classification models based on the scene features; the plurality of feature extraction models and classification models are trained using a training data set processed using a plurality of different masking strategies corresponding to the scene features; an extraction module, configured to input the facial image data into the plurality of feature extraction models to obtain a plurality of image features corresponding to the facial image; The recognition module is used to determine the classification recognition result corresponding to the facial image data based on the multiple image features and the classification model.

10. A face recognition system based on multi-model collaboration, characterized in that: The system comprises: a memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the face recognition method based on multi-model collaboration as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Multi-scene adaptive model fusion method and face recognition system

    CN113361488A

  • Facial expression recognition method in sheltered scene based on collaborative feature completion

    CN114821714A

  • Face recognition method and device based on rapid mask generation

    CN116665279A

  • Target attribute recognition method and apparatus, and model training method and apparatus

    WO2023246921A1

Cited By

  • Face recognition method and system based on cloud machine cooperation

    CN121527601A