Multi-modal face recognition method and system based on deep learning

Through a multimodal face recognition method based on deep learning, combined with sensor data of different modalities, feature extraction and fusion are performed, and the problem of single-modal face recognition being susceptible to the environment is solved, achieving higher recognition accuracy and robustness.

CN119942623AInactive Publication Date: 2025-05-06BEIJING ZHONGSHITONG TECH CO LTD

Patent Information

Application Number
CN202510436022.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing facial recognition technology has limitations in both single-modal and multi-modal fields. Single-modal facial recognition is susceptible to environmental and conditions, while multi-modal facial recognition still faces challenges in data integration and processing.

Method used

A multimodal face recognition method based on deep learning is adopted to collect data in real time through sensors of different modes, and perform denoising, graying, cropping and alignment processing. RGB image features are extracted using convolutional neural networks, and infrared image heat map features are extracted by CNN structures, and different modal data are fused at the network input layer, and the features are stitched or weighted summed and merged, and then classified through the fully connected layer.

Benefits of technology

It improves the accuracy and stability of face recognition, enhances the robustness and reliability of the recognition system, and effectively solves the problem of limited generalization ability of single-modal data in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942623A_ABST
    Figure CN119942623A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of face recognition, and particularly relates to a multi-modal face recognition method and system based on deep learning, and the method comprises the steps: carrying out the real-time collection of a target through employing sensors of different modals, and carrying out the denoising, graying, image cutting and alignment processing of a collected image; extracting features in the RGB image by using a convolutional neural network, and extracting heat map features from the infrared image by using a CNN structure; fusing the data of different modalities at an input layer of the network, merging the data into a multi-modal input for feature extraction, after features are independently extracted from each modal, splicing or performing weighted summation to merge the modalities, and classifying the modalities through a full connection layer; a deep neural network is used for training, deep semantic information is learned from the fused multi-modal features, a proper loss function is selected according to a task target, and fine tuning optimization is carried out according to a data set; the full connection layer or the SVM is adopted for final category prediction, a face recognition result is output, and the effect of efficiently recognizing the face is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of face recognition, and in particular to a multimodal face recognition method and system based on deep learning. Background Art

[0002] With the development of modern science and technology, face recognition technology plays an increasingly important role in the security field. Traditional single-modal face recognition technology usually relies on a single biometric feature (such as facial contour, eye iris, etc.) for identity recognition, which is limited to a certain extent by factors such as lighting, angle, expression, etc., and the recognition accuracy is limited. In recent years, the rise of deep learning technology has brought new breakthroughs in face recognition technology. Deep learning can effectively extract and analyze complex features in images, greatly improving the accuracy and stability of face recognition, but the generalization ability of single-modal data in complex scenarios is limited.

[0003] In order to solve the limitations of single-modal data, multimodal face recognition technology came into being. Multimodal face recognition can perform identity recognition more comprehensively and accurately by combining multiple biometric features such as face, voice, and behavior, thereby significantly improving the robustness and reliability of the recognition system. However, traditional multimodal face recognition methods still have certain challenges, such as how to effectively integrate data from different modalities, how to deal with inconsistencies between data, and how to improve recognition performance while ensuring efficiency. These problems limit the widespread application and development of multimodal face recognition technology. Therefore, it is crucial to design an efficient multimodal face recognition method.

[0004] In summary, although the current face recognition technology has made progress in both single-modality and multi-modality fields, both have certain limitations. Single-modality face recognition is easily affected by the environment and conditions, and although multi-modality face recognition can provide more recognition information, it still faces challenges in data integration and processing. Therefore, it is of great scientific significance and practical value to study a face recognition method that combines deep learning and can effectively integrate and analyze multi-modal data.

[0005] In response to the above-mentioned technical defects, a multimodal face recognition method and system solution based on deep learning is proposed. Summary of the invention

[0006] To solve the above problems, the present invention provides the following technical solutions: Multimodal face recognition method based on deep learning, including: Use sensors of different modalities to collect targets in real time, and perform denoising, grayscale, image cropping and alignment on the collected images; Use convolutional neural networks to extract features from RGB images, and use CNN structures to extract heat map features from infrared images; The data of different modalities are fused at the input layer of the network and merged into a multimodal input for feature extraction. After extracting features independently from each modality, they are merged by concatenation or weighted summation and then classified through the fully connected layer; Use deep neural networks for training, learn deeper semantic information from the fused multimodal features, select appropriate loss functions based on the task objectives, and perform fine-tuning and optimization based on the dataset; Use the fully connected layer or SVM to make the final category prediction and output the face recognition result.

[0007] Furthermore, the real-time acquisition uses an RGB camera and an infrared sensor to synchronously capture data to ensure time and space alignment and avoid mismatches between modalities due to device delays or motion blur; non-local mean denoising is used for RGB images, and adaptive median filtering is used to process thermal noise for infrared images; RGB images are grayscaled to reduce redundant information; MTCNN is used to detect facial key points, and after cropping, they are aligned to a standard size through affine transformation to ensure spatial consistency between modalities.

[0008] Furthermore, the use of a convolutional neural network to extract features from an RGB image includes adjusting the input RGB image to a fixed size required by the network, normalizing the image, normalizing the pixel values ​​to between 0 and 1, or subtracting the mean and dividing by the standard deviation to speed up training and improve the stability of the model. During the training process, data enhancement techniques are applied to increase the diversity of data and improve the generalization ability of the model. A 3x3 or 5x5 filter is used to perform a convolution operation on the input image to extract low-level features such as edge, texture, and color information. An activation function is applied after the convolution layer to introduce nonlinearity and enhance the expression ability of the network. By stacking multiple convolution layers, more advanced and abstract features are gradually extracted.

[0009] Furthermore, the use of convolutional neural networks to extract features from RGB images also includes applying a maximum pooling operation after the convolution layer to reduce the spatial dimension of the feature map, reduce the number of parameters, and improve the model's ability to resist overfitting. The average value within the pooling window is taken, which is also used to reduce the dimension. A 2x2 window and a stride of 2 are used to effectively reduce the size of the feature map. In deep networks, residual connections are used to help alleviate the gradient vanishing problem and improve the ability to train deep networks. Multi-scale features are extracted through filters of different sizes or convolution operations of different levels to adapt to targets of different sizes. An attention mechanism is introduced to allow the network to automatically focus on important areas in the image and enhance the ability to discriminate features.

[0010] Furthermore, fusing the data of different modalities at the input layer of the network includes collecting data of different modalities, such as RGB images and infrared images, adjusting the data of all modalities to the same size to ensure spatial alignment during fusion, and normalizing the data of each modality to have the same scale and distribution; The data of different modes are fused and calculated as follows: , in, is the concatenated input, For RGB images, For infrared images, is the splicing function; Data fusion is performed through weighted summation, and the calculation is as follows: , in, is the concatenated input, is an RGB image, For infrared images, , is the learning parameter, and + =1; Ensure that the input layer of the network can accept the fused multimodal data. Use the splicing method, the number of input channels is 4, use the convolutional neural network to extract the features of the fused data, select the appropriate loss function according to the task objectives, and optimize the weight parameters in the fusion operation through the back propagation algorithm during the training process; The loss function is calculated as follows: , in, is the cross entropy loss, is the true label, For output, is the target data; Evaluate the performance of the model on the validation set, adjust the hyperparameters to optimize the model, and deploy the trained model to practical applications to perform classification or recognition tasks of multimodal data.

[0011] Furthermore, the use of deep neural networks for training increases data diversity through rotation, flipping, and scaling techniques to prevent overfitting, standardizes data to the same scale, normalizes pixel values ​​to between 0 and 1, divides data into training sets, validation sets, and test sets, selects models based on data types and task requirements, sets the number of hidden layers, the number of neurons in each layer, and activation functions, uses a data loader to load data, divides data into small batches, and inputs each batch into the model for training to improve training efficiency, iterates training data, calculates losses, and updates parameters, and after each round of training, uses a validation set to evaluate model performance to prevent overfitting, loads the model in actual applications, performs predictions, and integrates the model into actual applications, such as image classification and speech recognition.

[0012] According to one aspect of the present invention, a multimodal face recognition system based on deep learning is provided, comprising: A data acquisition module, used to collect multi-modal facial data, including visible light images, infrared images, depth images, video streams or biometric data; A data preprocessing module is used to preprocess the collected multimodal data, including denoising, alignment, standardization or data enhancement; A feature extraction module, used to extract deep features from the preprocessed multimodal data, wherein the feature extraction is based on a deep learning model, including a convolutional neural network, a Transformer or a graph neural network; A multimodal data fusion module, used to fuse features of different modalities, wherein the fusion method includes attention mechanism, weighted fusion or feature alignment; A classification and recognition module, used to classify the fused features and perform face recognition, wherein the classification includes 1:N recognition, 1:1 verification or liveness detection based on face features; Data storage and management module, used to store and manage facial feature data, recognition results and related information; The result output module is used to output the recognition results, including but not limited to identity information, matching probability or alarm information.

[0013] Furthermore, the classification and recognition module includes further processing of the fused multimodal features, extracting more discriminative feature vectors to meet the needs of the classification model, standardizing the extracted features to ensure the comparability of features between different samples and improve the performance of the classification model, and using the trained classification model to classify the feature vectors. The model must be fully trained to ensure the generalization ability on multimodal data, calculate the probability of each sample belonging to each category, and provide a basis for subsequent decision-making. According to the classification results and probability values, make the final recognition decision, such as determining identity or verification, and output the recognition results, including identity information and matching probability, for use by users or systems.

[0014] According to one aspect of the present invention, there is provided a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned multimodal face recognition method based on deep learning when executing the computer program.

[0015] According to one aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of the multimodal face recognition method based on deep learning described above are implemented.

[0016] Compared with the prior art, the present invention has the following beneficial effects: 1. In the multimodal face recognition method based on deep learning of the present invention, the target is collected in real time by using sensors of different modalities, and the collected images are denoised, grayed, cropped and aligned; the features in the RGB image are extracted by using a convolutional neural network, and the heat map features are extracted from the infrared image by using a CNN structure; the data of different modalities are fused in the input layer of the network and merged into a multimodal input for feature extraction, and after each modality extracts features independently, they are merged by splicing or weighted summing, and then classified through a fully connected layer; a deep neural network is used for training, and deeper semantic information is learned from the fused multimodal features, a suitable loss function is selected according to the task objective, and fine-tuning and optimization are performed according to the data set; a fully connected layer or SVM is used for final category prediction, and the face recognition result is output, which has the effect of efficiently recognizing faces.

[0017] 2. In the multimodal face recognition method and system based on deep learning of the present invention, a data acquisition module is used to collect multimodal face data, including visible light images, infrared images, depth images, video streams or biometric data; a data preprocessing module is used to preprocess the collected multimodal data, including denoising, alignment, standardization or data enhancement; a feature extraction module is used to extract deep features from the preprocessed multimodal data, and the feature extraction is based on a deep learning model, including a convolutional neural network, a Transformer or a graph neural network; a multimodal data fusion module is used to fuse features of different modalities, and the fusion method includes an attention mechanism, weighted fusion or feature alignment; a classification and recognition module is used to classify the fused features and recognize faces, and the classification includes 1:N recognition based on face features, 1:1 verification or liveness detection; a data storage and management module is used to store and manage face feature data, recognition results and related information; a result output module is used to output recognition results, including but not limited to identity information, matching probability or alarm information, and has the effect of efficient and intelligent face recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to facilitate understanding by those skilled in the art, the present invention is further described below in conjunction with the accompanying drawings; Figure 1 Schematic diagram of the overall framework of the multimodal face recognition method based on deep learning of the present invention; Figure 2 It is a module schematic diagram of a multimodal face recognition system based on deep learning of the present invention; Figure 3 It is a schematic diagram of a computer device in the multimodal face recognition method based on deep learning of the present invention. DETAILED DESCRIPTION

[0019] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0020] like Figure 1-Figure 3 As shown, the present application provides a multimodal face recognition method based on deep learning, including: S1: Use sensors of different modalities to collect the target in real time, and perform denoising, grayscale, image cropping and alignment on the collected images; S2: Use convolutional neural networks to extract features from RGB images and use CNN structures to extract heat map features from infrared images; S3: The data of different modalities are fused at the input layer of the network and merged into a multimodal input for feature extraction. After extracting features independently from each modality, they are merged by concatenation or weighted summation, and then classified through the fully connected layer; S4: Use deep neural networks for training, learn deeper semantic information from the fused multimodal features, select appropriate loss functions according to the task objectives, and perform fine-tuning and optimization based on the dataset; S5: Use the fully connected layer or SVM to make the final category prediction and output the face recognition result.

[0021] Specifically, the real-time acquisition uses an RGB camera and an infrared sensor to synchronously capture data to ensure time and space alignment and avoid mismatch between modalities due to device delays or motion blur; non-local mean denoising is used for RGB images, and adaptive median filtering is used to process thermal noise for infrared images; RGB images are grayscaled to reduce redundant information; MTCNN is used to detect facial key points, and after cropping, they are aligned to a standard size through affine transformation to ensure spatial consistency between modalities.

[0022] Specifically, the use of a convolutional neural network to extract features from an RGB image includes adjusting the input RGB image to a fixed size required by the network, normalizing the image, normalizing the pixel values ​​to between 0 and 1, or subtracting the mean and dividing by the standard deviation to speed up training and improve the stability of the model. During the training process, data enhancement techniques are applied to increase the diversity of data and improve the generalization ability of the model. A 3x3 or 5x5 filter is used to perform a convolution operation on the input image to extract low-level features such as edge, texture, and color information. An activation function is applied after the convolution layer to introduce nonlinearity and enhance the expression ability of the network. By stacking multiple convolution layers, more advanced and abstract features are gradually extracted.

[0023] Specifically, the use of convolutional neural networks to extract features from RGB images also includes applying a maximum pooling operation after the convolution layer to reduce the spatial dimension of the feature map, reduce the number of parameters, and improve the model's ability to resist overfitting. The average value within the pooling window is taken, which is also used to reduce the dimension. A 2x2 window and a stride of 2 are used to effectively reduce the size of the feature map. In deep networks, residual connections are used to help alleviate the gradient vanishing problem and improve the ability to train deep networks. Multi-scale features are extracted through filters of different sizes or convolution operations of different levels to adapt to targets of different sizes. An attention mechanism is introduced to allow the network to automatically focus on important areas in the image and enhance the ability to discriminate features.

[0024] Specifically, fusing the data of different modalities at the input layer of the network includes collecting data of different modalities, such as RGB images and infrared images, adjusting the data of all modalities to the same size to ensure spatial alignment during fusion, and normalizing the data of each modality to have the same scale and distribution; The data of different modes are fused and calculated as follows: , in, is the concatenated input, is an RGB image, For infrared images, is the splicing function; Data fusion is performed through weighted summation, and the calculation is as follows: , in, is the concatenated input, For RGB images, For infrared images, , is the learning parameter, and + =1; Ensure that the input layer of the network can accept the fused multimodal data. Use the splicing method, the number of input channels is 4, use the convolutional neural network to extract the features of the fused data, select the appropriate loss function according to the task objectives, and optimize the weight parameters in the fusion operation through the back propagation algorithm during the training process; The loss function is calculated as follows: , in, is the cross entropy loss, is the true label, For output, is the target data; Evaluate the performance of the model on the validation set, adjust the hyperparameters to optimize the model, and deploy the trained model to practical applications to perform classification or recognition tasks of multimodal data.

[0025] In one embodiment, the specific operations are as follows: #Data preprocessing defpreprocess_data(rgb_image,infrared_image): # Adjust the size rgb_resized=resize(rgb_image,target_size) infrared_resized=resize(infrared_image,target_size) #Normalization rgb_normalized=(rgb_resized-rgb_mean) / rgb_std infrared_normalized=(infrared_resized-infrared_mean) / infrared_std returnrgb_normalized,infrared_normalized # Modal Fusion deffuse_modalities(rgb,infrared,w1,w2): # Ensure consistent sizes if rgb.shape != infrared.shape : raiseValueError("Modal data size is inconsistent") #Weighted sum fused=w1*rgb+w2*infrared return fused # Network structure class MultiModalNet(nn.Module): def __init__(self, input_channels): super(MultiModalNet, self).__init__() self.conv1 = nn.Conv2d(input_channels, 32, kernel_size = 3, padding = 1) self.relu = nn.ReLU() self.fc = nn.Linear(32 * H * W, num_classes) def forward(self, x): x = self.relu(self.conv1(x)) x = x.view(x.size(0), -1) x = self.fc(x) return x # Training process def train_model(train_loader, model, criterion, optimizer, epochs): for epoch in range(epochs): for rgb, infrared, labels in train_loader: # Preprocessing rgb_normalized, infrared_normalized = preprocess_data(rgb, infrared) # Fusion fused_input = fuse_modalities(rgb_normalized, infrared_normalized, w1, w2) # Forward propagation outputs = model(fused_input) loss = criterion(outputs, labels) # Backward propagation optimizer.zero_grad() loss.backward() optimizer.step() The data of different modalities can be fused at the input layer of the network through methods such as concatenation or weighted summation. The key steps include data preprocessing, fusion operation, network design and parameter optimization. By properly selecting fusion methods and optimizing parameters, the performance of the model can be effectively improved, especially in the classification and recognition tasks of multimodal data.

[0026] Specifically, the use of deep neural networks for training increases data diversity through rotation, flipping, and scaling techniques to prevent overfitting, standardizes data to the same scale, normalizes pixel values ​​to between 0 and 1, divides data into training sets, validation sets, and test sets, selects models based on data types and task requirements, sets the number of hidden layers, the number of neurons in each layer, and activation functions, uses a data loader to load data, divides data into small batches, and inputs each batch into the model for training to improve training efficiency, iterates training data, calculates losses, and updates parameters, and after each round of training, uses a validation set to evaluate model performance to prevent overfitting, loads the model in actual applications, performs predictions, and integrates the model into actual applications, such as image classification and speech recognition.

[0027] According to one aspect of the present invention, a multimodal face recognition system based on deep learning is provided, comprising: A data acquisition module, used to collect multi-modal facial data, including visible light images, infrared images, depth images, video streams or biometric data; A data preprocessing module is used to preprocess the collected multimodal data, including denoising, alignment, standardization or data enhancement; A feature extraction module, used to extract deep features from the preprocessed multimodal data, wherein the feature extraction is based on a deep learning model, including a convolutional neural network, a Transformer or a graph neural network; A multimodal data fusion module, used to fuse features of different modalities, wherein the fusion method includes attention mechanism, weighted fusion or feature alignment; A classification and recognition module, used to classify the fused features and perform face recognition, wherein the classification includes 1:N recognition, 1:1 verification or liveness detection based on face features; Data storage and management module, used to store and manage facial feature data, recognition results and related information; The result output module is used to output the recognition results, including but not limited to identity information, matching probability or alarm information.

[0028] Specifically, the classification and recognition module includes specific processing of the fused multimodal features, extracting more discriminative feature vectors to meet the needs of the classification model, standardizing the extracted features to ensure the comparability of features between different samples and improve the performance of the classification model, and using the trained classification model to classify the feature vectors. The model must be fully trained to ensure the generalization ability on multimodal data, calculate the probability of each sample belonging to each category, and provide a basis for subsequent decision-making. According to the classification results and probability values, make the final recognition decision, such as determining identity or verification, and output the recognition results, including identity information and matching probability, for use by users or systems.

[0029] In one embodiment, the specific operations are: Configure a visible light camera, an infrared camera and a depth sensor to collect visible light images, infrared images and depth images of the user's face respectively.

[0030] The collected multimodal data are transmitted to the system processing unit via USB or network interface.

[0031] Data preprocessing module: The visible light image is denoised and the median filter algorithm is used to remove image noise.

[0032] Enhance the contrast of infrared images to improve image quality.

[0033] Smooth the depth image to remove depth noise.

[0034] Feature extraction module: Use a pre-trained convolutional neural network (such as ResNet-50) to extract features from visible light images and obtain a 2048-dimensional feature vector.

[0035] A similar network architecture is used to extract features from infrared images and depth images respectively.

[0036] The feature vector of each modality is L2 regularized to ensure that the features are comparable.

[0037] Multimodal data fusion module: The attention mechanism is used to perform weighted fusion on the feature vectors of the three modalities. The specific steps are as follows: a. Calculate the attention weights of each modality feature vector, and the weights are generated by a fully connected layer.

[0038] b. Perform weighted summation on the feature vectors of each modality according to the attention weight to obtain the final fused feature vector.

[0039] The dimension of the fused feature vector is 2048, which retains the complementary information of each modality.

[0040] Classification and recognition module: The fused feature vector is input into the classifier, which is implemented based on support vector machine (SVM) or softmax classifier.

[0041] The classifier classifies the fused features and outputs a confidence score for the user identity.

[0042] Based on the confidence score, the system decides whether to release the user and records the recognition result.

[0043] Data storage and management module: The user's feature vector and recognition results are stored in the local database, and AES encryption technology is used to protect data security.

[0044] The database supports add, delete, modify and query operations to facilitate user management.

[0045] Result output module: The recognition result is output through the display or buzzer, such as "Authentication successful" or "Authentication failed".

[0046] At the same time, the system sends the recognition results to the cloud server for further analysis and recording.

[0047] Example 2 In one embodiment, the application of face recognition based on multimodal data in a payment system includes: Data acquisition module: Use the front camera of the mobile phone to collect visible light images and infrared images of the user's face.

[0048] Use ToF sensor to obtain depth image.

[0049] Data preprocessing module: Perform face recognition alignment on visible light images to ensure the accuracy of facial key points.

[0050] Perform contrast adjustment on infrared images to enhance facial features.

[0051] Perform noise reduction on the depth image to improve the accuracy of the depth information.

[0052] Feature extraction module: Use lightweight deep learning models such as MobileNetV2 to extract feature vectors of visible light, infrared, and depth images.

[0053] The feature vectors are normalized to ensure consistency.

[0054] Multimodal data fusion module: A Transformer-based fusion method is used to capture the long-range dependencies between different modalities.

[0055] The dimension of the fused feature vector is 1280, which has strong discrimination.

[0056] Classification and recognition module: Use a trained classification model (such as XGBoost or neural network) to classify the fused features.

[0057] The classification results are used to verify user identity and ensure payment security.

[0058] Data storage and management module: User feature vectors are stored in a cloud database and end-to-end encryption technology is used to protect privacy.

[0059] The database supports high concurrent access to ensure stable operation of the system.

[0060] Result output module: The recognition result is displayed on the mobile phone screen, such as "Authentication successful, transaction completed".

[0061] At the same time, the system generates a transaction record and sends it to the user's email or mobile phone text message.

[0062] The multimodal face recognition method and system based on deep learning of the present invention collects the target in real time by using sensors of different modes, and performs denoising, grayscale, image cropping and alignment processing on the collected images; uses convolutional neural network to extract features in RGB images, and uses CNN structure to extract heat map features from infrared images; fuses data of different modes at the input layer of the network and merges them into a multimodal input for feature extraction, and after each mode extracts features independently, merges them by splicing or weighted summation, and then classifies them through the fully connected layer; uses deep neural network for training, learns deeper semantic information from the fused multimodal features, selects appropriate loss function according to the task objective, and performs fine-tuning and optimization according to the data set; uses fully connected layer or SVM for final category prediction, outputs face recognition results, and has the effect of efficiently recognizing faces; uses data acquisition module to collect multimodal face data, including visible light images, infrared images, etc. External images, depth images, video streams or biometric data, data preprocessing module, used to preprocess the collected multimodal data, including denoising, alignment, standardization or data enhancement; feature extraction module, used to extract deep features from the preprocessed multimodal data, the feature extraction is based on a deep learning model, including a convolutional neural network, a Transformer or a graph neural network; a multimodal data fusion module, used to fuse the features of different modalities, the fusion method includes an attention mechanism, weighted fusion or feature alignment; a classification and recognition module, used to classify the fused features and perform face recognition, the classification includes 1:N recognition based on facial features, 1:1 verification or liveness detection; a data storage and management module, used to store and manage facial feature data, recognition results and related information; a result output module, used to output recognition results, including but not limited to identity information, matching probability or alarm information, with the effect of efficient and intelligent face recognition.

[0063] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned multimodal face recognition method based on deep learning when executing the computer program.

[0064] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the above-mentioned multimodal face recognition method based on deep learning are implemented.

[0065] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided in this application and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0066] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "includes a ..." does not exclude the presence of other identical elements in the process, device, article or method including the element.

[0067] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only specific implementation methods. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A multimodal face recognition method based on deep learning, characterized in that: include: Use sensors of different modalities to collect targets in real time, and perform denoising, grayscale, image cropping and alignment on the collected images; Use convolutional neural networks to extract features from RGB images, and use CNN structures to extract heat map features from infrared images; The data of different modalities are fused at the input layer of the network and merged into a multimodal input for feature extraction. After extracting features independently from each modality, they are merged by concatenation or weighted summation and then classified through the fully connected layer; Use deep neural networks for training, learn deeper semantic information from the fused multimodal features, select appropriate loss functions based on the task objectives, and perform fine-tuning and optimization based on the dataset; Use the fully connected layer or SVM to make the final category prediction and output the face recognition result.

2. The multimodal face recognition method based on deep learning according to claim 1, characterized in that: The real-time acquisition uses an RGB camera and an infrared sensor to synchronously capture data to ensure temporal and spatial alignment and avoid mismatches between modalities due to device delays or motion blur; non-local mean denoising is used for RGB images, and adaptive median filtering is used for infrared images to process thermal noise; Grayscale the RGB image to reduce redundant information; use MTCNN to detect facial key points, and align it to a standard size through affine transformation after cropping to ensure spatial consistency between modalities.

3. The multimodal face recognition method based on deep learning according to claim 1, characterized in that: The method of extracting features from RGB images using a convolutional neural network includes adjusting the input RGB image to a fixed size required by the network, normalizing the image, normalizing the pixel values ​​to between 0 and 1, or subtracting the mean and dividing by the standard deviation to speed up training and improve the stability of the model. During the training process, data enhancement technology is applied to increase the diversity of data and improve the generalization ability of the model. A 3x3 or 5x5 filter is used to perform a convolution operation on the input image to extract low-level features, including edge, texture and color information. An activation function is applied after the convolution layer to introduce nonlinearity and enhance the expression ability of the network. By stacking multiple convolution layers, high-level and abstract features are gradually extracted.

4. The multimodal face recognition method based on deep learning according to claim 3, characterized in that: The use of convolutional neural networks to extract features from RGB images also includes applying a maximum pooling operation after the convolution layer to reduce the spatial dimension of the feature map, reduce the number of parameters, and improve the model's ability to resist overfitting. The average value within the pooling window is taken, which is also used to reduce the dimension. A 2x2 window and a stride of 2 are used to effectively reduce the size of the feature map. In deep networks, residual connections are used to help alleviate the gradient vanishing problem and improve the ability to train deep networks. Multi-scale features are extracted through filters of different sizes or convolution operations of different levels to adapt to targets of different sizes. An attention mechanism is introduced to allow the network to automatically focus on important areas in the image and enhance the ability to discriminate features.

5. The multimodal face recognition method based on deep learning according to claim 1, characterized in that: The fusing of data of different modalities at the input layer of the network includes collecting data of different modalities, adjusting the data of all modalities to the same size through RGB images and infrared images, ensuring spatial alignment during fusion, and normalizing the data of each modality to have the same scale and distribution; The data of different modes are fused and calculated as follows: , in, is the concatenated input, For RGB images, For infrared images, is the splicing function; Data fusion is performed through weighted summation, and the calculation is as follows: , in, is the concatenated input, For RGB images, For infrared images, , is the learning parameter, and + =1; Ensure that the network input layer accepts the fused multimodal data, use the splicing method, the number of input channels is 4, use the convolutional neural network to extract the features of the fused data, select the loss function according to the task goal, and optimize the weight parameters in the fusion operation through the back propagation algorithm during the training process; The loss function is calculated as follows: , in, is the cross entropy loss, is the true label, For output, is the target data; Evaluate the performance of the model on the validation set, adjust the hyperparameters to optimize the model, and deploy the trained model to practical applications to perform classification or recognition tasks of multimodal data.

6. The multimodal face recognition method based on deep learning according to claim 5, characterized in that: The deep neural network is used for training to increase data diversity and prevent overfitting through rotation, flipping, and scaling techniques, standardize the data to the same scale, normalize the pixel values ​​to between 0 and 1, divide the data into training sets, validation sets, and test sets, select a model based on the data type and task requirements, set the number of hidden layers, the number of neurons in each layer, and the activation function, use a data loader to load data, divide the data into small batches, and input each batch into the model for training to improve training efficiency, iterate the training data, calculate the loss and update the parameters, use the validation set to evaluate the model performance after each round of training to prevent overfitting, load the model for prediction in actual applications, and integrate the model into actual applications, including image classification and speech recognition.

7. A multimodal face recognition system based on deep learning, characterized in that: include: A data acquisition module, used to collect multi-modal facial data, including visible light images, infrared images, depth images, video streams or biometric data; A data preprocessing module is used to preprocess the collected multimodal data, including denoising, alignment, standardization or data enhancement; Feature extraction module, used to extract deep features from preprocessed multimodal data. Feature extraction is based on deep learning models, including convolutional neural networks, Transformer or graph neural networks; Multimodal data fusion module, used to fuse features of different modalities. Fusion methods include attention mechanism, weighted fusion or feature alignment; The classification and recognition module is used to classify the fused features and perform face recognition. The classification includes 1:N recognition, 1:1 verification or liveness detection based on face features. Data storage and management module, used to store and manage facial feature data, recognition results and related information; The result output module is used to output the recognition results, including but not limited to identity information, matching probability or alarm information.

8. The multimodal face recognition system based on deep learning according to claim 7, characterized in that: The classification and recognition module includes further processing of the fused multimodal features, extracting feature vectors of discrimination to meet the needs of the classification model, standardizing the extracted features, ensuring the comparison of features between different samples, and improving the performance of the classification model. The feature vectors are classified using a trained classification model. The model needs to be fully trained to ensure the generalization ability on multimodal data, and calculate the probability of each sample belonging to each category to provide a basis for subsequent decision-making. According to the classification results and probability values, the final recognition decision is made to determine the identity or perform verification, and the recognition results, including identity information and matching probability, are output for use by users or systems.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the multimodal face recognition method based on deep learning described in any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multimodal face recognition method based on deep learning described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Face recognition method and system based on adaptive score fusion and deep learning

    CN106709477A

  • Multi-modal human face recognition method based on deep learning

    CN106909905A

  • Access control system based on face recognition and recognition method

    CN119672782A

Cited By

  • Multi-modal fusion protection device pressing plate state identification method and electronic equipment

    CN120198741A

  • Target multi-dimensional detection method based on deep learning multi-modal fusion technology

    CN120339645A

  • A multi-dimensional target detection method based on deep learning multimodal fusion technology

    CN120339645B

  • Face recognition method and device based on industrial internet of things, medium and electronic equipment

    CN120932279A

  • Target detection method and system combining router, gateway and camera

    CN121030429A