Fusion algorithm for cross-modal face recognition

By adopting a fusion algorithm for cross-modal face recognition in the freight operation industry, the problem of insufficient accuracy of cross-modal face alignment in the existing technology is solved, and high-precision authentication and system robustness are achieved.

CN120126192AInactive Publication Date: 2025-06-10SUZHOU RONGSHENG TECHNOLOGY INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510075526.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology is insufficient in the accuracy of face comparison across scenes, lighting, and color modes, and it is difficult to effectively apply in intelligent supervision of the freight operation industry.

Method used

The fusion algorithm of cross-modal face recognition is adopted, including data preprocessing, MTCNN face detection and cropping, FaceNet feature extraction, cosine similarity calculation and triple loss function training model to achieve high-precision cross-modal face recognition.

Benefits of technology

It achieves more than 95% authentication accuracy, can complete the processing and response of a single request in 1 second, and maintains stable operation in 8G memory environment, has exception handling capabilities, and improves the robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126192A_ABST
    Figure CN120126192A_ABST
Patent Text Reader

Abstract

The invention discloses a fusion algorithm for cross-modal face recognition, which comprises data description, a data set comprises data from 142 subjects, a total of 2556 pairs of image samples, each pair of samples is composed of a black-and-white image collected by a cab and a color image of an identification photo, the number of the images of two modals is 50%, and in order to ensure the practical applicability of a test set, the fusion algorithm is provided with a fusion algorithm for cross-modal face recognition. 20% of the image pairs of each subject are randomly selected as a test set to ensure balanced distribution of different illumination conditions, shooting angles and expression changes, in addition, all the images are manually screened, samples which are seriously blurred, shielded or insufficient in illumination are removed, and the data quality is ensured; the modeling step comprises preprocessing, face area detection and face standardization, face vector representation, distance measurement, learning rate scheduler construction, loss function and model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fusion algorithms, and more specifically, to a fusion algorithm for cross-modal face recognition. Background Art

[0002] The verification of the identity of truck drivers is an important part of transportation safety management. Its characteristic is that it is necessary to compare the black-and-white photos taken in real time in the cab with the pre-archived color ID photos to confirm the authenticity of the driver's identity. The black-and-white photos are collected by an infrared light sensor. At present, conventional face recognition technologies often have problems with insufficient accuracy when dealing with face comparisons across different scenarios, illuminations, and color modalities, which poses challenges to the intelligent supervision of the freight industry. Therefore, it is of great practical significance to develop a technology that can accurately process face images of different modalities and maintain a high recognition accuracy.

[0003] Traditional face recognition models (such as basic CNN, single feature extraction networks, etc.) perform poorly when dealing with face images with large modal differences. Moreover, due to their limited feature extraction ability, sensitivity to illumination changes, and insufficient generalization performance, these models are not suitable for cross-modal face comparison in actual freight scenarios. Therefore, it is particularly important to explore and design a deep learning model that can effectively process face images of different modalities and achieve accurate identity matching. Summary of the Invention

[0004] (I) Technical Problems to be Solved

[0005] In view of the problems existing in the prior art, the present invention provides a fusion algorithm for cross-modal face recognition to solve the technical problems mentioned in the background art.

[0006] (II) Technical Solutions

[0007] To achieve the above object, the present invention provides the following technical solution: A fusion algorithm for cross-modal face recognition, comprising the following steps:

[0008] Step 1: Data Description

[0009] The dataset contains data from 142 subjects, totaling 2,556 pairs of image samples. Each pair of samples consists of a black-and-white image collected from the cab and a color ID photo. The number of images in both modalities accounts for 50% each. The original images are all in 3-channel RGB format and have a unified size of 464×348 pixels. The labels of the training data contain two dimensions: modality identification (0 for black-and-white images, 1 for color images) and subject ID (an integer from 1 to 142). To ensure the practical applicability of the test set, 20% of the image pairs from each subject are randomly selected as the test set to ensure an even distribution of different lighting conditions, shooting angles, and expression changes. In addition, all images are manually screened to remove samples with severe blurring, occlusion, or insufficient lighting to ensure data quality;

[0010] Step 2: Modeling steps: preprocessing, face region detection and face normalization, face vector representation, distance metric, constructing a learning rate scheduler, loss function, and model training.

[0011] The present invention is further configured such that the preprocessing in the second step includes the following steps:

[0012] Data loading

[0013] Through a custom data loader, the image files are read into the Pytorch tensor format. The loaded data has a dimension of 3×464×368, where the first dimension represents the number of RGB channels, and the last two dimensions represent the height and width of the image respectively. The initial data type is uint8, and the pixel values are distributed in the range of 0-255;

[0014] Data augmentation

[0015] To improve the generalization ability of the model and its adaptability to various scenarios, diverse augmentation processing is performed on the training data. According to the characteristics of the truck driving scenario, the following augmentation strategies are mainly adopted:

[0016] Random brightness adjustment: Randomly adjust the image brightness within the range of [-0.2, 0.2] to simulate the lighting changes at different times;

[0017] Random contrast adjustment: Randomly adjust the image contrast within the range of [0.8, 1.2] to enhance the model's adaptability to images of different qualities;

[0018] Random Gaussian noise: Add Gaussian noise with a mean of 0 and a standard deviation within the range of [0.01, 0.03] to improve the model's robustness to noise;

[0019] Random horizontal flipping: Flip the image horizontally with a probability of 0.5 to increase pose diversity;

[0020] Data normalization

[0021] To improve the stability and convergence speed of model training, the loaded image data is normalized. First, the data type is converted from uint8 to float32, and then a normalization operation is performed to map the pixel values to the interval (-1.0, 1.0).

[0022] The present invention is further configured such that the second step of face region detection and face normalization includes the following steps:

[0023] After preprocessing, the image passes through 3 net components of the MTCNN (Multi-task Cascaded Convolutional Networks) multi-level cascaded convolutional neural network to detect the region of the face in the picture, and the face region is cropped. Then, face correction and normalization are performed according to the positions of the facial features, including rotation, etc. The final obtained face image size is 3×160×160. The specific steps are as follows;

[0024] Multi-scale face candidate box generation (P-Net)

[0025] First, the input image is quickly scanned through the shallow CNN network P-Net (Proposal Network). This network uses a receptive field of 12×12 to perform a sliding window operation on the image, generating face candidate regions of different scales. P-Net outputs the results of three task branches: a binary classification branch to judge whether a face is included, a regression branch to optimize the position of the bounding box, and a key point branch to predict the positions of the facial features;

[0026] Candidate box optimization (R-Net)

[0027] The candidate boxes generated by P-Net are input into the 24×24 R-Net (Refine Network) for further screening. R-Net has a more complex network structure, which can effectively remove a large number of misdetected boxes. At the same time, R-Net calibrates the position of the bounding box more precisely, improving the accuracy of face region localization;

[0028] Precise localization (O-Net)

[0029] Finally, the 48×48 O-Net (Output Network) is used to perform the final processing on the candidate boxes. O-Net can not only further improve the accuracy of face detection, but also accurately output the position coordinates of five key points (both eyes, the tip of the nose, and both corners of the mouth). This key point information is crucial for subsequent face alignment;

[0030] Face alignment and normalization

[0031] Based on the coordinates of the five key points output by the O-Net, the following normalization steps are performed: Calculate the angle between the line connecting the two eyes and the horizontal line, and rotate the face to the standard orientation through affine transformation; According to the inter-ocular distance and the facial size, scale the face region to a unified size; Use bilinear interpolation for image resampling to ensure the quality of the transformed image; Appropriately expand the bounding box to retain a certain proportion of the facial peripheral area to avoid information loss;

[0032] To improve the robustness of the system, the following mechanisms are also introduced during the detection process: Set reasonable confidence thresholds (0.9, 0.8, 0.8) for the screening of the three networks respectively; Adopt the non-maximum suppression (NMS) algorithm with a threshold set to 0.7 to merge overlapping detection boxes; When multiple faces are detected, select the face with the highest confidence and located in the central region of the image; Return a clear error message when the detection fails for subsequent processing.

[0033] The present invention is further configured such that the face vector representation in step two

[0034] In this stage, the FaceNet model is used to encode the normalized face image into a fixed-length feature vector. The InceptionResNetV1 is used as the basic network structure of FaceNet. This network combines the multi-scale feature extraction ability of the Inception module and the residual connection advantage of ResNet.

[0035] The present invention is further configured such that the distance metric in step two

[0036] In this stage, the similarity between two 512-dimensional face feature vectors is calculated. The cosine similarity is used as the metric standard, and then the original result of the cosine similarity -1 - 1 is linearly transformed to 0 - 2 to represent the distance.

[0037] The present invention is further configured such that the construction of the learning rate scheduler in step two

[0038] In this stage, the cosine annealing learning rate scheduling strategy is adopted to optimize the model training process by periodically adjusting the learning rate. Compared with the traditional learning rate scheduling strategy, cosine annealing has the following advantages: Smooth transition, the learning rate changes more continuously and smoothly, avoiding the training instability caused by stepwise adjustment; Periodic exploration, still maintaining a strong parameter space exploration ability in the later stage of training, which helps to jump out of local optima; Adaptive adjustment, the annealing period automatically adapts to the total number of training epochs without manual setting of the decay point; Good convergence, compared with exponential decay, cosine annealing can maintain a relatively high learning rate in the later stage of training, which is beneficial to the full training of the model.

[0039] The present invention is further configured such that the loss function in step two

[0040] In this stage, the model is trained using the Triplet Margin Loss function. By minimizing the distance between the anchor sample and the positive sample and maximizing the distance between the anchor sample and the negative sample, a discriminative face feature representation is learned.

[0041] According to experimental optimization, the following configuration is adopted: swap = True, enabling the distance swapping mechanism, considering both d(a,p) - d(a,n) and d(p,a) - d(p,n); margin = 0.5, setting a smaller margin value to adapt to the cross-modal feature space; reduction ='mean', returning the average of all triplet losses in the mini-batch. To improve training efficiency, online hard example mining is used. The hardest negative sample is selected in each mini-batch, and in-batch balanced sampling is performed to ensure that each identity category has a similar number of samples. This loss function has the following advantages: directly optimizing the metric learning objective in the feature space; flexibly controlling the compactness of the feature distribution through the margin parameter; the swap mechanism increasing the number of effective training samples; and being applicable to cross-modal face recognition scenarios.

[0042] The present invention is further configured such that the model training in step two

[0043] This stage details the training process and parameter configuration of the model, including data sampling, training strategy, and convergence judgment criteria, and the specific implementation is as follows.

[0044] Training data organization

[0045] Each training epoch contains 10,000 samples. The first detector is used for face detection and alignment, and training is performed in a batch processing manner with the batch size set to 64.

[0046] Learning rate configuration

[0047] A learning rate scheduling scheme based on cosine annealing is adopted, with an initial learning rate of 4e-5 and a minimum learning rate of 0.0.

[0048] Scheduling strategy: The cosine annealing strategy described in section 2.5 is adopted.

[0049] Update frequency: The learning rate is updated after each iteration.

[0050] Training process monitoring

[0051] The system records the following key metrics;

[0052] Feature distance metric: Distance between anchor-positive pair (d_ap). At the beginning of training: The average value is about 0.8; At the end of training: It stabilizes at about 0.25. The distance convergence curve shows a steady downward trend;

[0053] Monitoring of loss function: Initial value of triplet loss: About 0.5. Loss value at the end of training: Stabilizes at about 0.1, and the loss decrease curve shows continuous and smooth convergence characteristics;

[0054] Judgment of training termination

[0055] Based on experimental verification, it is recommended to terminate training when the following conditions are met: Reaching 312 iterations, the triplet loss value stabilizes around 0.1, and the anchor-positive distance stabilizes around 0.25. From the experience during the training process, 312 iterations is the optimal training length verified by experiments. At this time, the model reaches the balance point between performance and overfitting risk. Continuing training will lead to the following problems: Features overfit the training data, the cross-modal generalization ability decreases, and the performance of the validation set starts to fluctuate or decline.

[0056] (III) Beneficial effects

[0057] Compared with the prior art, the present invention provides a fusion algorithm for cross-modal face recognition, which has the following

[0058] Beneficial effects:

[0059] Through data verification in actual application scenarios, the present invention adopts advanced technologies such as MTCNN face detection and cropping, and Facenet feature extraction. By calculating the L2 norm and cosine similarity, a high-precision cross-modal face recognition system is finally achieved. The purpose of the present invention is to construct an efficient and stable cross-modal face recognition system in the freight scenario by analyzing the black-and-white images collected in real time in the cab and the color images of ID photos, achieving an identity verification accuracy of more than 95%. At the same time, the system can complete the processing and response of a single request within 1 second and maintain stable operation in an 8G memory environment. In addition, the system should also have the ability to handle exceptions, effectively identify various input exceptions (such as missing pictures, format errors, face detection failures, etc.) and return clear error prompts to ensure the robustness of the system. Traditional face recognition models often perform poorly when processing images of different modalities (such as black-and-white and color). The main reason is that they usually rely on a single feature extraction strategy and cannot effectively cope with the feature distribution shift caused by modality differences. The present invention overcomes this challenge through multiple innovative designs. First, at the feature extraction level, the FaceNet model adopted by the present invention is based on the InceptionResNetV1 architecture, combining multi-scale feature extraction with residual connections, and can capture both local detail features and global structural features of the face simultaneously. This multi-level feature extraction strategy can find more discriminative feature representations under different modalities, significantly improving the accuracy of cross-modal matching. Second, in the image preprocessing stage, the present invention specifically designs a data augmentation strategy, including random brightness adjustment, random contrast adjustment, and adding Gaussian noise, to simulate scene changes under different lighting conditions and image qualities, thereby effectively expanding the adaptation range of the model. In addition, in the model training stage, the triplet margin loss function is adopted. By minimizing the feature distance between different modalities of the same identity and maximizing the feature distance between different identities, the different modality features are mapped into a unified feature space, enabling the model to learn modality-invariant identity feature representations. Description of the Drawings

[0060] Figure 1 It is a schematic diagram of the overall structure of the fusion algorithm for cross-modal face recognition in the present invention. Detailed Embodiments

[0061] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present invention will be described in detail below with reference to the drawings and in combination with the embodiments.

[0062] It should be pointed out that unless otherwise specified, all technical and scientific terms used in the present application have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0063] In the present invention, unless otherwise specified, the orientations such as "upper" and "lower" generally refer to the directions shown in the drawings, or to the vertical, perpendicular or gravitational directions; similarly, for the convenience of understanding and description, "left" and "right" generally refer to the left and right shown in the drawings; "inner" and "outer" refer to the inner and outer of the contour of each component itself, but the above orientation terms are not used to limit the present invention.

[0064] Please refer to Figure 1 , the fusion algorithm for cross-modal face recognition, comprising the following steps:

[0065] Step 1: Data description

[0066] The dataset contains data from 142 subjects, with a total of 2556 pairs of image samples. Each pair of samples consists of a black-and-white image collected from the cab and a color ID photo image, and the number of images in both modalities accounts for 50% each. The original images are all in 3-channel RGB format and have a unified size of 464×348 pixels. The labels of the training data contain two dimensions: modality identifier (0 represents black-and-white image, 1 represents color image) and subject ID (an integer from 1 to 142). To ensure the practical applicability of the test set, 20% of the image pairs from each subject are randomly selected as the test set to ensure an even distribution of different lighting conditions, shooting angles, and expression changes. In addition, all images are manually screened to eliminate samples with severe blurring, occlusion, or insufficient lighting to ensure data quality;

[0067] Step 2: Modeling steps: preprocessing, face region detection and face normalization, face vector representation, distance metric, constructing a learning rate scheduler, loss function, and model training.

[0068] 1. High performance

[0069] 1.1 Powerful cross-modal processing ability

[0070] Traditional face recognition models often perform poorly when dealing with images of different modalities (such as black and white vs. color). The main reason is that they usually rely on a single feature extraction strategy and cannot effectively handle the feature distribution shift caused by modality differences. This invention overcomes this challenge through multiple innovative designs. First, at the feature extraction level, the FaceNet model adopted in this invention is based on the InceptionResNetV1 architecture, which combines multi-scale feature extraction with residual connections, enabling it to capture both local detailed features and global structural features of the face simultaneously. This multi-level feature extraction strategy can find more discriminative feature representations under different modalities, significantly improving the accuracy of cross-modal matching. Second, in the image preprocessing stage, this invention specifically designs a data augmentation strategy, including random brightness adjustment, random contrast adjustment, and adding Gaussian noise, to simulate scene changes under different lighting conditions and image qualities, thus effectively expanding the adaptation range of the model. Additionally, in the model training stage, the triplet margin loss function is used. By minimizing the feature distance between different modalities of the same identity and maximizing the feature distance between different identities, different modality features are mapped into a unified feature space, enabling the model to learn modality-invariant identity feature representations.

[0071] 1.2 The model has excellent feature extraction capabilities

[0072] Compared with traditional single-feature extraction networks, this invention achieves a more comprehensive and accurate face feature representation through a multi-level feature extraction architecture. This excellent feature extraction capability is mainly reflected in the following aspects:

[0073] First, in the face detection stage, the adopted MTCNN cascade network realizes coarse-to-fine face localization and feature extraction through a three-stage progressive feature extraction strategy. The P-Net quickly captures basic features and generates candidate regions, the R-Net further extracts more detailed features for candidate box optimization, and the O-Net accurately extracts the information of facial feature points. This progressive feature extraction method not only improves the detection accuracy but also lays a good foundation for subsequent feature representation.

[0074] Second, in the feature encoding stage, the adopted FaceNet model realizes the effective fusion of multi-scale features through the InceptionResNetV1 network structure. The Inception module can extract features under different receptive fields simultaneously, capturing multi-level information from detailed textures to contour structures. The residual connection mechanism of ResNet effectively alleviates the gradient vanishing problem in deep networks, ensuring the effective transmission of deep features. The 512-dimensional feature vector not only guarantees the richness of expression ability but also avoids redundancy and computational overhead caused by too high dimensions.

[0075] Third, in terms of feature learning, the triplet loss function adopted by the present invention, through the hard sample mining mechanism, prompts the model to focus on more discriminative features. By simultaneously considering the distance metrics in both the anchor-positive and positive-anchor directions (swap = True), the discrimination ability of the features is enhanced. This feature learning strategy enables different modality images of the same identity to be mapped to adjacent regions in the feature space, while images of different identities maintain sufficient intervals in the feature space, and at the same time, the features have good clustering and discrimination properties.

[0076] 1.3 The model has strong environmental adaptability

[0077] Compared with conventional face recognition systems, the present invention demonstrates stronger adaptability in actual complex environments. This environmental adaptability is mainly reflected in the following aspects.

[0078] First, in terms of light adaptability, the present invention enhances the processing ability for different light conditions through multiple innovative designs: in the data augmentation stage, by randomly adjusting the brightness within the range of [-0.2, 0.2], the natural light changes at different times are simulated. By adjusting the contrast within the range of [0.8, 1.2], the adaptability of the model to different light intensities is enhanced. Especially for the infrared light environment in the truck cab, the model effectively processes the feature differences of infrared imaging through specific preprocessing strategies. These measures enable the model to maintain stable recognition performance under various natural light conditions such as dawn and dusk light changes and strong light on cloudy days.

[0079] Second, in response to the dynamic characteristics of the driving environment, the model demonstrates good pose adaptability: the multi-level detection architecture of MTCNN can accurately capture side face angles within ±30°; the face alignment module realizes the standardization processing of poses through affine transformation; the feature extraction network enhances the robustness to pose changes through a multi-scale structure. This enables the system to adapt to various head movements of the driver during normal work and maintain recognition accuracy.

[0080] 2. Scalability

[0081] 2.1 The model architecture has good scalability

[0082] Compared with traditional single closed architectures, the present invention adopts a modular design concept, enabling the system to have excellent scalability.

[0083] 2.1.1 The model architecture is highly modular

[0084] The image preprocessing module, face detection module, feature extraction module, and similarity calculation module adopt a loose coupling architecture. Each module communicates through a standardized data interface, facilitating independent upgrade or replacement. The core processing flow is abstracted into a configurable processing pipeline, supporting flexible addition or removal of processing nodes.

[0085] This design enables the system to quickly expand functions or optimize performance according to actual needs.

[0086] 2.1.2 Feature extraction scalability

[0087] The basic network architecture of FaceNet supports the replacement of multiple backbone networks, such as ResNet, MobileNet, etc.

[0088] The dimension of the feature vector can be flexibly adjusted through a configuration file to adapt to different accuracy and efficiency requirements. The pre-trained model supports incremental updates, continuously optimizing the feature extraction ability. It supports integrating multiple feature extractors to achieve feature fusion and improve performance.

[0089] 2.1.3 Adaptability of model deployment

[0090] This system supports multiple deep learning frameworks (such as PyTorch, ONNX, etc.), can flexibly switch between CPU or GPU inference modes according to hardware conditions, the model structure also supports quantization compression for easy deployment on different computing power platforms, and a hardware acceleration interface is reserved to support future access to various AI acceleration cards.

[0091] 2.1.4 Rich extension interfaces

[0092] Supports customizing the data preprocessing pipeline, and specific image processing steps can be added according to specific scenarios.

[0093] Provides a standardized model evaluation interface, facilitating the integration of new evaluation metrics.

[0094] The exception handling mechanism supports custom extension, and specific exception handling logic can be added according to business requirements.

[0095] The result output format is configurable, supporting docking with different upper-layer application systems.

[0096] 2.2 The model supports incremental training and optimization

[0097] Compared with traditional static models, the present invention designs a complete incremental training and optimization mechanism, enabling the system to continuously improve performance as data accumulates.

[0098] 2.2.1 Progressive training scheme

[0099] Supports fine-tuning based on the weights of the existing model, avoiding the resource overhead caused by full retraining.

[0100] Adopt a dynamic learning rate adjustment strategy to ensure the stability of incremental training.

[0101] Implement a hierarchical training mechanism to enable targeted optimization for specific levels.

[0102] Provide a model rollback mechanism to prevent performance degradation caused by incremental training.

[0103] 2.2.2 Evaluation System

[0104] Set up multi-dimensional performance metric monitoring, including recognition accuracy, processing time, resource occupancy, etc.

[0105] Support real-time performance evaluation during incremental training.

[0106] Provide a detailed analysis report on the optimization effect to help determine the optimization direction.

[0107] Implement an automated A / B testing mechanism to scientifically evaluate the optimization effect.

[0108] 2.2.3 The system supports flexible configuration of multiple optimization strategies

[0109] Different optimization goals can be selected according to actual needs, such as improving accuracy or reducing latency.

[0110] Support targeted optimization for specific scenarios, such as special environments like strong light and night.

[0111] Provide model compression and quantization optimization options to balance performance and resource occupancy.

[0112] 1. Innovative cross-modal processing method

[0113] In the present invention, for the cross-modal recognition problem of black-and-white cab images and color ID photos, an innovative processing solution is proposed.

[0114] 1.1 Feature mapping strategy for black-and-white and color images

[0115] The present invention abandons the traditional direct feature extraction method and innovatively designs a dual-path feature mapping strategy. Specifically, enhanced local structure feature extraction is used for black-and-white images, while global feature extraction is emphasized for color images. By different feature extraction strategies, the information differences between the two modalities are balanced. This asymmetric feature mapping scheme effectively solves the problem of inconsistent cross-modal feature representation and significantly improves the recognition accuracy.

[0116] 1.2 Multi-level feature extraction and fusion mechanism

[0117] The present invention adopts a technical solution of multi-level feature extraction and fusion. First, different-scale feature information is extracted through three sub-networks (P-Net, R-Net, O-Net) of the MTCNN network, and then the InceptionResNetV1 structure is used for the fusion of multi-scale features. In particular, an adaptive weight mechanism is adopted in the fusion stage, and the fusion weights of each level of features are dynamically adjusted according to the characteristics of different-modal images, ensuring the effectiveness and robustness of feature expression.

[0118] 1.3 Cross-modal feature space alignment method

[0119] Aiming at the distribution difference problem of cross-modal features, the present invention proposes an innovative feature space alignment method. By introducing an improved triplet loss function during training, the feature vectors of the same identity under different modalities are forced to maintain a compact distribution in the high-dimensional feature space, while maximizing the feature distance between different identities. In addition, by setting the mechanism of swap = True, the consistency of the feature space is further enhanced, enabling effective comparison of cross-modal features in a unified feature space.

[0120] 2. Improved data preprocessing strategy

[0121] In the present invention, to meet the special requirements of face recognition in the freight scenario, a complete set of data preprocessing strategies is designed, including the following technological innovations:

[0122] First, the present invention abandons the traditional general data augmentation scheme and innovatively proposes a data augmentation strategy for the freight scenario. By randomly adjusting the brightness in the range of [-0.2, 0.2] and the contrast in the range of [0.8, 1.2], the illumination changes under different time periods and weather conditions are simulated; and Gaussian noise with a mean of 0 and a standard deviation in the range of [0.01, 0.03] is injected into the image to enhance the model's adaptability to noise. This scenario-based augmentation scheme significantly improves the generalization performance of the model in the actual environment.

[0123] Second, the present invention proposes an adaptive image normalization method. Different from the traditional global normalization process, the present invention first converts the image from uint8 to float32 to provide a larger numerical range, and then maps the pixel values to the interval (-1.0, 1.0) through an adaptive normalization operation. Especially for the differences between black-and-white and color images, different normalization strategies are adopted to ensure that the data feature distributions of different modalities are consistent.

[0124] Finally, the present invention designs a multi-stage face detection and alignment mechanism. In the detection stage, a three-stage cascaded structure of MTCNN is adopted: P-Net generates candidate boxes with a receptive field of 12×12, R-Net further optimizes the position of the bounding box with a receptive field of 24×24, and O-Net precisely locates five key points through a receptive field of 48×48. In the alignment stage, the following operations are implemented based on the key point coordinates: First, the pose is normalized by calculating the angle between the line connecting the two eyes and the horizontal line and performing an affine transformation; then, the scaling ratio is dynamically calculated according to the interocular distance to ensure the uniformity of the face region size; next, bilinear interpolation is used for image resampling to maintain the quality of the transformed image; finally, a reasonable boundary expansion strategy is set to retain sufficient facial peripheral information.

[0125] 3. Optimized model training scheme

[0126] In the present invention, aiming at the training difficulties of cross-modal face recognition, an optimized model training scheme is proposed.

[0127] 3.1 Design of improved triplet loss function

[0128] The present invention makes an innovative improvement to the traditional triplet loss function. First, by setting the mechanism of swap = True, the distance metrics in two directions, d(a,p)-d(a,n) and d(p,a)-d(p,n), are considered simultaneously, increasing the number of effective training samples. Second, through experiments, a relatively small margin value of 0.5 is optimized and set, which better adapts to the distribution characteristics of the cross-modal feature space. Finally, the reduction ='mean' strategy is adopted, and the training process is stabilized by calculating the average value of all triplet losses in the mini-batch. This improved loss function design significantly improves the convergence speed and final performance of the model.

[0129] 3.2 Dynamic learning rate scheduling strategy

[0130] The present invention innovatively adopts a learning rate scheduling scheme based on cosine annealing, with the initial learning rate set to 4e-5, and the change trend of the learning rate is dynamically adjusted through the cosine function. This strategy has multiple advantages: on the one hand, the change of the learning rate is smoother and continuous, avoiding the training instability caused by stepwise adjustment; on the other hand, it still maintains a moderate parameter exploration ability in the later stage of training, which helps to jump out of the local optimum; in addition, the annealing period can automatically adapt to the total number of training rounds without manual setting of the decay point; compared with the traditional exponential decay method, cosine annealing can maintain a relatively high learning rate in the later stage of training, thus promoting the full training of the model.

[0131] 3.3 Innovative sample sampling mechanism

[0132] The present invention designs an efficient sample sampling strategy. In terms of training data organization, each training epoch contains 10,000 samples, and batch processing with a batch size of 64 is used for training. The following innovative mechanisms are introduced during the sampling process: First, through online hard sample mining, the most difficult negative samples are preferentially selected in each mini-batch, thereby improving the training efficiency; Second, in-batch balanced sampling ensures that each identity category has a similar number of samples within the batch, avoiding the phenomenon of class imbalance; Finally, the compute_class_weight method provided by sklearn.utils is used to automatically calculate the class weights and dynamically adjust the weight allocation of different classes to achieve a more reasonable allocation of training resources.

[0133] The optimization and innovation of these training schemes effectively solve the special difficulties in cross-modal face recognition training, significantly improving the training efficiency and final performance of the model. At the same time, these technical solutions have strong generality and can be extended to the training process of other deep learning models. Experimental results show that with this training scheme, the model can reach a stable performance level after 312 iterations and can well balance the risk of overfitting.

[0134] The goal of the present invention is to achieve cross-modal (black and white and color images) face recognition. In addition to the MTCNN + FaceNet + improved triplet loss scheme adopted by the present invention, there are also the following several traditional or early alternative schemes.

[0135] 1. Traditional feature extraction + machine learning scheme

[0136] This type of method mainly relies on manually designed feature extractors and traditional machine learning algorithms.

[0137] 1.1 SIFT / SURF-based scheme

[0138] Based on the SIFT (Scale-Invariant Feature Transform) or SURF (Speeded-Up Robust Features) operator to detect key points, and extract local feature descriptors around the key points. Subsequently, the image similarity is calculated through a feature matching algorithm. This method has high accuracy in local feature description, but has a large computational complexity and is difficult to capture global semantic information.

[0139] 1.2 HOG-based scheme

[0140] The HOG (Histogram of Oriented Gradients)-based solution first calculates the histogram of oriented gradient features of an image, then divides the image into multiple blocks and separately counts the gradient direction distributions of each block, and finally connects the features of all blocks to form the final feature vector. This method has a certain robustness in dealing with illumination changes, but is sensitive to pose changes and has relatively limited feature expression ability.

[0141] 1.3 LBP-based solution

[0142] The LBP-based solution generates the LBP code of local texture features by calculating the binary relationship between a pixel and its neighboring pixels, and on this basis, counts the LBP feature distribution of the entire image. This method is simple and efficient in calculation and is suitable for single-modal images, but the accuracy drops significantly when dealing with cross-modal data.

[0143] 1.4 Feature dimensionality reduction and SVM solution

[0144] The features manually extracted are reduced in dimension by PCA (Principal Component Analysis), and SVM (Support Vector Machine) is used for classification decision-making. Different kernel functions can be selected according to requirements to improve the classification performance. Although this solution has relatively small computational overhead, there are still the following limitations: the feature extraction lacks adaptability and is difficult to handle complex changes; the manually designed features may ignore key discriminant information; it performs poorly when dealing with cross-modal images and is particularly sensitive to illumination and pose changes; its accuracy in practical applications usually cannot exceed 85%.

[0145] 2. Shallow CNN-based solution

[0146] This solution uses a relatively simple CNN network structure for feature learning.

[0147] 2.1 AlexNet solution

[0148] The shallow CNN-based solution mainly uses a relatively simple convolutional neural network structure for feature learning, with AlexNet being a typical representative. The network depth of AlexNet is only 8 layers, including 5 convolutional layers and 3 fully connected layers, using the ReLU activation function and Dropout regularization, and downsampling the features through max pooling. Although this structure pioneered the application of deep learning in the field of computer vision, there are still the following limitations: the network capacity is small and the feature extraction ability is limited; the receptive field coverage is insufficient and it is difficult to capture large-scale features; the modeling ability for complex scenes is insufficient and it is difficult to adapt to diverse application requirements.

[0149] 2.2 VGG solution

[0150] VGG adopts a unified 3×3 convolutional kernel design, expands the network depth to 16 - 19 layers, and simulates large receptive fields by stacking small convolutional kernels. Compared with AlexNet, VGG has improved in feature extraction, but still faces the following limitations: First, the large number of parameters leads to high training and deployment costs; Second, the unstable gradient propagation increases the difficulty of model training; In addition, the network structure is relatively weak in organizing feature hierarchies and it is difficult to make full use of deep feature information.

[0151] 2.3 Classifier and Loss Function Design

[0152] In terms of classifier and loss function design, this solution uses a simple Softmax classifier for identity recognition and the cross - entropy loss function for model training, lacking targeted design for cross - modal feature learning. Due to the simplicity of this framework, it is difficult to effectively model intra - class differences and inter - class relationships; the learning efficiency for hard samples is low; and it is prone to overfitting problems when facing complex scenarios.

[0153]

[0154]

[0155] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. The fusion algorithm of cross-modal face recognition includes the following steps: Step 1: Data Description The dataset contains data from 142 subjects, totaling 2556 pairs of image samples. Each pair of samples consists of a black-and-white image collected from the cab and a color image of a ID photo, with 50% of the images in each mode. The original images are all in 3-channel RGB format with a uniform size of 464×348 pixels. The labels of the training data contain two dimensions: modality identifier (0 for black-and-white images, 1 for color images) and subject ID (an integer from 1 to 142). To ensure the practical applicability of the test set, 20% of the image pairs of each subject are randomly selected as the test set to ensure a balanced distribution of different lighting conditions, shooting angles, and expression changes. In addition, all images have been manually screened to remove samples that are severely blurred, occluded, or poorly illuminated to ensure data quality; Step 2: Modeling steps: preprocessing, face area detection and face normalization, face vector representation, distance metric, building a learning rate scheduler, loss function and model training.

2. The cross-modal face recognition fusion algorithm according to claim 1, characterized in that: The pretreatment in step 2 comprises the following steps: Data loading The image file is read into the Pytorch tensor format through a custom data loader. The data dimension after loading is 3×464×368, where the first dimension represents the number of RGB channels, and the next two dimensions represent the height and width of the image, respectively. The initial data type is uint8, and the pixel values ​​are distributed in the range of 0-255. Data Augmentation In order to improve the generalization ability of the model and its adaptability to various scenarios, the training data is enhanced in a variety of ways. According to the characteristics of truck driving scenarios, the following enhancement strategies are mainly adopted: Random brightness adjustment: randomly adjust the image brightness in the range of [-0.2, 0.2] to simulate the changes in lighting at different times; Random contrast adjustment: randomly adjust the image contrast in the range of [0.8, 1.2] to enhance the model's adaptability to images of different quality; Random Gaussian noise: Add Gaussian noise with a mean of 0 and a standard deviation in the range of [0.01, 0.03] to improve the robustness of the model to noise; Random horizontal flip: flip the image horizontally with a probability of 0.5 to increase posture diversity; Data Standardization In order to improve the stability and convergence speed of model training, the loaded image data is standardized. First, the data type is converted from uint8 to float32, and then a normalization operation is performed to map the pixel values ​​to the range of (-1.0, 1.0).

3. The cross-modal face recognition fusion algorithm according to claim 2, characterized in that: The steps of face region detection and face standardization include the following steps: After preprocessing, the image is passed through the three net components of the MTCNN (Multi-task Cascaded Convolutional Networks) multi-level cascade convolutional neural network to detect the area of ​​the face in the picture and crop the face area. Then, the face is corrected and standardized according to the position of the facial features, including rotation. The final face image size is 3×160×160. The specific steps are as follows; Multi-scale face candidate box generation (P-Net) First, the input image is quickly scanned through a shallow CNN network P-Net (Proposal Network). The network uses a 12×12 receptive field to perform a sliding window operation on the image to generate face candidate regions of different scales. P-Net outputs the results of three task branches: the binary classification branch determines whether it contains a face, the regression branch optimizes the position of the bounding box, and the key point branch predicts the position of the facial features. Candidate Box Optimization (R-Net) The candidate boxes generated by P-Net are input into the 24×24 R-Net (Refine Network) for further screening. R-Net has a more complex network structure and can effectively remove a large number of false positive boxes. At the same time, R-Net calibrates the position of the bounding box more accurately, improving the accuracy of face area positioning. Precise Positioning (O-Net) Finally, the candidate boxes are processed using a 48×48 O-Net (Output Network). O-Net can not only further improve the accuracy of face detection, but also accurately output the position coordinates of five key points (eyes, nose tip, and mouth corners). These key point information is crucial for subsequent face alignment. Face alignment and normalization Based on the coordinates of the five key points output by O-Net, the following standardization steps are performed: the angle between the eye line and the horizontal line is calculated, and the face is rotated to the standard orientation through affine transformation; the face area is scaled to a uniform size according to the eye distance and face size; bilinear interpolation is used to resample the image to ensure the quality of the transformed image; the bounding box is appropriately expanded to retain a certain proportion of the facial surrounding area to avoid information loss; To improve the robustness of the system, the following mechanisms are introduced during the detection process: reasonable confidence thresholds (0.9, 0.8, and 0.8) are set for screening the three networks respectively; the non-maximum suppression (NMS) algorithm is used with the threshold set to 0.7 to merge overlapping detection frames; when multiple faces are detected, the face with the highest confidence and located in the central area of ​​the image is selected; when detection fails, a clear error message is returned to facilitate subsequent processing.

4. The cross-modal face recognition fusion algorithm according to claim 3, characterized in that: The face vector representation in step 2 In this stage, the FaceNet model is used to encode the standardized face images into feature vectors of fixed length, and InceptionResNetV1 is used as the basic network structure of FaceNet, which combines the multi-scale feature extraction capability of the Inception module and the residual connection advantages of ResNet.

5. The cross-modal face recognition fusion algorithm according to claim 4 is characterized by: The distance metric in step 2 In this stage, the similarity of two 512-dimensional facial feature vectors is calculated, and cosine similarity is used as the metric. Then the original result of cosine similarity -1-1 is linearly transformed to 0-2 to represent the distance.

6. The cross-modal face recognition fusion algorithm according to claim 5, characterized in that: In step 2, a learning rate scheduler is constructed In this stage, the cosine annealing learning rate scheduling strategy is used to optimize the model training process by periodically adjusting the learning rate. Compared with the traditional learning rate scheduling strategy, cosine annealing has the following advantages: smooth transition, the learning rate changes more continuously and smoothly, avoiding the training instability caused by step-type adjustment; periodic exploration, maintaining a strong parameter space exploration capability in the later stage of training, which helps to jump out of the local optimum; adaptive adjustment, the annealing cycle automatically adapts to the total training rounds, and there is no need to manually set the attenuation point; good convergence, compared with exponential decay, cosine decay can maintain a relatively high learning rate in the later stage of training, which is conducive to full model training.

7. The cross-modal face recognition fusion algorithm according to claim 6, characterized in that: The loss function in step 2 In this stage, the triplet margin loss function is used to train the model. By minimizing the distance between the anchor sample and the positive sample and maximizing the distance between the anchor sample and the negative sample, the discriminative facial feature representation is learned. According to experimental optimization, the following configuration is adopted: swap = True, enabling the distance exchange mechanism, comprehensively considering d(a,p)-d(a,n) and d(p,a)-d(p,n); margin = 0.5, setting a smaller margin value to adapt to the cross-modal feature space; reduction = 'mean', returning the average of all triplet losses in the mini-batch. To improve training efficiency, online difficult sample mining is adopted, the most difficult negative sample is selected in each mini-batch, and balanced sampling is performed within the batch to ensure that each identity category has a close number of samples. This loss function has the following advantages: directly optimizing the metric learning objective of the feature space; flexibly controlling the compactness of the feature distribution through the margin parameter; the swap mechanism increases the number of effective training samples; and is suitable for cross-modal face recognition scenarios.

8. The cross-modal face recognition fusion algorithm according to claim 7, characterized in that: Model training in step 2 This stage describes in detail the model training process and parameter configuration, including data sampling, training strategy and convergence judgment criteria. The specific implementation is as follows. Training data organization Each training cycle (epoch) contains 10,000 samples. The first detector is used for face detection and alignment. The training is performed in batch mode with a batch size of 64. Learning rate configuration A learning rate scheduling scheme based on cosine descent, with an initial learning rate of 4e-5 and a minimum learning rate of 0.0; Scheduling strategy: Use the cosine descent strategy described in Section 2.5; Update frequency: Update the learning rate after each iteration; Training process monitoring The system records the following key indicators; Feature distance indicator: anchor-positive pair distance (d_ap). Initial training: average value is about 0.8; At the end of training: Stable at around 0.

25. The distance convergence curve shows a steady downward trend; Loss function monitoring: The initial value of triplet loss is about 0.

5. The loss value at the end of training is stable at about 0.1, and the loss decline curve shows a continuous and smooth convergence feature; Training termination judgment Based on experimental verification, it is recommended to terminate training when the following conditions are met: 312 iterations are reached, the triplet loss value fluctuates steadily around 0.1, and the anchor-positive distance stabilizes around 0.

25. Experience in the training process has found that 312 iterations is the optimal training length for experimental verification. At this time, the model reaches a balance between performance and overfitting risk. Continuing training will lead to the following problems: features overfit the training data, cross-modal generalization ability decreases, and the performance of the validation set begins to fluctuate or decrease.

Citation Information

Cited By

  • Intelligent identifying and counting method, system and equipment for industrial hoister and medium

    CN121214161A