Method, system and equipment for generating face recognition fusion model
Through multimodal data fusion and timing adaptive modeling, the problem of degradation of recognition accuracy in complex environments by traditional face recognition technology is solved, and high accuracy recognition under factors such as lighting, angle and expression changes are achieved.
Patent Information
- Application Number
- CN202510381200.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional facial recognition technology has reduced recognition accuracy in complex environments, especially under the influence of light changes, angle changes, occlusions and expression changes, the recognition accuracy rate has dropped significantly.
By collecting historical face recognition image sets and sensing data, multi-modal data fusion is carried out, multi-layer network spatiotemporal features are integrated, model correlation matrix is constructed, and adaptive modal contribution weighting is carried out, face recognition timing adaptation model is established, and model evaluation and optimization is carried out.
It improves the robustness and accuracy of facial recognition, enhances the model's adaptability to complex environments, and ensures that high recognition performance is maintained in different scenarios.
Smart Images

Figure CN120356044A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly to a method, system and device for generating a face recognition fusion model. Background Art
[0002] As a kind of biometric recognition technology, face recognition has gradually become an important means of identity authentication due to its advantages such as non-contact, non-invasive, and high recognition accuracy. Traditional face recognition technologies mainly rely on methods based on feature engineering, such as principal component analysis (PCA), linear discriminant analysis (LDA), etc. These methods usually show good performance when dealing with standardized and clear images, but have many deficiencies in complex environments. Among them, a significant defect of traditional methods is their strong dependence on images. In practical applications, due to factors such as illumination changes, angle changes, occlusion, and expression changes, the performance of traditional methods is often limited. For example, the PCA method relies on the linear transformation of the main features in the image and has weak modeling ability for non-linear relationships, resulting in a significant decrease in accuracy when dealing with complex images. At the same time, the LDA method has good effects when the number of sample categories is small, but the recognition accuracy will also be affected when the data volume is large or the number of categories is large. Summary of the Invention
[0003] Based on this, it is necessary for the present invention to provide a method, system and device for generating a face recognition fusion model to solve at least one of the above technical problems.
[0004] To achieve the above object, a method for generating a face recognition fusion model includes the following steps:
[0005] Step S1: Collect a historical face recognition image set and historical device sensing data through a face recognition device, and perform multi-modal data fusion integration on the historical face recognition image set and historical device sensing data to obtain a multi-source fusion data set;
[0006] Step S2: Perform multi-layer network spatio-temporal feature integration on the multi-source fusion data set to obtain a high-dimensional feature spectrum;
[0007] Step S3: Perform modal information flow modeling based on the high-dimensional feature spectrum to obtain a modal correlation matrix, and perform adaptive modal contribution degree weighting on the modal correlation matrix to obtain a modal relationship matrix;
[0008] Step S4: Adjust the priority of the modal feature flow according to the modal relationship matrix to obtain an adaptive feature flow; perform face recognition time series adaptability modeling based on the adaptive feature flow to obtain a face recognition time series adaptation model;
[0009] Step S5: Evaluate the face recognition temporal adaptation model to obtain model accuracy evaluation data; establish an accuracy optimization matrix based on the model accuracy evaluation data, and use the accuracy optimization matrix to optimize the face recognition temporal adaptation model to obtain a face recognition fusion model.
[0010] In view of the problem that the recognition accuracy of traditional face recognition technology decreases in complex environments, the present invention improves the robustness and accuracy of recognition through multi-modal data fusion and temporal adaptability modeling. By introducing a historical face recognition image set and historical device sensing data, and performing multi-modal data fusion integration, it is possible to make full use of sensor information to enhance the environmental adaptability of recognition and effectively make up for the limitations of single image features. By adopting multi-layer network spatio-temporal feature integration, high-dimensional feature spectra can be extracted in different time and space dimensions, enhancing the adaptability of the model to changes in factors such as illumination, angle, and occlusion, and retaining key identity feature information. Through modal information flow modeling, the correlation between different modalities can be effectively analyzed, redundant feature interference can be avoided, and the information utilization efficiency can be improved. At the same time, adaptive modal contribution degree weighting can dynamically adjust the influence weights of different modalities, enabling the model to pay more attention to key features and improving the generalization ability of recognition. By adjusting the priority of modal feature flows based on the modal relationship matrix, the utilization order of feature flows can be optimized according to environmental and data feature changes, enabling the model to maintain a high recognition accuracy in different scenarios. Through temporal adaptability modeling, the recognition model can dynamically adapt to feature changes in the time dimension, improving the stability of long-term recognition. Through model evaluation and accuracy optimization matrix adjustment, refined optimization for different data distributions can be achieved, further improving the accuracy and real-time performance of face recognition. At the same time, the optimized model can more precisely adapt to recognition tasks in different environments, ensuring high recognition performance even when face features change significantly.
[0011] Optionally, step S1 is specifically as follows:
[0012] Step S11: Collect a historical face recognition image set and historical device sensing data through a face recognition device, and perform data preprocessing to obtain a historical face recognition image set to be analyzed and a historical device sensing data to be analyzed;
[0013] Step S12: Perform temporal synchronization and spatial feature mapping on the historical face recognition image set to be analyzed and the historical device sensing data to be analyzed to obtain a spatio-temporally aligned multi-modal data set;
[0014] Step S13: Perform principal component feature selection based on the spatio-temporally aligned multi-modal data set, and perform weighted fusion on the principal component features to obtain a multi-modal fusion data set;
[0015] Step S14: Perform data augmentation on the multi-modal fusion data set to obtain a multi-source fusion data set;
[0016] Step S15: Perform Monte Carlo data augmentation on the multi-source fusion dataset to obtain the multi-source fusion dataset.
[0017] The present invention collects historical image data and sensing data through a face recognition device and performs data preprocessing, which can effectively remove abnormal data and noise interference and maintain the consistency of data formats, providing clean input data for subsequent analysis. Secondly, by using time series synchronization and spatial feature mapping, data of different modalities are strictly aligned in the time dimension, and reasonable associations are established in space, ensuring more accurate feature fusion of image data and sensing data and avoiding model misjudgment caused by information misalignment. In the feature extraction stage, the principal component analysis method is used for feature selection to eliminate redundant features and only retain the most informative principal components. At the same time, combined with the feature weighted fusion technology, the distribution of the contribution degrees of each modality data is made more balanced, avoiding the dominant role of single modality features. To further improve the generalization ability of the model, a data augmentation method is introduced to generate simulated data under different scenarios and lighting conditions to enhance the robustness of the model. In addition, combined with the Monte Carlo data augmentation strategy, the data samples are further expanded to improve the diversity of training data, enabling the model to perform more stably in the face of different real application scenarios.
[0018] Optionally, step S13 is specifically:
[0019] Extract high-level face image features and sensing physical features from the spatio-temporally aligned multi-modal dataset, and perform standardization processing on the high-level face image features and sensing physical features to obtain a standardized high-level face image feature set and a standardized sensing physical feature set;
[0020] Perform face expression feature pattern clustering based on the standardized high-level face image feature set, evaluate the feature contribution degrees of the clustering results, and screen out the principal component features according to the feature contribution degree evaluation results to obtain the principal component high-level face image feature set;
[0021] Perform principal component feature selection based on the standardized sensing physical feature set to obtain the principal component sensing physical feature set;
[0022] Perform modal feature weighted fusion on the principal component high-level face image feature set and the principal component sensing physical feature set to obtain a multi-modal fusion feature set;
[0023] Perform non-linear optimization of the fusion features on the multi-modal fusion feature set to obtain the multi-modal fusion dataset.
[0024] After the alignment of multi-modal data, the present invention extracts high-level features and sensing physical features of face images and performs standardization processing to make the data of different modalities consistent in the numerical scale, eliminating the imbalance problem caused by different feature distributions, thereby improving the stability of subsequent calculations and the comparability of features. Secondly, in the processing of face image features, through the clustering of expression feature patterns, different categories of expression patterns can be effectively mined, and the feature contribution degree of the clustering results is evaluated. The principal component features that contribute the most to the recognition task are selected, redundant information is removed, and the calculation efficiency and the discriminant ability of the model are improved. At the same time, the principal component analysis is also used to reduce the dimension of the sensing physical features to ensure the optimal expression form of the sensing data, reduce the calculation complexity while retaining the most informative feature components. After that, the selected principal component features are weighted and fused in modal features, and a reasonable weighting mechanism is established between different modalities, so that the image features and sensing features can complement each other and avoid the deviation problem caused by the dominance of single-modal information in the recognition process. Finally, in order to further improve the separability of the multi-modal fusion features, a non-linear optimization method is used to optimize the fusion features to make them more in line with the input requirements of the deep learning model, improve the feature expression ability of the model, and finally obtain a high-quality multi-modal fusion data set.
[0025] Optionally, step S2 is specifically as follows:
[0026] Step S21: Use a preset multi-scale convolutional network to extract the image detail features and sensing trend features in the multi-source fusion data set, so as to obtain a multi-scale spatio-temporal feature set;
[0027] Step S22: Perform data feature adversarial optimization on the multi-scale spatio-temporal feature set to obtain a constrained optimization feature set;
[0028] Step S23: Perform dynamic spatio-temporal modeling on the constrained optimization feature set through a preset spatio-temporal relationship adaptive network, and adaptively adjust the contribution degree of the features at different time points and spatial regions to obtain an adaptive spatio-temporal feature set;
[0029] Step S24: Convert the adaptive spatio-temporal feature set into a graph structure to obtain a spatio-temporal graph feature set;
[0030] Step S25: Perform modal fusion and feature integration based on the spatio-temporal feature set to obtain a high-dimensional feature spectrum.
[0031] The present invention uses a multi-scale convolutional network to extract features from a multi-source fusion dataset, which can capture the detailed information in image data and simultaneously extract the changing trends of sensing data, thereby obtaining a more comprehensive spatio-temporal feature representation. The advantage of multi-scale feature extraction is that information at different scales can complement each other, improving the model's perception ability of local details and global trends, and ensuring stable recognition under factors such as face pose changes and lighting interference. Then, the extracted multi-scale spatio-temporal features are adversarially optimized. By constructing a constraint optimization mechanism, the features are made more robust, avoiding the interference of abnormal noise, and enhancing the separability and discriminability of the features. Next, a spatio-temporal relationship adaptive network is used to dynamically model the optimized features, adaptively adjusting the contribution degrees of the features at different time points and spatial regions, enabling the model to automatically adjust the weights according to the spatio-temporal changes of the data, and thus more accurately capturing the evolution laws of behavior patterns and face features. Further, the adaptive spatio-temporal feature set is converted into a graph structure. By constructing the topological relationship between nodes and edges, the correlation of spatio-temporal features is explicitly expressed, and then a more reasonable structured representation is provided for subsequent feature integration. Finally, based on the spatio-temporal graph features, modal fusion and feature integration are carried out to construct a high-dimensional feature spectrum, enabling multi-modal data to be fused at a higher level, ensuring that the complementary effect between different information sources is maximized, and improving the overall performance of face recognition.
[0032] Optionally, the modal information flow modeling in step S3 is specifically as follows:
[0033] Extract the face image feature flow and device sensing feature flow in the high-dimensional feature spectrum;
[0034] Evaluate the difference degree of the feature modal distribution of the high-dimensional feature spectrum, set the distribution deviation threshold to 0.15, and standardize the features in the face image feature flow and device sensing feature flow according to the distribution deviation threshold to obtain a standardized modal feature flow set;
[0035] Set the autoregressive order to 3, the moving average order to 2, calculate the smoothing window length to be 10s, and perform autoregressive moving average trend smoothing on the standardized modal feature flow set to obtain a smoothed modal time series;
[0036] Conduct cross-modal information coupling analysis based on the smoothed modal time series, calculate the mutual information gain of each modal feature flow, and set the number of discretization intervals for information gain calculation to 20, thereby constructing a modal mutual information gain matrix;
[0037] Set the clustering threshold to 0.6, perform hierarchical clustering based on the modal mutual information gain matrix, and conduct modal correlation modeling based on the hierarchical clustering results to generate a modal correlation matrix.
[0038] The characteristic modal distribution difference evaluation of the present invention can identify the statistical distribution deviation between different modal data, thereby ensuring that the data is in a relatively consistent standardized state before fusion. Setting the distribution deviation threshold to 0.15 can ensure data diversity while avoiding the stability of model training due to excessive distribution differences between modalities. Through standardization, the facial image feature stream and the device sensor feature stream can be compared and calculated on the same numerical scale, improving the effectiveness of cross-modal analysis. Autoregressive moving average trend smoothing can reduce noise in the data while retaining key time series features. The autoregressive order is set to 3, so that the model can use the information of the past 3 time steps for prediction, and the moving average order is 2, which can smooth out short-term random fluctuations, thereby avoiding trend misjudgment caused by short-term abnormal points. In addition, the smoothing window length is set to 10s, so that local changes in data in a short period of time can be effectively captured without losing the overall trend. This parameter combination ensures the stability of the time series, and is suitable for the dynamic analysis of facial image features and device sensor features, improving the system's anti-interference ability to short-term fluctuations. In the process of modal mutual information gain analysis, the mutual information metric can effectively evaluate the correlation strength between different modal feature flows, avoiding relying solely on traditional linear correlation analysis. The number of discretization intervals is set to 20, so that the data can be estimated at a finer granularity when calculating mutual information, thereby improving the accuracy of information gain calculation. Through this analysis, the characteristic modes with the most information contribution can be screened out, providing a basis for subsequent cross-modal fusion. Hierarchical clustering analysis is performed based on the modal mutual information gain matrix, and the clustering threshold is set to 0.6, so that highly correlated modes can be aggregated, while ensuring that low-correlation modes will not be merged incorrectly. Hierarchical clustering can explore potential feature correlation structures while retaining modal independence, providing a basis for modal correlation modeling. Through the final generated modal correlation matrix, the interaction mode of multimodal feature flows can be effectively optimized, the rationality of data fusion can be improved, and the adaptability and recognition accuracy of the model in complex environments can be improved.
[0039] Optionally, the adjustment of the modal feature flow priority described in step S4 is specifically:
[0040] The weight distribution of each modal characteristic flow is extracted from the modal relationship matrix and normalized to obtain a standardized modal weight set;
[0041] The modal time dependency is evaluated based on the standardized modal weight set, and the temporal influence of the modal characteristic flow is modeled to obtain the modal temporal weight matrix;
[0042] Perform local influence analysis on the modal timing weight matrix, evaluate the feature flow priority score based on the local influence, and obtain the initial modal feature flow priority ranking;
[0043] Combine the modal correlation matrix to perform dynamic optimization on the priority ranking of the initial modal feature stream, and generate an optimized priority ranking of the modal feature stream;
[0044] Based on the optimized modal feature stream ranking, perform normalization adjustment on the modal feature stream to obtain an adaptive feature stream.
[0045] The present invention extracts the weight distribution of each modal feature stream from the modal relationship matrix and performs normalization processing to make different modal features consistent in the numerical scale, ensuring the stability of subsequent calculations. On this basis, based on the standardized modal weight set, evaluate the time dependence of each feature stream, and establish a time series influence degree model to describe the contribution change of each modal feature stream at different time points, thereby constructing a modal time series weight matrix. This matrix can accurately express the importance change of each modal feature in time series, making the system have stronger adaptability in a time-varying environment. Subsequently, perform local influence degree analysis on the modal time series weight matrix, evaluate its contribution degree in different scenarios by quantifying the influence range of each feature stream, and calculate the priority score accordingly to generate the initial priority ranking of the modal feature stream. The initial ranking can reflect the importance of each feature stream, but still needs to be dynamically optimized in combination with the modal correlation matrix to ensure that feature fusion can take into account the information interaction between different modalities and improve the overall recognition effect. The optimized priority ranking of the modal feature stream can more accurately adjust the contributions of different features, enabling key modalities to play a greater role at the appropriate time, and suppressing the influence of non-key modalities reasonably. Finally, based on the optimized ranking, perform normalization adjustment on the modal feature stream to ensure that all modal feature streams are in the optimal fusion state at different time points and in different environments, thereby generating an adaptive feature stream. The adaptive feature stream can dynamically adjust the weights of modal features according to the environment, avoid the information redundancy problem brought by traditional static feature fusion methods, improve the utilization rate and effectiveness of data, and ultimately enhance the stability and recognition accuracy of the face recognition system in complex scenarios.
[0046] Optionally, the dynamic optimization is specifically as follows:
[0047] According to the initial priority ranking of the modal feature stream, perform division to obtain a high-priority modal feature stream and a low-priority modal feature stream;
[0048] Calculate the local correlation distribution of the low-priority modal feature stream in the modal correlation matrix. If the Pearson correlation coefficient with the high-priority modal feature stream is greater than or equal to 0.8, then reduce the downweight amplitude of the low-priority modal feature stream and maintain the contribution degree, thereby obtaining a high-correlation low-priority downweight ranking; if the Pearson correlation coefficient with the high-priority modal feature stream is less than 0.3, then perform exponential decay dynamic downweight on the reduced low-priority modal feature stream, thereby obtaining a low-correlation low-priority downweight ranking;
[0049] Calculate the global contribution of the high-priority modal feature flow in the modal correlation matrix. If the global contribution is greater than 0.8, perform Lagrangian fusion weight improvement to obtain the high-priority weighted ranking.
[0050] Perform sorting coupling on the high-correlation low-priority deweighted ranking, low-correlation low-priority deweighted ranking, and high-priority weighted ranking to obtain the optimized modal feature flow priority ranking.
[0051] The present invention divides according to the priority sorting of the initial modal feature stream, obtaining a high-priority modal feature stream and a low-priority modal feature stream. This division can ensure the hierarchy of information utilization and ensure that key features act on subsequent analysis first. When calculating the local correlation distribution of the low-priority modal feature stream in the modal correlation matrix, if the Pearson correlation coefficient between it and the high-priority modal feature stream is greater than or equal to 0.8, it means that although the overall contribution of this low-priority feature is small, there is a strong interaction with the key features in the local range. Therefore, reduce its downweighting amplitude and maintain the contribution degree, so that these features can still play a role under appropriate circumstances. The selection of the correlation threshold of 0.8 ensures that this type of feature has a strong linear correlation, thus preventing irrelevant features from being wrongly assigned high weights. When the Pearson correlation coefficient is less than 0.3, it indicates that the connection between the low-priority feature and the core feature is weak. Directly retaining it may lead to information redundancy or noise interference. Therefore, an exponential decay dynamic downweighting method is used for suppression, so that these features are gradually weakened during the optimization process, avoiding negative impacts on the final feature fusion. The exponential decay method is more in line with the actual situation of feature influence than the linear downweighting, making the role of low-correlation features show a rapid downward trend and further enhancing the robustness of the system to noise features. To further optimize the contribution degree of the high-priority modal feature stream, calculate its global contribution degree in the modal correlation matrix and set the threshold of 0.8 as the improvement standard. The selection of this threshold ensures that only the key features with a relatively high overall contribution degree are optimized in weight, rather than blindly increasing the influence of all high-priority features. For high-priority features with a contribution degree greater than 0.8, adopt the Lagrangian fusion weight improvement strategy to make them occupy a more important position in the final feature sorting, thus enhancing the global optimization effect of feature fusion. The Lagrangian weight adjustment method can effectively improve the self-adaptability of the weight, ensuring that key features can obtain greater influence without losing balance. Finally, perform sorting coupling on the high-correlation low-priority downweighting sorting, low-correlation low-priority downweighting sorting, and high-priority weighting sorting to ensure that the optimization strategies of different-priority feature streams can work together and improve the sorting rationality of the modal feature stream. Through this optimization mechanism, the system can flexibly adjust the weights of each modal feature according to the local correlation, global contribution degree, and dynamic change of the features, so that the final feature fusion can retain key information and avoid interference from invalid information, thereby improving the accuracy and stability of the overall recognition system.
[0052] Optionally, the face recognition time sequence adaptability modeling described in step S4 is specifically:
[0053] Perform time segmentation processing on the optimized modal feature stream priority sorting, set the time window to 15s for sliding feature extraction, and calculate the feature statistics within each time window to obtain a time-segmented feature mapping matrix;
[0054] Calculate the feature change rate between different time windows based on the time-segmented feature mapping matrix, and conduct feature change trend analysis to obtain the time-series feature change weight matrix;
[0055] Combine the preset face recognition accuracy index to screen the face recognition accuracy impact factors of the time-series feature change weight matrix, and obtain the face recognition accuracy impact feature stream;
[0056] Conduct hierarchical clustering analysis on the face recognition accuracy impact feature stream, identify the high-time-dependence feature stream, and assign a time decay factor [0.9, 0.99] to the high-time-dependence feature stream to generate the accuracy impact weight adjustment matrix;
[0057] Based on the accuracy impact weight adjustment matrix and the time-series feature change weight matrix, conduct cross-time-window feature correlation modeling to obtain the face recognition time-series adaptation model.
[0058] The present invention performs sliding feature extraction by setting the time window to 15 s, which can ensure capturing the feature change trend in a relatively short time, while avoiding the influence of local noise caused by too short a time window or the problem of feature blurring caused by too long a time window. The window length of 15 s ensures the stability of feature statistics and is suitable for the rhythm characteristics of face features changing over time, enabling the time-segmented feature mapping matrix to more accurately depict the feature distribution in different time periods. On this basis, the feature change rate between different time windows is calculated, and the feature change trend is analyzed to obtain the temporal feature change weight matrix. The construction of this matrix can effectively capture the feature fluctuations across time windows and provide data support for subsequent analysis. Modeling the feature change trend can avoid the limitations of single-time-window features, ensure that the system can make predictive adjustments based on historical information, and enhance the adaptability of the recognition system to time dynamics. Combining the preset face recognition accuracy index to screen the temporal feature change weight matrix, extracting the feature streams that have a greater impact on the recognition accuracy, thereby removing the features that contribute less or are irrelevant to the recognition effect, improving the calculation efficiency, and enhancing the generalization ability of the model. This screening process can prevent the redundant influence of low-contribution features and ensure that the system focuses on the key feature streams that truly affect the recognition accuracy. To further optimize the time dynamic adaptability of the recognition system, hierarchical clustering analysis is performed on the screened feature streams that affect the face recognition accuracy to identify the highly time-dependent feature streams. The highly time-dependent feature streams have a strong continuous influence in the temporal change. Therefore, by assigning a time decay factor [0.9, 0.99] to them, the influence of the features can gradually decay in the time dimension without affecting the overall trend due to short-term fluctuations. The decay factor is set between 0.9 and 0.99, which can ensure a smooth decay process and not overly weaken the historical information, thus taking into account the effectiveness and adaptability of historical information, enabling the system to maintain a high feature consistency in a short time and not accumulate outdated information over a long time span. Based on the accuracy influence weight adjustment matrix and the temporal feature change weight matrix, cross-time-window feature correlation modeling is performed to construct a face recognition temporal adaptation model. This model can fully consider the influence of time factors on the face feature streams and adaptively adjust the weights of the features, enabling the system to maintain a high recognition accuracy in a dynamic environment. The cross-time-window modeling method makes up for the limitations of traditional single-time-point feature analysis and enables the system to work more stably in complex time-changing scenarios.
[0059] Optionally, this specification also provides a generation system for a face recognition fusion model, which is used to execute the generation method of the face recognition fusion model as described above. The generation system for the face recognition fusion model includes:
[0060] The multi-modal data fusion module is used to collect the historical face recognition image set and historical device sensing data through the face recognition device, and perform multi-modal data fusion integration on the historical face recognition image set and historical device sensing data to obtain a multi-source fusion data set;
[0061] The spatio-temporal feature integration module is used to perform multi-layer network spatio-temporal feature integration on the multi-source fusion data set to obtain a high-dimensional feature spectrum;
[0062] The modal contribution degree weighting module is used to perform modal information flow modeling based on the high-dimensional feature spectrum to obtain a modal correlation matrix, and perform adaptive modal contribution degree weighting on the modal correlation matrix to obtain a modal relationship matrix;
[0063] The recognition time sequence modeling module is used to adjust the priority of the modal feature flow according to the modal relationship matrix to obtain an adaptive feature flow; perform face recognition time sequence adaptability modeling based on the adaptive feature flow to obtain a face recognition time sequence adaptation model;
[0064] The model optimization module is used to perform model evaluation on the face recognition time sequence adaptation model to obtain model accuracy evaluation data; establish an accuracy optimization matrix based on the model accuracy evaluation data, and use the accuracy optimization matrix to optimize the face recognition time sequence adaptation model to obtain a face recognition fusion model.
[0065] The generation system of the face recognition fusion model of the present invention, which can implement any generation method of the face recognition fusion model of the present invention, is used as the medium for coordinating the operations and signal transmissions between various modules to complete the generation method of the face recognition fusion model. The internal modules of the system cooperate with each other, thereby improving the accuracy and real-time performance of face recognition. At the same time, the optimized model can more accurately adapt to the recognition tasks in different environments, ensuring high recognition performance even when the face features change greatly.
[0066] Optionally, the present specification also provides a face recognition device, including a face recognition device main body, a power supply unit, and an electrical control unit. Among them, the power supply unit is installed inside the face recognition device main body. Inside the face device, there is a sensor group and a camera group electrically connected to the power supply unit. The electrical control unit is electrically connected to the power supply unit and is used to charge the face recognition device main body and control the face recognition device main body. The electrical control unit is used to execute a face recognition fusion model generation method as described in any one of the above. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objects, and advantages of the present invention will become more apparent:
[0068] Figure 1Schematic diagram of the steps of the method for generating the face recognition fusion model of the present invention;
[0069] Figure 2 Schematic diagram of the detailed steps of step S1 in the present invention;
[0070] Figure 3 Schematic diagram of the detailed steps of step S2 in the present invention;
[0071] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments
[0072] The technical method of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work belong to the scope of protection of the present invention.
[0073] In addition, the accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor methods and / or microcontroller methods.
[0074] It should be understood that although the terms "first", "second", etc. may be used here to describe each unit, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly the second unit can be called the first unit. The term "and / or" used here includes any and all combinations of one or more of the listed related items.
[0075] To achieve the above object, please refer to Figures 1 to 3 , the present invention provides a method for generating a face recognition fusion model, and the method includes the following steps:
[0076] Step S1: Collect a historical face recognition image set and historical device sensing data through a face recognition device, and perform multimodal data fusion integration on the historical face recognition image set and historical device sensing data to obtain a multi-source fusion data set;
[0077] In this embodiment, historical face recognition image sets and historical device sensing data are collected by a face recognition device. The collected historical face image sets include facial images under different time periods, different lighting conditions, angle changes, and expression changes, etc. The device sensing data includes sensor information such as temperature, humidity, acceleration, and pressure. These data are processed by a multi-modal data fusion integration technology, and a data alignment algorithm is used to synchronize the image sets and the sensing data in time. Specifically, the collected face images are preprocessed and then subjected to image enhancement (such as brightness and contrast adjustment) and noise removal processing to ensure the quality of the data set. The sensor data is standardized so that the data of different sensors have a unified scale, and temporal alignment is performed so that the image data and the sensor data at each moment can be effectively matched. The device sensing data is standardized to the interval of [0,1], and the accuracy of data alignment is 1ms. The size of the data set after fusion is 30 images and 5 sensor data values at each moment, forming a multi-modal data set that contains image features and environmental feature data related to face recognition, and forming a multi-modal data set for subsequent model training.
[0078] Step S2: Perform multi-layer network spatio-temporal feature integration on the multi-source fusion data set to obtain a high-dimensional feature spectrum;
[0079] In this embodiment, the multi-source fusion data set undergoes multi-layer network spatio-temporal feature integration. First, a convolutional neural network (CNN) is used to extract the spatial features of the face images. The convolutional layer of the network is set to 3 layers, the filter size is 3x3, the stride is 1, and the size of the max-pooling layer is 2x2. The ReLU activation function is used. Subsequently, a temporal model (such as LSTM) is used, with the number of LSTM layers set to 2 layers and the number of units in each layer being 128. A dropout layer is used to prevent overfitting to model the temporal dependence of the sensing data. During the processing, the image data undergoes multiple convolutional operations to obtain high-dimensional spatial features, and the sensor data, after being processed by a recurrent neural network (RNN), can capture the dynamic change features in time. For the fusion of spatio-temporal features, a joint learning strategy is used to obtain the influence weights of each feature at different time points through training. Finally, the spatial features of the face images and the temporal features of the sensing data are combined through a feature integration network, and the dimension of the features is extended from the input 32x32 image to a 1024-dimensional feature vector, obtaining a high-dimensional feature spectrum. This feature spectrum not only contains rich spatial information but also covers data features from environmental changes, and can provide more comprehensive inputs for subsequent face recognition.
[0080] Step S3: Perform modal information flow modeling based on the high-dimensional feature spectrum to obtain a modal correlation matrix, and perform adaptive modal contribution degree weighting on the modal correlation matrix to obtain a modal relationship matrix;
[0081] In this embodiment, based on the high-dimensional feature spectrum, modal information flow modeling is performed. First, by analyzing the correlation between different modalities, a modal correlation matrix is constructed. Each item in this matrix represents the degree of mutual influence between different modalities, and the Pearson correlation coefficient is used to measure the correlation between modalities. For example, if the correlation between image data and temperature and humidity sensing data is high, the correlation coefficient value can be set to 0.85. To make the model more flexible, the modal correlation matrix is further weighted by the adaptive modal contribution degree, and the influence of each modality is adjusted through the weighting coefficient. Specifically, a higher weight is assigned to the modality with a greater contribution degree, and vice versa, a lower weight is assigned. Finally, after the weighting process, a modal relationship matrix is obtained, which can reflect the comprehensive contribution degree of each modality and provide a basis for the subsequent adjustment of the feature flow priority. For example, if the Pearson correlation coefficient between the image modality and sensor data (such as temperature and humidity) is 0.85, then the corresponding value in the matrix is 0.85. The modal correlation matrix is in the following form: Next, through the adaptive modal contribution degree weighting, the elements in the matrix are weighted according to the importance of the modalities. For example, the weighting coefficient for the temperature and humidity sensor is 0.9, and the weighting coefficient for the image modality is 1.1. Finally, a modal relationship matrix is generated, which reflects the importance of the modalities and provides a basis for subsequent modeling.
[0082] Step S4: Adjust the priority of the modal feature flow according to the modal relationship matrix to obtain an adaptive feature flow; perform face recognition temporal adaptability modeling based on the adaptive feature flow to obtain a face recognition temporal adaptability model;
[0083] In this embodiment, according to the modal relationship matrix, the priority of the modal feature flow is adjusted. First, according to the contribution degree of each modality in the matrix, the preliminary priority of the modal feature flow is determined. For example, if the influence of face image data on the recognition result is greater than that of temperature and humidity data, its priority is higher. Then, based on this priority, temporal adaptability modeling is performed. This modeling process uses an adaptive algorithm to dynamically adjust the weights of each modality according to the changes in the real-time data stream in order to adapt to different environmental changes. For example, when the lighting conditions change, the priority of the image modality may be adjusted to be higher to improve the recognition accuracy. In this way, the model can perform efficient adaptive adjustment at different time periods and under different environmental conditions, improving the accuracy and real-time performance of face recognition. For example, the weight of the image modality is set to 0.6, and the weight of the temperature and humidity sensor data is 0.4. At this time, the image data has a higher priority in the temporal adaptability modeling. When the environmental lighting changes, to reduce the influence of environmental lighting, the weight of the image modality will dynamically increase to 0.8, while the weight of the temperature and humidity modality will decrease to 0.2. The form of the face recognition temporal adaptability model is specifically: Y t= softmax(W o H t ); where is the facial image feature stream at time t; is the device sensing feature stream at time t; W F and W S are the feature weight matrices of the image modality and the sensing modality respectively; A t is the adaptive attention weight vector at time t; H t is the hidden state of the LSTM; is the input feature after modality weighting; Y t is the probability distribution of the final recognized category; W o is the classification weight matrix; ⊙ is the element-wise multiplication; W A is the precision impact weight adjustment matrix; W R is the temporal feature change weight matrix.
[0084] Step S5: Evaluate the face recognition temporal adaptation model to obtain model precision evaluation data; establish a precision optimization matrix based on the model precision evaluation data, and use the precision optimization matrix to optimize the face recognition temporal adaptation model to obtain a face recognition fusion model.
[0085] In this embodiment, after the face recognition temporal adaptation model is established, model evaluation is performed. The specific evaluation methods include cross-validation and precision verification, and the precision evaluation data is obtained by comparing the model prediction results with the actual recognition results. For each cross-validation, the ratio of the training set to the test set is set to 80:20, and the model evaluation metrics include accuracy, precision, and recall, with the accuracy requirement being above 90%. For example, evaluation metrics such as recognition accuracy, misrecognition rate, and recall will be calculated. To further improve the model performance, a precision optimization matrix is established based on the evaluation results, and the model is optimized in combination with an optimization algorithm (such as a genetic algorithm or a gradient descent method). During the precision optimization process, the genetic algorithm is used to adjust the weights of each modality. The precision optimization matrix will include adjustment suggestions for each feature stream and weight parameter in the model. For example, first define the initial modality weight matrix as: W init = [w f w s = [0.6 0.4]; where w f represents the initial weight of the facial image feature stream. The value of this parameter determines the importance of the facial image features in the model. The initial value is 0.6, indicating that the model has a higher dependence on the facial image features initially. w sRepresents the initial weight of the device sensing feature stream. The value of this parameter is 0.4, indicating that the device sensing features have a relatively small impact on the model, but are still necessary to enhance the model's adaptability in complex environments. After model evaluation, the weights are adjusted according to the Precision Optimization Matrix (POM): Where Represents the optimized adjustment value for the facial image feature stream. This parameter is based on the model's evaluation metrics (such as accuracy, precision, etc.). By comparing the changes in model performance under different inputs, the weights of the facial image feature stream are dynamically adjusted. Assuming the adjustment value is +0.15, it means that when the facial image features contribute more to the recognition result, their weights are increased. Represents the optimized adjustment value for the device sensing feature stream. This parameter is also based on model evaluation. If the model depends too much or too little on device sensing features in certain situations, the impact of sensing data is optimized by adjusting this parameter. For example, if the sensing data contributes less to the recognition accuracy, the adjustment value may be negative, such as -0.15, to reduce the weight of the sensing features. The optimized modal weights are obtained: W opt = W init + POM = [0.75 0.25]; The optimized modal weight matrix is obtained by adjusting the initial modal weights through the Precision Optimization Matrix. For example, the initial weight matrix is W init = [0.6 0.4]. After optimization, by adding the optimization matrix POM = [+0.15 -0.15], the obtained optimized modal matrix W opt = [0.75 0.25]; If the model's performance degrades in a specific environment (such as low light), the POM will be further adjusted dynamically. The optimized recognition model can maintain a high accuracy rate in different scenarios. Different modal optimization schemes will be proposed in the optimization matrix. For example, the weight of the image modality is adjusted from 0.6 to 0.75, and at the same time, the weight of the sensor data is adjusted from 0.4 to 0.25. After optimization, the optimized face recognition fusion model is obtained, and its accuracy rate is increased by 4%. If some sensor data is found to have a small impact on the recognition accuracy, the optimization matrix will recommend reducing the weight of this data stream. After precision optimization, the model will generate a higher recognition accuracy, and finally, the optimized face recognition fusion model is obtained. This model can maintain a high accuracy rate in a changing environment and achieve fast response.
[0086] Optionally, step S1 is specifically as follows:
[0087] Step S11: Collect the historical face recognition image set and historical device sensing data through the face recognition device, and perform data preprocessing to obtain the historical face recognition image set to be analyzed and the historical device sensing data to be analyzed;
[0088] In this embodiment, historical face recognition image sets and historical device sensing data are collected by a face recognition device, and data preprocessing is first performed. For the face image sets, denoising processing is first carried out to remove interference factors in the images. The Gaussian blur algorithm is used to smooth the images to ensure that the image quality meets the requirements of subsequent processing. Then, histogram equalization processing is performed to enhance the image contrast and improve the effect of feature extraction. For the device sensing data, an interpolation method is used to fill in the missing data to ensure data integrity. At the same time, the z-score normalization technique is used to normalize the sensing data so that the mean of the data is 0 and the standard deviation is 1, enabling the data of each sensor to be compared and fused on the same scale. The obtained historical face recognition image sets and device sensing data sets can enter the subsequent analysis steps.
[0089] Step S12: Perform temporal synchronization and spatial feature mapping on the historical face recognition image set to be analyzed and the historical device sensing data to be analyzed to obtain a spatio-temporally aligned multimodal data set;
[0090] In this embodiment, temporal synchronization is performed on the historical face recognition image set to be analyzed and the historical device sensing data to ensure the temporal consistency between the image data and the sensor data. Specifically, for the face image data, each image records a timestamp, which can be compared with the timestamp of the sensor data. If there is a slight deviation in the timestamps, linear interpolation or spline interpolation methods are used for time alignment. For spatial feature mapping, based on a predefined spatial reference framework (such as a geographical location or a coordinate system), the sensor data and the image data are mapped to the same spatial dimension. During the spatial mapping process, principal component analysis (PCA) can be used for feature dimensionality reduction to ensure that the features of different modalities can be represented within the same spatial range, thereby generating a spatio-temporally aligned multimodal data set. This data set provides a unified spatio-temporal reference for subsequent feature extraction and fusion.
[0091] Step S13: Perform principal component feature selection based on the spatio-temporally aligned multimodal data set and perform weighted fusion on the principal component features to obtain a multimodal fusion data set;
[0092] In this embodiment, principal component analysis (PCA) is applied to perform feature selection on the spatio-temporally aligned multimodal dataset. For each data type (such as face image features, device sensing data features), PCA is used to reduce its dimension, retaining more than 95% of the principal component features to reduce noise and improve the efficiency of subsequent analysis. For example, the principal components of face images can represent key facial features (such as the positions of eyes, nose, facial contours, etc.), while the principal components of device sensing data can represent environmental change features (such as temperature, humidity, etc.). Subsequently, these principal component features are weighted and fused. The weighting coefficients are set based on the contribution degree of each modality. If image features have a greater impact on the recognition result in a specific situation, a higher weight can be assigned to the image features. For example, the weight of face image features is 0.7, and the weight of sensor data features is 0.3. The multimodal fusion dataset generated after weighted fusion provides rich information for subsequent data augmentation and expansion.
[0093] Step S14: Perform data augmentation on the multimodal fusion dataset to obtain a multi-source fusion dataset;
[0094] In this embodiment, data augmentation is performed on the multimodal fusion dataset to improve data diversity and the robustness of the model. Specifically, common data augmentation methods such as rotation, translation, and mirroring can be used to augment face images. For example, for each face image, the random rotation angle range is between ±30 degrees, and it can be randomly flipped horizontally or vertically; at the same time, the image can be randomly cropped to enhance the model's adaptability to different face angles and expressions. For device sensing data, the method of adding noise is used for augmentation. For example, the noise in the sensor data is added to the original data to simulate sensor errors and environmental interference. The augmented dataset can effectively improve the training effect of the model and enhance its adaptability to complex environments, thereby obtaining a multi-source fusion dataset.
[0095] Step S15: Perform Monte Carlo data expansion based on the multi-source fusion dataset to obtain a multi-source fusion dataset.
[0096] In this embodiment, Monte Carlo data expansion is performed based on the multi-source fusion dataset. The Monte Carlo method generates multiple data samples through random sampling to expand the existing dataset and enhance its diversity. First, a subset can be randomly selected from the existing multi-source fusion dataset, and new data points can be generated according to the data distribution of this subset. For example, in face image data, image sampling can be performed based on the color histogram and texture features of the image, while in sensor data, sampling can be performed according to the statistical characteristics of the data (such as mean, standard deviation, etc.). The expanded dataset can be used to train a more powerful model, help the model handle more possible input situations, and improve the generalization ability of the model. Suppose the temperature distribution of the device sensing data follows a normal distribution with a mean of 20°C and a standard deviation of 2°C, while the feature distribution of the face image data is a standard normal distribution with a mean of 0 and a standard deviation of 1. During the expansion process, the Monte Carlo method generates multiple new data points by randomly sampling from their respective data distributions, simulating scenarios under different environments and different device states. For example, 2000 temperature data points and 2000 image feature data points are generated, and finally an expanded multi-source fusion dataset is obtained. In this way, the diversity of the dataset is greatly increased, providing richer sample data for training the model, thereby improving the generalization ability of the model and its adaptability to abnormal situations.
[0097] Optionally, step S13 is specifically as follows:
[0098] Extract the high-level features of the face image and the sensing physical features from the spatio-temporally aligned multi-modal dataset, and perform standardization processing on the high-level features of the face image and the sensing physical features to obtain a standardized high-level face image feature set and a standardized sensing physical feature set;
[0099] In this embodiment, the high-level features of the face image are extracted from the spatio-temporally aligned multi-modal dataset. For this purpose, a convolutional neural network (CNN) is used for feature extraction, and a pre-trained ResNet-50 model is used to extract features from the face image. In specific operations, the image is scaled to 224x224 pixels and input into the ResNet-50 network to extract the high-level features therein, especially the detailed features of parts such as facial contours, eyes, nose, and mouth. In addition, the physical features of the sensor (such as temperature, humidity, acceleration, pressure, etc.) are also extracted and standardized. Suppose the temperature data range of the sensor is [18°C, 25°C], and the z-score standardization method is used to calculate the mean and standard deviation of this dataset, and then it is normalized. Through these two steps, a standardized "high-level face image feature set" and a "sensing physical feature set" are obtained, which respectively contain the facial features extracted by the CNN model and the standardized information of the sensor data, and are ready for subsequent processing.
[0100] Perform face expression feature pattern clustering based on the high-level feature set of the standardized face image, evaluate the feature contribution degree of the clustering result, and screen the principal component features according to the evaluation result of the feature contribution degree, so as to obtain the high-level feature set of the principal component face image;
[0101] In this embodiment, face expression feature pattern clustering is performed based on the high-level feature set of the standardized face image. Specifically, first, the K-means clustering algorithm is used to classify the high-level features of the face image, and the number of clusters is set to 5, that is, the images are divided into five different expression categories (such as smiling, serious, surprised, angry, calm). By calculating the cluster center of each category and evaluating the feature contribution degree of each cluster, for example, by calculating the variance or information gain of the features within the cluster to evaluate the contribution of each feature to the expression pattern. According to this contribution degree evaluation, select the features with higher contribution degrees, such as selecting the eye features that can best distinguish between smiling and angry expressions, and based on this result, screen out the principal component features, so as to obtain the "high-level feature set of the principal component face image". These features retain the components that can best represent the face expression information.
[0102] Perform principal component feature selection based on the standardized sensing physical feature set to obtain the principal component sensing physical feature set;
[0103] In this embodiment, principal component analysis (PCA) is performed on the standardized sensing physical feature set. Assume that the original sensor data set contains features such as temperature, humidity, and acceleration. Through PCA analysis, select the principal components that retain more than 90% of the variance. In the specific operation, the PCA algorithm first calculates the feature covariance matrix, then obtains the principal components through eigenvalue decomposition, and finally retains the first 3 principal components (assuming they respectively correspond to the change patterns of temperature, humidity, and acceleration) to obtain the "principal component sensing physical feature set". This feature set highlights the most critical part of the sensor data by reducing the dimension and removing redundant information, providing a basis for subsequent feature fusion.
[0104] Perform modal feature weighted fusion on the high-level feature set of the principal component face image and the principal component sensing physical feature set to obtain the multi-modal fusion feature set;
[0105] In this embodiment, the high-level feature set of the principal component face image and the principal component sensing physical feature set are subjected to modal feature weighted fusion. To achieve weighted fusion, first determine the importance of each modality (face image and sensor data) in the overall recognition task. Assume that the contribution degree of the face image to the final decision is 0.7, and the contribution degree of the sensor data is 0.3. Then, during the fusion process of the face feature set and the sensing feature set, weight the two types of features according to the weight coefficients 0.7 and 0.3 respectively. Specifically, assume that the high-level feature set of the face image F = F1, F2, F3, …, F n, the sensor data feature set S = S1, S2, S3, …, S m , the feature set after weighted fusion is represented as: F fusion = 0.7·F + 0.3·S; the “multi-modal fusion feature set” obtained after weighted fusion contains the comprehensive features of facial expressions and environmental data, making full use of the advantages of both modalities.
[0106] Perform non-linear optimization on the multi-modal fusion feature set to obtain the multi-modal fusion data set.
[0107] In this embodiment, non-linear optimization processing is performed on the obtained multi-modal fusion feature set to improve the identification ability of the model. Specifically, the gradient boosting algorithm can be used to optimize the fusion feature set. The model transforms the features through non-linear functions (such as sigmoid, ReLU, etc.) to better adapt to complex pattern recognition tasks. Assuming that the non-linear optimization function used is the activation function ReLU, the obtained multi-modal fusion features are input into a trained classification model, and its parameters are optimized through training, so that the final model can still maintain a high accuracy in the face of complex environments such as different illuminations, expression changes, and sensor noises. After this non-linear optimization, the “multi-modal fusion data set” is obtained, which can better handle the complex changes and challenges in practical applications.
[0108] Optionally, step S2 is specifically:
[0109] Step S21: Use a preset multi-scale convolutional network to extract the image detail features and sensing trend features in the multi-source fusion data set, so as to obtain the multi-scale spatio-temporal feature set;
[0110] In this embodiment, a multi-scale convolutional neural network (CNN) is used to extract the image detail features and sensing trend features in the multi-source fusion data set. To achieve this goal, a multi-scale convolutional network containing multiple convolutional layers is first adopted, and each convolutional layer uses different convolutional kernel sizes (for example, 3x3, 5x5, 7x7) to extract features at different scales. The image detail features are mainly extracted by smaller convolutional kernels (such as 3x3) to obtain local detail information, such as facial contours, eyes, mouth, etc.; while the sensing trend features (such as temperature, humidity, pressure change trends) are extracted by larger convolutional kernels (such as 7x7) to obtain long-term dependencies. This process will generate feature maps at multiple scales, corresponding to details, local information, and trend changes respectively. Taking a face image as an example, through the multi-scale convolutional network, the obtained multi-scale spatio-temporal feature set will include two types of information: detail features (such as facial structure) and trend features (such as image brightness changes). These features can provide rich spatio-temporal information for subsequent optimization and modeling processes.
[0111] Step S22: Perform data feature adversarial optimization on the multi-scale spatio-temporal feature set to obtain a constrained optimized feature set;
[0112] In this embodiment, data feature adversarial optimization is performed on the multi-scale spatio-temporal feature set. The main purpose is to enhance the robustness of the feature set through adversarial training, especially its performance in the face of noise and interference. Specifically, the architecture of a generative adversarial network (GAN) can be adopted, and a generator and a discriminator are constructed during the training process. The task of the generator is to generate "fake" features similar to the original spatio-temporal features, while the discriminator needs to distinguish which are real features and which are generated fake features. In this way, the features are forced to make more significant differentiations during the training process, making them more adaptable to different environmental changes (such as illumination, sensor noise, etc.). During this process, the goal of feature adversarial optimization is to improve the expression ability of features in a changing environment, such as the recognition accuracy of image detail information under illumination changes and the sensing trend under temperature changes. The optimized feature set is called the "constrained optimized feature set", which can effectively cope with different data interferences and improve the robustness of the model.
[0113] Step S23: Perform dynamic spatio-temporal modeling on the constrained optimized feature set through a preset spatio-temporal relationship adaptive network, adaptively adjust the contribution degrees of features at different time points and spatial regions, and obtain an adaptive spatio-temporal feature set;
[0114] In this embodiment, dynamic spatio-temporal modeling is performed on the constrained optimized feature set through a preset spatio-temporal relationship adaptive network. This network is designed as a spatio-temporal neural network (STNN) that can process features that change dynamically in time and space. Specifically, the network architecture can include a temporal convolutional network (TCN) and a spatial adaptive network (SAN). The TCN module is used to capture the long-term dependencies of time series data, while the SAN module adaptively adjusts the feature contribution degrees of spatial regions, especially when the contributions of different regions (such as different parts of an image or sensor acquisition points) to the overall feature are different. For example, in some cases, sensor data in a specific region (such as the region where a temperature sensor is close to a heat source) may have a greater impact on the model output, while in other regions, the contribution is smaller. The spatio-temporal relationship adaptive network can dynamically adjust the contribution degrees of each feature at different time points and spatial regions according to the actual scenario, and finally generate an "adaptive spatio-temporal feature set".
[0115] Step S24: Convert the adaptive spatio-temporal feature set into a graph structure to obtain a spatio-temporal graph feature set;
[0116] In this embodiment, a graph convolutional network (GCN) is used to construct a graph structure. Each feature point is regarded as a node in the graph, and the relationship between nodes is represented by an adjacency matrix. In graph convolution, spatio-temporal features will be embedded into the nodes of the graph, and the edges of the graph represent the relationships between nodes. For example, in a spatio-temporal graph based on a sensor network, the edges between sensor nodes may represent distance or similarity. To better express these spatio-temporal relationships, a self-attention mechanism can be used to enhance the model's attention to key spatio-temporal nodes, enabling the model to focus more on important regions and time periods, thereby obtaining a more effective spatio-temporal graph feature set. These graph features will subsequently provide a basis for the generation of high-dimensional feature spectra.
[0117] Step S25: Perform modality fusion and feature integration based on the spatio-temporal feature set to obtain a high-dimensional feature spectrum.
[0118] In this embodiment, modality fusion and feature integration are performed based on the spatio-temporal graph feature set to obtain a "high-dimensional feature spectrum". For this purpose, a multi-modal fusion algorithm can be adopted to combine features from different modalities (such as images, sensor data, etc.) through a weight learning strategy. Specifically, during implementation, first, the features of different modalities are weighted and fused through an adaptive weighting strategy, such as using weighted average or the depth fusion method in a convolutional neural network (CNN). Then, principal component analysis (PCA) is used to reduce the dimension of the fused high-dimensional features to obtain the final high-dimensional feature spectrum. This high-dimensional feature spectrum contains a comprehensive representation of multi-modal information and retains the key information that can best describe the target task. For example, assume that in multi-modal fusion, the weight of the image modality is 0.6 and the weight of the sensor data modality is 0.4. The fused feature set can more comprehensively express the changes in the scene, including the comprehensive features of facial expressions and environmental perception.
[0119] Optionally, the modality information flow modeling described in step S3 is specifically:
[0120] Extract the face image feature flow and device sensing feature flow in the high-dimensional feature spectrum;
[0121] In this embodiment, features in the time series are extracted through a pre-trained neural network or statistical analysis method. The face image feature flow focuses on extracting the spatio-temporal changes of facial expressions and face features, while the device sensing feature flow focuses on the changes in environmental data or physical characteristics. For example, for a multi-modal system (synchronization of face recognition and sensor data), the extracted feature flows may include the changes in facial expressions and the trend of temperature changes. This process generates two feature flow sets: the face image feature flow and the device sensing feature flow, and subsequent analysis will be based on these two feature flow sets.
[0122] Evaluate the difference degree of the characteristic mode distribution of the high-dimensional feature spectrum, set the distribution deviation threshold to 0.15, and standardize the features in the face image feature stream and the device sensing feature stream according to the distribution deviation threshold to obtain a standardized modal feature stream set;
[0123] In this embodiment, the difference degree of the characteristic mode distribution is evaluated, and the feature standardization process is carried out according to the distribution deviation threshold. First, for the face image feature stream and the device sensing feature stream, a distribution difference degree evaluation method, such as the Kullback-Leibler divergence (KL divergence), is used to measure the distribution difference degree between the two. If the difference degree is greater than the set deviation threshold (0.15), it indicates that the distribution difference between the two feature streams is large and needs to be standardized. In specific implementation, standardization methods such as Z-score standardization (mean normalization) or Min-Max standardization (scaled to the [0,1] interval) can be used to make the two types of features comparable on the same scale. For example, if the mean of the face image feature stream is 100 and the standard deviation is 15, and the mean of the sensing data is 50 and the standard deviation is 10, after standardization, the standardized feature streams of the two will have the same distribution range and numerical scale, ensuring that the subsequent analysis can effectively fuse these two types of features.
[0124] Set the autoregressive order to 3, the moving average order to 2, calculate the smoothing window length to be 10s, and perform autoregressive moving average trend smoothing on the standardized modal feature stream set to obtain a smoothed modal time series;
[0125] In this embodiment, autoregressive moving average trend smoothing is performed on the standardized modal feature stream set. Specifically, set the autoregressive order to 3, the moving average order to 2, and the smoothing window length to 10 seconds, and use a combination of the autoregressive (AR) model and the moving average (MA) model (ARMA model) to smooth the feature stream. The autoregressive order 3 means using the current value and the values of the previous three time points to predict future values; the moving average order 2 means using the current value and the residuals of the previous two time points to calculate the predicted value. Under this model, the value at each time point is obtained by weighted summation of the previous values. For a given sensor data stream, such as temperature data, after ARMA smoothing, sudden change values and short-term fluctuations can be eliminated, thereby obtaining a smoothed trend time series and reducing noise interference. After the smoothing process, the obtained smoothed modal time series can provide a more stable feature stream for subsequent analysis and modeling.
[0126] Perform cross-modal information coupling analysis based on the smoothed modal time series, calculate the mutual information gain of each modal feature stream, and set the number of discretization intervals for information gain calculation to 20, thereby constructing a modal mutual information gain matrix;
[0127] In this embodiment, cross-modal information coupling analysis is performed based on smooth modal time series, and the mutual information gain of the modal feature flow is calculated. First, the number of discretization intervals is set to 20, and the mutual information gain calculation method is used to evaluate the information correlation degree between different modalities. The mutual information gain is a measure of the dependence relationship between two variables. Specifically, when calculating, the continuous modal feature flow is discretized into 20 intervals, and then the mutual information between each pair of modalities is calculated to evaluate the degree of information sharing between them. For example, if there is a strong mutual information gain between the facial image feature flow and the sensor data (such as temperature change), it indicates that there is a high correlation between the two, and complementary information can be provided. The calculated mutual information gain matrix will reflect the mutual relationship between different modalities, thereby providing a basis for subsequent modal clustering and modeling. The elements of the mutual information gain matrix represent the mutual information between each pair of modal feature flows. The form of the matrix is as follows: where MI(f i , f j ) represents the mutual information amount between modality f i and modality f j . The calculation method is to discretize each modal feature flow and calculate the conditional entropy and joint entropy of each pair of modal feature flows, and then calculate the mutual information. Assume that after discretization, each modal feature flow f i and f j become the discretized probability distributions P(f i ) and P(f j ), then the calculation formula for mutual information is: The mutual information gain matrix will reflect the degree of association between each pair of modal feature flows. When the mutual information gain value between a pair of modalities is relatively high, it means that they share more information and there is a strong correlation between them.
[0128] Set the clustering threshold to 0.6, perform hierarchical clustering based on the modal mutual information gain matrix, and perform modal correlation modeling based on the hierarchical clustering result to generate a modal association matrix.
[0129] In this embodiment, a clustering algorithm is used for hierarchical clustering, and modal correlation modeling is performed based on the hierarchical clustering results. First, a clustering threshold of 0.6 is set, and the modal feature streams are grouped using a hierarchical clustering algorithm (such as agglomerative hierarchical clustering). The key to the clustering process is to determine the similarity between different modalities by calculating the similarity values in the mutual information gain matrix. Setting the clustering threshold to 0.6 means that only when the similarity between modalities exceeds this threshold will they be grouped into the same class. After clustering, according to the modal feature streams of each class, a modal correlation matrix is further constructed to clarify the correlation degree and influence between different modalities. For example, if a certain clustering group contains two modalities, "facial image expression" and "temperature sensor", and their correlation degree is high, it indicates that these two modalities are strongly correlated in this context. On this basis, the modal correlation matrix will reflect the mutual influence and correlation between different modalities, providing a basis for subsequent decision-making and model optimization. The modal correlation matrix is used to represent the strength of the correlation between modal feature streams. Assuming there are N modal feature streams, the modal correlation matrix can be obtained after hierarchical clustering, and the form of the matrix is as follows: Where C(f i ,f j ) represents the correlation metric value between modality f i and modality f j , and a correlation coefficient (such as Pearson correlation coefficient or Spearman correlation coefficient) can be used to calculate it: Where Cov(f i ,f j ) represents the covariance between modality f i and modality f j , and σ(f i ) and σ(f j ) are the standard deviations of modality f i and modality f j respectively.
[0130] Optionally, the adjustment of the priority of the modal feature streams described in step S4 is specifically as follows:
[0131] Extract the weight distribution of each modal feature stream from the modal relationship matrix and perform normalization to obtain a standardized modal weight set;
[0132] In this embodiment, the correlation information between different modal feature streams has been obtained from the modal correlation matrix C. For example, modality f1 is a facial image feature stream, modality f2 is a temperature sensor data stream, and modality f3 is a humidity sensor data stream. By summing each column in the modal relationship matrix, the overall contribution degree of each modality is obtained. Then, the weights are normalized using the following formula: W i is the weight of modality f i , representing modality fi Contribution to the entire system. The standardized modal weight set W is obtained through normalization. Assume the normalized weight set is: W = (0.4 0.35 0.25); this means the weight of f1 (face image feature stream) is 0.4, the weight of f2 (temperature sensor data stream) is 0.35, and the weight of f3 (humidity sensor data stream) is 0.25.
[0133] Evaluate the modal time dependence based on the standardized modal weight set, and model the temporal influence degree of the modal feature stream to obtain the modal temporal weight matrix;
[0134] In this embodiment, based on the obtained standardized modal weight set W, the time dependence of the modal is evaluated based on the time series data of the modal feature stream. Through time series analysis methods (such as autoregressive model AR, moving average MA, or long short-term memory network LSTM), the influence degree of each modal feature stream at different time sequences is calculated. For example, set the temporal weight w1(t) of modal f1 (face image feature stream) at time point t, the temporal weight w2(t) of modal f2 (temperature sensor data stream) at time point t, and the temporal weight w3(t) of modal f3 (temperature sensor data stream) at time point t. Through temporal modeling, the temporal weight matrix is obtained: Assume t1, t2,..., t N are different time points, and the matrix T shows the temporal influence degree of each modal at different time points. According to this matrix, the dynamic changes of the temporal feature stream can be captured and further used for optimization.
[0135] Conduct a local influence degree analysis on the modal temporal weight matrix, and evaluate the priority score of the feature stream according to the local influence degree to obtain the initial modal feature stream priority ranking;
[0136] In this embodiment, when conducting a local influence degree analysis on the temporal weight matrix T, the weighted average method is adopted, and the temporal influence degree of each modal is evaluated according to the weight of each modal at a specific time point. For example, set the local influence degree as the weighted average weight of each modal within a continuous time period, and the calculation is as follows: where, L i is the local influence degree of modal f i , N is the total number of time periods, and T(f i , t) is the weight of modal f i at time point t. By calculating the local influence degree of all modals and comparing them with other modals, the priority ranking of the initial modal feature stream is obtained. For example, assume the following priority scores are obtained after calculation: Priority ranking = {f1, f2, f3}; this means that f1 (face image feature stream) has the highest influence degree in terms of time sequence, so its priority is the highest.
[0137] Combine the modal relevance matrix to dynamically optimize the initial modal feature stream priority ranking and generate an optimized modal feature stream priority ranking;
[0138] In this embodiment, the priority ranking of the initial modal feature stream is dynamically optimized by combining the modal relevance matrix C. Specifically, we adjust the ranking results according to the correlation and mutual influence degree between modalities. For example, in the relevance matrix, if there is a high correlation between modality f1 and f2, their priority order can be adjusted to ensure that their feature streams receive more attention during the fusion process. The optimization process can be carried out through algorithms (such as genetic algorithms or particle swarm optimization), and finally an optimized priority ranking is generated.
[0139] Based on the optimized modal feature stream ranking, perform normalization adjustment on the modal feature stream to obtain an adaptive feature stream.
[0140] In this embodiment, after optimizing the modal feature stream ranking, a normalization adjustment method is used to adjust each modal feature stream to make its contribution to the system more balanced at different time points and spatial regions. Specifically, a normalization factor k i is set to adjust the weight of each modal feature stream so that the total weight sum of all modalities is 1: where k i is the normalization factor of modality f i , and W i is the weight of modality f i . The adjusted modal feature stream obtains a more reasonable weight distribution through adaptive adjustment to form the final adaptive feature stream.
[0141] Optionally, the dynamic optimization is specifically:
[0142] According to the initial modal feature stream priority ranking, perform partitioning to obtain a high-priority modal feature stream and a low-priority modal feature stream;
[0143] In this embodiment, after calculating the optimized modal feature stream priority ranking, the ranking result is partitioned, and a priority partitioning threshold T p = 0.6 is set, that is, when the local influence degree L i of the feature stream is greater than or equal to T p , this modal feature stream is classified as a high-priority modal feature stream. If L i is less than T p , it is classified as a low-priority modal feature stream. Set the modal feature stream set as M = {m1, m2,..., m n}, where the local influence degree set of the feature stream is L = {l1, l2,..., l n}, then the high-priority modal feature stream set and the low-priority modal feature stream set are respectively expressed as: MH = {m i | L i ≥ T p}, M L = {m j | L i < T p}};
[0144] Calculate the local correlation distribution of the low - priority modal feature stream in the modal correlation matrix. If the Pearson correlation coefficient with the high - priority modal feature stream is greater than or equal to 0.8, then reduce the down - weighting amplitude of the low - priority modal feature stream and maintain the contribution degree, so as to obtain the down - weighting sorting of low - priority with high correlation; if the Pearson correlation coefficient with the high - priority modal feature stream is less than 0.3, then perform exponential decay dynamic down - weighting on the low - priority modal feature stream, so as to obtain the down - weighting sorting of low - priority with low correlation;
[0145] In this embodiment, based on the modal correlation matrix C, calculate the Pearson correlation coefficient between the low - priority modal feature stream M L and the high - priority modal feature stream M H : where ρ ij represents the correlation between m i and m j in the time series. When ρ ij ≥ 0.8, it is determined that the low - priority modal feature stream and the high - priority modal feature stream have a strong correlation. Therefore, reduce its down - weighting amplitude to keep its contribution degree above 0.7. The adjusted feature stream weight is calculated as follows: W L is the original feature stream weight, and then perform a secondary sorting on the evaluation feature stream priority scores according to the adjusted feature stream weight to obtain the down - weighting sorting of low - priority with high correlation. If ρ ij < 0.3, then the low - priority modal feature stream and the high - priority modal feature stream have a low correlation and need to be dynamically down - weighted by exponential decay. Then perform a secondary sorting on the evaluation feature stream priority scores according to the adjusted feature stream weight to obtain the down - weighting sorting of low - priority with low correlation. Its weight adjustment function is set as: where W L is the original feature stream weight, λ = 0.05 is the exponential decay factor, and t is the time step of the feature stream in the observation window.
[0146] Calculate the global contribution degree of the high - priority modal feature stream in the modal correlation matrix. If the global contribution degree is greater than 0.8, then perform Lagrangian fusion weight increase to obtain the high - priority weighted sorting;
[0147] In this embodiment, for the high - priority modal feature stream MH Perform global contribution evaluation. Define the global contribution C i as where R ij represents the correlation coefficient between modality i and modality j, and W j is the current weight of modality j. If C i > 0.8, then the high local influence modality feature stream makes a greater contribution to the overall system. It is necessary to optimize the weight through Lagrangian fusion to enhance its influence, and then perform secondary priority sorting to obtain a high-priority weighted sorting. Set the optimization weight adjustment formula as follows: where λ L = 0.1 is the Lagrangian adjustment coefficient, and L(W) is the optimization objective function, which is set as: The optimized weight can enhance the global contribution of the high local influence modality feature stream, thereby improving the rationality of the feature stream sorting.
[0148] Couple the sorting of high-correlation low-priority deweighting, low-correlation low-priority deweighting, and high-priority weighting to obtain the optimized modality feature stream priority sorting.
[0149] In this embodiment, after completing the deweighting of the low local influence modality feature stream and the weighted optimization of the high local influence modality feature stream, it is necessary to couple the three types of sorting to finally form the optimized modality feature stream local influence sorting. A weighted fusion strategy can be used for sorting optimization, and the sorting is performed in descending order according to the local influence to obtain the finally optimized modality feature stream local influence sorting.
[0150] Optionally, the face recognition time series adaptability modeling described in step S4 is specifically:
[0151] Perform time segmentation processing on the optimized modality feature stream priority sorting. Set the time window to 15s for sliding feature extraction, and calculate the feature statistics within each time window to obtain the time-segmented feature mapping matrix;
[0152] In this embodiment, during the process of optimizing the modality feature stream priority sorting, in order to improve the stability and calculation efficiency of the time series features, a sliding time window method is used for time segmentation processing. Set the time window length to 15s and the sliding step to 5s to ensure the continuity of the time series data and the dynamics of the feature stream. Assume that the modality feature stream set is represented as: M = {m1, m2,..., m n}; within each time window, calculate the statistical features of each modality feature stream, including mean, standard deviation, skewness, kurtosis, maximum value, minimum value, etc. The statistical features constitute the time-segmented feature mapping matrix: Among them, rows represent different statistical features, such as mean, standard deviation, etc., and columns represent different modal feature stream data. This matrix will be used in subsequent steps to analyze the change trend of the feature stream and perform optimization processing in combination with the time dimension.
[0153] Calculate the feature change rate between different time windows based on the time-segmented feature mapping matrix, and conduct feature change trend analysis to obtain the time-series feature change weight matrix;
[0154] In this embodiment, in order to evaluate the change of the feature stream between different time windows, based on the time-segmented feature mapping matrix F k , calculate the feature change rate between time window t k and t k+1 . Set the time interval to 15s, and define the feature change rate matrix where r ij represents the change rate of the i-th statistical feature to the j-th modal feature stream within the time window t k . For feature streams with drastic changes, higher dynamic adjustment weights will be assigned. Further, based on R k calculate the feature change trend and establish the time-series feature change weight matrix where w rij is comprehensively calculated from the feature change rate and the time window statistical features.
[0155] Combine the preset face recognition accuracy index to screen the face recognition accuracy impact factors of the time-series feature change weight matrix to obtain the face recognition accuracy impact feature stream;
[0156] In this embodiment, in order to ensure the adaptability of the feature stream sorting optimization to the face recognition task, in combination with the face recognition accuracy index, screen the impact factors of the time-series feature change weight matrix W R to identify the feature streams that contribute more to the recognition accuracy. Set a contribution threshold of 0.75, screen the feature stream set M P , and construct the face recognition accuracy impact feature matrix: where p ij represents the feature stream that has an important impact on the face recognition accuracy after screening. For feature streams below the threshold, lower priority weights will be assigned.
[0157] Conduct hierarchical clustering analysis on the face recognition accuracy impact feature stream, identify the high-time-dependence feature stream, and assign a time decay factor [0.9, 0.99] to the high-time-dependence feature stream to generate the accuracy impact weight adjustment matrix;
[0158] In this embodiment, through the hierarchical clustering method for M PPerform time-dependent analysis to identify high-time-dependency feature streams that vary stably within multiple time windows. For these feature streams, set the time decay factor λ i , and form the precision impact weight adjustment matrix W A : where the time decay factor λ i is set to a range of [0.9, 0.99], which is used to dynamically adjust the feature streams with strong time-dependency, enabling them to maintain relatively high weights on long time scales while reducing the impact of short-term sharp fluctuations on the recognition results.
[0159] Based on the precision impact weight adjustment matrix and the temporal feature change weight matrix, conduct cross-time-window feature correlation modeling to obtain the face recognition temporal adaptation model.
[0160] In this embodiment, based on the precision impact weight adjustment matrix W A and the temporal feature change weight matrix W R , optimize the feature stream fusion method under different time windows on the basis of the face recognition temporal adaptation model (FRTAM), enabling the model to adaptively adjust the modality fusion weights and improve the time robustness and precision stability of face recognition. The specific calculation of LSTM modeling time-dependency is as follows: represents the feature representation at the current time step and conducts temporal modeling by combining historical information. Incorporate the precision impact weight adjustment matrix for W A adaptive modality attention calculation: Calculate the importance weights of the image feature stream and the sensing feature stream, and normalize them through softmax. At the same time, introduce W A to dynamically adjust the contribution degrees of different modalities to the recognition precision. Combine the temporal feature change weight matrix W R for feature fusion to adjust the impact of feature changes between different time windows and ensure that the time dimension information can be effectively utilized in the recognition task. Final classification decision: Y t = softmax(W o H t ); where is the face image feature stream at time t; is the device sensing feature stream at time t; W F and W S are the feature weight matrices of the image modality and the sensing modality respectively; A t is the adaptive attention weight vector at time t; H t is the hidden state of LSTM; X t is the input feature after modality weighting; Y t is the probability distribution of the final recognized category; Wo is the classification weight matrix; ⊙ is the element-wise multiplication; W A is the weight adjustment matrix for accuracy influence; W R is the weight matrix for temporal feature changes.
[0161] Optionally, this specification also provides a generation system for a face recognition fusion model, which is used to execute the method for generating a face recognition fusion model as described above. The generation system for the face recognition fusion model includes:
[0162] A multi-modal data fusion module, which is used to collect a historical face recognition image set and historical device sensing data through a face recognition device, and perform multi-modal data fusion integration on the historical face recognition image set and historical device sensing data to obtain a multi-source fusion data set;
[0163] A spatio-temporal feature integration module, which is used to perform multi-layer network spatio-temporal feature integration on the multi-source fusion data set to obtain a high-dimensional feature spectrum;
[0164] A modal contribution degree weighting module, which is used to perform modal information flow modeling based on the high-dimensional feature spectrum to obtain a modal correlation matrix, and perform adaptive modal contribution degree weighting on the modal correlation matrix to obtain a modal relationship matrix;
[0165] An identification time series modeling module, which is used to adjust the priority of the modal feature flow according to the modal relationship matrix to obtain an adaptive feature flow; perform face recognition time series adaptability modeling based on the adaptive feature flow to obtain a face recognition time series adaptation model;
[0166] A model optimization module, which is used to perform model evaluation on the face recognition time series adaptation model to obtain model accuracy evaluation data; establish an accuracy optimization matrix based on the model accuracy evaluation data, and use the accuracy optimization matrix to optimize the face recognition time series adaptation model to obtain a face recognition fusion model.
[0167] Optionally, this specification also provides a face recognition device, which includes a face recognition device main body, a power supply unit, and an electrical control unit. Among them, the power supply unit is installed inside the face recognition device main body. Inside the face recognition device, there is a sensor group and a camera group that are electrically connected to the power supply unit. The electrical control unit is electrically connected to the power supply unit. The electrical control unit is used to charge the face recognition device main body and control the face recognition device main body. The electrical control unit is used to execute a method for generating a face recognition fusion model as described in any one of the above.
[0168] Therefore, from any perspective, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to encompass all changes that fall within the meaning and scope of the equivalent elements of the application documents within the present invention.
[0169] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for generating a face recognition fusion model, characterized in that It includes the following steps: Step S1: Collect the historical face recognition image set and historical device sensing data through a face recognition device, and perform multi-modal data fusion integration on the historical face recognition image set and historical device sensing data to obtain a multi-source fusion data set; Step S2: Perform multi-layer network spatio-temporal feature integration on the multi-source fusion data set to obtain a high-dimensional feature spectrum; Step S3: Based on the high-dimensional feature spectrum, perform modal information flow modeling to obtain a modal correlation matrix, and perform adaptive modal contribution degree weighting on the modal correlation matrix to obtain a modal relationship matrix; Step S4: Adjust the priority of the modal feature flow according to the modal relationship matrix to obtain an adaptive feature flow; Based on the adaptive feature flow, perform face recognition time series adaptability modeling to obtain a face recognition time series adaptation model; Step S5: Evaluate the face recognition time series adaptation model to obtain model accuracy evaluation data; establish an accuracy optimization matrix based on the model accuracy evaluation data, and use the accuracy optimization matrix to optimize the face recognition time series adaptation model to obtain a face recognition fusion model.
2. The method for generating a face recognition fusion model according to claim 1, wherein Specifically, step S1 is as follows: Step S11: Collect the historical face recognition image set and historical device sensing data through a face recognition device, and perform data preprocessing to obtain the historical face recognition image set to be analyzed and the historical device sensing data to be analyzed; Step S12: Perform time series synchronization and spatial feature mapping on the historical face recognition image set to be analyzed and the historical device sensing data to be analyzed to obtain a spatio-temporally aligned multi-modal data set; Step S13: Based on the spatio-temporally aligned multi-modal data set, perform principal component feature selection, and perform weighted fusion on the principal component features to obtain a multi-modal fusion data set; Step S14: Perform data augmentation on the multi-modal fusion data set to obtain a multi-source fusion data set; Step S15: According to the multi-source fusion data set, perform Monte Carlo data expansion to obtain a multi-source fusion data set.
3. The method for generating a face recognition fusion model according to claim 2, wherein Specifically, step S13 is as follows: Extract the high-level face image features and sensing physical features from the spatio-temporally aligned multi-modal data set, and perform standardization processing on the high-level face image features and sensing physical features to obtain a standardized high-level face image feature set and a standardized sensing physical feature set; Based on the standardized high-level face image feature set, perform face expression feature pattern clustering, evaluate the feature contribution degree of the clustering result, and screen the principal component features according to the feature contribution degree evaluation result to obtain a principal component high-level face image feature set; Based on the standardized sensing physical feature set, perform principal component feature selection to obtain a principal component sensing physical feature set; Perform modal feature weighted fusion on the principal component high-level face image feature set and the principal component sensing physical feature set to obtain a multi-modal fusion feature set; Perform fusion feature non-linear optimization on the multi-modal fusion feature set to obtain a multi-modal fusion data set.
4. The method for generating a face recognition fusion model according to claim 1, wherein Specifically, step S2 is as follows: Step S21: Use a preset multi-scale convolutional network to extract the image detail features and sensing trend features in the multi-source fusion data set to obtain a multi-scale spatio-temporal feature set; Step S22: Perform data feature adversarial optimization on the multi-scale spatio-temporal feature set to obtain a constrained optimization feature set; Step S23: Dynamically model the spatio-temporal relationship of the constrained optimization feature set through a preset spatio-temporal relationship adaptive network, adaptively adjust the contribution degrees of features at different time points and spatial regions, and obtain an adaptive spatio-temporal feature set; Step S24: Convert the adaptive spatio-temporal feature set into a graph structure to obtain a spatio-temporal graph feature set; Step S25: Perform modal fusion and feature integration based on the spatio-temporal feature set to obtain a high-dimensional feature spectrum.
5. The method for generating a face recognition fusion model according to claim 1, wherein The modal information flow modeling described in Step S3 is specifically: Extract the face image feature flow and device sensing feature flow in the high-dimensional feature spectrum; Evaluate the difference degree of feature modal distribution of the high-dimensional feature spectrum, set the distribution deviation threshold to 0.15, and standardize the features in the face image feature flow and device sensing feature flow according to the distribution deviation threshold to obtain a standardized modal feature flow set; Set the autoregressive order to 3, the moving average order to 2, calculate the smoothing window length to be 10s, and perform autoregressive moving average trend smoothing on the standardized modal feature flow set to obtain a smoothed modal time series; Conduct cross-modal information coupling analysis according to the smoothed modal time series, calculate the mutual information gain of each modal feature flow, and set the number of discretization intervals for information gain calculation to 20, so as to construct a modal mutual information gain matrix; Set the clustering threshold to 0.6, perform hierarchical clustering based on the modal mutual information gain matrix, and perform modal correlation modeling based on the hierarchical clustering result to generate a modal correlation matrix.
6. The method for generating a face recognition fusion model according to claim 1, wherein The specific adjustment of the modal feature flow priority described in Step S4 is: Extract the weight distribution of each modal feature flow from the modal relationship matrix and normalize it to obtain a standardized modal weight set; Evaluate the modal time dependence based on the standardized modal weight set and perform a temporal influence degree modeling on the modal feature flow to obtain a modal temporal weight matrix; Conduct a local influence degree analysis on the modal temporal weight matrix, evaluate the feature flow priority score according to the local influence degree, and obtain an initial modal feature flow priority ranking; Combined with the modal correlation matrix, perform dynamic optimization on the initial modal feature flow priority ranking to generate an optimized modal feature flow priority ranking; Perform modal feature flow normalization adjustment based on the optimized modal feature flow ranking to obtain an adaptive feature flow.
7. The method for generating a face recognition fusion model according to claim 6, wherein The so-called dynamic optimization is specifically: Perform division according to the initial modal feature flow priority ranking to obtain a high-priority modal feature flow and a low-priority modal feature flow; Calculate the local correlation distribution of the low-priority modal feature flow in the modal correlation matrix. If the Pearson correlation coefficient with the high-priority modal feature flow is greater than or equal to 0.8, then reduce the downweight amplitude of the low-priority modal feature flow and maintain the contribution degree, so as to obtain a high-correlation low-priority downweight ranking; if the Pearson correlation coefficient with the high-priority modal feature flow is less than 0.3, then perform exponential decay dynamic downweight on the reduced low-priority modal feature flow, so as to obtain a low-correlation low-priority downweight ranking; Calculate the global contribution degree of the high-priority modal feature flow in the modal correlation matrix. If the global contribution degree is greater than 0.8, then perform Lagrangian fusion weight increase to obtain a high-priority weighted ranking; Perform sorting coupling on the high-correlation low-priority deweighting sorting, low-correlation low-priority deweighting sorting, and high-priority weighting sorting to obtain the optimized modal feature stream priority sorting.
8. The method for generating a face recognition fusion model according to claim 1, wherein The specific process of the face recognition timing adaptability modeling described in step S4 is as follows: Perform time segmentation processing on the optimized modal feature stream priority sorting, set the time window to 15s for sliding feature extraction, and calculate the feature statistics within each time window to obtain the time-segmented feature mapping matrix; Calculate the feature change rate between different time windows based on the time-segmented feature mapping matrix, and perform feature change trend analysis to obtain the timing feature change weight matrix; Combine the preset face recognition accuracy index to screen the face recognition accuracy impact factors of the timing feature change weight matrix to obtain the face recognition accuracy impact feature stream; Perform hierarchical clustering analysis on the face recognition accuracy impact feature stream, identify the high-time-dependency feature stream, and assign a time decay factor [0.9, 0.99] to the high-time-dependency feature stream to generate the accuracy impact weight adjustment matrix; Perform cross-time-window feature correlation modeling based on the accuracy impact weight adjustment matrix and the timing feature change weight matrix to obtain the face recognition timing adaptation model.
9. A generation system for a face recognition fusion model, characterized in that, For implementing the method for generating the face recognition fusion model as described in claim 1, the face recognition fusion model generation system includes: A multimodal data fusion module, which is used to collect the historical face recognition image set and historical device sensing data through a face recognition device, and perform multimodal data fusion integration on the historical face recognition image set and historical device sensing data to obtain a multi-source fusion data set; A spatio-temporal feature integration module, which is used to perform multi-layer network spatio-temporal feature integration on the multi-source fusion data set to obtain a high-dimensional feature spectrum; A modal contribution degree weighting module, which is used to perform modal information flow modeling based on the high-dimensional feature spectrum to obtain a modal correlation matrix, and perform adaptive modal contribution degree weighting on the modal correlation matrix to obtain a modal relationship matrix; An identification timing modeling module, which is used to adjust the modal feature stream priority according to the modal relationship matrix to obtain an adaptive feature stream; perform face recognition timing adaptability modeling based on the adaptive feature stream to obtain a face recognition timing adaptation model; A model optimization module, which is used to perform model evaluation on the face recognition timing adaptation model to obtain model accuracy evaluation data; establish an accuracy optimization matrix based on the model accuracy evaluation data, and use the accuracy optimization matrix to optimize the face recognition timing adaptation model to obtain a face recognition fusion model.
10. A face recognition device, characterized in that, It includes a face recognition device main body, a power supply unit, and an electrical control unit. Among them, the power supply unit is installed inside the face recognition device main body. Inside the face device, there is a sensor group and a camera group that are electrically connected to the power supply unit. The electrical control unit is electrically connected to the power supply unit. The electrical control unit is used to charge the face recognition device main body and control the face recognition device main body. The electrical control unit is used to execute the method for generating a face recognition fusion model as described in any one of claims 1-8.
Citation Information
Cited By
Face recognition method and system based on cross-modal face feature fusion
CN120853241A
Object identification method and system based on multi-modal biological characteristic dynamic weight
CN121330782A
A multi-modal biometric dynamic weight object recognition method and system
CN121330782B