Biological feature recognition and verification method and system based on multi-model collaborative decision

By employing a multi-model collaborative decision-making method, multimodal semantic representations are parsed and generated in real time. Combined with dynamic weight allocation and virtual feature generation, the accuracy and privacy protection issues of biometric recognition in complex environments are resolved, enabling continuous iterative upgrades of the recognition strategy and support for user feedback.

CN121744286APending Publication Date: 2026-03-27DONGGUAN YOULIAN YINUO BIOTECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing biometric recognition technologies suffer from low accuracy in complex environments, lack the ability to dynamically adjust weights, exhibit large errors in scenarios with missing modalities, have low conflict resolution efficiency, insufficient privacy protection strategies, and inadequate user feedback support, making it difficult to achieve continuous iterative upgrades of recognition strategies.

Method used

A multi-model collaborative decision-making method is adopted. By parsing multimodal biometrics in real time, a unified multimodal semantic representation is generated. Combined with dynamic weight allocation of environmental perception and modal complementarity analysis, virtual features are generated. Edge-cloud collaborative encrypted transmission is used, a local decision-making module is deployed to resolve conflicts, and the recognition strategy is optimized based on user feedback.

Benefits of technology

It achieves efficient and secure biometric recognition in complex environments, reduces recognition latency, improves recognition accuracy, balances privacy and convenience, and supports continuous iteration and upgrading of recognition strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121744286A_ABST
    Figure CN121744286A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of biometric feature recognition, and discloses a biometric feature recognition verification method based on multi-model collaborative decision, which comprises the steps of real-time multi-modal original data, extraction of core attributes, generation of multi-modal semantic representation based on environmental perception dynamic weight distribution and modal complementarity analysis, and verification of biometric feature recognition. Generating a virtual feature through attention weight migration when the mode is missing; when a multi-modal recognition conflict occurs, a conflict source is positioned through a four-layer detection system, association is analyzed in combination with a graph neural network, and a strategy is generated in a classified manner and executed by an edge node; after the conflict is solved, a target edge node is selected in combination with the biological characteristic data and the geographic position of the user, and a privacy protection strategy is dynamically configured; and obtaining user biological feature input, analyzing quality information to generate decision features, determining candidate results through deep reinforcement learning adaptive fusion and outputting the candidate results, and collecting user feedback to update an environment perception model and a decision fusion strategy. According to the invention, continuous iterative upgrading of the identification strategy can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of biometric recognition, and in particular to a biometric recognition verification method and system based on multi-model collaborative decision-making. Background Technology

[0002] Currently, multimodal biometric recognition technology has been widely applied in fields such as financial security, smart access control, and medical authentication, improving recognition reliability by integrating data from multiple modalities such as face, fingerprint, iris, and voiceprint. However, existing technologies mostly adopt static strategies in weight allocation, which cannot dynamically adjust according to real-time environmental parameters such as lighting and noise, as well as the quality of each modality feature. This leads to a significant decrease in recognition accuracy in complex environments. At the same time, in scenarios with missing modalities, traditional imputation methods are prone to introducing large errors, making it difficult to meet the needs of practical applications.

[0003] Furthermore, existing conflict resolution mechanisms have significant shortcomings. They mostly focus on single-dimensional detection and processing, failing to comprehensively locate conflicts across multiple levels, including data, features, decisions, and semantics. They also lack efficient conflict correlation analysis methods, resulting in low conflict resolution efficiency and slow response. In terms of privacy protection, they mostly adopt a uniform strategy without dynamically adapting to differences in user identity levels and application scenarios. This makes it difficult to balance privacy security with ease of identification, and user feedback provides insufficient support for model optimization, hindering continuous iterative upgrades of the identification strategy.

[0004] As can be seen from the above, how to achieve continuous iterative upgrades of the recognition strategy remains to be solved. Summary of the Invention

[0005] To achieve continuous iterative upgrades of the identification strategy, this application provides a biometric identification verification method based on multi-model collaborative decision-making.

[0006] Firstly, this application provides a biometric identification and verification method based on multi-model collaborative decision-making, employing the following technical solution: A biometric identification and verification method based on multi-model collaborative decision-making includes: The system performs real-time analysis of raw data from different biometric modalities, extracts the core attributes of biometrics, and generates a unified multimodal semantic representation. The biometric modalities include face, fingerprint, iris, and voiceprint, and the core attributes include feature point distribution, texture features, and spectral features. The multimodal semantic representation establishes a semantic mapping relationship between cross-modal biometrics based on a dynamic weight allocation mechanism of environmental perception and modal complementarity analysis rules. For modality-deficient scenarios, virtual features are generated through attention weight transfer. The standardized data is transmitted to the local secure processing module through an edge-cloud collaborative encrypted transmission mechanism. Based on the biometric attribute information and real-time environmental parameters in the standardized data, a conflict resolution strategy is dynamically generated when a multimodal recognition conflict occurs. The conflict resolution strategy is executed by a local decision module deployed on an edge computing node. The local decision module locates the conflict source through a four-layer conflict detection system, combines graph neural network modeling to analyze conflict correlation, responds quickly to sudden conflicts, and processes periodic conflicts in batches. The four-layer conflict detection system includes data level, feature level, decision level, and semantic level. After the conflict is resolved, the biometric data content after the conflict is resolved is obtained. In the computing architecture composed of edge computing nodes and local security modules, the target edge node is selected based on the biometric data content and the user's geographical location information, and the privacy protection processing strategy of the target edge node is dynamically configured. The system acquires user biometric input, parses corresponding biometric quality information from the input, generates corresponding multi-dimensional decision features based on the biometric quality information and feature data in the configured privacy protection strategy, determines a candidate recognition result set based on the multi-dimensional decision features through an adaptive fusion mechanism driven by deep reinforcement learning, outputs corresponding verification results based on the candidate recognition result set, acquires user feedback on the verification results, and updates the environmental perception model and decision fusion strategy based on the feedback.

[0007] Optionally, the dynamic weight allocation mechanism for environmental perception also includes: The weight allocation problem is modeled as a Markov decision process, with real-time environmental parameters, the quality of each modality feature, and historical recognition accuracy as the state space. The real-time environmental parameters include light intensity, noise decibels, temperature, and humidity, while the quality of each modality feature includes face clarity, fingerprint integrity, and voiceprint signal-to-noise ratio. The action space is defined by the weight adjustment range of four modalities: face, fingerprint, iris, and voiceprint. The weight adjustment range of each modality is within the specified range and the sum of the weights of all modalities is a specific set value. A reward function is constructed based on recognition accuracy, response latency, and device power consumption, with each factor having a different weight. Adaptive weight optimization is achieved through a deep Q-network.

[0008] Optionally, in the process of generating virtual features through attention weight transfer, the method further includes: By using a multi-head cross-modal semantic attention module, the semantic associations between existing modalities and missing modalities are mined. The semantic associations include the correlation between face age and voiceprint frequency band, and the correlation between iris structure and fingerprint minutiae distribution. Based on the attention weight distribution of existing modalities, structural and texture features of missing modalities are generated; and the error rate of virtual features is controlled within a specific range.

[0009] Optionally, in the four-layer conflict detection system, the method also includes: At the data level, a Gaussian mixture model is used to detect modal integrity and data quality, and to identify incomplete fingerprint collection and disconnected voiceprint data. Feature consistency is detected at the feature level through a dynamic threshold of cross-modal feature cosine similarity. The dynamic threshold of cross-modal feature cosine similarity is within a specific range and is adjusted according to environmental parameters. The decision-making level detects conflicts based on the confidence difference of each modality recognition result. When the confidence difference reaches a specific standard, it is judged as a conflict. Semantic-level verification of cross-modal semantic contradictions is achieved using a biometric knowledge graph, which includes association rules for age, gender, and physiological characteristics.

[0010] Optionally, when modeling and analyzing conflict associations using graph neural networks, the method further includes: Biometric data, features, and decision results are respectively used as data nodes, feature nodes, and decision nodes in the graph structure; The edge weights are determined by the strength of the association between nodes, and the strength of the association between nodes is dynamically assigned based on data quality and semantic relevance. The influence of nodes is calculated through graph convolution operations to locate conflict source nodes. Nodes whose influence ranking is within a specific range are identified as conflict sources, and the conflict location time is controlled within a specific time range.

[0011] Optionally, in the process of dynamically configuring the privacy protection processing strategy for the target edge node, the method further includes: Based on the user identity level adaptation strategy, a basic privacy strategy is adopted for ordinary users, which is to perform local encrypted storage only at the edge node. Enhanced privacy policies are adopted for high-privilege users, including financial account holders and medical practitioners. These enhanced privacy policies include blockchain notarization, homomorphic encryption, and real-time uploading of data access logs. Furthermore, the response time for triggering privacy policies for high-privilege users is controlled within a specific time range.

[0012] Optionally, after generating the candidate recognition result set, the deep reinforcement learning-driven adaptive fusion mechanism incorporates user historical feedback preference correction. The method further includes: Based on user feedback on verification results in similar scenarios in the past, the feedback included users' preference for fingerprint verification and rejection of blurry face verification. The ranking weights of the candidate identification result set are adjusted a second time to increase the user satisfaction of the adjusted candidate set to a certain percentage or above, and the time taken for the second adjustment is controlled within a certain time range.

[0013] Secondly, this application provides a biometric identification and verification system based on multi-model collaborative decision-making, which adopts the following technical solution: A biometric identification and verification system based on multi-model collaborative decision-making includes: The multimodal feature parsing and encryption module performs real-time parsing of raw data from different biometric modalities, extracts the core attributes of biometric features, and generates a unified multimodal semantic representation. Biometric modalities include face, fingerprint, iris, and voiceprint, and core attributes include feature point distribution, texture features, and spectral features. The multimodal semantic representation establishes a semantic mapping relationship between cross-modal biometric features based on an environment-aware dynamic weight allocation mechanism and modality complementarity analysis rules. For modality-deficient scenarios, virtual features are generated through attention weight transfer. The standardized data is transmitted to the local secure processing module through an edge-cloud collaborative encrypted transmission mechanism. The dynamic conflict resolution module identifies and generates conflict resolution strategies based on biometric attribute information and real-time environmental parameters in the standardized data when multimodal conflict identification occurs. The conflict resolution strategies are executed by a local decision-making module deployed on an edge computing node. The local decision-making module locates the conflict source through a four-layer conflict detection system, analyzes the conflict correlation by combining graph neural network modeling, responds quickly to sudden conflicts, and processes periodic conflicts in batches. The four-layer conflict detection system includes data level, feature level, decision level, and semantic level. The edge node adaptation and protection module acquires the biometric data content after conflict resolution. In the computing architecture composed of edge computing nodes and local security modules, it selects target edge nodes based on biometric data content and user geographic location information, and dynamically configures the privacy protection processing strategy of the target edge nodes. The feature verification model update module acquires user biometric input, parses the corresponding biometric quality information from the biometric input, generates corresponding multi-dimensional decision features based on the biometric quality information and feature data in the configured privacy protection strategy, determines a candidate recognition result set based on the multi-dimensional decision features through a deep reinforcement learning-driven adaptive fusion mechanism, outputs the corresponding verification result based on the candidate recognition result set, acquires user feedback information on the verification result, and updates the environmental perception model and decision fusion strategy based on the feedback information.

[0014] Thirdly, this application provides a biometric identification and verification system based on multi-model collaborative decision-making, which adopts the following technical solution: A biometric identification and verification system based on multi-model collaborative decision-making includes a processor, wherein the processor runs a program of any one of the above-described biometric identification and verification methods based on multi-model collaborative decision-making.

[0015] Fourthly, this application provides a storage medium, which adopts the following technical solution: A storage medium storing a program for the biometric identification and verification method based on multi-model collaborative decision-making as described in any one of the above.

[0016] In summary, this application includes at least one of the following beneficial technical effects: By employing edge-cloud collaborative encrypted transmission and dynamic privacy protection strategy adaptation (deploying encryption schemes differently based on user identity levels), core data processing and privacy protection are completed locally at edge nodes. This avoids the security risks of remote transmission of sensitive data while ensuring recognition response efficiency through lightweight edge computing, achieving a precise balance between privacy security and convenience. Simultaneously, relying on a user feedback mechanism, user preferences and suggestions for correction based on verification results are fed back into the model. Combined with a deep reinforcement learning-driven adaptive fusion mechanism, the environmental perception model and decision fusion strategy are continuously optimized, providing direct and accurate user demand support for the iteration of recognition strategies.

[0017] On the one hand, the dynamic weight allocation mechanism constructs a Markov decision process based on real-time environmental parameters, modal quality, and historical accuracy. It continuously optimizes modal weights through a deep Q-network, while the four-layer conflict detection system and graph neural network continuously explore conflict correlation patterns to improve conflict resolution strategies. On the other hand, it introduces user historical feedback preferences to correct the ranking of candidate recognition results, transforming user acceptance / rejection scenario-based feedback into a basis for model optimization. Combined with feedback information after each verification, it updates the decision logic, forming a closed-loop iteration of "data collection - model decision - user feedback - parameter optimization," ensuring that the recognition strategy always adapts to environmental changes and user needs, and achieving continuous upgrades. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a biometric identification and verification method based on multi-model collaborative decision-making, according to an exemplary embodiment.

[0019] Figure 2 This is a structural block diagram of a biometric identification and verification system based on multi-model collaborative decision-making, according to an exemplary embodiment. Detailed Implementation

[0020] The embodiments of this application are described in detail below, and examples of the embodiments are shown in the accompanying drawings.

[0021] In the description of this specification, the references to "certain embodiments," "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples" indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0022] This application discloses a biometric identification and verification method based on multi-model collaborative decision-making, referring to... Figure 1 ,include.

[0023] The S100 performs real-time analysis of raw data from different biometric modalities, extracts the core attributes of biometrics, and generates a unified multimodal semantic representation. Biometric modalities include face, fingerprint, iris, and voiceprint, and core attributes include feature point distribution, texture features, and spectral features. The multimodal semantic representation establishes a semantic mapping relationship between cross-modal biometrics based on a dynamic weight allocation mechanism of environmental perception and modal complementarity analysis rules. For modality-deficient scenarios, virtual features are generated through attention weight transfer. The standardized data is transmitted to the local secure processing module through an edge-cloud collaborative encrypted transmission mechanism.

[0024] S100 is the core data preprocessing and transmission stage for multimodal biometric recognition and verification. It is executed step-by-step according to the logic of "data parsing → feature extraction → semantic unification → missing data compensation → encrypted transmission," and the specific process is as follows: Real-time parsing of multimodal raw data: The system synchronously receives raw data from four biometric modalities: face, fingerprint, iris, and voiceprint. Through a dedicated data parsing interface, it performs format standardization processing on the data of different modalities—for example, converting the raw pixel data of face images into a processable matrix format, converting the raw audio signal of voiceprint into a digital signal, and parsing the collected data of fingerprints and iris into vector data with recognizable feature points, ensuring that raw data from different sources and in different formats have a unified processing basis.

[0025] Accurate extraction of core attributes: For the parsed standardized raw data, modality-specific algorithms are used to extract core attributes: for face data, feature point distribution (such as the coordinates of key points such as the corners of the eyes and mouth) is extracted; for fingerprint data, texture features (such as the direction of ridges and the location of bifurcation points) are extracted; for iris data, texture features (such as the distribution of iris ring texture) are extracted; and for voiceprint data, spectral features (such as the distribution of fundamental frequency and harmonics) are extracted, providing core feature support for the subsequent establishment of cross-modal semantic association.

[0026] Unified generation of multimodal semantic representation: Based on an environment-aware dynamic weight allocation mechanism, dynamic weights are assigned to four modalities by combining real-time environmental parameters (such as illumination and noise) and the quality of each modality feature (such as facial clarity and fingerprint integrity). For example, the weights of fingerprint and iris modalities are increased when there is insufficient illumination, and the weights of facial and iris modalities are increased when there is a lot of noise. At the same time, according to the modal complementarity analysis rules, semantic associations between different modalities are mined (such as the correlation between facial age and voiceprint frequency band), and a semantic mapping relationship of cross-modal biometric features is established. The extracted multimodal core attributes are integrated into a unified semantic representation to eliminate the heterogeneity of different modal features.

[0027] Virtual feature compensation for missing modalities: If data for a certain modality is missing (such as fingerprint acquisition failure or voiceprint not being recognized), virtual features are generated through attention weight transfer technology: using a multi-head cross-modal semantic attention module, the semantic association between the existing effective modalities and the missing modalities is mined. Based on the attention weight distribution of the existing modalities, the structural and texture features of the missing modalities are simulated to ensure that even if there are missing modalities, a complete multimodal semantic representation can still be formed, avoiding the decrease in recognition accuracy caused by incomplete data.

[0028] Data standardization and edge-cloud collaborative encrypted transmission: The generated unified multimodal semantic representation is subjected to data standardization processing (such as normalization and noise reduction) to eliminate differences in data dimensions; then, through the edge-cloud collaborative encrypted transmission mechanism, after the local preliminary encryption is completed at the edge node, the standardized data is securely transmitted to the local secure processing module to avoid leakage of sensitive biometric data during transmission and to ensure data privacy and security.

[0029] By accurately analyzing multimodal biometric data, extracting core attributes, unifying semantic fusion, and compensating for missing features, the heterogeneity and incompleteness of different biometric modalities are effectively addressed. Furthermore, dynamic weight allocation adapts to complex environmental changes, ensuring the integrity, consistency, and validity of the input data. Edge-cloud collaborative encrypted transmission constructs a privacy and security barrier during data preprocessing, providing high-quality, high-security core data support for subsequent S200 conflict detection, S300 privacy protection configuration, and S400 identification and verification. Modal missing feature compensation also enhances the system's adaptability to complex data acquisition scenarios, laying a solid foundation for the overall accuracy, stability, and privacy security of identification and verification.

[0030] S200 dynamically generates conflict resolution strategies when multimodal recognition conflicts occur, based on biometric attribute information and real-time environmental parameters in standardized data. The conflict resolution strategies are executed by the local decision-making module deployed on the edge computing node. The local decision-making module locates the conflict source through a four-layer conflict detection system, combines graph neural network modeling to analyze conflict correlations, responds quickly to sudden conflicts, and processes periodic conflicts in batches. The four-layer conflict detection system includes data level, feature level, decision level, and semantic level.

[0031] S200 is the core component of conflict resolution in multimodal biometric recognition. It is executed step-by-step according to the logic of "conflict trigger judgment → four-layer detection to locate the conflict source → graph neural network analysis and association → classification and strategy generation → edge node execution," as detailed below: Conflict Trigger Condition Judgment: After receiving standardized data transmitted from S100, the local decision module synchronously reads the biometric attribute information (such as facial feature point distribution, fingerprint texture, etc.) and real-time environmental parameters (such as light intensity, noise decibels, temperature and humidity), and cross-validates the multimodal recognition results. When the recognition results of different modalities contradict each other (such as facial recognition passing but fingerprint recognition failing, iris and voiceprint recognition results not matching), it is determined to be a multimodal recognition conflict, triggering the conflict resolution process.

[0032] A four-layer conflict detection system locates the source of conflict: The local decision-making module uses a four-layer detection system—data-level, feature-level, decision-level, and semantic-level—to systematically investigate and locate the root cause of conflict. Data level: Gaussian mixture model is used to detect the integrity and quality of data of each modality, and to identify data-level problems such as incomplete fingerprint collection, disconnected voiceprint data, and blurred face images; Feature level: Calculate the cosine similarity of cross-modal features, detect feature consistency through dynamic threshold (adjusted in real time with environmental parameters), and investigate feature extraction deviations caused by environmental interference; Decision level: Compare the confidence differences of the recognition results of each modality. When the difference reaches a preset standard (such as the confidence difference between two modalities exceeding 0.3), it is judged as a conflict at the decision level. Semantic level: By leveraging a biometric knowledge graph that includes association rules for age, gender, and physiological characteristics, cross-modal semantic contradictions (such as a large difference between the age inferred from facial recognition and the age inferred from voiceprint) are verified, and semantic-level conflicts are located.

[0033] Graph neural network modeling and analysis of conflict associations: Biometric data, extracted features, and decision results of each modality are mapped to data nodes, feature nodes, and decision nodes in a graph structure, respectively; the association strength between nodes (dynamically assigned based on data quality and semantic relevance, such as higher association strength between high-quality data nodes and corresponding feature nodes) is used as the edge weight to construct a conflict association graph; the influence of each node is calculated through graph convolution operations, and nodes whose influence ranking is within a preset range (such as the top 30%) are identified as core conflict sources, clarifying the propagation path and association relationship of the conflict.

[0034] Classification-based conflict resolution strategy generation: Based on conflict source type and correlation analysis results, and combined with conflict occurrence frequency, strategies are generated according to the classification. Sudden conflicts (such as facial feature recognition deviation caused by a single dust interference): A rapid response strategy is adopted to directly adjust the weight allocation or feature extraction parameters of the corresponding modality to achieve immediate correction; Periodic conflicts (such as repeated voiceprint data disconnections within a specific time period): A batch processing strategy is adopted to optimize the sensor acquisition parameters and environmental adaptation algorithms for that period, thereby resolving the problem of repeated conflicts at its source.

[0035] Local execution of strategies on edge nodes: The local decision-making module deployed on the edge computing node directly executes the generated conflict resolution strategy without uploading it to the cloud for processing, reducing response latency; after the strategy is executed, the processing result is fed back in real time to confirm whether the conflict has been resolved. If it has not been resolved, it returns to the conflict detection stage for re-analysis.

[0036] The S200 achieves precise conflict localization in multimodal recognition through a four-layer conflict detection system. Combined with a graph neural network, it clearly outlines the conflict correlation logic, avoiding the drawbacks of the traditional "one-size-fits-all" approach to conflict handling. By classifying and processing sudden and periodic conflicts, it ensures both rapid resolution of immediate problems and root cause optimization of recurring problems. The local execution strategy at edge nodes significantly reduces conflict processing latency, effectively improving the stability and reliability of multimodal recognition. At the same time, accurate conflict source localization and correlation analysis provide targeted problem-solving basis for subsequent optimization of the S300 privacy protection strategy and iteration of the S400 model, further strengthening the robustness of the entire recognition and verification system.

[0037] After conflict resolution, the S300 acquires the biometric data content after conflict resolution. In the computing architecture composed of edge computing nodes and local security modules, it selects target edge nodes based on the biometric data content and user geographical location information, and dynamically configures the privacy protection processing strategy of the target edge nodes.

[0038] S300 is the node adaptation and privacy protection configuration stage for multimodal biometric recognition. It is executed step-by-step according to the logic of "processed data acquisition → target edge node screening → dynamic configuration of privacy policies → policy effectiveness verification," as detailed below: Step 1, Data Acquisition After Conflict Resolution: The system receives the conflict resolution result output by S200 and simultaneously acquires the biometric data content after conflict correction, including complete multimodal semantic representation, conflict processing records (such as conflict source type, correction parameters), and other core data; at the same time, it collects the user's real-time geographical location information (such as latitude and longitude, and base station information in the area) through the positioning module to ensure the accuracy of subsequent node selection.

[0039] Step 2, Intelligent Selection of Target Edge Nodes: Under the computing architecture of edge computing nodes and local security modules working together, a node selection evaluation system is constructed: the core evaluation indicators are the complexity of biometric data (such as the number of modalities and the size of the data), the physical distance between the user's geographical location and each edge node, and the current load of the edge node (such as CPU utilization and remaining processing resources); the suitability of each edge node is calculated through a weighted scoring algorithm, and the node with the highest suitability is selected as the target edge node to ensure low latency and high stability of data processing.

[0040] Step 3, Dynamic Configuration of Privacy Protection Policies: Based on the sensitivity of biometric data and user identity level, dynamically match privacy protection policies for target edge nodes. For regular users: a basic privacy policy is adopted, which only performs local encrypted storage at the target edge node, restricts data transmission across nodes, and ensures basic privacy and security; High-privilege users: For groups such as financial account holders and medical practitioners, enhanced privacy policies are enabled, including blockchain notarization (ensuring data is tamper-proof), homomorphic encryption (supporting data processing in an encrypted state), and real-time uploading of data access logs to the local security module (enabling full traceability of operations).

[0041] Step 4, Policy Activation and Real-time Verification: After receiving the configuration command, the target edge node automatically loads the corresponding privacy protection algorithm and rules to complete the policy deployment; the local security module monitors the policy activation status in real time, verifies key indicators such as the effectiveness of the encryption algorithm and the rationality of access control permissions. If a configuration anomaly occurs (such as encryption failure or incorrect permission allocation), the reconfiguration process is immediately triggered to ensure that the privacy protection policy is accurately implemented.

[0042] By combining biometric data characteristics with user geographic location to select the optimal target edge node, the S300 achieves precise matching of computing resources, significantly reducing subsequent data processing latency and improving system response efficiency. Its dynamic privacy protection strategy configuration based on user identity levels not only meets the basic privacy needs of ordinary users but also ensures the security of sensitive data for high-privilege users through high-strength encryption technology, effectively balancing privacy security and service convenience. Simultaneously, the collaborative architecture between edge nodes and local security modules constructs a full-process privacy protection barrier, providing a secure and reliable processing environment for the S400's identification and verification, further enhancing the overall system's privacy protection capabilities and scenario adaptability.

[0043] S400: Obtain user biometric input, parse the corresponding biometric quality information from the biometric input, generate corresponding multi-dimensional decision features based on the biometric quality information and feature data in the configured privacy protection strategy, determine the candidate recognition result set based on the multi-dimensional decision features through an adaptive fusion mechanism driven by deep reinforcement learning, output the corresponding verification result based on the candidate recognition result set, obtain user feedback information on the verification result, and update the environmental perception model and decision fusion strategy based on the feedback information.

[0044] S400 is the core decision-making and iterative optimization stage of multimodal biometric recognition verification. It is executed step by step according to the logic of "biometric input acquisition → quality information parsing → decision feature generation → candidate result determination → verification result output → feedback collection → model strategy update". The specific process is as follows: Step 1, User Biometric Input Acquisition: The system receives biometric input actively submitted by the user in real time, covering data collection from four modalities: face, fingerprint, iris, and voiceprint (such as face images captured by the camera, fingerprint patterns collected by the fingerprint sensor, and voiceprint audio recorded by the microphone). At the same time, the system performs preliminary format verification on the input data and removes invalid data (such as completely blurry face images and fingerprint data without valid patterns) to ensure that the input data has basic processing value.

[0045] Step 2, Biometric Quality Information Analysis: For valid input data, a modality-specific quality assessment algorithm is used to analyze the corresponding biometric quality information: for face data, the clarity and occlusion level are assessed; for fingerprint data, the integrity and texture contrast are assessed; for iris data, the texture clarity and noise interference level are assessed; and for voiceprint data, the signal-to-noise ratio and audio stability are assessed. The quality level of each modality feature is determined by quantitative scoring (e.g., 0-100 points), providing a quality basis for subsequent feature fusion.

[0046] Step 3, Multi-dimensional Decision Feature Generation: Combining the parsed biometric quality information with the feature data in the privacy protection policy configured on the S300 (such as encrypted core feature fields and feature ranges adapted to permissions), feature selection and fusion are performed: core features of high-quality modalities are retained first, noise reduction and optimization are performed on low-quality modal features, and sensitive and redundant fields are removed in strict accordance with privacy policy requirements; through feature normalization and dimensional compression processing, multi-dimensional decision features that balance recognition accuracy and privacy security are generated.

[0047] Step 4, Determining the candidate recognition result set: Based on multi-dimensional decision features, a deep reinforcement learning-driven adaptive fusion mechanism is initiated: with recognition accuracy, response speed, and user experience as optimization objectives, the fusion weights of each modality feature are dynamically adjusted through a deep reinforcement learning agent (e.g., increasing the weight of high-quality fingerprint features and decreasing the weight of low-quality voiceprint features); the fused features are matched and recognized, and the top N (e.g., N=3) recognition results with the highest similarity are selected to form the candidate recognition result set.

[0048] Step 5, Verification Result Output: Prioritize the candidate recognition result set (from high to low according to feature matching similarity), and output the final verification result to the user, including the recognition pass / fail conclusion, the candidate result ranking and key matching basis (such as "fingerprint feature matching degree 98%, face recognition passed"); at the same time, provide feedback entry (such as button, pop-up window) to allow users to confirm, reject or correct the verification result.

[0049] Step 6, User Feedback Acquisition: Collect user feedback on the verification results in real time, including approval / disapproval of the recognition conclusion, suggestions for revising the candidate result ranking, and evaluation of the recognition effect of specific modalities (such as "reject blurry face verification" and "prioritize fingerprint verification"). The feedback information is structured and transformed into labeled data that can be used for model optimization (such as the label "low-quality face → user rejection").

[0050] Step 7, Model and Strategy Iterative Update: Using the structured user feedback data as the core optimization basis, the environmental perception model and decision fusion strategy are updated in reverse: the threshold parameters of quality assessment for each modality in the environmental perception model are adjusted (e.g., improving the sensitivity of quality scoring in face occlusion scenarios), and the weight allocation logic of the deep reinforcement learning fusion mechanism is optimized (e.g., increasing the fusion weight of user preference modalities); the updated model and strategy take effect immediately and are used in subsequent recognition and verification processes to achieve continuous iteration.

[0051] The S400 generates decision features by combining biometric quality assessment with privacy protection strategies, ensuring both the accuracy of identification and verification while protecting privacy. Its deep reinforcement learning-driven adaptive fusion mechanism dynamically adapts to quality differences across different modalities, improving the reliability of identification results in complex scenarios. The core user feedback closed-loop design effectively addresses the shortcomings of traditional identification models that rely on data-driven approaches and lack sufficient user demand support. By transforming real-time user feedback into the driving force for model optimization, it enables continuous iteration and upgrading of the environmental perception model and decision fusion strategy. This allows the identification strategy to constantly adapt to user preferences and scenario changes, while simultaneously outputting clear verification results and feedback channels. This balances identification convenience and user experience, providing core support for the long-term optimization of the entire multimodal biometric identification system.

[0052] Based on the scheme in the embodiments of this application, that is, the case of a biometric identification and verification method based on multi-model collaborative decision-making, the following description is provided: Taking the biometric verification scenario of a VIP user (high-privilege user) at a bank's offline branch handling a large-amount transfer as an example, the user needs to complete identity verification through multimodal recognition of face, fingerprint, and voiceprint. The on-site environment suffers from insufficient lighting and significant background noise. Furthermore, the user's fingerprints are worn down due to prolonged use (potentially leading to incomplete collection). Traditional recognition solutions either overemphasize privacy and security by employing complex encryption processes, resulting in recognition response delays exceeding 10 seconds; or they prioritize convenience by simplifying verification steps, but at the risk of sensitive biometric data leakage. Moreover, users have repeatedly reported issues such as "false positives in blurry face recognition and low success rates in fingerprint recognition," and the lack of model optimization has consistently failed to address these problems. This fully exposes the technical pain points of an imbalance between privacy and convenience, and insufficient support from user feedback.

[0053] In the above scenario, the continuous iterative upgrade of the recognition strategy is achieved through a closed-loop mechanism of "user feedback - model optimization - effect verification": After the user submits face, fingerprint, and voiceprint input, the S100 stage improves the weight of fingerprint and voiceprint through dynamic weight allocation based on environmental perception, and generates missing clear virtual facial features through attention weight transfer, ensuring data security through edge-cloud collaborative encrypted transmission; the S200 stage locates "incomplete fingerprint acquisition (data-level conflict)" and "excessive difference in confidence between face and voiceprint recognition (decision-level conflict)" through a four-layer conflict detection system, and after combining graph neural network analysis and correlation, generates "optimized fingerprint acquisition parameters". The system employs a sudden conflict resolution strategy of "number + adjustment of voiceprint fusion weights"; the S300 uses a privacy-enhancing strategy of blockchain notarization and homomorphic encryption based on the user's high-privilege identity configuration, while selecting the nearest and lowest-load edge node to reduce processing latency; after the S400 outputs the verification result, the user submits the opinion "fingerprint recognition still failed, I hope to use voiceprint verification first" through the feedback entry. The system structures this feedback into the label "low-quality fingerprint → user prefers voiceprint", updates the fingerprint quality evaluation threshold in the environmental perception model in reverse, and optimizes the deep reinforcement learning fusion mechanism to improve the fusion weight of voiceprint modalities, completing the first strategy iteration.

[0054] When the user subsequently conducted business again, the iterated model was able to accurately adapt to their biometric characteristics and preferences: in the S100 stage, voiceprint data was prioritized and weighted, while in the S400 stage, high-quality voiceprint features were used directly for identification and verification. The recognition response latency was reduced to within 3 seconds, and privacy and security were continuously guaranteed through enhanced strategies, significantly improving user satisfaction. Simultaneously, the user's feedback data was incorporated into the model training sample library, helping the system develop general optimization strategies for users with similar "fingerprint wear and voiceprint preferences," enabling continuous upgrades to the recognition strategy. This mechanism, driven by user feedback and combining multi-stage data accumulation with dynamic model updates, completely solves the problem of insufficient user feedback support in traditional solutions, achieving an organic unity of privacy and security, ease of identification, and iterative strategy upgrades.

[0055] In this embodiment of the application, the dynamic weight allocation mechanism for environmental perception, taking multimodal user recognition in two scenarios as an example (Scenario 1: Office, normal lighting, low noise; Scenario 2: Subway station, dim lighting, high noise), the execution process of the dynamic weight allocation mechanism for environmental perception further includes: Step 1: Problem Modeling (Markov Decision Process Transformation) The weight allocation problem between the two scenarios is transformed into a Markov decision process, defining the state space: In Scenario 1, the real-time environmental parameters are 500 lux illumination, 30 dB noise, and 25°C / 40% temperature and humidity. The quality of each modality feature is 90 points for face clarity, 85 points for fingerprint integrity, and 80 points for voiceprint signal-to-noise ratio. The historical recognition accuracy is 92% for face, 95% for fingerprint, 90% for iris, and 88% for voiceprint. In Scenario 2, the real-time environmental parameters are 100 lux illumination, 70 dB noise, and 28°C / 60% temperature and humidity. The quality of each modality feature is 60 points for face clarity, 80 points for fingerprint integrity, and 50 points for voiceprint signal-to-noise ratio. The historical recognition accuracy is the same as in Scenario 1. These parameters together constitute the state basis for decision-making, reflecting the actual situation of the current environment and modality.

[0056] Step 2: Spatial Definition (Action Space Determination) Define the weight adjustment range for the four modalities of face, fingerprint, iris, and voiceprint as [0.1, 0.4], and strictly ensure that the sum of the weights of the four modalities is 1. For example, in scenario 1, the adjustment space for face and fingerprint can be biased towards higher values. In scenario 2, due to poor lighting and high noise, the adjustment space for face and voiceprint needs to be narrowed, while the adjustment space for iris can be widened, ensuring that the weight allocation conforms to both modal characteristics and the sum constraint.

[0057] Step 3: Reward Function Construction (Optimization Target Quantization) A reward function is constructed based on recognition accuracy (weight 0.6), response latency (0.2), and device power consumption (0.2). In Scenario 1, priority is given to ensuring high accuracy (e.g., target ≥95%) and low latency (≤2 seconds). In Scenario 2, due to environmental interference, a slight decrease in accuracy is allowed (≥90%), but latency (≤3 seconds) and device power consumption for iris recognition must be strictly controlled (avoiding excessive use of high-power sensors). The optimization priorities for different scenarios are clarified through the weighting of these metrics.

[0058] Step 4: Weight Optimization (Deep Q-Network Output) The state space data from both scenarios are input into the deep Q-network. The network learns the optimization direction of the reward function and adaptively outputs the optimal weights: In Scenario 1, the output weights are: face 0.3, fingerprint 0.3, iris 0.2, and voiceprint 0.2 (utilizing high-resolution faces and complete fingerprints to improve efficiency); in Scenario 2, the output weights are: face 0.1, fingerprint 0.3, iris 0.4, and voiceprint 0.2 (reducing the weights of faces and voiceprints, which are affected by lighting and noise, while increasing the weight of iris to ensure accuracy). This ultimately achieves dynamic adaptation of weights under different environments.

[0059] Through the above process, the system can accurately adjust the multimodal weights according to changes in the scene, balancing efficiency and energy consumption while ensuring recognition effect, and significantly improving the recognition robustness in complex environments.

[0060] In this embodiment, the process of generating virtual features through attention weight transfer is executed step-by-step according to the logic of "semantic association mining → virtual feature generation → error rate control", as follows: Semantic association mining: Through the multi-head cross-modal semantic attention module, we can deeply explore the intrinsic semantic associations between existing effective modalities and missing modalities. For example, we can identify the correlation between facial age and voiceprint frequency band (such as the decrease in voiceprint fundamental frequency corresponding to age) and the correlation between iris structure and fingerprint minutiae distribution (such as the positive correlation between iris texture complexity and fingerprint bifurcation density), providing semantic basis for virtual feature generation.

[0061] Virtual feature generation: Based on the attention weight distribution of the existing modalities (i.e. the degree of attention the model pays to the features in the existing modalities that are related to the missing modalities), simulate and generate structural features (such as the outline of a face) and texture features (such as the ridge texture of a fingerprint) of the missing modalities, ensuring that the generated features are semantically consistent with the features of the existing modalities.

[0062] Error rate control: After generating virtual features, the error rate between virtual features and real features (historically collected valid data of the same modality) is controlled within a specific range (e.g., ≤5%) through methods such as feature matching degree verification, to ensure the reliability of virtual features.

[0063] By accurately mining cross-modal semantic associations and controlling errors, high-quality virtual features are generated in scenarios with missing modalities, effectively solving the problem of incomplete data, ensuring the integrity and accuracy of multimodal semantic representation, and providing reliable feature support for subsequent recognition and verification.

[0064] In this embodiment, the four-layer conflict detection system is executed layer by layer according to the logic of "data-level detection → feature-level detection → decision-level detection → semantic-level detection", as follows: Data-level detection: Gaussian mixture models are used to analyze the raw data of each modality, with a focus on detecting data integrity and quality. For example, by fitting the fingerprint data distribution through the model, incomplete fingerprint collection caused by sensor failure can be identified; the temporal continuity of voiceprint data is analyzed to determine whether there is a disconnection in voiceprint data, thus eliminating potential conflicts at the data level from the source.

[0065] Feature-level detection: Calculate the cosine similarity of cross-modal features (such as the distribution of facial feature points and iris texture features), set a dynamic threshold within a specific range (such as 0.6-0.8) (this threshold is adjusted according to environmental parameters, such as appropriately reducing the threshold when noise increases), and determine that the features are inconsistent when the similarity is lower than the threshold, thus locating the feature extraction deviation problem caused by environmental interference.

[0066] Decision-level detection: Compare the confidence levels of the recognition results of each modality (e.g., 90% confidence level for face recognition and 60% confidence level for fingerprint recognition). When the difference in confidence levels reaches a specific standard (e.g., difference ≥ 30%), it is determined to be a decision-level conflict, identifying contradictions in results caused by differences in the stability of modality recognition.

[0067] Semantic-level detection: Calling a biometric knowledge graph containing rules for association of age, gender, and physiological characteristics (such as "the fundamental frequency of a 20-year-old male's voiceprint is usually higher than that of a 50-year-old male" and "the complexity of iris texture is positively correlated with the density of fingerprint ridges") to verify cross-modal semantic logic. For example, if the face recognition age is 20 years old but the voiceprint inference age is 60 years old, the knowledge graph will mark the semantic contradiction and locate the deep semantic conflict.

[0068] By employing a four-layer progressive detection approach, combined with technologies such as Gaussian mixture models, dynamic thresholds, confidence analysis, and knowledge graphs, we have achieved precise location of conflict sources across all dimensions, from data to semantics. This significantly improves the accuracy and comprehensiveness of conflict detection, providing a precise basis for generating subsequent conflict resolution strategies.

[0069] In this embodiment of the application, when modeling and analyzing conflict associations using graph neural networks, the process is executed step-by-step according to the logic of "node definition → edge weight assignment → influence calculation and conflict source location → time control," as follows: Node definition: The core elements in the biometric identification process are mapped to nodes in a graph structure—raw biometric data (such as fingerprint images and voiceprint audio) are used as data nodes, extracted features (such as facial feature points and iris textures) are used as feature nodes, and the identification conclusions of each modality (such as "pass" and "fail") are used as decision nodes, thus constructing the basic graph structure for conflict correlation analysis.

[0070] Edge weight assignment: The edge weights of the graph structure are based on the strength of the association between nodes. The association strength is dynamically assigned based on data quality and semantic relevance. For example, the association strength between a high-quality fingerprint data node and its corresponding fingerprint feature node is assigned 0.9 (strong association), and the association strength between a low-quality face data node and its voiceprint decision node is assigned 0.3 (weak association), thus quantifying the dependency between nodes.

[0071] Influence calculation and conflict source localization: The graph structure is aggregated by graph convolution operation to calculate the influence of each node on the overall recognition result (e.g., if the influence value of a certain feature node is 0.8, it indicates that its contribution to the conflict is high); the nodes whose influence ranking is in a specific range (e.g., the top 20%) are identified as the core conflict sources, and the key elements that cause the conflict are identified.

[0072] Time control: Throughout the modeling and analysis process, the conflict location time is strictly controlled within a specific time range (e.g., ≤500 milliseconds) by optimizing the complexity of the graph structure and limiting the number of nodes, thus ensuring real-time performance.

[0073] By using graph neural networks to perform structured modeling and quantitative analysis of conflict correlations, the core conflict sources can be accurately located. At the same time, the location time can be controlled to meet real-time requirements, providing an efficient and reliable basis for the targeted generation of subsequent conflict resolution strategies.

[0074] In this embodiment, the process of dynamically configuring the privacy protection processing strategy for the target edge node is explained in more detail, and is executed step by step according to the logic of "user identity level identification → differentiated strategy configuration → response time control", as follows: User identity level identification: Determine the identity level through user account information, business permissions and other dimensions to distinguish between ordinary users and high-privilege users (high-privilege users include financial account holders, medical practitioners and other groups involving sensitive information), and provide a basis for policy adaptation.

[0075] Differentiated policy configuration: Implement corresponding privacy policies based on identity level—for ordinary users, adopt a basic privacy policy, perform local encrypted storage only at the target edge node, restrict data transmission across nodes, and simplify the process while ensuring basic privacy; for high-privilege users, adopt an enhanced privacy policy, enable blockchain notarization (to ensure data is tamper-proof), homomorphic encryption (to support data processing in encrypted state), and upload data access logs to the local security module in real time (to achieve full traceability), and strengthen the protection of sensitive data.

[0076] Response time control: Enhanced privacy policies for high-privilege users strictly control the policy trigger response time within a specific time range (e.g., ≤2 seconds) by optimizing encryption algorithm efficiency and preloading policy modules, so as to avoid a decline in user experience due to the complexity of security measures.

[0077] By configuring differentiated privacy policies based on user identity levels, the system not only meets the needs of ordinary users for convenient identification, but also protects the privacy and security of high-privilege users through enhanced policies. At the same time, it controls response time to balance security and efficiency, achieving precise privacy protection and scenario adaptability.

[0078] In this embodiment, after the deep reinforcement learning-driven adaptive fusion mechanism generates a set of candidate recognition results, an explanation of the execution process for correcting user historical feedback preferences is introduced. This process is executed step-by-step according to the logic of "historical feedback extraction → secondary adjustment of ranking weights → satisfaction and time consumption control," as follows: User history feedback extraction: The system retrieves user feedback data on verification results in similar scenarios (such as similar environments and similar business), such as "prioritize fingerprint verification", "reject blurry face verification", "suggestions for correcting voiceprint recognition results", etc., to clarify the user's preference in specific scenarios.

[0079] Secondary adjustment of ranking weights: Based on the extracted historical feedback, the ranking weights of the candidate recognition result set generated by deep reinforcement learning are optimized a second time. For example, if a user has repeatedly given preference to fingerprint verification, the weight of fingerprint feature matching in the ranking is increased, so that fingerprint-related candidate results are ranked higher. If a user has rejected blurry face verification, the weight of candidate results corresponding to low-quality face features is reduced, and the ranking order is adjusted.

[0080] Satisfaction and time consumption control: By adjusting the weights, we ensure that the matching degree between the corrected candidate set and user preferences is improved, so that user satisfaction reaches a certain percentage (e.g., ≥90%). At the same time, by optimizing the adjustment algorithm (e.g., preloading user preference parameters), we strictly control the time consumption of the secondary adjustment within a specific time range (e.g., ≤100 milliseconds) to avoid affecting the recognition response speed.

[0081] By incorporating users' historical feedback preferences to make secondary corrections to the candidate recognition results, the adaptability of the recognition results to users' habits is improved, ensuring high user satisfaction. At the same time, the time consumption control balances personalized optimization and recognition efficiency, further enhancing the user orientation and iterative effectiveness of the recognition strategy.

[0082] This application discloses a biometric identification and verification system based on multi-model collaborative decision-making, referring to... Figure 2 ,include: The multimodal feature parsing and encryption module 001 performs real-time parsing of raw data from different biometric modalities, extracts the core attributes of biometric features, and generates a unified multimodal semantic representation. Biometric modalities include face, fingerprint, iris, and voiceprint, and core attributes include feature point distribution, texture features, and spectral features. The multimodal semantic representation establishes a semantic mapping relationship between cross-modal biometric features based on a dynamic weight allocation mechanism of environmental perception and modal complementarity analysis rules. For modality-deficient scenarios, it generates virtual features through attention weight transfer. The standardized data is transmitted to the local secure processing module through an edge-cloud collaborative encrypted transmission mechanism. The dynamic conflict resolution module 002, based on the biometric attribute information in the standardized data and real-time environmental parameters, dynamically generates conflict resolution strategies when multimodal conflict identification occurs. The conflict resolution strategies are executed by the local decision-making module deployed on the edge computing node. The local decision-making module locates the conflict source through a four-layer conflict detection system, combines graph neural network modeling to analyze conflict correlation, responds quickly to sudden conflicts, and processes periodic conflicts in batches. The four-layer conflict detection system includes data level, feature level, decision level, and semantic level. The edge node adaptation and protection module 003, after conflict resolution, acquires the biometric data content after conflict resolution. In the computing architecture composed of edge computing nodes and local security modules, it selects target edge nodes based on biometric data content and user geographic location information, and dynamically configures the privacy protection processing strategy of target edge nodes. The feature verification model update module 004 acquires the user's biometric input, parses the corresponding biometric quality information based on the biometric input, generates corresponding multi-dimensional decision features based on the biometric quality information and the feature data in the configured privacy protection strategy, determines the candidate recognition result set based on the multi-dimensional decision features through a deep reinforcement learning-driven adaptive fusion mechanism, outputs the corresponding verification result based on the candidate recognition result set, acquires the user's feedback information on the verification result, and updates the environmental perception model and decision fusion strategy based on the feedback information.

[0083] This application also discloses a biometric identification and verification system based on multi-model collaborative decision-making, including a processor, wherein the processor runs a program of the biometric identification and verification method based on multi-model collaborative decision-making described in any one of the above-mentioned embodiments.

[0084] This application also discloses a storage medium storing the program of the biometric identification and verification method based on multi-model collaborative decision-making as described in any one of the above embodiments.

[0085] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A biometric identification verification method based on multi-model collaborative decision, characterized in that, Comprise: Real-time analysis of raw data from different biometric modalities, extraction of core attributes of biometric features and generation of unified multi-modal semantic representation, biometric modalities include face, fingerprint, iris and voiceprint, core attributes include feature point distribution, texture feature, spectral feature; The multi-modal semantic representation establishes a semantic mapping relationship between cross-modal biometrics based on a dynamic weight distribution mechanism and a modal complementarity analysis rule, and generates virtual features through attention weight transfer for modal missing scenarios, and the standardized data is transmitted to the local security processing module through the edge-cloud collaborative encryption transmission mechanism; Based on the biometric attribute information in the standardized data and the real-time environmental parameters, a conflict resolution strategy is dynamically generated when multi-modal recognition conflict occurs, which is executed by the local decision module deployed in the edge computing node. The local decision module locates the conflict source through a four-layer conflict detection system, analyzes the conflict association combined with graph neural network modeling, and responds quickly to sudden conflicts and processes periodic conflicts in batches. The four-layer conflict detection system includes data level, feature level, decision level and semantic level. After conflict resolution, the biometric feature data content after conflict resolution is obtained, and based on the biometric feature data content combined with user geographic location information, the target edge node is selected, and the privacy protection processing strategy of the target edge node is dynamically configured in the computing architecture composed of edge computing node and local security module; Obtain user biometric feature input, based on the biometric feature quality information parsed in the biometric feature input, generate corresponding multi-dimensional decision features based on the biometric feature quality information combined with the feature data in the configured privacy protection strategy, determine the candidate recognition result set based on the multi-dimensional decision features through the deep reinforcement learning driven adaptive fusion mechanism, and output the corresponding verification result based on the candidate recognition result set. Get user feedback information on the verification result, and update the environment perception model and decision fusion strategy based on the feedback information.

2. The multi-model based co-decision making biometric identification verification method according to claim 1, characterized in that, In the dynamic weight distribution mechanism of environment perception, the method further comprises: Model the weight distribution problem as a Markov decision process, with real-time environmental parameters, feature quality of each modality and historical recognition accuracy as state space, wherein the real-time environmental parameters include light intensity, noise decibel, temperature and humidity, and the feature quality of each modality includes face sharpness, fingerprint integrity and voiceprint signal-to-noise ratio; The weight adjustment interval of the four modalities of face, fingerprint, iris and voiceprint is taken as the action space, and the weight adjustment interval of each modality is within the specified range and the sum of all modality weights is a specific set value; The reward function is constructed with recognition accuracy, response delay and device power consumption, wherein the recognition accuracy, response delay and device power consumption occupy different weight proportions; weight adaptive optimization is realized through deep Q network.

3. The multi-model based co-decision making biometric identification verification method according to claim 2, characterized in that, In the process of attention weight transfer to generate virtual features, the method further comprises: By the multi-head cross-modal semantic attention module, the semantic association between the existing modalities and the missing modalities is excavated, including the correlation between face age and voiceprint frequency band, and the correlation between iris structure and fingerprint minutiae point distribution. Based on the attention weight distribution of the existing modalities, the structural features and texture features of the missing modalities are generated, and the error rate of the virtual features is controlled within a certain range.

4. The multi-model based co-decision making biometric identification verification method of claim 3, wherein, In the four-layer conflict detection system, the method further includes: The Gaussian mixture model is used to detect the modal integrity and data quality at the data level, and to identify incomplete fingerprint collection and disconnected voiceprint data; At the feature level, the feature consistency is detected by the cross-modal feature cosine similarity dynamic threshold, which is in a certain range and is adjusted according to environmental parameters; At the decision level, the conflict is detected according to the confidence difference of each modality recognition result, and when the confidence difference reaches a certain standard, it is determined as a conflict; At the semantic level, the cross-modal semantic contradiction is checked by means of the biometric knowledge graph, which contains age, gender, and physiological feature association rules.

5. The multi-model based co-decision making biometric identification verification method according to claim 4, characterized in that, When analyzing the conflict association by graph neural network modeling, the method further includes: The biometric data, features, and decision results are taken as data nodes, feature nodes, and decision nodes in the graph structure respectively; The edge weight between nodes is based on the dynamic assignment of data quality and semantic correlation; The influence of the node is calculated by graph convolution operation, and the conflict source node is located, and the nodes with influence ranking in a certain range are determined as the conflict source; and the conflict positioning time is controlled within a certain time range.

6. The multi-model based co-decision making biometric identification verification method according to claim 5, characterized in that, In the process of dynamically configuring the privacy protection processing strategy of the target edge node, the method further includes: Based on the user identity level, the basic privacy strategy is adopted for ordinary users, which is only local encryption storage at the edge node; The enhanced privacy strategy is adopted for high-privilege users, including financial account holders and medical practitioners, which includes blockchain storage, homomorphic encryption, and real-time uploading of data access logs; and the privacy policy trigger response time of high-privilege users is controlled within a certain time range.

7. The multi-model based co-decision making biometric identification verification method according to claim 6, characterized in that, After the deep reinforcement learning driven adaptive fusion mechanism generates the candidate recognition result set, the user historical feedback preference is introduced for correction, and the method further includes: Based on the user's feedback on the verification results of similar scenarios in the past, the feedback includes the user's preference for fingerprint verification and rejection of fuzzy face verification; The sorting weight of the candidate recognition result set is adjusted again; the user satisfaction of the adjusted candidate set is improved to a certain percentage or more; and the time consumption of the secondary adjustment is controlled within a certain time range.

8. A multi-model collaborative decision based biometric identification verification system characterized in that, It includes: The multi-modal feature analysis encryption module analyzes raw data from different biological feature modalities in real time, extracts core attributes of biological features, and generates unified multi-modal semantic representations. The biological feature modalities include face, fingerprint, iris, and voiceprint, and the core attributes include feature point distribution, texture feature, and spectral feature. The multi-modal semantic representation establishes a semantic mapping relationship between cross-modal biological features based on a dynamic weight distribution mechanism of environmental perception and a modal complementarity analysis rule, and generates virtual features through attention weight migration for modal missing scenarios. The standardized data is transmitted to the local security processing module through an edge-cloud collaborative encryption transmission mechanism. The conflict recognition dynamic resolution module dynamically generates a conflict resolution strategy based on biological feature attribute information in the standardized data and real-time environmental parameters when a multi-modal recognition conflict occurs. The conflict resolution strategy is executed by a local decision module deployed in an edge computing node. The local decision module locates the conflict source through a four-layer conflict detection system, analyzes conflict associations through graph neural network modeling, responds quickly to sudden conflicts, and processes periodic conflicts in batches. The four-layer conflict detection system includes data level, feature level, decision level, and semantic level. The edge node adaptation protection module obtains biological feature data content after conflict resolution. In the computing architecture composed of the edge computing node and the local security module, the target edge node is selected based on the biological feature data content and user geographic location information, and the privacy protection processing strategy of the target edge node is dynamically configured. The feature verification model update module obtains user biological feature input, analyzes corresponding biological feature quality information based on the biological feature input, generates corresponding multi-dimensional decision features based on the biological feature quality information and the configured privacy protection strategy, determines a candidate recognition result set based on the multi-dimensional decision features through a deep reinforcement learning driven adaptive fusion mechanism, and outputs corresponding verification results based on the candidate recognition result set. Feedback information of the user on the verification results is obtained, and the environmental perception model and the decision fusion strategy are updated based on the feedback information.

9. A multi-model collaborative decision based biometric identification verification system characterized in that, A processor running a program of the multi-model collaborative decision based biological feature recognition verification method according to any one of claims 1-7.

10. A storage medium, characterized by A storage storing a program of the multi-model collaborative decision based biological feature recognition verification method according to any one of claims 1-7.