Identity authentication method based on multiple modes and storage medium
By combining multimodal identity authentication methods with various biometric data for identity authentication, this approach addresses the shortcomings of existing security robots in terms of identity verification and privacy protection. It achieves efficient and secure identity verification and privacy protection, adapting to changes in different environments and personnel.
Patent Information
- Application Number
- CN202511003984.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-11
AI Technical Summary
Existing security robots in home and small commercial environments have shortcomings in identity verification and privacy protection. Single-modal recognition is not reliable enough, and they lack multimodal fusion, dynamic patrol and proactive interaction capabilities. Privacy data is easily leaked and there is a lack of convenient cancellation mechanisms.
It adopts a multimodal identity authentication method, combining multiple biometric data for identity authentication. Through data collection, comparison with a preset database, and interactive authentication, it ensures the accuracy and security of identity verification, manages privacy data through a local database, and supports dynamic patrol routes and proactive interaction.
It improves the accuracy and flexibility of identity authentication, reduces the false recognition rate, enhances the security of patrol areas, reduces human intervention, and ensures the security and management efficiency of privacy data.
Smart Images

Figure CN120930118A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and specifically to a multimodal identity authentication method and storage medium. Background Technology
[0002] Security patrol robots and monitoring devices are increasingly emerging in current home and small business environments, such as home patrol robots with cameras or robotic vacuum cleaners with facial recognition capabilities. However, these existing technologies still have many shortcomings in terms of authentication and privacy protection.
[0003] First, many security systems rely solely on a single biometric identification method, such as facial recognition or voiceprint recognition, for identity verification. This single-modal recognition approach is unreliable in practical applications. Facial recognition requires the person being checked to face the camera and provide a clear facial image, while voiceprint recognition often requires reciting a fixed password, failing to achieve non-intrusive real-time identity verification. If the person being checked is uncooperative or objective conditions are unfavorable, these technologies often fail, resulting in high false recognition and false negative rates. Furthermore, although some patrol robots have begun to integrate multiple identification methods, such as facial recognition, gait recognition, and RFID identification, multimodal fusion has not yet been fully applied in home and small commercial patrol applications, failing to provide a complete solution for multimodal identity verification in lightweight devices.
[0004] Secondly, traditional patrol and security robots primarily rely on passive monitoring, lacking proactive interaction mechanisms. When strangers intrude, the robots lack the ability to actively verify their identities, and reliance on remote monitoring by personnel may result in missed opportunities for timely intervention. Furthermore, many commercial security robots patrol along preset trajectories, lacking dynamic patrol routes, leading to blind spots or the risk of malicious evasion. Existing technologies fail to adequately consider dynamically adjusting patrol routes based on environmental conditions and historical data to improve patrol coverage and unpredictability.
[0005] Finally, existing technologies also have shortcomings in terms of privacy protection. Sensitive personal biometric data is involved in the identity verification process, and many systems upload this data to cloud servers for processing, posing a risk of privacy breaches. If the cloud database is compromised, users' private information may be leaked. Furthermore, even after a user stops using the device, their biometric data remains stored in the system, lacking a convenient data logout mechanism, making it difficult to meet the high security requirements of home and small business environments. Summary of the Invention
[0006] One objective of this invention is to provide a multimodal identity authentication method and storage medium to address the technical problems of significant shortcomings in traditional technologies regarding active identity verification, multimodal fusion, dynamic patrol, and privacy protection.
[0007] In a first aspect, embodiments of the present invention provide a multimodal authentication method applied to at least one electronic device in a security system, the method comprising: Within the preset patrol area, data is collected on the detected persons to be identified, and the collected data corresponding to the persons to be identified is obtained; Based on the collected data and the preset database, the person to be identified is authenticated and identified to obtain the identification result; If the identification result is "not identified", then the person to be identified is authenticated through interaction to obtain the identity verification result of the person to be identified.
[0008] In a second aspect, a computer device is provided, the computer device including a memory and a processor, the memory being connected to the processor, the processor being configured to execute one or more computer programs stored in the memory, the processor, when executing the one or more computer programs, causing the computer device to implement the multimodal authentication method as described in the first aspect.
[0009] In a third aspect, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the multimodal authentication method as described in the first aspect.
[0010] In the above embodiment based on the multimodal identity authentication method, computer equipment, and storage medium, data is first collected from detected individuals to be identified within a preset patrol area to obtain the collected data corresponding to the individuals to be identified. Next, based on the collected data and a preset database, identity authentication is performed on the individuals to be identified to obtain an identification result. Finally, if the identification result is a failure, the individuals to be identified are authenticated interactively to obtain the identity verification result. This embodiment, by first performing identity authentication through data collection and comparison with a preset database, can effectively improve the accuracy of identity authentication for individuals to be identified and reduce the possibility of misidentification. When the identification result fails, further interactive authentication is performed to ensure that only verified personnel can obtain target operation instructions, thereby enhancing the security of the patrol area. Further interactive authentication of the individuals to be identified can handle complex identity verification scenarios, adapt to changes in different environments and personnel, improve flexibility and accuracy, and reduce manual intervention, thus increasing efficiency. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram of the framework of a security system according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a multimodal identity authentication method according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a multimodal identity authentication device according to an embodiment of the present invention. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.
[0014] It should be noted that, unless otherwise specified, the various features in the embodiments of this invention can be combined with each other, all of which are within the protection scope of this invention. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. Moreover, the terms "first," "second," and "third" used in this invention do not limit the data or execution order, but only distinguish identical or similar items with essentially the same function and effect.
[0015] Please see Figure 1 , Figure 1 This is a structural diagram of a security system. Figure 1 In the security system 10, there is an electronic device 20, which includes at least one processor 201 and a memory 202.
[0016] Among them, electronic device 20 can be a mobile patrol device that can perform security patrols, such as a patrol robot or patrol dog, etc., and is not limited to one specific device.
[0017] The processor 201 is configured to support the security system in performing the corresponding functions of the multimodal authentication method described in the above method embodiments. The processor 201 can be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The hardware chip can be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0018] Specifically, the processor 201 may include a transmitting card, a receiving card, and a driver chip.
[0019] The memory 202 is used to store program code and common storage components, including graph databases, vector databases, MySQL, Redis, and MQ, to ensure data storage and management. The memory 202 may include volatile memory (VM), such as random access memory (RAM); it may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); or it may include a combination of the above types of memory.
[0020] See Figure 2 , Figure 2 This is a flowchart illustrating a multimodal authentication method provided in an embodiment of the present invention, applied to at least one electronic device in a security system. The method includes the following steps: S10. Within the preset patrol area, data is collected on the detected persons to be identified, and the collected data corresponding to the persons to be identified is obtained.
[0021] The preset patrol area is a specific geographical area, typically monitored by security systems or electronic devices (such as robots or cameras). The boundaries and extent of this area are pre-defined based on security requirements and site layout. This can be achieved, for example, by delineating a polygonal area on a map or by setting virtual boundaries.
[0022] Specifically, these preset patrol zones are designed to limit the activity range of electronic devices, ensuring they only operate within authorized or monitored areas, thereby improving task efficiency and security. For example, in a large warehouse, the preset patrol zones might be Warehouse Area A, Warehouse Area B, or an office area; in a smart home, they might be the living room, hallway, and study.
[0023] The detected individuals refer to people identified by electronic devices within their preset patrol areas using their own sensors (mainly cameras). These individuals have not yet been identified, and the detection process is usually completed by computer vision algorithms, such as face detection algorithms.
[0024] Data acquisition refers to collecting biometric or behavioral data of individuals to be identified through various sensors and devices, such as facial images, gait, and voice.
[0025] The collected data can be biometric or behavioral data of the person to be identified, such as facial images, gait, and voice.
[0026] As can be seen, this embodiment can effectively collect data on the personnel to be identified, ensuring the safety and efficiency of management within the patrol area.
[0027] S20. Based on the collected data and the preset database, perform identity authentication and identification on the person to be identified to obtain the identification result.
[0028] The preset database refers to the database, local database, or database that is periodically updated locally and synchronized with the cloud, that is ultimately formed after the above-mentioned upload process and can be used by the robot for real-time identity authentication queries. It contains encrypted perturbation feature information of all target personnel.
[0029] The preset database may include, but is not limited to, each person's ID number and corresponding digital templates such as facial feature vectors and voiceprint features.
[0030] For example, a cloud database instance contains a table or set of encrypted perturbation facial feature data packets for all registered family members, such as "father," "mother," and "child," which is a preset database.
[0031] Identity authentication refers to the process of confirming the identity of a person by comparing the data collected from them with target data in a pre-set database. Identity authentication not only determines whether someone is a known person in the database, but may also assess the credibility of that identity.
[0032] Optionally, the preset database is dynamically maintained and updated based on actual conditions. When changes in the appearance or voice characteristics of an authorized person (e.g., change in hairstyle, wearing glasses, hoarseness, etc.) lead to a decrease in recognition confidence, the preset database will be updated. After confirming the person's identity (through administrator confirmation or multimodal verification), the newly collected feature data is added to the person's template entry, enriching the template library. For example, when an electronic device encounters its owner multiple times during patrols, facial features under different lighting and angles can be accumulated, ensuring correct recognition, to improve the owner's template in the database, making subsequent recognition more accurate and robust. Conversely, when an authorized person no longer needs permissions (e.g., employee resignation, visitor permissions expiring), their template can be removed or deregistered from the preset database via control commands, ensuring the real-time effectiveness of the template library.
[0033] Furthermore, the specific process of removing or deregistering a template from the preset database via control commands can involve selecting the user ID to be deregistered through an electronic device control terminal or a companion mobile app, and issuing a deregistration command. The command can be sent to the template management server or broadcast directly to all electronic devices.
[0034] Upon receiving the instruction, each electronic device immediately searches for the template data of the corresponding person ID in the preset database, calls the secure deletion function to remove all biometric data of that person from the database, and overwrites the storage to prevent recovery (e.g., writing random data to overwrite storage sectors multiple times). After local deletion is complete, the device returns a success notification. Furthermore, upon receiving the deregistration instruction or notifications from each device, the preset server also deletes the stored template data package for that person and any backups. If the server retains aggregated statistical information related to that person, according to the differential privacy policy, such statistics do not affect the privacy of individual individuals and require no special handling; however, if plaintext logs involve that ID, they are also cleaned up. The preset server collects deletion success notifications from all devices to verify that each device has performed the deletion operation. If a device does not respond, the preset server will retry sending the deregistration instruction or force execution upon its next online session to ensure no omissions. After all is completed, the server sends a confirmation message to the administrator stating "The biometric template of person XXX has been completely deleted," and can generate a digitally signed deregistration report (including execution time, a list of involved devices, etc.) for auditing purposes.
[0035] The identification result can be either "identified successfully" or "identified unsuccessfully".
[0036] The specific implementation process of S20 can be found in S201-S203, and will not be repeated here.
[0037] As can be seen, this embodiment can efficiently and accurately authenticate the identities of personnel to be identified within the patrol area, ensuring the effective execution of security management.
[0038] S30. If the identification result is that the identification fails, the person to be identified is authenticated through interaction to obtain the identity verification result of the person to be identified.
[0039] Interactive authentication of the person to be identified refers to the process of verifying identity through interaction between a person and an electronic device after automatic identification fails. This typically involves the user actively cooperating to complete additional verification steps.
[0040] The identity verification result is the final conclusion reached after the interactive authentication and identification process. The identity verification result clearly indicates whether the identity of the person to be identified has been confirmed. The identity verification result includes successful verification (identity confirmed) or verification failure (identity not confirmed).
[0041] Optionally, if a recognition failure is detected, the system automatically switches to interactive authentication mode. The electronic device may issue voice prompts through a speaker or display text / image prompts on the screen, posing questions / instructions to the person being recognized that require an answer or action. For example: a voice prompt: "Hello, the system cannot recognize you. Who in your family has blonde hair?" Or, the electronic device may display a dynamically changing digital verification code on its screen, requiring the person being recognized to read it aloud. Or, a specific action instruction: "Please raise your right hand and make an 'OK' gesture in front of your chest."
[0042] Furthermore, the person being identified will respond accordingly after hearing or seeing the instructions. For example, they might answer, "Yes, my sister has blonde hair," or read out the verification code "6789," or complete the specified gesture.
[0043] Furthermore, if it is a voice question and answer, the answer can be converted into text and compared with the preset correct answer (such as "Zhang San", "my sister", etc.); if it is reading a verification code, the captured voice content is compared with the verification code displayed on the screen; if it is an action command, the system captures the action of the person to be identified through the camera and compares it with the preset action template.
[0044] As can be seen, this embodiment improves security and reliability through an interactive authentication mechanism, especially in cases where automated recognition may result in misjudgment or failure to recognize (such as in low light or when the face is obscured), providing an effective supplementary verification method.
[0045] This embodiment improves the accuracy of identity verification by first collecting data and comparing it with a preset database, thus reducing the possibility of misidentification. If the identification result fails, further interactive authentication is performed to ensure that only verified personnel can obtain target operation instructions, thereby enhancing the security of the patrol area. Further interactive authentication of the personnel to be identified can handle complex identity verification scenarios, adapt to changes in different environments and personnel, improve flexibility and accuracy, reduce manual intervention, and increase efficiency.
[0046] Optionally, after the step of performing authentication and identification on the person to be identified through interaction if the identification result is a failure, and obtaining the identity verification result of the person to be identified, the method further includes: generating a target operation instruction for the person to be identified based on the identity verification result.
[0047] The target operation instruction refers to the specific, executable command generated based on the identity verification result, that is, what action the electronic device needs to perform on the person to be identified. The content of the instruction depends entirely on the verification result and the preset behavior rules.
[0048] The preset behavior rules can be as follows: Rule 1: If the identity verification result is "authentication passed", then generate a "allow access" command; Rule 2: If the identity verification result is "authentication passed" and the person is a visitor, then generate a "notify owner" command; Rule 3: If the identity verification result is "authentication failed", then generate a "deny access" command; Rule 4: If the identity verification result is "authentication failed" and the system is in high alert mode, then generate an "issue alarm" command.
[0049] For example, if the identity verification confirms that the person is an authorized visitor or resident (e.g., the person answered security questions correctly or their biometrics matched), the electronic device will play a confirmation message such as "Identity verified, thank you for your cooperation." The event will be logged and archived, and regular patrols will resume.
[0050] Alternatively, if verification fails (the other party cannot provide reasonable identification information, or their biometrics do not match), the electronic device will determine that they are a suspicious person. At this point, it will initiate an alert process, issuing a warning message via speaker such as, "Warning! You have entered a monitored area. Please identify yourself immediately, or security will be notified." Simultaneously, warning lights will illuminate and a buzzer will sound. If the other party remains uncooperative and behaves abnormally, the electronic device can activate a remote alarm program, sending real-time audio and video to the owner / security personnel and requesting support. While ensuring its own safety, the electronic device may attempt to track their movements or prevent them from entering sensitive areas (e.g., patrolling the doorway to block them).
[0051] As can be seen, this embodiment achieves proactive security protection, which not only ensures a friendly experience for normal visitors (polite interaction and fast passage) but also deters potential intruders, greatly improving the intelligence and proactivity of security systems in homes and small commercial venues.
[0052] S201. In one embodiment, the collected data includes at least two types of collected data. The step of performing identity authentication and identification on the person to be identified based on the collected data and a preset database to obtain an identification result includes: extracting features from each of the at least two types of collected data to obtain at least two sets of first feature data; performing feature matching processing on each set of first feature data according to the preset database to obtain multiple matching values; if any one of the multiple matching values is greater than or equal to a preset matching threshold, then the identification result is determined to be successful; or, if all the multiple matching values are less than the preset matching threshold, then the identification result is determined to be unsuccessful.
[0053] The collected data includes at least two types of data, which refer to data collected from the person to be identified that belong to different categories or sources. For example, it could be a facial image and gait video clip, or a facial image and voice clip, or even gait and voice. These at least two types of collected data can reflect the characteristics of an individual from different dimensions.
[0054] In this embodiment, the feature extraction process involves extracting features independently for each type of collected data. Specifically, facial features are extracted from facial data, gait features from gait data, and speech features from speech data.
[0055] Specifically, specialized algorithms for different data types are invoked. For example, a face recognition model is used to extract facial feature vectors, a gait analysis algorithm is used to extract gait parameters, and a speech recognition / voiceprint model is used to extract speech feature vectors. For each type of data (such as face image A, gait video B), the corresponding feature extraction algorithm is run to obtain its respective feature representation, resulting in at least two sets of feature data (corresponding to at least two types of collected data), which constitute the first feature data.
[0056] Here, at least two sets of first feature data refer to feature representations obtained from at least two different types of collected data through a feature extraction step. For example, one set is a facial feature vector, and the other set is a gait feature vector.
[0057] The feature matching process involves extracting each set of first feature data (such as facial features or gait features) from the person to be identified and comparing it with corresponding feature data stored in a pre-defined database. The purpose of feature matching is to calculate similarity or the degree of matching. Specifically, executing the comparison algorithm may involve calculating the distance (such as Euclidean distance or cosine distance) or similarity score between two sets of feature vectors.
[0058] In practice, for each set of first feature data (e.g., facial features to be identified), the corresponding feature data of all registered individuals (e.g., facial features of all registered individuals) are found in a pre-set database, and a similarity score is calculated by comparing them one by one. This process is repeated for each set of third feature data. Therefore, if there are three sets of first feature data (face, gait, and voice), three rounds of matching are required.
[0059] The matching value refers to the similarity score or matching value obtained through feature matching processing, which is used to determine the recognition result.
[0060] For example, if there are two types of collected data (face and gait), then two sets of matching values will be obtained. For face feature matching, the similarity with Zhang San is 0.95 and with Li Si is 0.30; for gait feature matching, the similarity with Zhang San is 0.85 and with Li Si is 0.25.
[0061] The preset matching threshold is a pre-defined boundary value used to determine whether a matching result is sufficiently high. A match is considered successful only when the calculated matching value reaches or exceeds this threshold. The preset matching threshold needs to be adjusted based on the specific application scenario, feature type, and algorithm performance to balance recognition accuracy and false recognition rate.
[0062] For example, the gait characteristics of a person to be identified currently show an 80% similarity to "Person A" and a 65% similarity to "Person B" in the database. It is tentatively considered most likely to be A, but the confidence level is low. Simultaneously, if clear speech is captured, voiceprint features are extracted within a few hundred milliseconds and the best match is found in the voiceprint template library, for example, a 90% match with "Person A's" voiceprint. At this point, gait and voiceprint recognition results are obtained separately. If both the gait and voiceprint modules indicate that the target is most likely "Person A," and their respective match rates exceed the threshold, then the fusion result confidently identifies the target as Person A (e.g., a family member, authorized user). If the two results are inconsistent, for example, the gait is more like "A" while the voiceprint is more like "B," then several possibilities are considered: Case 1: If the person has never faced the camera directly, facial recognition is difficult, but the voiceprint is clear, then the voiceprint result may be more reliable, and the fusion module tends to consider the target as B; Case 2: If the voice acquisition is not clear enough but the gait characteristics are obvious, then the gait result tends to be A; Case 3: If both have low confidence, then no definitive conclusion is drawn, and the identity is marked as uncertain.
[0063] As can be seen, this embodiment no longer relies on a single type of biometric data, but further combines data collected from at least two different sources or types (such as face and gait, or face and voice) for comprehensive judgment, thereby improving the accuracy and robustness of recognition and reducing recognition failures caused by interference from a single feature (such as face being occluded, poor lighting, or speaking too softly).
[0064] S202. In one embodiment, the collected data includes at least two types of collected data. The step of performing identity authentication and identification on the person to be identified based on the collected data and a preset database to obtain an identification result includes: extracting features from each of the at least two types of collected data to obtain at least two sets of first feature data; inputting each set of third feature data from the at least two sets of third feature data into a preset discrimination model for data fusion processing to obtain a probability value output by the preset discrimination model; if the probability value is greater than or equal to a preset probability threshold, then the identification result is determined to be successful; or, if the probability value is less than the preset probability threshold, then the identification result is determined to be unsuccessful.
[0065] Among them, at least two sets of first feature data are consistent with the description in S201, and will not be repeated here.
[0066] The pre-defined discrimination model is a machine learning or deep learning-based model designed to fuse and analyze multiple sets of feature data to output a probability value for identity recognition. This model is pre-trained and can comprehensively assess the authenticity of an identity based on multimodal features. Common model types may include Support Vector Machines (SVM), Logistic Regression, and Neural Networks.
[0067] Data fusion processing refers to the process of processing multiple sets of features within the pre-defined discrimination model. The model combines input features from different data sources (such as facial features and gait features) to obtain a unified probability value that reflects the credibility of identity matching.
[0068] The probability value refers to the numerical value output by the discrimination model, which represents the likelihood of successful identity authentication, and is usually between 0 and 1.
[0069] The preset probability threshold refers to a pre-defined probability limit used to determine whether identity authentication is successful. For example, the preset probability threshold can be set to 0.8 or 0.9. The preset probability threshold can be set according to the actual application scenario and security requirements.
[0070] S201 and S202 can be executed in parallel, one of them can be executed, or they can be executed sequentially.
[0071] As can be seen, this embodiment uses a multimodal data fusion and discrimination model to integrate multiple biometric information, improve the accuracy and robustness of recognition, and effectively prevent the risk of misjudgment caused by single feature recognition.
[0072] S203. In one embodiment, before inputting each set of first feature data from the at least two sets of first feature data into a preset discrimination model for data fusion processing to obtain the probability value output by the preset discrimination model, the method further includes: obtaining a preset feature standard; detecting whether each set of first feature data satisfies the corresponding preset feature standard to obtain a detection result corresponding to each set of first feature data; if any detection result does not satisfy the corresponding preset feature standard, then re-collecting data on the person to be identified; or, if at least two sets of detection results satisfy the corresponding preset feature standard, then inputting each set of first feature data from the at least two sets of first feature data into a preset discrimination model for data fusion processing to obtain the probability value output by the preset discrimination model.
[0073] Among them, preset feature standards are used as benchmarks or conditions for evaluating the quality of collected data. Preset feature standards can be quality standards for specific types of data (such as facial images, voiceprints, etc.) to ensure that the data is suitable for further processing and recognition. These standards may include image sharpness, lighting conditions, sound signal strength, etc., and are not limited to a single factor here.
[0074] Furthermore, the preset feature standard can be the preset feature standard corresponding to the first feature data or the preset feature standard corresponding to other data; no single limitation is imposed here. Specifically, the acquired facial image can be input into a preset quality judgment model to determine whether the image quality is up to standard. The feature extraction network of this quality model is different from the feature extraction network of the facial recognition model.
[0075] Specifically, quality standards are obtained for different feature types (such as face and voiceprint). For example, if the currently collected image is a face, the "Face Quality Standard" module is called, which includes parameters such as sharpness and occlusion detection; if the collected data is voiceprint data, the "Voiceprint Quality Standard" module is called, which includes parameters such as speech duration and noise threshold.
[0076] In the process of detecting whether each set of first feature data meets the corresponding preset feature standard, each set of first feature data is compared with the preset feature standard. The detection process can use a specific quality assessment model or algorithm to evaluate whether the data meets the standard. For example, for facial images, a dedicated image quality assessment model may be used to evaluate the image's sharpness and lighting conditions.
[0077] The detection result refers to the comparison result of each group of first feature data with the preset feature standard, which is divided into "satisfied" or "not satisfied".
[0078] The preset discrimination model is an algorithm used to fuse multiple sets of personality feature data, and outputs a probability value for identity matching (e.g., 0.9 indicates a high matching degree). The preset discrimination model can be a multimodal fusion model (e.g., face + voiceprint joint recognition) or a dynamic weighted model (adjusting weights according to the reliability of each modality), etc., and is not limited to one specific model here.
[0079] Specifically, if any of the detection results does not meet the corresponding preset feature standard, a re-collection is triggered, prompting the user to adjust (e.g., "Please remove the obstruction" or "Please speak again").
[0080] Specifically, if at least two sets of the detection results meet the corresponding preset feature criteria, each set of first feature data from the at least two sets of first feature data is input into a preset discrimination model for data fusion processing. For example, if the user's face image and voiceprint data are both qualified, the two sets of data are input into the "multimodal fusion model" respectively, and the model outputs a probability value (such as 0.95), indicating a high matching degree.
[0081] As can be seen, this embodiment ensures the quality of data used for identity recognition by limiting the preset feature standards, thereby improving the reliability and accuracy of the recognition.
[0082] In one embodiment, if the identification result is a failure, the person to be identified is authenticated interactively to obtain the identity verification result. This includes: outputting a preset query statement to the person to be identified and collecting the biometric data of the person to be identified; receiving the response data of the person to be identified in response to the preset query statement; performing authentication processing on the response data and the biometric data respectively to obtain a first authentication result corresponding to the response data and a second authentication result corresponding to the biometric data; if the first authentication result and / or the second authentication result conforms to preset authentication rules, the identity verification result of the person to be identified is determined to be successful; or, if neither the first authentication result nor the second authentication result conforms to the preset authentication rules, the identity verification result of the person to be identified is determined to be unsuccessful.
[0083] The preset inquiry statements are pre-defined questions used to question the person to be identified. Only someone who truly knows or is related to the target's identity can answer these preset questions accurately. They can be based on public information (such as "Do you know Zhang San?") or private information (such as "What is Zhang San's birthday?"). The purpose is to verify identity through a person's knowledge and memory.
[0084] Biometric data refers to physiological or behavioral characteristics collected from the person to be identified. These can include voiceprints, facial micro-expressions, irises (if observed at close range), gait (if the person is moving), etc.
[0085] The response data refers to the verbal answers given by the person to be identified to preset questions. This includes both the text content and the audio signal of the response.
[0086] The authentication process is the process of verifying identity by combining response data and biometric data through specific algorithms and rules.
[0087] The first authentication result refers to the result obtained after authenticating the response data (content). For example, "content matches", "response is reasonable", "content does not match", or "response is suspicious".
[0088] The second authentication result refers to the result obtained after authenticating the biometric data. For example, "voiceprint match", "high similarity", "voiceprint mismatch", or "low similarity".
[0089] Among them, the preset authentication rules are predefined judgment criteria used to determine how the first and second authentication results should be combined to ultimately determine that the identity verification is successful. The rules can take many forms.
[0090] For example, Rule A ("AND-OR" logic): if either the first result or the second result passes, the authentication is considered successful; Rule B ("AND" logic): both the first and second results must pass for the authentication to be considered successful; Rule C (weighted logic): different weights are assigned according to the importance of the first and second results, and the overall score must reach a threshold to be considered successful.
[0091] For example, the electronic device plays a pre-set inquiry in a polite tone through a speaker, such as, "Hello, I am the security patrol robot for this area. Are you a resident of this building? Please let me know if you need any assistance." Simultaneously, it listens for the respondent's reply. If the respondent provides identification information (such as name and apartment number), the electronic device converts the speech to text and searches for a match in the authorized list; or if the respondent directly states their intention to be a visitor, it judges based on keywords. If the respondent does not respond or attempts to avoid the interaction, the robot can remind them again: "For security reasons, please answer my question or contact the homeowner." During the conversation, the electronic device combines the acquired speech content and voiceprint characteristics to attempt to confirm identity again. For example, if a visitor states their name, the electronic device searches for the corresponding voiceprint / face template in the background database (possibly through a visitor list provided by the homeowner via a security gateway) or a pre-set database to verify its authenticity. If the visitor has pre-registered their face or voiceprint, direct comparison is possible. If not, they can be asked to look directly at the door camera for a few seconds to capture a clear face for real-time comparison. Throughout the interaction, the robot continuously monitors the actions of the person to be identified: if the person suddenly approaches the robot quickly or attempts to force their way in, the electronic device immediately interrupts normal interaction and enters alert mode.
[0092] As can be seen, the dual verification method combining inquiry and biometrics in this embodiment is more reliable than single automatic identification or single inquiry verification, effectively preventing impersonation and deception, and significantly improving security.
[0093] In one embodiment, after outputting a preset query statement to the person to be identified and collecting the biometric data of the person to be identified, the method further includes: performing authentication processing on the biometric data to obtain a third authentication result corresponding to the biometric data; if no response data of the person to be identified to the preset query statement is collected and / or the third authentication result does not conform to the preset authentication rule, then the identity verification result of the person to be identified is determined to be authentication failure.
[0094] The third authentication result can be "match" or "pass" (indicating that the biometrics are highly consistent with the target identity), "not match" or "fail" (indicating that the biometrics are inconsistent with the target identity).
[0095] Among them, the preset authentication rules are the standards and conditions set to judge the authentication results, ensuring the accuracy of identity verification.
[0096] As can be seen, this embodiment improves security and accuracy by using interactive authentication to treat failure to interact effectively (no response) or biometric mismatch as a high-risk signal and directly rejecting the application.
[0097] In one embodiment, after determining that the identity verification result of the person to be identified is successful when the first authentication result and / or the second authentication result conforms to the preset authentication rules, the method further includes: obtaining the current biometric data of the person to be identified; updating the biometric data of the person to be identified to the preset database to obtain the updated preset database.
[0098] Biometric data refers to data related to an individual's biological characteristics, used for identity verification. This includes, but is not limited to, physical characteristics (such as hairstyle, whether glasses are worn, and facial features) and voice characteristics (such as timbre, tone of voice, and degree of hoarseness).
[0099] The pre-set database refers to a database of biometric data for all known individuals, used for comparison and identification. This database needs to be updated regularly to maintain the accuracy and validity of the data.
[0100] In the process of acquiring the biometric data of the person to be identified, the latest biometric features of the person to be identified are collected by sensors (such as cameras and microphones) or extracted from existing records.
[0101] During the process of updating the biometric data of the person to be identified to the preset database, the newly collected biometric data is compared with existing data in the database. If the identification result confirms the identity, the record in the database is updated to reflect the current biometric status. The update process may involve replacing old data or adding new entries to ensure the accuracy and completeness of the database information.
[0102] Specifically, the system initially identifies the person in the database by comparing their name, ID, or face; if the database already contains the person's features (such as an old hairstyle image), the new data is used to overwrite or supplement it (such as storing both old and new hairstyles); the modified data is then written back to the database to ensure that subsequent authentication uses the latest features.
[0103] As can be seen, by continuously updating biometric data in this embodiment, it is possible to dynamically adapt to changes in a person's appearance and voice, thus maintaining the efficiency and accuracy of identity recognition.
[0104] In one embodiment, before collecting data on detected persons to be identified within a preset patrol area to obtain the collected data corresponding to the persons to be identified, the method further includes: collecting a target feature dataset for each of at least one target person; performing differential privacy processing on the target feature dataset of each target person to obtain a perturbation dataset; encrypting the perturbation dataset to obtain an encrypted target data packet; and uploading the target data packet to an initial database in a preset server to obtain the preset database.
[0105] In this context, "target individuals" refers to people whose identities are known in advance and whom electronic devices wish to identify. Target individuals are typically authorized users, family members, employees, etc. Their identity information is known and is entered into the system.
[0106] For example, in a family setting, the target individuals might be the father, mother, and children at home; in an office setting, they might be employees authorized to enter a specific area.
[0107] The target feature dataset is a set of raw biometric data associated with the identity of each target person, such as facial images, fingerprints, and voice samples. For example, if 10 clear facial photos of the target person "A" are collected from different angles and under different lighting conditions, these 10 photos constitute his "target feature dataset".
[0108] Differential privacy processing involves adding carefully calculated noise to raw data during analysis or publication, ensuring that the inclusion of a single target individual's data in the dataset does not imperceptibly affect the final analysis results. The purpose of differential privacy processing is to protect individual privacy, making it impossible to precisely trace a specific individual even when data is aggregated or analyzed.
[0109] For example, suppose some statistical features (such as the mean of facial keypoint coordinates, texture feature statistics of facial regions, etc.) are extracted from the target feature dataset. Differential privacy processing adds tiny, randomly generated noise values to these statistical features. For example, a small random offset is added to the mean of a keypoint coordinate.
[0110] The perturbation dataset is the target feature dataset after differential privacy processing (adding noise). It retains the general pattern and information of the original data, but the specific details have been obscured to protect privacy. For example, adding noise to the statistical features extracted from the target feature dataset of "A" yields his "perturbation dataset".
[0111] Encryption is the process of converting a perturbed dataset into ciphertext using encryption algorithms (such as AES, RSA, etc.) and keys. This ensures that even if the data is intercepted during storage or transmission, it cannot be easily read or understood, thus providing confidentiality protection.
[0112] For example, using a secure symmetric encryption algorithm and key, the perturbation dataset of "A" can be encrypted into a binary data block that cannot be directly read.
[0113] The encrypted target data packet is a data unit containing the disturbance characteristics of a single target person after encryption. It is a structured and secure storage unit.
[0114] The initial database in the preset server is an initial version of the reference database used by electromechanical devices for identity authentication. It is stored on a designated server (which may be a cloud server or a local server) and contains encrypted perturbation signature data packets of all registered target personnel.
[0115] The default server can be a template management server, a cloud server, a local server, or, in the case of a multi-electronic device deployment within a home network, a main electronic device can act as the server to avoid sending data to the external network.
[0116] Specifically, the default server can be a template management server, i.e., a secure server deployed on a user's home LAN or in the cloud, used to relay and store encrypted biometric templates between multiple devices. This server has access control and encrypted computing capabilities, and only authorized devices can access or update the template data within it.
[0117] For example, a database instance on a cloud server might contain a table or collection specifically designed to store encrypted, perturbed facial feature data packets for all family members (target individuals).
[0118] Optionally, homomorphic encryption or secure multi-party computation techniques can be used to further enhance the privacy of cloud processing. For example, when cloud-based collaborative identification is required, homomorphic encryption allows the server to complete the matching calculation without decrypting the user's biometrics. The identification conclusion is obtained directly after the calculation result is decrypted, and the server cannot see the plaintext facial data throughout the entire process.
[0119] Optionally, other patrol robots periodically or in real-time query the preset server for template updates. When a new template package is detected, it is downloaded and decrypted using its respective encryption module. If the decryption key matches, the new biometric template and its ID are obtained. The template is then added to the local template library (if a new person is added) or replaces the old version (if personnel information is updated). After the local library update is complete, the device sends a confirmation receipt to the preset server. Upon receiving successful update confirmations from all relevant devices, the preset server can mark the synchronization package as deployed. If a device is offline for an extended period, the server can temporarily store the update until it comes online to ensure eventual consistency.
[0120] As can be seen, this embodiment ensures that the preset database used for identification contains privacy-protected and encrypted data, thereby enhancing the security of the entire system and the level of user privacy protection.
[0121] In one embodiment, performing differential privacy processing on the target feature dataset of each target person to obtain a perturbed dataset includes: injecting random noise into each target feature data in the target feature dataset of each target person; and performing noise-adding processing on each target feature data according to the random noise to obtain a perturbed dataset.
[0122] In this context, target feature data refers to a single data item within the target feature dataset. Depending on the specific composition of the dataset, it could be a face image, a feature vector, a set of keypoint coordinates, etc.
[0123] For example, in the feature dataset of B, a certain 128-dimensional feature vector or a face photo taken from a specific angle is the target feature data.
[0124] Random noise refers to numerical values generated randomly according to a specific probability distribution (such as a Laplace distribution or a Gaussian distribution). These values are random, but their distribution characteristics are deterministic. The purpose of noise is to "mask" potentially private individual information contained in a single data point.
[0125] For example, a random number that follows a Laplace distribution may have values such as -0.1, +0.3, -0.05, +0.2, etc.
[0126] Injecting random noise refers to the process of performing mathematical operations (usually addition) on the generated random noise values and the original target feature data. The random noise injection process needs to ensure that the noise addition is systematic and its magnitude is controlled to guarantee a level of privacy protection (usually measured by a privacy budget). In practice, the "injection" of noise is usually not a physical mixing, but a mathematical superposition.
[0127] In this context, noise addition refers to the process of performing specific mathematical operations to combine injected random noise with the original target feature data, generating a new, noisy version of the data. For example, adding a random noise value to each value in a feature vector.
[0128] The perturbation dataset refers to the target feature dataset after the aforementioned noise-adding process. Each data item (target feature data) in the perturbation dataset has been added with random noise, thus differing from the original dataset, but still reflecting certain statistical characteristics of the target population as a whole. For example, the perturbation dataset for user B contains his feature vectors or images that have already been subjected to random noise.
[0129] Optionally, a privacy budget value is set before injecting random noise into each target feature data in the target feature dataset for each target person. The privacy budget value is a non-negative real number. The smaller the value, the stronger the privacy protection (the greater the added noise), but the more severe the data distortion; the larger the value, the weaker the privacy protection and the higher the data fidelity. The privacy budget value needs to be weighed according to the application scenario's requirements for privacy and accuracy, and no single limit is set here.
[0130] Furthermore, choose a noise distribution (such as a Laplace distribution). Calculate the sensitivity of the data, which is the maximum possible change in the result of a statistical query when a single data record (such as a person's characteristic data) changes. For adding noise to a single characteristic data point, the sensitivity is typically 1 or some fixed value. Calculate the scale parameter of the noise based on the privacy budget and sensitivity.
[0131] Furthermore, based on the selected noise distribution (such as a Laplace distribution) and the calculated scale parameter, one or more random noise values are generated. If the target feature data is a vector (e.g., 128-dimensional), it is typically necessary to generate an independent noise value for each dimension of the vector. The generated random noise values are then added to the original target feature data to obtain the perturbed data.
[0132] Therefore, the process continues until all target feature data of the current target person has been processed with noise and stored to obtain a perturbation dataset for that target person. Further, this continues until all pre-defined target person feature datasets have undergone differential privacy processing, generating their respective perturbation datasets. Finally, the perturbation datasets generated by all target persons are integrated, or each target person's perturbation dataset is labeled with their identity and stored, ultimately forming a perturbation dataset for subsequent encryption and uploading.
[0133] As can be seen, this embodiment can continue to use data for effective identification and analysis while protecting personal privacy.
[0134] Optionally, the step of performing identity authentication on the person to be identified based on the collected data and a preset database to obtain an identification result includes: performing feature analysis on the collected data to obtain feature types and corresponding second feature data; using the feature type as a query identifier, querying the corresponding target feature type in the preset database; retrieving the third feature data corresponding to the feature type in the preset database; comparing the third feature data with the second feature data to obtain a target feature value; if the target feature value is greater than or equal to a preset feature threshold, then determining that the identification result is successful; or, if the target feature value is less than the preset feature threshold, then determining that the identification result is unsuccessful.
[0135] In this embodiment, the collected data can be a single data set, such as collected facial data. This typically refers to the raw image region or intermediate result used for feature extraction after preliminary processing (such as face detection and alignment).
[0136] Feature analysis involves using algorithms to extract key features from collected data, forming feature data that can be used for comparison.
[0137] Feature type refers to the label or classification that identifies the type of extracted features. It defines which biometric or behavioral features we use for identity recognition. Examples include facial features, gait features, voice features, fingerprint features, iris features, etc., as long as it is a single type of feature data used for identity recognition.
[0138] The second feature data is a numerical representation obtained after feature extraction from the collected data (such as a face), which is usually a high-dimensional feature vector.
[0139] The target feature type is an identifier of the feature type that needs to be matched when searching in the preset database. It corresponds to the second feature type, ensuring that the database search yields reference data of the same type as the current analysis result.
[0140] The third feature data is reference feature data of the same type as the second feature data, retrieved from a pre-set database. The third feature data is a feature vector pre-stored for target individuals with known identities (such as family members or authorized visitors).
[0141] The feature comparison process involves comparing the second and third feature data, calculating their similarity or difference, and applying similarity calculation algorithms. Commonly used algorithms include cosine similarity or Euclidean distance.
[0142] The target feature value refers to the similarity score or matching value obtained by comparing the second feature data and the third feature data.
[0143] The preset feature threshold is a pre-defined numerical limit used to determine whether the comparison result meets the standard for successful recognition. The preset feature threshold needs to be carefully adjusted according to the specific application scenario, algorithm performance, and acceptable error rates (such as false recognition rate (FAR) and false rejection rate (FRR).
[0144] For example, facial feature values are extracted and compared one by one with facial templates in a pre-set database to calculate similarity. If the similarity is higher than a pre-set threshold and a match is found with an existing identity ID, the person is initially determined to be an authorized individual. If the facial comparison does not yield a high-confidence result, other features such as voice recognition are attempted: the electronic device politely requests the other party to say a pre-set phrase or give a short answer through a speaker, and after acquiring the sound, voiceprint features are extracted and searched for a match in a pre-set template library. If no authorized identity is found again, the person is marked as having an unknown identity.
[0145] As can be seen, this embodiment can effectively identify and verify the identity of the person to be identified based on the collected data, ensuring security and accuracy.
[0146] It should be noted that in the above embodiments, there is no necessarily a certain order between the steps. Those skilled in the art can understand from the description of the embodiments of this application that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in turn, etc.
[0147] As another aspect of the embodiments of this application, this application provides a multimodal identity authentication device. The multimodal identity authentication device can be a software module, which includes several instructions stored in a memory. A processor can access the memory and execute the instructions to complete the multimodal identity authentication method described in the various embodiments above.
[0148] See Figure 3 , Figure 3 This is a schematic diagram of a multimodal identity authentication device provided in an embodiment of this application. Figure 2 As shown, the multimodal authentication device 300 includes: The data acquisition unit 301 is used to collect data on the detected person to be identified within a preset patrol area, and obtain the data corresponding to the person to be identified. The identification unit 302 is used to perform identity authentication and identification on the person to be identified based on the collected data and the preset database, and obtain the identification result; The identification unit 302 is further configured to, if the identification result is that the identification fails, perform authentication identification on the person to be identified through interaction to obtain the identity verification result of the person to be identified.
[0149] This embodiment improves the accuracy of identity verification by first collecting data and comparing it with a preset database, thus reducing the possibility of misidentification. If the identification result fails, further interactive authentication is performed to ensure that only verified personnel can obtain target operation instructions, thereby enhancing the security of the patrol area. Further interactive authentication of the personnel to be identified can handle complex identity verification scenarios, adapt to changes in different environments and personnel, improve flexibility and accuracy, reduce manual intervention, and increase efficiency.
[0150] In one embodiment, the collected data includes at least two types of collected data. In the process of performing identity authentication and identification on the person to be identified based on the collected data and a preset database to obtain an identification result, the identification unit 302 is further configured to: extract features from each of the at least two types of collected data to obtain at least two sets of first feature data; perform feature matching processing on each set of first feature data according to the preset database to obtain multiple matching values; if any one of the multiple matching values is greater than or equal to a preset matching threshold, then the identification result is determined to be successful; or, if all the multiple matching values are less than the preset matching threshold, then the identification result is determined to be unsuccessful.
[0151] In one embodiment, the collected data includes at least two types of collected data. In the process of performing identity authentication and identification on the person to be identified based on the collected data and a preset database to obtain an identification result, the identification unit 302 is further configured to: extract features from each of the at least two types of collected data to obtain at least two sets of first feature data; input each set of first feature data into a preset discrimination model for data fusion processing to obtain a probability value output by the preset discrimination model; if the probability value is greater than or equal to a preset probability threshold, then the identification result is determined to be successful; or, if the probability value is less than the preset probability threshold, then the identification result is determined to be unsuccessful.
[0152] In one embodiment, before inputting each of the at least two sets of first feature data into a preset discrimination model for data fusion processing to obtain the probability value output by the preset discrimination model, the identification unit 302 is further configured to: acquire a preset feature standard; detect whether each set of first feature data satisfies the corresponding preset feature standard, and obtain a detection result corresponding to each set of first feature data; if any detection result does not satisfy the corresponding preset feature standard, then re-collect data on the person to be identified; or, if at least two sets of detection results satisfy the corresponding preset feature standard, then input each of the at least two sets of first feature data into the preset discrimination model for data fusion processing to obtain the probability value output by the preset discrimination model.
[0153] In one embodiment, in the step of verifying the person to be identified through interaction if the identification result is a failure, to obtain the identity verification result of the person to be identified, the identification unit 302 is further configured to: output a preset query statement to the person to be identified and collect the biometric data of the person to be identified; receive the response data of the person to be identified in response to the preset query statement; perform authentication processing on the response data and the biometric data respectively to obtain a first authentication result corresponding to the response data and a second authentication result corresponding to the biometric data; if the first authentication result and / or the second authentication result conform to the preset authentication rules, then determine that the identity verification result of the person to be identified is verified as successful; or, if neither the first authentication result nor the second authentication result conforms to the preset authentication rules, then determine that the identity verification result of the person to be identified is unsuccessful.
[0154] In one embodiment, after outputting a preset query statement to the person to be identified and collecting the biometric data of the person to be identified, the identification unit 302 is further configured to: perform authentication processing on the biometric data to obtain a third authentication result corresponding to the biometric data; if no response data of the person to be identified to the preset query statement is collected and / or the third authentication result does not conform to the preset authentication rule, then the identity verification result of the person to be identified is determined to be authentication failure.
[0155] In one embodiment, after determining that the identity verification result of the person to be identified is successful when the first authentication result and / or the second authentication result conforms to the preset authentication rules, the identification unit 302 is further configured to: acquire the current biometric data of the person to be identified; update the biometric data of the person to be identified to the preset database to obtain the updated preset database.
[0156] In one embodiment, before collecting data on the detected person to be identified within the preset patrol area to obtain the collected data corresponding to the person to be identified, the collection unit 301 is further configured to: collect a target feature dataset for each of at least one target person; perform differential privacy processing on the target feature dataset of each target person to obtain a perturbation dataset; encrypt the perturbation dataset to obtain an encrypted target data packet; and upload the target data packet to an initial database in a preset server to obtain the preset database.
[0157] In one embodiment, in the process of performing differential privacy processing on the target feature dataset of each target person to obtain a perturbed dataset, the acquisition unit 301 is further configured to: inject random noise into each target feature data in the target feature dataset of each target person; and perform noise-adding processing on each target feature data according to the random noise to obtain a perturbed dataset.
[0158] It should be noted that the above-described multimodal identity authentication device can execute the multimodal identity authentication method provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects of the method. Technical details not described in detail in the embodiments of the multimodal identity authentication device can be found in the multimodal identity authentication method provided in the embodiments of this application.
[0159] This application also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the multimodal authentication method as described in the foregoing embodiments.
[0160] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0161] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.
Claims
1. A multimodal identity authentication method, characterized in that, The method includes: Within the preset patrol area, data is collected on the detected persons to be identified, and the collected data corresponding to the persons to be identified is obtained; Based on the collected data and the preset database, the person to be identified is authenticated and identified to obtain the identification result; If the identification result is "not identified", then the person to be identified is authenticated through interaction to obtain the identity verification result of the person to be identified.
2. The method according to claim 1, characterized in that, The collected data includes at least two types of collected data. The process of identifying the person to be identified based on the collected data and a preset database, and obtaining the identification result, includes: Feature extraction is performed on each of the at least two types of collected data to obtain at least two sets of first feature data. Based on the preset database, feature matching processing is performed on each group of the first feature data to obtain multiple matching values; If any one of the multiple matching values is greater than or equal to a preset matching threshold, then the recognition result is determined to be successful; or, If all of the multiple matching values are less than the preset matching threshold, then the identification result is determined to be a failure.
3. The method according to claim 1, characterized in that, The collected data includes at least two types of collected data. The process of identifying the person to be identified based on the collected data and a preset database, and obtaining the identification result, includes: Feature extraction is performed on each of the at least two types of collected data to obtain at least two sets of first feature data. Each of the at least two sets of first feature data is input into a preset discrimination model for data fusion processing to obtain the probability value output by the preset discrimination model; If the probability value is greater than or equal to a preset probability threshold, then the recognition result is determined to be successful; or, If the probability value is less than the preset probability threshold, then the identification result is determined to be a failure.
4. The method according to claim 3, characterized in that, Before inputting each of the at least two sets of first feature data into a preset discrimination model for data fusion processing to obtain the probability value output by the preset discrimination model, the method further includes: Obtain preset feature standards; Detect whether each group of first feature data satisfies the corresponding preset feature standard, and obtain the detection result corresponding to each group of first feature data; If any of the detection results does not meet the corresponding preset feature standard, then the data of the person to be identified is collected again; or, If at least two sets of the detection results satisfy the corresponding preset feature criteria, then each set of first feature data from the at least two sets of first feature data is input into a preset discrimination model for data fusion processing to obtain the probability value output by the preset discrimination model.
5. The method according to claim 1, characterized in that, If the identification result is a failure, then the person to be identified is authenticated and identified interactively to obtain the identity verification result of the person to be identified, including: Output preset query statements to the person to be identified and collect the biometric data of the person to be identified; Receive the response data of the person to be identified in response to the preset query statement; The response data and the biometric data are respectively processed for authentication to obtain a first authentication result corresponding to the response data and a second authentication result corresponding to the biometric data. If the first authentication result and / or the second authentication result conform to the preset authentication rules, then the identity verification result of the person to be identified is determined to be successful; or, If neither the first authentication result nor the second authentication result meets the preset authentication rules, then the identity verification result of the person to be identified is determined to be authentication failure.
6. The method according to claim 5, characterized in that, After outputting a preset query to the person to be identified and collecting the biometric data of the person to be identified, the method further includes: The biometric data is processed for authentication to obtain a third authentication result corresponding to the biometric data. If no response data of the person to be identified to the preset query is collected and / or the third authentication result does not conform to the preset authentication rule, then the identity verification result of the person to be identified is determined to be authentication failure.
7. The method according to claim 5, characterized in that, After determining that the identity verification result of the person to be identified is successful when the first authentication result and / or the second authentication result conforms to the preset authentication rules, the method further includes: Obtain the biometric data of the person to be identified; The biometric data of the person to be identified is updated in the preset database to obtain the updated preset database.
8. The method according to claim 1, characterized in that, Before collecting data on the detected person to be identified within the preset patrol area to obtain the collected data corresponding to the person to be identified, the method further includes: Collect a dataset of target features for each target person in at least one target person; Differential privacy processing is performed on the target feature dataset of each target person to obtain a perturbed dataset; The disturbed dataset is encrypted to obtain the encrypted target data packet; The target data packet is uploaded to the initial database in the preset server to obtain the preset database.
9. The method according to claim 8, characterized in that, The differential privacy processing of the target feature dataset for each target person to obtain a perturbed dataset includes: Inject random noise into each target feature data in the target feature dataset of each target person; The random noise is used to add noise to each target feature data to obtain a perturbed dataset.
10. A storage medium, characterized in that, The storage medium stores a computer program, which includes program instructions that, when executed by a processor, cause the processor to perform the multimodal authentication method as described in any one of claims 1-9.