A Multimodal Digital Human Authentication Method and System

By dynamically adjusting speech differences and biometric weights in multimodal biometric recognition, the problem that multimodal biometric technology cannot distinguish verified personnel in special circumstances is solved, and the security and accuracy of identity verification are improved.

CN120068036BActive Publication Date: 2025-08-01INSPUR SMART SUPPLY CHAIN TECH (SHANDONG) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510549980.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-01
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing multimodal biometric technology cannot accurately distinguish different verification personnel under special circumstances, resulting in insufficient security of identity verification.

Method used

By presetting multiple verification voice texts, obtain the difference in voice data between identity people and other biological humans, select the most different voice text as the verification text, and dynamically adjust the weight of biological characteristics according to the voice difference, and combine multiple biological characteristics information for identity verification.

Benefits of technology

Improve the security of identity verification, reduce the possibility of being unable to distinguish between verification personnel in extreme cases, and enhance the discernment ability of different verification personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068036B_ABST
    Figure CN120068036B_ABST
Patent Text Reader

Abstract

This application relates to the field of digital human identity authentication technology, and particularly to a multi-modal digital human identity authentication method and system. The method collects multiple biometric information of the verification personnel. The identity voice information is the voice data when the verification personnel reads the verification text. The multiple biometric information is input into a preset verification model to obtain the probability of passing the verification of the multiple biometric information. Weights are assigned to the multiple biometric information according to the difference degree, where the difference degree is proportional to the weight of the probability of passing the verification of the identity voice information. The weighted value after weighting the probability of passing the verification of the multiple biometric information is obtained, and the weighted value of the probability of passing the verification of the multiple biometric information is used as the final verification probability. The verification personnel are verified based on the final verification probability. This application has the effect of facilitating the distinction and identification of verification personnel in extreme situations and improving the security of verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of digital human identity verification, and particularly to a multi-modal digital human identity verification method and system. Background Art

[0002] With the development of artificial intelligence and virtual reality technologies, digital humans (virtual humans) have gradually become important identity representatives in the digital world. Digital humans can interact with humans in various ways, simulating the behaviors and emotions of real people. In this context, digital human identity verification has become an important means to ensure the security of the interaction between the virtual world and the real world. To ensure the credibility and security of digital humans in various scenarios, it is necessary to verify the user identities of digital humans. Traditional user identity verification methods include password verification, ID number verification, etc. These methods are prone to leakage, leading to problems such as identity theft and fraud. With the development and progress of digital technologies, biometric technologies have gradually been applied to security verification. Common biometric technologies include fingerprint recognition, face recognition, etc. Due to the difficulty of forging biometric features, this technology can improve the security of identity verification to a certain extent. However, there are also some problems in the application of biometric recognition technologies. For example, during fingerprint verification, if the user's fingerprint is dirty or damaged, it may cause changes in the fingerprint during user verification, ultimately resulting in fingerprint verification failure; or during face recognition verification, the user's facial data may be affected by factors such as light, angle, and makeup, ultimately resulting in inaccurate user verification. Therefore, multi-modal biometric technologies have received more attention from industry technicians. Multi-modal biometric technologies mainly combine the verification of multiple biometric features. For example, the voice information and facial information of the verification personnel can be collected, and the voice information, facial information, and preset verification information are compared, and the final verification result of the verification personnel is obtained by synthesizing the two verification results.

[0003] In the process of calculating the final comprehensive verification result of multi-modal biometric technologies, the weights of each biometric information are the same. This method cannot make accurate judgments for some special situations. For example, in a multi-modal biometric system, multiple biometric technologies are collected, and only one of them has a large difference. However, due to the relatively small proportion of this biometric feature, the impact on the final verification result is also relatively small, and ultimately it may lead to the situation where the biometric system cannot distinguish between the two verification personnel. Therefore, the security of multi-modal biometric technologies in related technologies needs to be further improved. Summary of the Invention

[0004] To solve the problem of being unable to distinguish different verification personnel in special situations and improve the security of identity verification, the present application provides a multi-modal digital human identity verification method and system.

[0005] In a first aspect, the present application provides a multimodal digital human identity verification method, adopting the following technical solutions:

[0006] A multimodal digital human identity verification method presets multiple verification voice texts, obtains the voice data of the identity person reading different verification voice texts as the verification analysis voice, and obtains the voices of a preset number of biological persons reading different verification voice texts as the comparison analysis voice; compares the differences between the verification analysis voice and the comparison analysis voice to obtain the difference degree, sums the multiple difference degrees corresponding to the same verification voice text as the voice difference, and takes the verification voice text with the largest voice difference as the verification text.

[0007] Collect multiple biometric information of the verification personnel. The multiple biometric information includes: identity voice information, face information, and fingerprint information. The identity voice information is the voice data when the verification personnel read the verification text. Input the multiple biometric information into a preset verification model to obtain the probabilities of the multiple biometric information passing the verification; divide weights for the multiple biometric information according to the difference degree, where the difference degree is proportional to the weight of the probability of the identity voice information passing the verification; obtain the weighted value after weighting the probabilities of the multiple biometric information passing the verification, take the weighted value of the probabilities of the multiple biometric information passing the verification as the final verification probability, and verify the verification personnel based on the final verification probability.

[0008] In the present application, the differences between the identity person and other biological persons reading different verification voice texts are calculated. If the voice difference corresponding to one of the verification voice texts is larger, it indicates that the voice data of the identity person when reading this verification voice text is easier to distinguish. Therefore, selecting this verification voice text for subsequent verification of the verification personnel can improve the ability to distinguish the verification personnel, reduce the situation where the verification personnel cannot be distinguished, and improve the security of identity verification. At the same time, in the present application, according to the size of the voice difference, weights are set for multiple biometric features when comprehensively verifying the identity of the verification personnel subsequently, so that when the voice difference is large, more attention can be paid to the probability of the voice information passing the verification, thereby further improving the ability to distinguish different verification personnel. Compared with the traditional method where the weights of various biometric information are the same when verifying different digital humans, the biometric information in the present application is dynamically changed when verifying the identity of different digital humans, which is more in line with the unique biometric features of the identity person, such as voice features, and improves the ability to distinguish whether the verification personnel is the identity person.

[0009] Optionally, the step of comparing the differences between the verification analysis voice and the comparison analysis voice to obtain the difference degree includes: obtaining the frequency domain difference and time domain difference between the verification analysis voice and the comparison analysis voice, and taking the sum of the normalized result of the time domain difference and the normalized result of the frequency domain difference as the difference degree between the verification analysis voice and the comparison analysis voice.

[0010] Compare the features of the verification analysis speech and the comparative analysis speech from multiple dimensions of time-domain difference and frequency-domain difference, so as to improve the accuracy of the difference degree in subsequent calculations.

[0011] Optionally, the steps for obtaining the frequency-domain difference between the verification analysis speech and the comparative analysis speech include: obtaining the Mel-frequency cepstral coefficients of the verification analysis speech and the comparative analysis speech, and calculating the distance between the two Mel-frequency cepstral coefficients, and taking this distance as the frequency-domain difference.

[0012] The distance between the Mel-frequency cepstral coefficients of the verification analysis speech and the comparative analysis speech reflects the frequency-domain difference between the verification speech and the comparative analysis speech, and realizes the calculation of the frequency-domain difference.

[0013] Optionally, obtain the Mel-frequency cepstral coefficients of the verification analysis speech and the comparative analysis speech, and calculate the distance between the two Mel-frequency cepstral coefficients; obtain the first-order difference results of the two Mel-frequency cepstral coefficients, and obtain the distance between the two first-order difference results; take the sum of the two distances as the frequency-domain difference.

[0014] In this method, by combining the differences in the first-order difference results of the Mel-frequency cepstral coefficients and considering the dynamic differences between the verification analysis speech and the comparative analysis speech, the accuracy of the frequency-domain difference calculation is improved.

[0015] Optionally, the steps for obtaining the time-domain difference between the verification analysis speech and the comparative analysis speech include: obtaining multiple time-domain features of the verification analysis speech and the comparative analysis speech, calculating the differences of the time-domain features of the verification analysis speech and the comparative analysis speech, and taking the sum of the differences of the multiple time-domain features as the time-domain difference.

[0016] Obtain multiple time-domain features of the verification analysis speech and the comparative analysis speech, and compare the differences in the time-domain features of the verification analysis speech and the comparative analysis speech, so as to obtain the time-domain difference between the verification analysis speech and the comparative analysis speech.

[0017] Optionally, the steps for obtaining the time-domain difference between the verification analysis speech and the comparative analysis speech include: obtaining multiple time-domain features of the verification analysis speech and the comparative analysis speech, calculating the differences of the time-domain features of the verification analysis speech and the comparative analysis speech, calculating the squared results of the differences of the time-domain features, taking this result as the amplified difference, and taking the sum of the amplified differences of the multiple time-domain features as the time-domain difference.

[0018] In this method, the square value of the time-domain feature difference is calculated based on the calculation of the time-domain feature difference, so as to improve the sensitivity of the finally calculated time-domain difference to the time-domain feature difference between the verified analysis speech and the comparative analysis speech.

[0019] Optionally, multiple time-domain features include: zero-crossing rate, energy, and fundamental frequency.

[0020] Optionally, the calculation formula for the weight of the probability that the identity voice information passes verification is:

[0021] ; where represents the weight of the probability that the identity voice information passes verification; represents the voice difference corresponding to the verified text.

[0022] Optionally, the method for verifying the verifier based on the final verification probability includes: setting a passing threshold, and in response to the final verification probability being greater than the passing threshold, determining that the verifier passes the verification.

[0023] In a second aspect, the present application provides a multi-modal digital human identity verification system, adopting the following technical solution:

[0024] A multi-modal digital human identity verification system includes: a processor and a memory, and the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a multi-modal digital human identity verification method according to the above is implemented.

[0025] The beneficial effect is: generating a computer program for the above multi-modal digital human identity verification method and storing it in the memory to be loaded and executed by the processor, so that a system is made according to the memory and the processor, which is convenient to use.

[0026] The present application has the following technical effects: In the present application, different verification texts are set for each digital human, and weights are determined for multiple biometric feature information participating in subsequent verification based on the differences in voice data when the identity person and other biological persons read the verification texts. Therefore, different digital humans perform identity verification based on different weights, which can reduce the situation where the identity of the verifier cannot be verified in extreme cases (multiple biometric features are similar and one biometric information has a large difference). Description of the Drawings

[0027] Figure 1 is a flowchart of a multi-modal digital human identity verification method according to an embodiment of the present application. Detailed Embodiment <s

[0028] An embodiment of this application discloses a multi-modal digital human identity verification method and system, which enables the identity person and other biological persons to read the same set of verification voice texts, obtains the difference degree when the identity person and other biological persons read the same section of voice, and selects the verification voice text with the largest difference degree as the verification text. Therefore, different identity persons correspond to different verification texts when verifying their respective identities.

[0029] When subsequently verifying the identity of the verification personnel, multiple biometric features of the verification personnel are collected and input into a preset model to obtain the passing probabilities of each biometric feature. Subsequently, weights are set for different biometric information according to the difference degree, and the final verification probability of the verification personnel is calculated by combining the passing probabilities of multiple biometric information after weighting. The verification result is obtained according to the final verification probability of the verification personnel, and the identity verification of the verification personnel is completed.

[0030] During the identity verification process of the verification personnel, the weights of the passing probabilities of the biometric information of different verification personnel are different, so it is possible to reduce the situation where the verification system cannot distinguish the verification personnel in special cases.

[0031] Refer to Figure 1 , the multi-modal digital human identity verification method includes steps S1-S4.

[0032] S1: Preset multiple verification voice texts, obtain the voice data of the identity person reading different verification voice texts as verification analysis voices, and obtain the voices of a preset number of biological persons reading different verification voice texts as comparison analysis voices;

[0033] The identity person refers to the owner of the digital human (virtual human), and the identity person corresponds to the digital human one by one and is unique. Other biological persons refer to any other natural persons except the identity person. Multiple other biological persons are selected. For example, 100 or 1000 other biological persons can be selected. In this embodiment, 100 other biological persons are preset. The specific quantity can be determined according to the actual situation.

[0034] The verification voice text is text or numbers, and the length of the verification voice text can be set by itself according to the actual situation. Exemplarily, the verification voice text can be: "24859", "35559", "The sun along the mountain bows", "Innovative and enterprising", etc.

[0035] The identity person reads the verification voice text, and the voice data is collected during the reading process to obtain a set of verification analysis voices.

[0036] The other biological persons read this set of verification voices, and at the same time, the voice data during the reading process is collected to obtain a set of comparison analysis voices.

[0037] One verified speech corresponds to one verified analysis speech and multiple comparison analysis speeches.

[0038] S2: Compare the differences between the verified analysis speech and the comparison analysis speeches. The sum of the differences of multiple degrees corresponding to the same verified speech is the speech difference, and the verified speech text with the largest speech difference is used as the verified text.

[0039] When different people read the same verified speech text, there will be certain differences, such as in speech rate, intonation, loudness, etc. Therefore, the difference degree between the verified analysis speech and the comparison analysis speech can be obtained by calculating the time-domain difference and frequency-domain difference between the verified analysis speech and the comparison analysis speech.

[0040] In one embodiment, the steps of obtaining the frequency-domain difference between the verified analysis speech and the comparison analysis speech include: obtaining the Mel frequency cepstral coefficients of the verified analysis speech and the comparison analysis speech, and calculating the distance between the two Mel frequency cepstral coefficients, and taking this distance as the frequency-domain difference.

[0041] Specifically, the calculation formula for the frequency-domain difference is: ; in the formula, represents the frequency-domain difference between the verified analysis speech and one of the comparison analysis speeches; represents the Mel frequency cepstral coefficients of the verified analysis speech; represents the Mel frequency cepstral coefficients of the comparison analysis speech; represents (Dynamic Time Warping) distance.

[0042] The greater the distance between the Mel frequency cepstral coefficients of the identity person and other biological persons when they say the same verified speech text, the greater the difference between the two sound data, and thus the greater the frequency-domain difference between the two speeches.

[0043] In another embodiment, the steps of obtaining the frequency-domain difference between the verified analysis speech and the comparison analysis speech include: obtaining the Mel frequency cepstral coefficients of the verified analysis speech and the comparison analysis speech, and calculating the distance between the two Mel frequency cepstral coefficients; obtaining the first-order difference results of the two Mel frequency cepstral coefficients, and obtaining the distance between the two first-order difference results; taking the sum of the two distances as the frequency-domain difference.

[0044] Specifically, the calculation formula for the frequency-domain difference is:

[0045] ; in the formula, Indicates the frequency domain difference between the verification analysis speech and one of the comparative analysis speeches; Indicates the Mel-frequency cepstral coefficients of the verification analysis speech; Indicates the Mel-frequency cepstral coefficients of the comparative analysis speech; Indicates the first-order difference of the Mel-frequency cepstral coefficients of the verification analysis speech; Indicates the first-order difference of the Mel-frequency cepstral coefficients of the comparative analysis speech; Indicates Distance.

[0046] In this formula The static characteristics of the frequency domain of the verification analysis speech and the comparative analysis speech are considered; The dynamic characteristics of the frequency domain of the verification analysis speech and the comparative analysis speech, combining the differences in dynamic characteristics and static characteristics, thus improving the accuracy of calculating the frequency domain difference.

[0047] In one embodiment, the steps of obtaining the time domain difference between the verification analysis speech and the comparative analysis speech include: obtaining multiple time domain characteristics of the verification analysis speech and the comparative analysis speech, calculating the differences of the respective time domain characteristics of the verification analysis speech and the comparative analysis speech, and taking the sum of the differences of the multiple time domain characteristics as the time domain difference.

[0048] Extract multiple time domain characteristics of the verification analysis speech and the comparative analysis speech from the waveform signals of the verification analysis speech and the comparative analysis speech. The time domain characteristics extracted in this embodiment include: zero-crossing rate, energy, fundamental frequency. In other embodiments, other characteristic values can be extracted according to actual situations.

[0049] Specifically, the calculation formula for the time domain difference is: ; In the formula, Indicates the time domain difference between the verification analysis speech and one of the comparative analysis speeches; Indicates the total number of time domain characteristics; Indicates the th time domain characteristic value of the verification analysis speech; Indicates the th time domain characteristic value of the comparative analysis speech.

[0050] In another embodiment, the steps of obtaining the time domain difference between the verification analysis speech and the comparative analysis speech include: obtaining multiple time domain characteristics of the verification analysis speech and the comparative analysis speech, calculating the differences of the respective time domain characteristics of the verification analysis speech and the comparative analysis speech, calculating the squared results of the differences of the respective time domain characteristics, taking this result as the amplified difference, and taking the sum of the amplified differences of the multiple time domain characteristics as the time domain difference.

[0051] Specifically, the calculation formula for the time domain difference is: ; In the formula, Indicates the time-domain difference between the verified analysis voice and one of the comparative analysis voices; Indicates the total number of time-domain features; Indicates the th time-domain feature value of the verified analysis voice; Indicates the th time-domain feature value of the comparative analysis voice.

[0052] In this formula, the squared results of the differences of each time-domain feature are calculated, which can improve the sensitivity of the final feature difference to the change of the time-domain feature difference. For example, there is a set of data, , , , when changes, , then the impact of the final time-domain difference is , which is more sensitive to the change.

[0053] For an identity person and any other biological person, there is a time-domain difference and a frequency-domain difference. The time-domain difference and the frequency-domain difference are normalized to obtain the normalized result of the time-domain difference and the normalized result of the frequency-domain difference, and then summed to obtain the difference degree between the identity person and any other biological person.

[0054] For any speech verification text, there is an identity person and multiple other biological persons, so there are multiple difference degrees. The multiple difference degrees are added to obtain the speech difference of the speech verification text.

[0055] Each verification speech text corresponds to a speech difference. The larger the speech difference, the greater the difference in the voice data between the identity person and other biological persons when reading the verification speech text. Therefore, the larger the speech difference, the easier it is to distinguish the identity person and other biological persons when reading the speech text, improving the accuracy of subsequent verification. Therefore, the verification speech text with the largest speech difference is selected as the verification text.

[0056] S3: Collect multiple biometric information of the verification personnel. The multiple biometric information includes: identity voice information, face information, fingerprint information. The identity voice information is the voice data when the verification personnel read the verification text. Input the multiple biometric information into a preset verification model to obtain the probability that the multiple biometric information passes the verification;

[0057] The verifier refers to the person who requests to log in to the digital human system. The process of requesting to log in to the digital human system can be understood as the process of verifying whether a natural person is an identity person. In this process, multiple biometric information of the verifier is collected, such as identity voice information, face information, fingerprint information, iris information, etc.; in this embodiment, the identity voice information, face information, and fingerprint information of the verifier are collected. Among them, the identity voice information refers to the verifier reading the verification text.

[0058] Input the collected identity voice information, face information, and fingerprint information into their respective preset verification models to obtain the probability of passing the verification for each.

[0059] Among them, the voice verification model can be DeepSpeech. The face verification model can be FaceNet or VGG-Face. The fingerprint verification model can be DeepPrint.

[0060] S4: Divide weights for multiple biometric information according to the difference degree, where the difference degree is proportional to the weight of the probability of the identity voice information passing the verification, obtain the weighted value after weighting the probabilities of multiple biometric information passing the verification, use the weighted value of the probabilities of multiple biometric information passing the verification as the final verification probability, and verify the verifier based on the final verification probability.

[0061] In one embodiment, the calculation formula for the weight of the probability of the identity voice information passing the verification is:

[0062] ; in the formula, represents the weight of the probability of the identity voice information passing the verification; represents the voice difference corresponding to the verification text.

[0063] In the formula, the greater the voice difference, it indicates that when reading the same verification text, the difference between the voice data of the identity person and that of other people is greater. Therefore, the attention to the voice data of the verifier can be improved. It can also be understood that the greater the voice difference, the higher the confidence level of the probability of the identity voice verification information obtained in the voice verification model passing the verification. Therefore, when the voice difference becomes larger, the weight of the probability of the identity voice information passing the verification becomes larger.

[0064] In another embodiment, the calculation formula for the weight of the probability of the identity voice information passing the verification is:

[0065] ; in the formula, represents the weight of the probability of the identity voice information passing the verification; is a hyperbolic function; represents the voice difference corresponding to the verification text.

[0066] In this formula, through the hyperbolic function, The value range of is controlled between 0 and 1, so as to effectively suppress the weight of the probability that the identity voice information passes the verification when the voice difference is small.

[0067] The calculation formula for the weight of the probability that the face information passes the verification is: ; In the formula, represents the weight of the probability that the identity voice information passes the verification; represents the weight of the probability that the face information passes the verification; represents the weight of the probability that the identity voice information passes the verification.

[0068] The calculation formula for the weight of the probability that the fingerprint information passes the verification is: ; represents the weight of the probability that the fingerprint information passes the verification; represents the weight of the probability that the identity voice information passes the verification.

[0069] Obtain the weighted value after weighting the probabilities that multiple biometric information pass the verification, and use the weighted value of the probabilities that multiple biometric information pass the verification as the final verification probability, and verify the verifier based on the final verification probability.

[0070] Specifically, the calculation formula for the final verification probability is:

[0071] ; In the formula, represents the final verification probability; represents the weight of the probability that the fingerprint information passes the verification; represents the weight of the probability that the face information passes the verification; represents the weight of the probability that the identity voice information passes the verification; represents the probability that the identity voice information passes the verification; represents the probability that the fingerprint information passes the verification; represents the probability that the face information passes the verification.

[0072] Verify the verifier based on the final verification probability, set a passing threshold, and when the response is greater than the passing threshold for the final verification probability, determine that the verifier passes the verification. In this embodiment, the passing threshold is set to 0.9, and in other embodiments, it can be adjusted according to the actual situation.

[0073] The embodiment of the present application also discloses a multi-modal digital human identity verification system, including a processor and a memory, and the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a multi-modal digital human identity verification method and system method according to the present application are implemented.

[0074] The above system further includes other components well-known to those skilled in the art, such as a communication bus and a communication interface. Their settings and functions are known in the art, so they will not be elaborated here.

[0075] The above are all preferred embodiments of this application. The protection scope of this application is not limited thereby. Therefore, all equivalent changes made according to the structure, shape, and principle of this application shall be covered within the protection scope of this application.

Claims

1. A multi-modal digital human authentication method, characterized in that, Including the steps: preset multiple verification voice texts, obtain the voice data of the identity person reading different verification voice texts as the verification analysis voice, and obtain the voices of a preset number of biological persons reading different verification voice texts as the comparison analysis voice; Compare the differences between the verification analysis voice and the comparison analysis voice to obtain the difference degree, including: obtaining the frequency domain difference and the time domain difference between the verification analysis voice and the comparison analysis voice, and taking the sum of the normalized result of the time domain difference and the normalized result of the frequency domain difference as the difference degree between the verification analysis voice and the comparison analysis voice; taking the sum of multiple difference degrees corresponding to the same verification voice as the voice difference, and taking the verification voice text with the largest voice difference as the verification text; The steps for obtaining the frequency-domain difference between the verification analysis speech and the comparison analysis speech include: obtaining the Mel-frequency cepstral coefficients of the verification analysis speech and the comparison analysis speech, and calculating the distance between the two Mel-frequency cepstral coefficients; obtaining the first-order difference results of the two Mel-frequency cepstral coefficients, and obtaining the distance between the two first-order difference results; taking the sum of the two distances as the frequency-domain difference; The steps of obtaining the time domain difference between the verification analysis voice and the comparison analysis voice include: obtaining multiple time domain features of the verification analysis voice and the comparison analysis voice, calculating the differences between the time domain features of the verification analysis voice and the comparison analysis voice, calculating the square results of the differences between the time domain features, taking this result as the amplified difference, and taking the sum of the amplified differences of multiple time domain features as the time domain difference; Collect multiple biometric information of the verification personnel, and the multiple biometric information includes: identity voice information, face information, fingerprint information, where the identity voice information is the voice data when the verification personnel read the verification text. Input the multiple biometric information into a preset verification model to obtain the probability of the multiple biometric information passing the verification; divide weights for the multiple biometric information according to the difference degree, where the difference degree is proportional to the weight of the probability of the identity voice information passing the verification, obtain the weighted value after weighting the probabilities of the multiple biometric information passing the verification, take the weighted value of the probabilities of the multiple biometric information passing the verification as the final verification probability, and verify the verification personnel based on the final verification probability; The calculation formula for the weight of the probability that the identity voice information passes verification is as follows: ; In the formula, represents the weight of the probability that the identity voice information passes verification; represents the voice difference corresponding to the verification text.

2. The multimodal digital human identity authentication method according to claim 1, characterized in that, Another method for obtaining the frequency-domain difference between the verification analysis speech and the comparison analysis speech includes: obtaining the Mel-frequency cepstral coefficients of the verification analysis speech and the comparison analysis speech, and calculating the distance between the two Mel-frequency cepstral coefficients, and using this distance as the frequency-domain difference.

3. A multimodal digital human identity authentication method according to claim 1, characterized in that, Another method of obtaining the time domain difference between the verification analysis voice and the comparison analysis voice includes: obtaining multiple time domain features of the verification analysis voice and the comparison analysis voice, calculating the differences between the time domain features of the verification analysis voice and the comparison analysis voice, and taking the sum of the differences of the multiple time domain features as the time domain difference.

4. A multimodal digital human identity authentication method according to claim 1, characterized in that, The multiple time domain features include: zero crossing rate, energy, fundamental frequency.

5. A multimodal digital human identity authentication method according to claim 1, characterized in that: The method of verifying the verification personnel based on the final verification probability includes: setting a passing threshold, and in response to the final verification probability being greater than the passing threshold, determining that the verification personnel pass the verification.

6. A multi-modal digital human authentication system, characterized in that, Including: A processor and a memory, the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the multi-modal digital human identity verification method according to any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Intelligent automobile remote unlocking method and device based on multi-modal biological characteristics

    CN118470836A

  • Method of processing a speech signal for speaker recognition and electronic apparatus implementing same

    US20190244612A1