Multi-mode digital human identity verification method and system

By calculating the degree of pronunciation difference between identity humans and other biological humans when reading and verifying pronunciation texts, and determining the weight of biometric information based on the degree of difference, the problem that multimodal biometric technology cannot accurately distinguish verified personnel under special circumstances is solved, and the security of identity verification is improved.

CN120068036AActive Publication Date: 2025-05-30INSPUR SMART SUPPLY CHAIN TECH (SHANDONG) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510549980.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

Existing multimodal biometric technology cannot accurately distinguish different verification personnel under special circumstances, resulting in insufficient security of identity verification.

Method used

By presetting multiple verification pronunciation texts, obtain the sound data of the identities and other biological humans reading these pronunciation texts, calculate the pronunciation difference degree, and determine the weight of the biometric information based on the difference degree, and combine multiple biometric information for identity verification.

Benefits of technology

It improves the security of identity verification, reduces the situation where the identity of a verified person cannot be verified in extreme cases, and enhances the ability to distinguish different verified persons.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068036A_ABST
    Figure CN120068036A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of digital human identity authentication, in particular to a multi-mode digital human identity authentication method and system. The method comprises the steps that multiple pieces of biological characteristic information of a verification person are collected, identity voice information is voice data when the verification person reads verification characters, the multiple pieces of biological characteristic information are input into a preset verification model, and the probability that the multiple pieces of biological characteristic information pass verification is obtained; dividing weights for the multiple pieces of biological characteristic information according to the difference degree, obtaining weighted values of the multiple pieces of biological information after the probability of passing the verification is weighted, and taking the weighted values of the probability of passing the verification of the multiple pieces of biological information as final verification probabilities, and verifying the verification personnel based on the final verification probability. The method has the advantages that verification personnel can be distinguished and recognized conveniently under extreme conditions, and verification safety is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of digital human identity verification, and in particular, to a multi-modal digital human identity verification method and system. Background Art

[0002] With the development of artificial intelligence and virtual reality technologies, digital humans (virtual humans) have gradually become important identity representatives in the digital world. Digital humans can interact with humans in various ways, simulating the behaviors and emotions of real people. In this context, digital human identity verification has become an important means to ensure the security of the interaction between the virtual world and the real world. To ensure the credibility and security of digital humans in various scenarios, it is necessary to verify the user identities of digital humans. Traditional user identity verification methods include password verification, ID number verification, etc. These methods are prone to leakage, leading to problems such as identity theft and fraud. With the development and progress of digital technologies, biometric technologies have gradually been applied to security verification. Common biometric technologies include fingerprint recognition, face recognition, etc. Due to the difficulty of forging biometric features, this technology can improve the security of identity verification to a certain extent. However, there are also some problems in the application process of biometric recognition technologies. For example, in the process of fingerprint verification, if the user's fingerprint is dirty or damaged, it may cause changes in the fingerprint during user verification, ultimately resulting in fingerprint verification failure; or in the process of face recognition verification, the user's facial data may be affected by factors such as light, angle, and makeup, ultimately resulting in inaccurate user verification. Therefore, multi-modal biometric technologies have received more attention from industry technicians. Multi-modal biometric technologies mainly combine the verification of multiple biometric features. For example, the voice information and facial information of the verifier can be collected, and the voice information, facial information, and preset verification information are compared, and the final verification result of the verifier is obtained by synthesizing the two verification results.

[0003] In the process of calculating the final comprehensive verification result of multi-modal biometric technologies, the weights of each biometric information are the same. This method cannot make accurate judgments for some special situations. For example, in a multi-modal biometric system, multiple biometric technologies are collected, and only one of them has a large difference. However, due to the small proportion of this biometric feature, its impact on the final verification result is relatively small, and ultimately it may lead to the situation where the biometric system cannot distinguish between the two verifiers. Therefore, the security of multi-modal biometric technologies in related technologies needs to be further improved. Summary of the Invention

[0004] To solve the problem of being unable to distinguish different verifiers in special situations and improve the security of identity verification, this application provides a multi-modal digital human identity verification method and system.

[0005] In a first aspect, the present application provides a multimodal digital human identity verification method, which adopts the following technical solutions: A multimodal digital human identity verification method, which presets multiple verification voice texts, obtains the voice data of the identity person reading different verification voice texts as the verification analysis voice, and obtains the voices of a preset number of biological persons reading different verification voice texts as the comparison analysis voice; compares the differences between the verification analysis voice and the comparison analysis voice, obtains the difference degree, takes the sum of multiple difference degrees corresponding to the same verification voice text as the voice difference, and takes the verification voice text with the largest voice difference as the verification text; Collects multiple biometric information of the verification personnel. The multiple biometric information includes: identity voice information, face information, and fingerprint information. Among them, the identity voice information is the voice data when the verification personnel reads the verification text. Inputs the multiple biometric information into a preset verification model to obtain the probability of the multiple biometric information passing the verification; divides weights for the multiple biometric information according to the difference degree, where the difference degree is directly proportional to the weight of the probability of the identity voice information passing the verification; obtains the weighted value after weighting the probabilities of the multiple biometric information passing the verification, takes the weighted value of the probabilities of the multiple biometric information passing the verification as the final verification probability, and verifies the verification personnel based on the final verification probability.

[0006] In the present application, the differences between the identity person and other biological persons reading different verification voice texts are calculated. If the voice difference corresponding to one of the verification voice texts is larger, it means that the voice data of the identity person when reading this voice verification text is easier to distinguish. Therefore, selecting this verification voice text can improve the ability to distinguish the verification personnel during subsequent verification of the verification personnel, reduce the situation where the verification personnel cannot be distinguished, and improve the security of identity verification. At the same time, in the present application, according to the size of the voice difference, weights are set for multiple biometric features when comprehensively verifying the identity of the verification personnel subsequently, so that when the voice difference is large, more attention can be paid to the probability of the voice information passing the verification, thereby further improving the ability to distinguish different verification personnel. Compared with the traditional method where the weights of various biometric information are the same during verification of different digital humans, in the present application, the biometric information is dynamically changed during identity verification of different digital humans, which is more in line with the unique biometric features of the identity person, such as voice features, and improves the ability to distinguish whether the verification personnel is the identity person.

[0007] Optionally, the step of comparing the differences between the verification analysis voice and the comparison analysis voice and obtaining the difference degree includes: obtaining the frequency domain difference and time domain difference between the verification analysis voice and the comparison analysis voice, and taking the sum of the normalized result of the time domain difference and the normalized result of the frequency domain difference as the difference degree between the verification analysis voice and the comparison analysis voice.

[0008] Compare the characteristics of the verification analysis speech and the comparative analysis speech from multiple dimensions of time-domain differences and frequency-domain differences, so as to improve the accuracy of the difference degree in subsequent calculations.

[0009] Optionally, the steps for obtaining the frequency-domain difference between the verification analysis speech and the comparative analysis speech include: obtaining the Mel-frequency cepstral coefficients of the verification analysis speech and the comparative analysis speech, and calculating the distance between the two Mel-frequency cepstral coefficients, and taking this distance as the frequency-domain difference.

[0010] The distance between the Mel-frequency cepstral coefficients of the verification analysis speech and the comparative analysis speech reflects the frequency-domain difference between the verification speech and the comparative analysis speech, and realizes the calculation of the frequency-domain difference.

[0011] Optionally, obtain the Mel-frequency cepstral coefficients of the verification analysis speech and the comparative analysis speech, and calculate the distance between the two; obtain the first-order difference results of the two Mel-frequency cepstral coefficients, and obtain the distance between the two first-order difference results; take the sum of the two distances as the frequency-domain difference.

[0012] In this method, by combining the differences in the first-order difference results of the Mel-frequency cepstral coefficients and considering the dynamic differences between the verification analysis speech and the comparative analysis speech, the accuracy of the frequency-domain difference calculation is improved.

[0013] Optionally, the steps for obtaining the time-domain difference between the verification analysis speech and the comparative analysis speech include: obtaining multiple time-domain features of the verification analysis speech and the comparative analysis speech, calculating the differences of the respective time-domain features of the verification analysis speech and the comparative analysis speech, and taking the sum of the differences of the multiple time-domain features as the time-domain difference.

[0014] Obtain multiple time-domain features of the verification analysis speech and the comparative analysis speech, and compare the differences in the time-domain features of the verification analysis speech and the comparative analysis speech, so as to obtain the time-domain difference between the verification analysis speech and the comparative analysis speech.

[0015] Optionally, the steps for obtaining the time-domain difference between the verification analysis speech and the comparative analysis speech include: obtaining multiple time-domain features of the verification analysis speech and the comparative analysis speech, calculating the differences of the respective time-domain features of the verification analysis speech and the comparative analysis speech, calculating the squared results of the differences of the respective time-domain features, taking this result as the amplified difference, and taking the sum of the amplified differences of the multiple time-domain features as the time-domain difference.

[0016] In this method, the square value of the time-domain feature difference is calculated based on the calculation of the time-domain feature difference, so as to improve the sensitivity of the finally calculated time-domain difference to the time-domain feature difference between the verified speech and the comparative speech.

[0017] Optionally, multiple time-domain features include: zero-crossing rate, energy, and fundamental frequency.

[0018] Optionally, the calculation formula for the weight of the probability that the identity speech information passes verification is: ; in the formula, represents the weight of the probability that the identity speech information passes verification; represents the speech difference corresponding to the verified text.

[0019] Optionally, the method for verifying the verifier based on the final verification probability includes: setting a passing threshold, and in response to the final verification probability being greater than the passing threshold, determining that the verifier passes the verification.

[0020] In a second aspect, the present application provides a multimodal digital human identity verification system, adopting the following technical solution: A multimodal digital human identity verification system includes: a processor and a memory, and the memory stores computer program instructions, which when executed by the processor implement a multimodal digital human identity verification method according to the above.

[0021] The beneficial effect is: generating a computer program for the above multimodal digital human identity verification method and storing it in the memory to be loaded and executed by the processor, so that, making a system according to the memory and the processor is convenient to use.

[0022] The present application has the following technical effects: in the present application, different verification texts are set for each digital human, and weights are determined for multiple biometric information participating in the subsequent verification based on the difference in the voice data when the identity person and other biological persons read the verification texts. Therefore, different digital humans perform identity verification based on different weights, and further, the situation where the identity of the verifier cannot be verified in extreme cases (multiple biometric features are similar and one biometric information has a large difference) can be reduced. Description of the Drawings

[0023] Figure 1 is the flowchart of a multimodal digital human identity verification method according to an embodiment of the present application. Detailed Embodiments

[0024] The embodiments of the present application disclose a multimodal digital human identity verification method and system, which enable the identity person and other biological persons to read the same set of verification voice texts, obtain the difference degree when the identity person and other biological persons read the same section of voice, and select the verification voice text with the largest difference degree as the verification text. Therefore, different identity persons correspond to different verification texts when verifying their respective identities.

[0025] When subsequently verifying the identity of the verification personnel, multiple biometric features of the verification personnel are collected and input into a preset model to obtain the probabilities of passing verification for each biometric feature. Subsequently, weights are set for different biometric information according to the difference degree, and the final verification probability of the verification personnel is calculated by combining the probabilities of passing verification weighted by multiple biometric information. The verification result is obtained based on the final verification probability of the verification personnel, and the identity verification of the verification personnel is completed.

[0026] During the identity verification process of the verification personnel, the weights of the probabilities of passing verification of the biometric information of different verification personnel are different. Therefore, the situation where the verification system cannot distinguish the verification personnel in special cases can be reduced.

[0027] Refer to Figure 1 , the multimodal digital human identity verification method includes steps S1 - S4.

[0028] S1: Preset multiple verification voice texts, obtain the voice data of the identity person reading different verification voice texts as verification analysis voices, and obtain the voices of a preset number of biological persons reading different verification voice texts as comparison analysis voices; The identity person refers to the owner of the digital human (virtual human). The identity person corresponds to the digital human one by one and is unique. Other biological persons refer to any other natural persons except the identity person. Multiple other biological persons are selected. For example, 100 or 1000 other biological persons can be selected. In this embodiment, 100 other biological persons are preset. The specific quantity can be determined according to the actual situation.

[0029] The verification voice text is text or numbers, and the length of the verification voice text can be set by itself according to the actual situation. Exemplarily, the verification voice text can be: "24859", "35559", "The sun along the mountain bows", "Innovative and enterprising", etc.

[0030] The identity person reads the verification voice text, and the voice data is collected during the reading process to obtain a set of verification analysis voices.

[0031] The other biological persons read this set of verification voices, and at the same time, the voice data during the reading process is collected to obtain a set of comparison analysis voices.

[0032] One verification voice text corresponds to one verification analysis voice and multiple comparison analysis voices.

[0033] S2: Compare and verify the differences between the verification analysis speech and the comparison analysis speech. The sum of the differences of multiple degrees corresponding to the same verification speech is the speech difference, and the verification speech text with the largest speech difference is used as the verification text. When different people read the same verification speech text, there will be certain differences, such as in speech rate, intonation, loudness, etc. Therefore, by calculating the time-domain difference and frequency-domain difference between the verification analysis speech and the comparison analysis speech, the difference degree between the verification analysis speech and the comparison analysis speech can be obtained.

[0034] In one embodiment, the steps of obtaining the frequency-domain difference between the verification analysis speech and the comparison analysis speech include: obtaining the Mel Frequency Cepstral Coefficients (MFCCs) of the verification analysis speech and the comparison analysis speech, and calculating the distance between the two MFCCs, and taking this distance as the frequency-domain difference.

[0035] Specifically, the calculation formula for the frequency-domain difference is: ; where represents the frequency-domain difference between the verification analysis speech and one of the comparison analysis speeches; represents the Mel Frequency Cepstral Coefficients of the verification analysis speech; represents the Mel Frequency Cepstral Coefficients of the comparison analysis speech; represents (Dynamic Time Warping, DTW) distance.

[0036] The greater the distance between the Mel Frequency Cepstral Coefficients of the identity person and other biological persons when they say the same verification speech text, the greater the difference between the two sound data, and thus the greater the frequency-domain difference between the two speeches.

[0037] In another embodiment, the steps of obtaining the frequency-domain difference between the verification analysis speech and the comparison analysis speech include: obtaining the Mel Frequency Cepstral Coefficients of the verification analysis speech and the comparison analysis speech, and calculating the distance between the two MFCCs; obtaining the first-order difference results of the two MFCCs, and obtaining the distance between the two first-order difference results; and taking the sum of the two distances as the frequency-domain difference.

[0038] Specifically, the calculation formula for the frequency-domain difference is: ; where represents the frequency-domain difference between the verification analysis speech and one of the comparison analysis speeches; represents the Mel Frequency Cepstral Coefficients of the verification analysis speech; Denote the Mel Frequency Cepstral Coefficients (MFCCs) for comparative analysis of speech; Denote the first-order difference of the Mel Frequency Cepstral Coefficients for verification analysis of speech; Denote the first-order difference of the Mel Frequency Cepstral Coefficients for comparative analysis of speech; Denote Distance.

[0039] In this formula The static characteristics of the frequency domain of the verification analysis speech and the comparative analysis speech are considered; The dynamic characteristics of the frequency domain of the verification analysis speech and the comparative analysis speech, combining the differences in dynamic characteristics and static characteristics, thereby improving the accuracy of calculating the frequency domain differences.

[0040] In one embodiment, the steps of obtaining the time domain differences between the verification analysis speech and the comparative analysis speech include: obtaining multiple time domain characteristics of the verification analysis speech and the comparative analysis speech, calculating the differences of the respective time domain characteristics of the verification analysis speech and the comparative analysis speech, and taking the sum of the differences of the multiple time domain characteristics as the time domain difference.

[0041] Extract multiple time domain characteristics of the verification analysis speech and the comparative analysis speech from the waveform signals of the verification analysis speech and the comparative analysis speech. The time domain characteristics extracted in this embodiment include: zero-crossing rate, energy, fundamental frequency. In other embodiments, other characteristic values can be extracted according to actual situations.

[0042] Specifically, the calculation formula for the time domain difference is: ; where Denote the time domain difference between the verification analysis speech and one of the comparative analysis speeches; Denote the total number of time domain characteristics; Denote the th time domain characteristic value of the verification analysis speech; Denote the th time domain characteristic value of the comparative analysis speech.

[0043] In another embodiment, the steps of obtaining the time domain differences between the verification analysis speech and the comparative analysis speech include: obtaining multiple time domain characteristics of the verification analysis speech and the comparative analysis speech, calculating the differences of the respective time domain characteristics of the verification analysis speech and the comparative analysis speech, calculating the squared results of the differences of the respective time domain characteristics, taking this result as the amplified difference, and taking the sum of the amplified differences of the multiple time domain characteristics as the time domain difference.

[0044] Specifically, the calculation formula for the time domain difference is: ; where Denote the time domain difference between the verification analysis speech and one of the comparative analysis speeches; Denote the total number of time domain characteristics; indicating the th time-domain eigenvalue for verifying and analyzing speech; indicating the th time-domain eigenvalue for comparative analysis of speech.

[0045] In this formula, the squared results of the differences of each time-domain feature are calculated, which can improve the sensitivity of the final feature difference to the change of the time-domain feature difference. For example, there is a set of data, , , , when changes, , then the impact of the final time-domain difference is , and it is more sensitive to the change.

[0046] For an identity person and any other biological person, there is a time-domain difference and a frequency-domain difference. The time-domain difference and the frequency-domain difference are normalized to obtain the normalized results of the time-domain difference and the frequency-domain difference, and then they are summed to obtain the difference degree between the identity person and any other biological person.

[0047] For any speech verification text, there is an identity person and multiple other biological persons, so there are multiple difference degrees. The multiple difference degrees are added to obtain the speech difference of the speech verification text.

[0048] Each verification speech text corresponds to a speech difference. The larger the speech difference, the greater the difference in voice data between the identity person and other biological persons when reading the verification speech text. Therefore, the larger the speech difference, the easier it is to distinguish the identity person and other biological persons when reading the speech text, improving the accuracy of subsequent verification. Therefore, the verification speech text with the largest speech difference is selected as the verification text.

[0049] S3: Collect multiple biometric information of the verification personnel. The multiple biometric information includes: identity voice information, face information, fingerprint information, where the identity voice information is the voice data when the verification personnel read the verification text. Input the multiple biometric information into a preset verification model to obtain the probability that the multiple biometric information passes the verification; The verification personnel refer to the personnel who request to log in to the digital human system. The process of requesting to log in to the digital human system can be understood as the process of verifying whether a natural person is an identity person. In this process, multiple biometric information of the verification personnel is collected, such as identity voice information, face information, fingerprint information, iris information, etc.; in this embodiment, the identity voice information, face information, and fingerprint information of the verification personnel are collected. The identity voice information refers to the verification personnel reading the verification text.

[0050] Input the collected identity voice information, face information, and fingerprint information into their respective preset verification models to obtain the probabilities of passing the verification for each.

[0051] Among them, the voice verification model can be DeepSpeech. The face verification model can be FaceNet or VGG-Face. The fingerprint verification model can be DeepPrint.

[0052] S4: Divide the weights for multiple biometric information according to the degree of difference, where the degree of difference is directly proportional to the weight of the probability of the identity voice information passing the verification. Obtain the weighted value after weighting the probabilities of multiple biometric information passing the verification, and use the weighted value of the probabilities of multiple biometric information passing the verification as the final verification probability, and verify the verifier based on the final verification probability.

[0053] In one embodiment, the calculation formula for the weight of the probability of the identity voice information passing the verification is: ; In the formula, represents the weight of the probability of the identity voice information passing the verification; represents the voice difference corresponding to the verification text.

[0054] In the formula, the greater the voice difference, it indicates that when reading the same verification text, the difference between the voice data of the identity person and that of others is greater. Therefore, it can improve the attention to the voice data of the verifier. It can also be understood that the greater the voice difference, the higher the confidence level of the probability of the identity voice verification information obtained in the voice verification model passing the verification. Therefore, when the voice difference becomes larger, the weight of the probability of the identity voice information passing the verification becomes larger.

[0055] In another embodiment, the calculation formula for the weight of the probability of the identity voice information passing the verification is: ; In the formula, represents the weight of the probability of the identity voice information passing the verification; is a hyperbolic function; represents the voice difference corresponding to the verification text.

[0056] In this formula, the hyperbolic function is used to control the value range of between 0 and 1, so that when the voice difference is small, the weight of the probability of the identity voice information passing the verification can be effectively suppressed.

[0057] The calculation formula for the weight of the probability of the face information passing the verification is: ; In the formula, represents the weight of the probability of the identity voice information passing the verification; represents the weight of the probability of the face information passing the verification; The weight representing the probability that the identity voice information passes the verification.

[0058] The calculation formula for the weight of the probability that the fingerprint information passes the verification is: ; represents the weight of the probability that the fingerprint information passes the verification; The weight representing the probability that the identity voice information passes the verification.

[0059] Obtain the weighted value after weighting the probabilities that multiple biometric information passes the verification, use the weighted value of the probabilities that multiple biometric information passes the verification as the final verification probability, and verify the verifier based on the final verification probability.

[0060] Specifically, the calculation formula for the final verification probability is: ; In the formula, represents the final verification probability; represents the weight of the probability that the fingerprint information passes the verification; represents the weight of the probability that the face information passes the verification; represents the weight of the probability that the identity voice information passes the verification; represents the probability that the identity voice information passes the verification; represents the probability that the fingerprint information passes the verification; represents the probability that the face information passes the verification.

[0061] Verify the verifier based on the final verification probability, set a passing threshold, and when the response is that the final verification probability is greater than the passing threshold, determine that the verifier passes the verification. In this embodiment, the passing threshold is set to 0.9, and in other embodiments, it can be adjusted according to actual situations.

[0062] The embodiment of the present application also discloses a multimodal digital human identity verification system, including a processor and a memory. The memory stores computer program instructions, and when the computer program instructions are executed by the processor, a multimodal digital human identity verification method and system method according to the present application are implemented.

[0063] The above system also includes other components well-known to those skilled in the art such as a communication bus and a communication interface, and their settings and functions are known in the art, so they will not be elaborated here.

[0064] The above are all the preferred embodiments of the present application. The protection scope of the present application is not limited by this. Therefore, all equivalent changes made according to the structure, shape, and principle of the present application should be covered within the protection scope of the present application.

Claims

1. A multimodal digital human identity authentication method, characterized in that: The method comprises the following steps: presetting a plurality of verification voice characters, obtaining sound data of an identity person reading different verification voice characters as verification analysis voice, obtaining a preset number of sounds of biological persons reading different verification voice characters as comparative analysis voice; comparing the difference between the verification analysis voice and the comparative analysis voice to obtain the difference degree, taking the sum of multiple difference degrees corresponding to the same verification voice as the voice difference, and taking the verification voice character with the largest voice difference as the verification character; Collect multiple biometric information of the verification personnel, which include: identity voice information, facial information, and fingerprint information, where the identity voice information is the sound data when the verification personnel reads the verification text, input the multiple biometric information into the preset verification model, and obtain the probability that the multiple biometric information passes the verification; assign weights to the multiple biometric information according to the difference, where the difference is proportional to the weight of the probability that the identity voice information passes the verification, obtain the weighted value of the weighted probability of the multiple biometric information passing the verification, use the weighted value of the probability that the multiple biometric information passes the verification as the final verification probability, and verify the verification personnel based on the final verification probability.

2. A multimodal digital human identity authentication method according to claim 1, characterized in that: The steps of comparing the difference between the verification analysis speech and the comparative analysis speech and obtaining the difference degree include: obtaining the frequency domain difference and time domain difference between the verification analysis speech and the comparative analysis speech, and taking the sum of the normalized result of the time domain difference and the normalized result of the frequency domain difference as the difference degree between the verification analysis speech and the comparative analysis speech.

3. A multi-modal digital human identity authentication method according to claim 2, characterized in that: The steps of obtaining the frequency domain difference between the verification analysis speech and the comparative analysis speech include: obtaining the Mel frequency cepstral coefficients of the verification analysis speech and the comparative analysis speech, calculating the difference between the two Mel frequency cepstral coefficients, and calculating the difference between the two Mel frequency cepstral coefficients. distance, the Distance as frequency domain difference.

4. A multi-modal digital human identity authentication method according to claim 2, characterized in that: The steps of obtaining the frequency domain difference between the verification analysis speech and the comparative analysis speech include: obtaining the Mel frequency cepstral coefficients of the verification analysis speech and the comparative analysis speech, calculating the difference between the two Mel frequency cepstral coefficients, and calculating the difference between the two Mel frequency cepstral coefficients. Distance; Get the first-order difference results of two Mel-frequency cepstral coefficients, and get the distance between the two first-order difference results distance; The sum of the distances is taken as the frequency domain difference.

5. A multi-modal digital human identity authentication method according to claim 2, characterized in that: The step of obtaining the time domain difference between the verification analysis speech and the comparative analysis speech includes: obtaining multiple time domain features of the verification analysis speech and the comparative analysis speech, calculating the difference between the various time domain features of the verification analysis speech and the comparative analysis speech, and taking the sum of the differences of the multiple time domain features as the time domain difference.

6. A multi-modal digital human identity authentication method according to claim 2, characterized in that: The step of obtaining the time domain difference between the verification analysis speech and the comparative analysis speech includes: obtaining multiple time domain features of the verification analysis speech and the comparative analysis speech, calculating the difference between the various time domain features of the verification analysis speech and the comparative analysis speech, calculating the square result of the difference between each time domain feature, taking the result as the amplified difference, and taking the sum of the amplified differences of multiple time domain features as the time domain difference.

7. A multi-modal digital human identity authentication method according to claim 6, characterized in that: Multiple time domain features include: zero crossing rate, energy, and fundamental frequency.

8. A multi-modal digital human identity authentication method according to claim 1, characterized in that: The calculation formula for the weight of the probability that the identity voice information passes the verification is: ; In the formula, A weight representing the probability that the identity voice information is verified; Indicates the pronunciation difference corresponding to the verification text.

9. A multi-modal digital human identity authentication method according to claim 1, characterized in that: The method for verifying a verification person based on a final verification probability includes: setting a pass threshold, and in response to the final verification probability being greater than the pass threshold, determining that the verification person has passed the verification.

10. A multimodal digital human identity authentication system, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a multimodal digital human identity authentication method according to any one of claims 1-9 is implemented.

Citation Information

Patent Citations

  • Intelligent automobile remote unlocking method and device based on multi-modal biological characteristics

    CN118470836A

  • Voice content recognition method and system

    CN118865951A

  • Non-inductive payment system and method based on adaptive multi-modal fusion and behavior prediction

    CN119539815A

  • Speech recognition method and device

    KR1019980076309A

  • Method of processing a speech signal for speaker recognition and electronic apparatus implementing same

    US20190244612A1