Multi-mode voiceprint identity verification method and device with enhanced security

By acquiring and fusing acoustic features and multimodal voiceprint information, and determining the phoneme loss tag and feature interaction deviation, accurate and reliable identity authentication in multimodal voiceprint authentication is achieved, solving the problem of instability in identity authentication under environmental and condition changes in the prior art.

CN120378113APending Publication Date: 2025-07-25MINAMI ACOUSTICS LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510332604.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing multimodal voiceprint authentication methods are insufficient in complex environments and conditions, making it difficult to provide stable identity authentication under different environments and conditions.

Method used

By obtaining the acoustic feature information of the voice data segment to be detected, collecting multi-modal voiceprint information, performing deviation sorting and feature fusion, determining phoneme loss tags and feature interaction deviations, and cross-verification combined with audio identification attributes, improving the accuracy and reliability of identity verification.

Benefits of technology

It improves the adaptability of the authentication system under different environments and conditions, enhances the accuracy and robustness of identity authentication, reduces interference from changes in ambient noise and voice quality, and ensures stable identity authentication in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378113A_ABST
    Figure CN120378113A_ABST
Patent Text Reader

Abstract

The invention provides a security-enhanced multi-modal voiceprint identity verification method and device, and relates to the technical field of voiceprint identity verification, and the method comprises the steps: determining a fusion voiceprint sequence of voiceprint features corresponding to an individual identity after multi-modal feature fusion, comparing the fusion voiceprint sequence with a voiceprint template of the individual identity, and obtaining a verification result of the individual identity; obtaining a phoneme loss label of the voiceprint of the individual identity in the synchronous voice sample; according to the voiceprint verification identifier, determining a feature interaction deviation of a voiceprint corresponding to the individual identity during voiceprint recognition, and further determining a voiceprint intention track of the voiceprint of the individual identity in the synchronous voice sample according to the feature interaction deviation and the audio recognition attribute; and according to the phoneme loss label and the voiceprint intention track, carrying out cross verification on the multi-mode voiceprint identity in the mobile voice service. According to the method and the device, the individual identity can be accurately and reliably verified under the fusion of the multi-modal voiceprint information, so that the adaptability of a verification system in different environments and conditions is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of voiceprint identity authentication. More specifically, this application relates to a method and device for authenticating multi-modal voiceprint identities with enhanced security. Background Art

[0002] Voiceprint identity authentication is a biometric recognition technology that confirms an individual's identity by analyzing their voice characteristics (such as pitch, timbre, speech rate, formants, etc.). Different from traditional passwords, fingerprints, or facial recognition, voiceprint identity authentication confirms identity through unique biometric characteristics in the voice signal, with advantages such as non-contact, convenience, and privacy protection. In this process, first, a high-quality microphone captures the individual's voice sample, extracts the voiceprint feature information therein, and then compares these features with the stored voiceprint templates to evaluate the similarity to confirm the identity.

[0003] However, in existing methods for authenticating multi-modal voiceprint identities with enhanced security, identity authentication usually relies on a single voiceprint feature, making it vulnerable to the influence of environmental noise, emotional changes, and fluctuations in the user's physiological state. This results in insufficient verification accuracy and robustness, leading to a decline in the reliability of the verification results, and thus it is difficult to provide stable identity authentication in complex actual application scenarios. Therefore, how to accurately and reliably verify an individual's identity under the fusion of multi-modal voiceprint information to improve the adaptability of the verification system in different environments and conditions is a difficult problem faced by the industry. Summary of the Invention

[0004] This application provides a method and device for authenticating multi-modal voiceprint identities with enhanced security, which can accurately and reliably verify an individual's identity under the fusion of multi-modal voiceprint information to improve the adaptability of the verification system in different environments and conditions.

[0005] In a first aspect, this application provides a method for authenticating multi-modal voiceprint identities with enhanced security, and the authentication method includes the following steps:

[0006] Obtain the acoustic feature information of the voice data segment to be detected in the mobile voice service with enhanced security, perform attribute recognition on the acoustic feature information, and obtain the audio identification attributes corresponding to the individual identity in the acoustic feature information;

[0007] Collect the multi-modal voiceprint information of the individual identity, perform deviation sorting on the multi-modal voiceprint information, obtain the fused voiceprint sequence after the fusion of the voiceprint features corresponding to the individual identity in the multi-modal features, compare the fused voiceprint sequence with the voiceprint template of the individual identity, and obtain the phoneme loss label of the voiceprint of the individual identity in the synchronous voice sample;

[0008] Obtain the voiceprint verification identifier in the multi-modal voiceprint corresponding to the individual identity, determine the feature interaction deviation of the voiceprint corresponding to the individual identity during voiceprint recognition based on the voiceprint verification identifier, and then determine the voiceprint intention trajectory of the voiceprint corresponding to the individual identity in the synchronous voice sample from the feature interaction deviation and the audio identification attribute;

[0009] Perform cross-verification on the multi-modal voiceprint identities in the mobile voice service according to the phoneme loss label and the voiceprint intention trajectory.

[0010] In this embodiment, performing attribute recognition on the acoustic feature information to obtain the audio identification attribute corresponding to the individual identity in the acoustic feature information specifically includes:

[0011] Determine the voice unit corresponding to the individual identity in the acoustic feature information according to the acoustic feature information;

[0012] Determine the acoustic attribute corresponding to the individual identity in the acoustic feature information according to each voice unit;

[0013] Determine the audio identification attribute corresponding to the individual identity in the acoustic feature information through the acoustic attribute.

[0014] In this embodiment, performing deviation sorting on the multi-modal voiceprint information to obtain the fused voiceprint sequence after multi-modal feature fusion of the voiceprint features corresponding to the individual identity specifically includes:

[0015] Determine the vocal data segment in the multi-modal voiceprint information;

[0016] Determine the voiceprint feature corresponding to the individual identity through the vocal data segment;

[0017] Determine the voiceprint configuration table according to the voiceprint feature;

[0018] Align and fuse the voiceprint configuration table to obtain the fused voiceprint sequence after multi-modal feature fusion of the voiceprint features corresponding to the individual identity.

[0019] In this embodiment, comparing the fused voiceprint sequence with the voiceprint template of the individual identity to obtain the phoneme loss label of the voiceprint of the individual identity in the synchronous voice sample specifically includes:

[0020] Determine the voiceprint template of the individual identity;

[0021] Determine the target voice intention corresponding to the voiceprint of the individual identity according to the voiceprint template;

[0022] Use the target voice intention to screen the fused voiceprint sequence to obtain the phoneme loss label of the voiceprint of the individual identity in the synchronous voice sample.

[0023] In this embodiment, the feature interaction deviation of the voiceprint corresponding to the individual identity determined according to the voiceprint verification identifier during voiceprint recognition specifically includes:

[0024] Determine the recognition similarity of the voiceprint verification identifier during voiceprint recognition;

[0025] Extract the voiceprint interaction rule of the voiceprint corresponding to the individual identity during voiceprint recognition based on the recognition similarity;

[0026] Determine the feature interaction deviation of the voiceprint corresponding to the individual identity during voiceprint recognition through the voiceprint interaction rule.

[0027] In this embodiment, determining the voiceprint intention trajectory of the voiceprint of the individual identity in the synchronous voice sample from the feature interaction deviation and the audio identification attribute specifically includes:

[0028] Determine the context information of the voiceprint of the individual identity in the synchronous voice sample according to the feature interaction deviation;

[0029] Determine the voiceprint confidence corresponding to the voiceprint of the individual identity according to the audio identification attribute;

[0030] Determine the voiceprint intention trajectory of the voiceprint of the individual identity in the synchronous voice sample through the context information and the voiceprint confidence.

[0031] In this embodiment, the voiceprint intention trajectory refers to the dynamic path of the voice intention of an individual changing with time in the synchronous voice sample.

[0032] In this embodiment, the feature interaction deviation refers to the difference between the feature change and the voiceprint template of an individual voiceprint in different situations.

[0033] In this embodiment, the phoneme loss label represents the deviation between the phonemes pronounced by an individual and the phonemes in the target voice intention in the synchronous voice sample.

[0034] In a second aspect, the present application provides a verification device for enhancing the security of multi-modal voiceprint identity, which is used to execute a verification method for enhancing the security of multi-modal voiceprint identity. The verification device includes:

[0035] An attribute recognition module, configured to obtain the acoustic feature information of the voice data segment to be detected in the mobile voice service for enhancing security, perform attribute recognition on the acoustic feature information, and obtain the audio identification attribute corresponding to the individual identity in the acoustic feature information;

[0036] A voiceprint comparison module, configured to collect multimodal voiceprint information of an individual identity, perform deviation sorting on the multimodal voiceprint information to obtain a fused voiceprint sequence after multimodal feature fusion of the voiceprint features corresponding to the individual identity, compare the fused voiceprint sequence with the voiceprint template of the individual identity, and obtain a phoneme loss label of the voiceprint of the individual identity in the synchronous speech sample;

[0037] A feature identification module, configured to obtain a voiceprint verification identifier in the multimodal voiceprint corresponding to the individual identity, determine a feature interaction deviation of the voiceprint corresponding to the individual identity during voiceprint recognition according to the voiceprint verification identifier, and further determine a voiceprint intention trajectory of the voiceprint corresponding to the individual identity in the synchronous speech sample from the feature interaction deviation and the audio identification attribute;

[0038] A voiceprint verification module, configured to perform cross-verification on the multimodal voiceprint identities in the mobile voice service according to the phoneme loss label and the voiceprint intention trajectory.

[0039] The technical solution provided by the embodiments disclosed in this application has the following beneficial effects:

[0040] By obtaining the acoustic feature information of the voice data segment to be detected in the mobile voice service with enhanced security, performing attribute recognition on the acoustic feature information to obtain the audio identification attribute corresponding to the individual identity in the acoustic feature information; collecting the multimodal voiceprint information of the individual identity, performing deviation sorting on the multimodal voiceprint information to obtain a fused voiceprint sequence after multimodal feature fusion of the voiceprint features corresponding to the individual identity, comparing the fused voiceprint sequence with the voiceprint template of the individual identity, and obtaining a phoneme loss label of the voiceprint of the individual identity in the synchronous speech sample; obtaining the voiceprint verification identifier in the multimodal voiceprint corresponding to the individual identity, determining the feature interaction deviation of the voiceprint corresponding to the individual identity during voiceprint recognition according to the voiceprint verification identifier, and further determining the voiceprint intention trajectory of the voiceprint corresponding to the individual identity in the synchronous speech sample from the feature interaction deviation and the audio identification attribute; performing cross-verification on the multimodal voiceprint identities in the mobile voice service according to the phoneme loss label and the voiceprint intention trajectory.

[0041] It can be seen that in this application, first, by extracting the acoustic feature information of the voice data segment, the audio recognition attributes of an individual in a specific environment can be accurately captured, thereby improving the recognition accuracy of the system for the individual identity and reducing the interference of environmental noise and voice quality changes on identity verification; by collecting and sorting the deviations of multi-modal voiceprint information, the voiceprint features from different modalities (such as voice, emotion, speech rate, etc.) can be fused to enhance the system's ability to identify an individual's identity, making identity verification more accurate and reliable in various scenarios; by comparing the fused voiceprint sequence with the voiceprint template, the phoneme loss of an individual in the synchronous voice sample can be accurately identified, thereby improving the system's ability to recognize subtle voice differences and enhancing the accuracy of identity verification; by obtaining the voiceprint verification identifier and analyzing the feature interaction deviation, the changes in an individual's identity in different contexts can be more accurately recognized, thereby enhancing the robustness of the voiceprint recognition system and ensuring stable identity verification even in complex environments or under emotional changes; by cross-verifying the phoneme loss label with the voiceprint intention trajectory, different feature information can be effectively integrated to further improve the accuracy and reliability of voiceprint identity verification and ensure accurate identification of identity under multi-modal conditions.

[0042] In summary, the technical solution adopted in this application can accurately and reliably verify an individual's identity under the fusion of multi-modal voiceprint information, so as to improve the adaptability of the verification system in different environments and conditions. Brief Description of the Drawings

[0043] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, form a part of this application, and do not constitute a limitation on the embodiments of the present invention. In the drawings:

[0044] Figure 1 is a flowchart of a method for verifying an enhanced security multi-modal voiceprint identity provided by this application;

[0045] Figure 2 is an exemplary flowchart of a fused voiceprint sequence after multi-modal feature fusion of the voiceprint features corresponding to an individual identity provided by this application;

[0046] Figure 3 is an exemplary flowchart of determining the feature interaction deviation during voiceprint recognition of the voiceprint corresponding to an individual identity provided by this application;

[0047] Figure 4 is a module structure diagram of a device for verifying an enhanced security multi-modal voiceprint identity provided by this application. Detailed Description of the Embodiments

[0048] To make the objectives, technical solutions, and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings. The illustrative embodiments of the present invention and their descriptions are only used to explain the present invention and do not limit the present invention. It should be noted that the present invention has been in the actual R & D and use stage.

[0049] Embodiment 1. To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners. Refer to Figure 1 As shown, this figure is an exemplary flowchart of a method for verifying a multi-modal voiceprint identity with enhanced security according to the present embodiment of the present application. The verification method includes the following steps:

[0050] In step S1, in a mobile voice service with enhanced security, the acoustic feature information of the voice data segment to be detected is obtained, and the attribute recognition is performed on the acoustic feature information to obtain the audio recognition attribute corresponding to the individual identity in the acoustic feature information.

[0051] Specifically, the acquisition of the acoustic feature information of the voice data segment to be detected in a mobile voice service with enhanced security can be implemented in the following manner: First, the collected voice signal is preprocessed, including: using Wiener filter for noise reduction and normalization processing. Then, the voice signal is divided into several short-time frames, that is: each frame is 20 - 30 milliseconds, and the frame shift is 10 milliseconds to capture the stationary characteristics of the voice in a short time. For each frame, the time-domain signal is converted into a frequency-domain signal through short-time Fourier transform (STFT), and the spectrum information is extracted. Next, the Mel filter bank is used to map the spectrum to the Mel scale to simulate the perception sensitivity of the human ear to different frequencies. Subsequently, the logarithmic transformation is performed on the energy spectrum of the Mel frequency and the discrete cosine transform (DCT) is calculated. Finally, the Mel frequency cepstral coefficients (MFCC) are obtained, and the Mel frequency cepstral coefficients can be used as the acoustic feature information of the voice data segment to be detected, which will not be elaborated here.

[0052] It should be noted that the acoustic feature information of the voice data segment to be detected in this application refers to the parametric description of the physical and perceptual characteristics of the voice in the frequency domain, time domain, and perceptual domain.

[0053] In this embodiment, the attribute recognition of the acoustic feature information to obtain the audio recognition attribute corresponding to the individual identity in the acoustic feature information can be implemented by the following steps:

[0054] Determine the voice unit corresponding to the individual identity in the acoustic feature information according to the acoustic feature information;

[0055] Determine the acoustic attribute corresponding to the individual identity in the acoustic feature information according to each voice unit;

[0056] Determine the audio identification attribute corresponding to the individual identity in the acoustic feature information through the acoustic attribute.

[0057] In specific implementation, first, process the collected speech signal. First, use noise reduction technology (such as frequency domain spectral subtraction) to remove background noise, and then perform normalization processing on the signal amplitude to ensure the consistency of the signal during analysis. Then, divide the speech signal into short-time frames (for example, each frame is 20 milliseconds, and the frame shift is 10 milliseconds). Apply windowing processing (such as Hamming window) to each frame to reduce spectral leakage. Use a trained speech model (such as an acoustic model constructed based on a speech dataset) to decode the acoustic features, and match the features within the short-time frame with known speech units (such as phonemes or morphemes). Through frame-by-frame analysis, combined with the speech structure rules in the model, predict the speech units. Next, analyze the periodic amplitude change of the speech frame to identify the fundamental period. By observing the repeated pattern of the signal, find the most significant frequency component in each frame, which represents the pitch of the speaker. Analyze the time distribution of the speech units, count the number of phonemes or morphemes completed within a certain time, and calculate the speaking speed. Through spectral analysis, find the frequency bands where the energy of the speech signal is most concentrated. These frequency bands are called formants, which reflect the modulation characteristics of the vocal organs (such as the oral cavity and nasal cavity) on the sound. The distribution and position of the formants can reflect the vocal tract characteristics of an individual. Finally, combine multiple acoustic attributes such as pitch, speaking speed, and formant position to form a complete acoustic feature description, such as pitch range, speaking speed index, and formant frequency band distribution. Use a machine learning classifier (such as a model based on a deep neural network) to analyze the acoustic attribute data. This model compares the acoustic attributes with the pre-labeled identity tags to identify the identity attributes reflected in the audio, that is, obtain the audio identification attribute corresponding to the individual identity.

[0058] It should be noted that in this application, a speech unit represents the smallest constituent unit of a speech signal, an acoustic attribute represents a key parameter describing speech pronunciation characteristics, and an audio identification attribute represents identity-related information derived from the acoustic attribute.

[0059] In step S2, collect the multi-modal voiceprint information of the individual identity, perform deviation ranking on the multi-modal voiceprint information, obtain the fused voiceprint sequence after multi-modal feature fusion of the voiceprint feature corresponding to the individual identity, compare the fused voiceprint sequence with the voiceprint template of the individual identity, and obtain the phoneme loss label of the voiceprint of the individual identity in the synchronous speech sample.

[0060] In specific implementation, the acquisition of multi-modal voiceprint information of an individual identity can be achieved in the following manner, that is: First, record the target voice sample, collect the audio through a high-quality microphone device to ensure signal clarity, and at the same time combine noise reduction technologies for specific scenarios (such as beamforming) to reduce environmental interference. Second, extract multi-modal voiceprint features, including traditional Mel Frequency Cepstral Coefficients (MFCCs) to represent pronunciation characteristics, pitch and formant frequencies to reflect vocal tract structures, as well as dynamic features such as speech rate and energy distribution. Further, obtain emotional features (such as tone and intonation) and pronunciation habits through an acoustic perception model. Finally, integrate multi-source data, construct a multi-modal feature vector of the individual identity, and store it in a database. Reading the multi-modal voiceprint information of the individual identity from this database will not be elaborated here.

[0061] It should be noted that in this application, multi-modal voiceprint information refers to the voice data that fuses sound data from multiple audio sources or feature dimensions, comprehensively representing the characteristic information of an individual identity.

[0062] Preferably, in this embodiment, referring to Figure 2 As shown, this figure is an exemplary flowchart of the fused voiceprint sequence after multi-modal feature fusion of the voiceprint features corresponding to an individual identity in the embodiment of this application. In this embodiment, the deviation sorting of the multi-modal voiceprint information is performed to obtain the fused voiceprint sequence after multi-modal feature fusion of the voiceprint features corresponding to an individual identity, which can be specifically achieved by the following steps:

[0063] First, in step S21, determine the voice data segment in the multi-modal voiceprint information;

[0064] Next, in step S22, determine the voiceprint features corresponding to the individual identity through the voice data segment;

[0065] Then, in step S23, determine the voiceprint configuration table according to the voiceprint features;

[0066] Finally, in step S24, align and fuse the voiceprint configuration tables to obtain the fused voiceprint sequence after multi-modal feature fusion of the voiceprint features corresponding to the individual identity.

[0067] In specific implementation, first, by using a voice activity detection algorithm (such as an energy threshold-based method), it is determined which parts of the audio segment contain valid vocalization data. The energy spectrum of the audio signal is calculated to distinguish between silent segments and speech segments, ensuring that only the parts containing valid vocalization are processed. Through short-time analysis of the audio signal, the valid speech segments are located, and the irrelevant noise parts are removed to obtain the vocalization data segment. Then, audio features such as Mel Frequency Cepstral Coefficients (MFCC), Pitch, etc. are extracted from each vocalization data segment. By applying the Short-Time Fourier Transform (STFT) or Mel filters, spectral features are extracted, and then logarithmic compression and Discrete Cosine Transform (DCT) are applied to generate MFCC. Combining with dynamic features such as emotion and speech rate, unique voiceprint features of an individual are formed. For example, based on different vocal tract features and pronunciation methods, the audio patterns related to the individual's identity, that is, the voiceprint features, are recognized. Then, the extracted voiceprint features are matched with the pre-stored voiceprint templates, and the unique feature patterns of the individual are recognized through classification or clustering algorithms (such as Support Vector Machine SVM or K-means clustering). Based on the voiceprint features, a voiceprint configuration table is created, which records the specific values of various acoustic attributes (such as pitch, speech rate, etc.) of the individual and their changes in different situations, that is, the voiceprint configuration table is obtained. Finally, the voiceprint features from different sources (such as speech samples collected in different environments and different situations) are time-aligned. For example, through Dynamic Time Warping (DTW) or other alignment algorithms, the feature sequences from different modalities are aligned to ensure that they can accurately correspond in the same time dimension. The aligned voiceprint features are weighted and fused. The fusion method can be simple weighted averaging or feature-level fusion through a trained neural network to retain the useful information of each modality and suppress the influence of noise. The result after fusion is the final voiceprint sequence corresponding to the individual identity, that is, the fused voiceprint sequence after multi-modal feature fusion of the voiceprint features.

[0068] It should be noted that in this application, the vocalization data segment refers to the audio segment containing valid pronunciation extracted from the speech signal; the voiceprint feature refers to the audio feature extracted by analyzing the physiological and behavioral characteristics in the individual's speech signal; the voiceprint configuration table refers to the table recording the voiceprint feature information of the individual in different vocalization situations; the fused voiceprint sequence refers to integrating the information from multiple modalities (such as voiceprint features collected under different devices or different environments) into a unified feature sequence.

[0069] In this embodiment, comparing the fused voiceprint sequence with the voiceprint template of the individual identity to obtain the phoneme loss label of the voiceprint of the individual identity in the synchronous speech sample can be specifically implemented in the following way, that is:

[0070] Determine the voiceprint template of the individual identity;

[0071] Determine the target voice intention corresponding to the voiceprint of the individual identity according to the voiceprint template;

[0072] Use the target voice intention to screen the fused voiceprint sequence to obtain the phoneme loss label of the voiceprint of the individual identity in the synchronous voice sample.

[0073] In specific implementation, first, based on multiple collected individual voice samples, extract the typical voiceprint features of the individual (such as MFCC, pitch, formant, etc.) to construct the voiceprint template of the individual; then, by analyzing the voiceprint template and voice samples of the individual, determine the target voice intention. This process compares the voice features in the voiceprint template with the features in the voice to be recognized to evaluate the matching degree between the voice content and the target voice intention. Common methods include using a deep learning-based speech recognition model or a voice intention recognition model for intention prediction. Finally, through the matching degree with the target voice intention, the part related to the target intention in the fused voiceprint sequence is screened out. After screening out the relevant voiceprint features, further analyze the difference between the phonemes (i.e., basic speech units) in the fused voiceprint sequence and the phonemes included in the target voice intention. The phoneme loss label identifies the deviation between the actual voiceprint and the expected phonemes in the synchronous voice sample, reflecting the gap between the individual's pronunciation and the template. Finally, calculate these deviations to generate the phoneme loss label.

[0074] It should be noted that in this application, the voiceprint template is a standardized expression of the voiceprint features of the individual identity, usually composed of the features in multiple voice samples, representing the unique voice features of the individual; the target voice intention represents the voice content or purpose expressed by the individual in a specific voice interaction, usually used to describe the intention of their speech (for example, "verify identity" or "greet"); the synchronous voice sample represents; the phoneme loss label represents the deviation between the phonemes of the individual's pronunciation and the phonemes in the target voice intention in the synchronous voice sample, used to measure the accuracy of the voice.

[0075] In step S3, obtain the voiceprint verification identifier in the multi-modal voiceprint corresponding to the individual identity, determine the feature interaction deviation of the voiceprint corresponding to the individual identity during voiceprint recognition according to the voiceprint verification identifier, and then determine the voiceprint intention trajectory of the voiceprint of the individual identity in the synchronous voice sample from the feature interaction deviation and the audio identification attribute.

[0076] In specific implementation, obtaining the voiceprint verification identifier in the multi-modal voiceprint corresponding to the individual identity can be achieved by the following method, that is: extracting voiceprint features of multiple modalities from the collected voice data, such as MFCC, pitch, speech rate, etc., fusing the features of different modalities to form a multi-modal voiceprint representation of the individual. Using a voiceprint matching algorithm (such as support vector machine, neural network) to compare the fused features with the known individual template to determine the verification identifier. This identifier represents the verification result of the individual identity, usually a unique mark or label, to confirm whether the feature matches the pre-stored template, thereby confirming the identity, that is, obtaining the voiceprint verification identifier in the multi-modal voiceprint.

[0077] It should be noted that in this application, the voiceprint verification identifier represents a unique identity identifier extracted from multiple audio sources or modalities by analyzing the voiceprint features of an individual.

[0078] Preferably, in this embodiment, referring to Figure 3 As shown, this figure is an exemplary flowchart of the feature interaction deviation when identifying the voiceprint corresponding to the individual identity in the embodiment of this application. In this embodiment, determining the feature interaction deviation when identifying the voiceprint corresponding to the individual identity according to the voiceprint verification identifier can be specifically implemented by the following steps:

[0079] First, in step S31, determine the recognition similarity when the voiceprint verification identifier is used for voiceprint recognition;

[0080] Then, in step S32, extract the voiceprint interaction rule when the voiceprint corresponding to the individual identity is used for voiceprint recognition based on the recognition similarity;

[0081] Finally, in step S33, determine the feature interaction deviation when the voiceprint corresponding to the individual identity is used for voiceprint recognition through the voiceprint interaction rule.

[0082] In specific implementation, first, the extracted multi-modal voiceprint features are compared with the voiceprint verification identifier. Common similarity calculation methods include Euclidean distance, cosine similarity, or similarity evaluation based on deep learning models. For example, the similarity score between the voiceprint features and the verification identifier is calculated through a neural network to quantify their proximity in the feature space, that is, the recognition similarity is obtained; then, after obtaining the recognition similarity, by analyzing the relationship between different voiceprint verification identifiers and similarity scores, voiceprint interaction rules are constructed. These rules are usually based on the behavior of an individual's voiceprint under different conditions (such as pitch change, speech rate change, etc.). For example, an individual may have a large pitch fluctuation under a tense state, while the pitch change is small under a calm state, and then machine learning or pattern recognition algorithms (such as decision trees, SVMs, or neural networks) are used to extract voiceprint interaction rules from historical data to indicate how voiceprint features interact and change in different situations. Finally, the extracted voiceprint interaction rules are used to analyze the feature changes of an individual's voiceprint in different situations. By calculating the deviation between the target voice sample and the individual voiceprint template under specific rules (for example, the deviation of pitch or speech rate), the "feature interaction deviation" is obtained.

[0083] It should be noted that in this application, the recognition similarity refers to the matching degree between the input voiceprint features and the voiceprint verification identifier during the voiceprint verification process, which is used to measure the accuracy of recognition; the voiceprint interaction rule refers to the law of how an individual's voiceprint features (such as pitch, speech rate, etc.) interact and change in different pronunciation situations, which is used to describe the voiceprint change pattern of an individual in different situations; the feature interaction deviation refers to the difference between the feature change of an individual's voiceprint in different situations and the voiceprint template.

[0084] In this embodiment, determining the voiceprint intention trajectory of the voiceprint of an individual identity in the synchronous voice sample from the feature interaction deviation and the audio identification attribute can be specifically implemented by the following steps:

[0085] Determine the context information of the voiceprint of an individual identity in the synchronous voice sample according to the feature interaction deviation;

[0086] Determine the voiceprint confidence corresponding to the voiceprint of an individual identity according to the audio identification attribute;

[0087] Determine the voiceprint intention trajectory of the voiceprint of an individual identity in the synchronous voice sample through the context information and the voiceprint confidence.

[0088] In specific implementation, first, according to the feature interaction deviation, the system analyzes the changes in an individual's voiceprint in different situations. For example, when the individual's mood fluctuates, the environmental noise changes, or the speech rate increases, the features of the voiceprint will change accordingly. The feature interaction deviation reflects these changes and provides clues about the environment where the speech is located. Through these changes, the context or situation of the speech can be inferred to form context information. Through pattern recognition algorithms (such as hidden Markov models or deep learning models), context information about the speech can be extracted from the feature interaction deviation, such as whether the speech is in a noisy environment or whether the individual is speaking in an emotionally excited state. Then, the audio recognition attribute provides the confidence level of the voiceprint recognition result, which is often quantified by calculating the matching degree between the voiceprint and the individual template. For example, the confidence level of the recognition result can be evaluated by the cosine similarity with the template or a distance metric (such as Euclidean distance). Finally, during the conversation, the speech intention of the individual may change from a simple greeting to a more complex request. By combining the context information (such as the speech situation) with the voiceprint confidence level (the accuracy of recognition), the system can track the evolution of the individual's speech intention. Then, algorithms based on sequence data analysis (such as long short-term memory network LSTM) are used to comprehensively process the context and confidence level data to generate an accurate voiceprint intention trajectory.

[0089] It should be noted that in this application, the context information refers to the environmental or situational information where an individual is located in a specific speech sample; the voiceprint confidence level represents a measure of the matching degree between the voiceprint of an individual's identity and the voiceprint template, reflecting the confidence of the recognition system in confirming this identity; the voiceprint intention trajectory refers to the dynamic path of the voice intention of an individual in a synchronous speech sample changing over time.

[0090] In step S4, cross-validation is performed on the multi-modal voiceprint identity in the mobile voice service according to the phoneme loss label and the voiceprint intention trajectory.

[0091] In specific implementation, multiple voice sample data are collected, and a phoneme loss label and a voiceprint intention trajectory are extracted for each sample. At this time, the system constructs an individual multi-modal voiceprint identity model based on multi-modal voiceprint information (such as audio, voice emotion, speech rate, etc.). First, the phoneme loss label is matched with the voiceprint intention trajectory. The phoneme loss label provides the accuracy of the speech content, while the voiceprint intention trajectory provides the context information of the pronunciation. By comparing the phoneme loss label with the corresponding intention trajectory, the system can check the matching degree of the individual identity between multiple modalities. Then, using a cross-validation algorithm (such as k-fold cross-validation), the data set is divided into multiple subsets. Validation is performed on each subset to ensure the stability and accuracy of the model. Cross-validation can help the system identify whether it can stably identify an individual's identity under different conditions and find the biases that may affect identity verification. By calculating the matching degree of each validation, the system can obtain an overall identity verification score. If the phoneme loss is small and the intention trajectory is consistent with the template, it is considered that the identity matching degree is high, that is, the cross-validation of the multi-modal voiceprint identity in the mobile voice service is completed.

[0092] Thus, it can be seen that in this application, first, by extracting the acoustic feature information of the voice data segment, the audio identification attributes of an individual in a specific environment can be accurately captured, thereby improving the system's recognition accuracy of the individual identity and reducing the interference of environmental noise and speech quality changes on identity verification; through the collection and deviation ranking of multi-modal voiceprint information, the voiceprint features from different modalities (such as voice, emotion, speech rate, etc.) can be fused, enhancing the system's ability to identify an individual's identity and making identity verification more accurate and reliable in various scenarios; by comparing the fused voiceprint sequence with the voiceprint template, the phoneme loss of an individual in the synchronous voice sample can be accurately identified, further improving the system's ability to recognize subtle voice differences and enhancing the accuracy of identity verification; by obtaining the voiceprint verification identifier and analyzing the feature interaction deviation, the changes of an individual's identity in different contexts can be more accurately identified, thereby enhancing the robustness of the voiceprint recognition system and ensuring stable identity verification even in complex environments or emotional changes; by cross-verifying the phoneme loss label with the voiceprint intention trajectory, different feature information can be effectively integrated, further improving the accuracy and reliability of voiceprint identity verification and ensuring accurate identity recognition under multi-modal conditions.

[0093] In summary, the technical solution adopted in this application can accurately and reliably verify an individual's identity under the fusion of multi-modal voiceprint information, so as to improve the adaptability of the verification system in different environments and conditions.

[0094] Embodiment 2, this application provides a verification device for multi-modal voiceprint identity with enhanced security. Refer to Figure 4 As shown in the figure, which is a schematic diagram of the verification device according to this embodiment of this application, the verification device includes:

[0095] An attribute recognition module 100 is configured to obtain acoustic feature information of a voice data segment to be detected in a mobile voice service with enhanced security, perform attribute recognition on the acoustic feature information, and obtain an audio identification attribute corresponding to an individual identity in the acoustic feature information;

[0096] A voiceprint comparison module 200 is configured to collect multimodal voiceprint information of an individual identity, perform deviation sorting on the multimodal voiceprint information, obtain a fused voiceprint sequence after multimodal feature fusion of the voiceprint features corresponding to the individual identity, compare the fused voiceprint sequence with a voiceprint template of the individual identity, and obtain a phoneme loss label of the voiceprint of the individual identity in a synchronous voice sample;

[0097] A feature recognition module 300 is configured to obtain a voiceprint verification identifier in the multimodal voiceprint corresponding to an individual identity, determine a feature interaction deviation of the voiceprint corresponding to the individual identity during voiceprint recognition according to the voiceprint verification identifier, and further determine a voiceprint intention trajectory of the voiceprint of the individual identity in a synchronous voice sample from the feature interaction deviation and the audio identification attribute;

[0098] A voiceprint verification module 400 is configured to perform cross-verification on the multimodal voiceprint identities in the mobile voice service according to the phoneme loss label and the voiceprint intention trajectory.

[0099] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for verifying a multi-modal voiceprint identity with enhanced security, characterized in that, The verification method includes the following steps: Obtain the acoustic feature information of the voice data segment to be detected in the mobile voice service with enhanced security, perform attribute recognition on the acoustic feature information, and obtain the audio identification attributes corresponding to the individual identity in the acoustic feature information; Collect the multi-modal voiceprint information of the individual identity, perform deviation ranking on the multi-modal voiceprint information, obtain the fused voiceprint sequence after multi-modal feature fusion of the voiceprint features corresponding to the individual identity, compare the fused voiceprint sequence with the voiceprint template of the individual identity, and obtain the phoneme loss label of the voiceprint of the individual identity in the synchronous voice sample; Obtain the voiceprint verification identifier in the multi-modal voiceprint corresponding to the individual identity, determine the feature interaction deviation of the voiceprint corresponding to the individual identity during voiceprint recognition according to the voiceprint verification identifier, and further determine the voiceprint intention trajectory of the voiceprint of the individual identity in the synchronous voice sample from the feature interaction deviation and the audio identification attributes; Cross-verify the multi-modal voiceprint identities in the mobile voice service according to the phoneme loss label and the voiceprint intention trajectory.

2. The verification method for a multi-modal voiceprint identity with enhanced security as claimed in claim 1, wherein, Performing attribute recognition on the acoustic feature information to obtain the audio identification attributes corresponding to the individual identity in the acoustic feature information specifically includes: Determine the voice units corresponding to the individual identity in the acoustic feature information according to the acoustic feature information; Determine the acoustic attributes corresponding to the individual identity in the acoustic feature information according to each voice unit; Determine the audio identification attributes corresponding to the individual identity in the acoustic feature information through the acoustic attributes.

3. The verification method for a multi-modal voiceprint identity with enhanced security as claimed in claim 1, wherein, Performing deviation ranking on the multi-modal voiceprint information to obtain the fused voiceprint sequence after multi-modal feature fusion of the voiceprint features corresponding to the individual identity specifically includes: Determine the vocal data segment in the multi-modal voiceprint information; Determine the voiceprint features corresponding to the individual identity through the vocal data segment; Determine the voiceprint configuration table according to the voiceprint features; Align and fuse the voiceprint configuration tables to obtain the fused voiceprint sequence after multi-modal feature fusion of the voiceprint features corresponding to the individual identity.

4. The verification method for a multi-modal voiceprint identity with enhanced security according to claim 1, wherein Comparing the fused voiceprint sequence with the voiceprint template of the individual identity to obtain the phoneme loss label of the voiceprint of the individual identity in the synchronous voice sample specifically includes: Determine the voiceprint template of the individual identity; Determine the target voice intention corresponding to the voiceprint of the individual identity according to the voiceprint template; Use the target voice intention to screen the fused voiceprint sequence to obtain the phoneme loss label of the voiceprint of the individual identity in the synchronous voice sample.

5. The verification method of a multi-modal voiceprint identity for enhancing security according to claim 1, wherein, Determining the feature interaction deviation of the voiceprint corresponding to the individual identity during voiceprint recognition according to the voiceprint verification identifier specifically includes: Determine the recognition similarity of the voiceprint verification identifier during voiceprint recognition; Extract the voiceprint interaction rule of the voiceprint corresponding to the individual identity during voiceprint recognition based on the recognition similarity; Determine the feature interaction deviation of the voiceprint corresponding to the individual identity during voiceprint recognition through the voiceprint interaction rule.

6. The verification method for a multi-modal voiceprint identity with enhanced security as described in claim 1, characterized in that, Determining the voiceprint intention trajectory of the voiceprint of the individual identity in the synchronous voice sample from the feature interaction deviation and the audio identification attributes specifically includes: Determine the context information of the voiceprint of the individual identity in the synchronous voice sample according to the feature interaction deviation; The voiceprint confidence level corresponding to the voiceprint for determining the individual identity based on the audio identification attribute; Determine the voiceprint intention trajectory of the voiceprint for the individual identity in the synchronous voice sample based on the context information and the voiceprint confidence level.

7. The verification method for a multi-modal voiceprint identity with enhanced security according to claim 1, characterized in that, The voiceprint intention trajectory refers to the dynamic path of the voice intention of an individual changing over time in the synchronous voice sample.

8. The verification method of a multi-modal voiceprint identity with enhanced security according to claim 1, characterized in that, The feature interaction deviation refers to the difference between the feature change and the voiceprint template of an individual's voiceprint in different situations.

9. The verification method for a multi-modal voiceprint identity with enhanced security as claimed in claim 1, wherein, The phoneme loss label represents the deviation between the phonemes pronounced by an individual and the phonemes in the target voice intention in the synchronous voice sample.

10. A verification device for multi-modal voiceprint identity with enhanced security, which is used to execute a verification method for multi-modal voiceprint identity with enhanced security according to any one of claims 1 to 9, characterized in that, The verification device includes: An attribute recognition module, configured to obtain the acoustic feature information of the voice data segment to be detected in the mobile voice service with enhanced security, perform attribute recognition on the acoustic feature information, and obtain the audio identification attribute corresponding to the individual identity in the acoustic feature information; A voiceprint comparison module, configured to collect the multimodal voiceprint information of the individual identity, perform deviation ranking on the multimodal voiceprint information, obtain the fused voiceprint sequence after the multimodal feature fusion of the voiceprint features corresponding to the individual identity, compare the fused voiceprint sequence with the voiceprint template of the individual identity, and obtain the phoneme loss label of the voiceprint of the individual identity in the synchronous voice sample; A feature identification module, configured to obtain the voiceprint verification identifier in the multimodal voiceprint corresponding to the individual identity, determine the feature interaction deviation of the voiceprint corresponding to the individual identity during voiceprint recognition according to the voiceprint verification identifier, and further determine the voiceprint intention trajectory of the voiceprint corresponding to the individual identity in the synchronous voice sample from the feature interaction deviation and the audio identification attribute; A voiceprint verification module, configured to perform cross-verification on the multimodal voiceprint identities in the mobile voice service according to the phoneme loss label and the voiceprint intention trajectory.

Citation Information

Cited By

  • Government affair form submission system and method based on multi-modal verification

    CN121811888A

  • Government form submission system and method based on multi-modal verification

    CN121811888B