Speech recognition authentication method and system based on multi-modal features and dynamic evaluation

Through a speech recognition authentication method with multimodal features and dynamic evaluation, integrating voiceprint, semantic and behavioral features, and combining with a risk assessment model, it solves the shortcomings of the existing technology such as the decreased accuracy in multi-noise scenarios and the static authentication strategy, and achieves improved security and reliability in different risk scenarios.

CN120748413AActive Publication Date: 2025-10-03JIANGSU VARIABLE SUPERCOMP TECH

Patent Information

Application Number
CN202511202874.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-10-03
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Existing voice recognition authentication technology has reduced accuracy in high-noise scenarios, weak anti-attack capabilities, and static authentication strategies cannot adapt to dynamic risk scenarios, resulting in high false rejection rates and security vulnerabilities.

Method used

It adopts multimodal features and dynamic evaluation methods, integrates voiceprint, semantic, behavioral characteristics and environmental adaptation mechanisms, combines with quantitative risk assessment models, adjusts authentication thresholds in real time, and enforces two-factor authentication in high-risk scenarios.

Benefits of technology

It significantly enhances the security and reliability of authentication, can dynamically adjust the authentication method under different risk scenarios, and improves the recognition accuracy and anti-attack capabilities in multi-noise environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748413A_ABST
    Figure CN120748413A_ABST
Patent Text Reader

Abstract

The invention discloses a speech recognition and authentication method and system based on multi-modal features and dynamic evaluation in the technical field of speech recognition and authentication, and the method comprises the steps: collecting an original speech signal of a user through a microphone, and carrying out the preprocessing of the original speech signal, and obtaining the preprocessing speech data; and extracting feature data of the preprocessed voice data by adopting a multi-dimensional feature hierarchical extraction technology, and injecting a multi-source noise sample into an acoustic feature space of a voiceprint feature model based on an initial training stage established by the voiceprint feature model to construct an anti-noise mixed voiceprint map. Through integrating voiceprint, semantics, behavior characteristics and an environment adaptation mechanism, an authentication threshold is adjusted in real time according to dynamic risk assessment, meanwhile, a risk scoring model is utilized to calculate a comprehensive risk value, authentication modes of different levels are started according to risk scenes of different degrees, and two-factor authentication is forcibly implemented for high-risk scenes. And the authentication security and reliability can be obviously enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition and authentication, and in particular to a speech recognition and authentication method and system based on multimodal features and dynamic evaluation. Background Art

[0002] Voice recognition authentication is a technology that verifies identity by analyzing user voice features. Prior art includes a voiceprint recognition method disclosed in Chinese Patent Application No. CN201910281641.1. This method receives a voice signal to be recognized from an unknown user, extracts the frame voiceprint features of each frame, calculates its posterior probability, classifies the frame voiceprint features based on the posterior probability, generates a model to be recognized and a voiceprint recognition model, and determines the user's identity based on model similarity. Although this method can improve the accuracy of text-independent voice signal recognition, especially in terms of recognition efficiency of short text-independent voice signals, it mainly focuses on single acoustic feature processing and model similarity comparison, and fails to fully consider comprehensive factors such as multimodal feature fusion, dynamic risk assessment, and anti-attack capabilities.

[0003] Existing voiceprint authentication is vulnerable and has the following shortcomings: (1) Existing voiceprint authentication is fragile. In a pure laboratory environment, the equal error rate (EER) based on MFCC features is about 8.2%. However, in actual noisy scenarios (SNR ≤ 20dB), the EER deteriorates to 15.7%, and the false acceptance rate of synthetic speech attacks (such as WaveGAN generation) is as high as 23.5%. (2) The fixed threshold of the existing static voice authentication strategy cannot adapt to dynamic risk scenarios such as remote login, resulting in an increased false rejection rate or security vulnerabilities.

[0004] (3) The existing voice single-factor authentication method has high security risks in scenarios where it is low in security, easy to crack, has a high risk of privacy leakage, and is difficult to defend against attacks. Summary of the Invention

[0005] The purpose of the present invention is to provide a voice recognition authentication method and system based on multimodal features and dynamic evaluation. By integrating voiceprints, semantics, behavioral characteristics and environmental adaptation mechanisms, combined with quantitative risk assessment models, dynamic risk assessment is performed, and the authentication threshold is adjusted in real time. Different levels of authentication methods can be enabled according to different risk scenarios. At the same time, for high-risk scenarios, two-factor authentication is enforced, significantly enhancing the security and reliability of authentication.

[0006] To achieve the above object, the present invention provides the following technical solutions: In a first aspect, a speech recognition authentication method based on multimodal features and dynamic evaluation is provided, comprising: Using a microphone to collect the user's original voice signal, preprocessing the original voice signal to obtain preprocessed voice data; Extracting feature data of the pre-processed speech data using a multi-dimensional feature hierarchical extraction technique, wherein the feature data includes feature alignment benchmarks, prosodic behavior, and semantics of the pre-processed speech data; During the initial training phase based on the voiceprint feature model, multi-source noise samples are injected into the acoustic feature space of the voiceprint feature model to construct a noise-resistant hybrid voiceprint map. An Adv-GAN (Adversarial Generative Adversarial Network) discriminator is deployed to perform feature deconstruction on deepfake speech generated by the WaveGAN (Waveform Generative Adversarial Network) in the discriminator based on a joint spectrum-time domain identification mechanism. A dynamic adaptation mechanism for the acoustic environment is also established simultaneously. Based on the user's mobile terminal device and preset scoring rules, the user's fingerprint information, geographic location and historical behavior logs are collected, and a risk scoring model is used to identify the user's voice signal recognition risk to obtain a voice recognition risk score; Multimodal feature fusion and model training based on noise-resistant MFCC acoustic features, prosodic behavior and semantics; When the authentication scenario is judged to be a low-risk scenario, only voice authentication is required; when the authentication scenario is judged to be a medium-risk scenario, two-factor authentication is enabled for enhanced verification; when the authentication scenario is judged to be a high-risk scenario, the second factor authentication is forcibly enabled, where the second factor authentication is fingerprint or dynamic password authentication; In high-risk scenarios, liveness detection, feature matching, and second-factor authentication tasks are assigned to multi-core processors for parallel execution. When authentication is successful, access is authorized and an encrypted token is generated, and the login time, device information, risk score, and authentication results are stored in the security log.

[0007] As a further solution of the present invention, a microphone is used to collect the original voice signal of the user, and the original voice signal is preprocessed to obtain preprocessed voice data, including: Read the original speech signal frame by frame and remove repeated or invalid segments to obtain valid speech data, store the valid speech data as a standardized data set, convert the standardized data set into a WAV format file and output it; Using the RNNoise deep learning model to eliminate ambient noise in the WAV format file to obtain a de-noised audio signal, filtering the non-speech audio segment of the de-noised audio signal to obtain a frequency domain speech signal, wherein the non-speech audio segment includes mechanical noise and / or sudden noise; The frequency domain speech signal is divided into short-time frame speech signals by using a speech segmentation method, and speech start and end points of the short-time frame speech signals are detected based on short-time energy and zero-crossing rate to obtain framed speech data, wherein the short-time frame speech signals are speech signals with a frame length not exceeding 20 ms. Speech segments with insufficient speech length or too low signal-to-noise ratio in the framed speech data are eliminated to obtain preprocessed speech data, wherein the speech segment with insufficient speech length is a speech segment with a time length of less than 0.5 seconds, and the speech segment with too low signal-to-noise ratio is a speech segment with a signal-to-noise ratio of less than 15 dB.

[0008] As a further solution of the present invention: the method of extracting feature data of the pre-processed speech data using a multi-dimensional feature hierarchical extraction technology includes: Obtaining a feature alignment benchmark for the preprocessed speech data; analyzing prosodic behavior of the preprocessed speech data; The semantics of the pre-processed speech data is understood.

[0009] As a further solution of the present invention: the obtaining of feature alignment reference includes: Extracting short-term energy envelope prosodic behavior markers of the preprocessed speech data using a noise-resistant MFCC and fundamental frequency trajectory fluctuation analysis method including first-order and / or second-order dynamic difference; generating a behavior-enhanced voiceprint of the prosodic behavior marker based on time-frequency domain joint coding to obtain a feature alignment benchmark for the preprocessed speech data, wherein the prosodic behavior marker includes a syllable starting point and / or an energy transition gradient; The analyzing the prosodic behavior of the pre-processed speech data includes: Segmenting the speech stream of the pre-processed speech data into independent syllables based on an energy envelope mutation detection algorithm; Calculating a dynamic coefficient of variation of the speaking rate of the independent syllable, and quantifying pronunciation stability based on the dynamic coefficient of variation of the speaking rate, wherein the dynamic coefficient of variation of the speaking rate includes a standard deviation and / or a mean; An LSTM classifier is used to analyze the distribution of abnormal pauses in the independent syllables, and the energy spectrum of the second-order derivative fluctuation of the fundamental frequency of the independent syllables (0-20 Hz) is extracted to characterize the naturalness of the fundamental frequency change, thereby obtaining fundamental frequency change data. An abnormal pause is a speech segment with a pause duration greater than 300 milliseconds, and the energy spectrum of the second-order derivative fluctuation of the fundamental frequency is an energy spectrum with a frequency of 0-20 Hz. Based on the self-attention mechanism, the features of the ±20% area of ​​the speech rate mutation in the fundamental frequency change data are weightedly fused, and a 128-dimensional prosodic feature vector is generated to detect abnormal behaviors of unnatural speech rate jumps, where abnormal behaviors include robotic voices.

[0010] The understanding of the semantics of the pre-processed speech data includes: Based on the BERT-wwm (BERT-based natural language processing) model, the speech text of the preprocessed speech data is input into the BERT-wwm model to extract the 768-dimensional [CLS] semantic vector and construct a dialogue logic graph; A dual-channel strategy is used to detect sensitive words.

[0011] As a further solution of the present invention: the construction of the dialogue logic graph includes: The nodes in the 768-dimensional [CLS] semantic vector are defined as named entities, and the BiLSTM-CRF (bidirectional long short-term memory network and conditional random field) model is used to identify the named entities. Based on the co-occurrence frequency of the semantic vectors and the correlation with BERT (pre-trained language model), the edge weight is calculated using a threshold of cosine similarity greater than 0.7 as a strong correlation judgment criterion.

[0012] As a further solution of the present invention: the dual-channel strategy for detecting sensitive words includes: The first channel is set as an exact matching phase, and in the exact matching phase, the words are strictly compared with a preset word library, wherein the preset word library includes standard sensitive words and variant combinations of the standard sensitive words; The second channel is set as the fuzzy matching link, and the Sentence-BERT (an improved model based on BERT) model is used to calculate the semantic similarity between the speech text and the sensitive concepts.

[0013] As a further solution of the present invention: the synchronous establishment of the acoustic environment dynamic adaptation mechanism includes: Set the KL divergence threshold of the spectral feature to ≤ 0.15, compare the KL divergence of the spectral features of the registered voiceprint and the real-time sound scene, train the spectral correlation decision model, and obtain the target spectral correlation decision model; A target spectrum correlation decision model is used to detect the credibility mapping and forgery traces of acoustic features, where the forgery traces include artificial harmonic distortion components.

[0014] As a further solution of the present invention: the method of collecting the user's fingerprint information, geographic location, and historical behavior log based on the user's mobile terminal device and preset scoring rules, and using a risk scoring model to identify the user's voice signal identification risk includes: Set the scoring rules as: When a user uses a device for the first time or has not bound the device, the device unfamiliarity score is increased; When a device logs in from a different location or is triggered by a high-risk IP, the geographic location anomaly score is increased; When high-frequency voice acquisition attempts fail or unusual operations are detected, the behavioral deviation score is increased. High-frequency voice acquisition refers to more than three acquisitions per minute. Unusual operations include accessing sensitive functions between 1:00 AM and 3:00 AM. Set the weight of each mode to: The weight of device unfamiliarity is 30%, the weight of geographic location anomaly is 40%, and the weight of behavioral deviation is 30%. Using a risk scoring model: , calculate the recognition risk of the speech signal, where, For the comprehensive risk score, Score the device unfamiliarity. is the geographic location anomaly score, is the behavioral deviation score, , , All are weights, 0.3, 0.4, and 0.3 respectively; when When the risk is low, it is judged as a low-risk scenario; when When the risk is too high, it is judged as a medium-risk scenario; when It is determined to be a high-risk scenario.

[0015] As a further solution of the present invention: the multimodal feature fusion and model training based on noise-resistant MFCC acoustic features, prosodic behavior and semantics includes: The system integrates noise-resistant MFCC acoustic features, prosodic and semantic features, and uses the DTW algorithm in the multimodal Transformer network and the cross-modal attention mechanism to align the acoustic features and prosodic timing of the speech. The attention mechanism is introduced to dynamically assign weights to each modality based on risk assessment results, and a learnable parameter matrix is ​​used to adaptively adjust the weights of each modality to obtain fused features. The fused features are input into SVM (kernel switching support vector machine), and RBF (radial basis function) kernel is used to process nonlinear data; When the kernel processes nonlinear data in a sparse dataset, it switches to a linear kernel to effectively prevent overfitting. A multi-task learning framework is used to integrate the losses of voiceprint recognition, semantic analysis, and behavior prediction, and the F1-score indicator is used to optimize weight distribution to achieve overall performance balance.

[0016] In a second aspect, the present invention provides a speech recognition and authentication system based on multimodal features and dynamic evaluation, which is applied to the speech recognition and authentication method based on multimodal features and dynamic evaluation as described in the above scheme. The system includes a data preprocessing module, a feature extraction module, an adaptation mechanism establishment module, a risk scoring module, a model training module, a scenario authentication module, an execution module, and an authentication pass module: The data preprocessing module is configured to use a microphone to collect the user's original voice signal, preprocess the original voice signal, and obtain preprocessed voice data; The feature extraction module is configured to extract feature data of the pre-processed speech data using a multi-dimensional feature hierarchical extraction technique, wherein the feature data includes feature alignment benchmarks, prosodic behavior, and semantics of the pre-processed speech data; The adaptation mechanism establishment module is configured to inject multi-source noise samples into the acoustic feature space of the voiceprint feature model during the initial training phase, construct a noise-resistant hybrid voiceprint spectrum, deploy the Adv-GAN discriminator, perform feature deconstruction on the deep fake speech generated by WaveGAN in the discriminator based on the spectrum-time domain joint identification mechanism, and simultaneously establish a dynamic adaptation mechanism for the acoustic environment; The risk scoring module is configured to collect the user's fingerprint information, geographic location and historical behavior log based on the user's mobile terminal device and preset scoring rules, identify the user's voice signal recognition risk, and obtain a voice recognition risk score; The model training module is configured to perform multimodal feature fusion and model training based on noise-resistant MFCC acoustic features, prosodic behavior and semantics; The scenario authentication module is configured to, when determining that the authentication scenario is a low-risk scenario, require only voice authentication; when determining that the authentication scenario is a medium-risk scenario, enable two-factor authentication for enhanced verification; when determining that the authentication scenario is a high-risk scenario, forcibly enable the second factor authentication, wherein the second factor authentication is fingerprint or dynamic password authentication; The execution module is configured to distribute liveness detection, feature comparison, and second factor verification tasks to multi-core processors for parallel execution in high-risk scenarios; The authentication module is configured to authorize access and generate an encrypted token when authentication is passed, and store login time, device information, risk score and authentication results in a security log.

[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. In the present invention, by breaking through the traditional single-modal dependence, the rhythmic behavior layer and the semantic layer are innovatively integrated based on the acoustic layer to construct a three-dimensional feature system. Among them, the rhythmic layer segments syllables and calculates the speaking rate variation coefficient through energy envelope mutation detection, combines LSTM to analyze abnormal pauses and fundamental frequency fluctuation energy spectrum, and extracts 128-dimensional time series features to quantify natural pronunciation characteristics; the semantic layer uses BERT-wwm to generate context vectors, combines exact matching with Sentence-BERT fuzzy detection to block semantic attacks, further aligns multimodal time series through a dynamic time warping algorithm, uses a cross-modal attention mechanism to achieve semantic-guided acoustic focusing, and dynamically allocates weights based on risk scores, ultimately generating a comprehensive speech representation with significantly enhanced anti-forgery capabilities.

[0018] 2. In this invention, by integrating voiceprint, semantics, behavioral characteristics and environmental adaptation mechanism, a formulaic risk scoring model is introduced for the first time in voice authentication. Quantitative judgment of risk scenarios can ensure the convenience of voice risk identification in low-risk scenarios and enhance security in high-risk scenarios. It can adjust the authentication threshold in real time based on dynamic risk assessment and enable different levels of authentication methods according to different levels of risk scenarios. At the same time, for high-risk scenarios, two-factor authentication is enforced to significantly enhance authentication security and reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a framework diagram of the method of the present invention; Figure 2 A diagram showing the steps of the method of the present invention; Figure 3 It is a system module diagram of the present invention.

[0020] In the figure: 1. Data preprocessing module; 2. Feature extraction module; 3. Adaptation mechanism establishment module; 4. Risk scoring module; 5. Model training module; 6. Scenario authentication module; 7. Execution module; 8. Authentication module. DETAILED DESCRIPTION

[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0022] Example: See also Figure 1-Figure 2 In an embodiment of the present invention, a speech recognition authentication method based on multimodal features and dynamic evaluation is provided, comprising the following steps: S1: Use a microphone to collect the user's original voice signal, preprocess the original voice signal to obtain preprocessed voice data; S2: extracting feature data of the pre-processed speech data using a multi-dimensional feature hierarchical extraction technique, wherein the feature data includes feature alignment benchmarks, prosodic behavior, and semantics of the pre-processed speech data; S3: In the initial training phase based on the voiceprint feature model, multi-source noise samples are injected into the acoustic feature space of the voiceprint feature model to construct a noise-resistant hybrid voiceprint map. The Adv-GAN discriminator is deployed to perform feature deconstruction on the deep fake speech generated by WaveGAN in the discriminator based on the spectrum-time domain joint identification mechanism, and a dynamic adaptation mechanism for the acoustic environment is simultaneously established. S4: Based on the user's mobile terminal device and preset scoring rules, the user's fingerprint information, geographic location and historical behavior log are collected, and a risk scoring model is used to identify the user's voice signal recognition risk, thereby obtaining a voice recognition risk score to identify the voice signal recognition risk; S5: Multimodal feature fusion and model training based on noise-resistant MFCC acoustic features, prosodic behavior and semantics; S6: Determine the scenario level based on the risk scoring results. When the authentication scenario is judged to be a low-risk scenario, the cosine similarity threshold is set to 0.80, and only voice authentication is required; when When the authentication scenario is judged to be a medium-risk scenario, the cosine similarity threshold is set to 0.85, and two-factor authentication is enabled for enhanced verification; when When the authentication scenario is judged as a high-risk scenario, the cosine similarity threshold is set to 0.90, and the second factor authentication is forcibly enabled. The second factor authentication is fingerprint or dynamic password authentication; S7: In high-risk scenarios, liveness detection, feature comparison, and second-factor authentication tasks are assigned to multi-core processors for parallel execution. S8: When authentication is successful, access is authorized and an encrypted token is generated, and the login time, device information, risk score, and authentication results are stored in the security log.

[0023] In this embodiment, the risk scoring model is:

[0024] in: To create a comprehensive risk score; Score device unfamiliarity; Score for geographic location anomaly; score for behavioral deviance; , , All are weights, which are 0.3, 0.4, and 0.3 respectively.

[0025] In this embodiment, low-risk scenarios are scenarios with a score ≤ 30, and the low-risk scenario threshold is set to 0.80 (cosine similarity); medium-risk scenarios are scenarios with a score > 30 and ≤ 70, and the medium-risk scenario threshold is set to 0.85 (cosine similarity); high-risk scenarios are scenarios with a score > 70, and the high-risk scenario threshold is set to 0.90; In this embodiment, liveness detection includes fingerprint detection and iris detection.

[0026] In this embodiment, the signal-to-noise ratio of the multi-source noise sample is 20dB~35dB; In this embodiment, the invalid segment includes a silent segment.

[0027] In this embodiment, step S1 is used to complete the normalization process of the original speech signal.

[0028] In this embodiment, feature deconstruction is to detect false resonance peaks.

[0029] In this embodiment, the multimodal Transformer network is a Transformer-based deep learning model, and the attention mechanism includes semantically guided acoustic focusing).

[0030] In this embodiment, dynamic time warping (DTW) technology is used to align acoustic features and rhythmic timing to ensure the synchronization of multimodal data. In cross-modal fusion, an attention mechanism (such as semantically guided acoustic focusing) is introduced to enhance the interaction between different modalities. The weights of each modality are dynamically allocated based on the risk assessment results, and adaptive adjustment is achieved through a learnable parameter matrix to improve the robustness of the model. In the kernel switching support vector machine (SVM) after fusion of feature input, the radial basis function (RBF) kernel is used by default to process nonlinear data, but it automatically switches to a linear kernel in sparse data set scenarios to effectively prevent overfitting problems. The multi-task learning framework integrates losses such as voiceprint recognition, semantic analysis, and behavior prediction, and uses the F1-score indicator to optimize the task weight distribution to achieve overall performance balance.

[0031] Preferably, a microphone is used to collect the user's original voice signal, and the original voice signal is preprocessed to obtain preprocessed voice data, including: Read the original speech signal frame by frame and remove duplicate or invalid segments to obtain valid speech data, store the valid speech data as a standardized data set, convert the standardized data set into a WAV format file and output it; The RNNoise deep learning model is used to remove ambient noise from a WAV file to obtain a de-noised frequency spectrum. The non-speech audio segments of the de-noised frequency spectrum are filtered to obtain a frequency-domain speech signal. The non-speech audio segments include mechanical noise and / or sudden noise. The frequency domain speech signal is divided into short-time frame speech signals by using a speech segmentation method. The speech start and end points of the short-time frame speech signal are detected based on the short-time energy and zero-crossing rate to obtain framed speech data. The short-time frame speech signal is a speech signal with a frame length of no more than 20ms. Speech segments with insufficient speech length or too low signal-to-noise ratio are eliminated from the framed speech data to obtain preprocessed speech data, wherein speech segments with insufficient speech length are speech segments with a time length of less than 0.5 seconds, and speech segments with too low signal-to-noise ratio are speech segments with a signal-to-noise ratio of less than 15 dB.

[0032] Preferably, a multi-dimensional feature hierarchical extraction technique is used to extract feature data of the pre-processed speech data, including: Obtain feature alignment benchmarks for preprocessed speech data; Analyze the prosodic behavior of preprocessed speech data; Understand the semantics of pre-processed speech data.

[0033] Preferably, obtaining a feature alignment benchmark includes: The noise-resistant MFCC method including second-order dynamic difference and fundamental frequency trajectory fluctuation analysis method are used to extract the short-term energy envelope prosodic behavior markers of the preprocessed speech data; Generate a behavior-enhanced voiceprint based on prosodic behavior markers based on time-frequency domain joint coding to obtain a feature alignment benchmark for preprocessed speech data, wherein the prosodic behavior markers include syllable onsets and / or energy transition gradients; Analyze the prosodic behavior of preprocessed speech data, including: Segment the speech stream of preprocessed speech data into independent syllables based on energy envelope mutation detection algorithm; Calculating a dynamic coefficient of variation of speech rate for an independent syllable, and quantifying pronunciation stability based on the dynamic coefficient of variation of speech rate, wherein the dynamic coefficient of variation of speech rate includes a standard deviation and / or a mean; An LSTM classifier was used to analyze the distribution of abnormal pauses in independent syllables. The energy spectrum of the second-order derivative fluctuation of the fundamental frequency of independent syllables (0-20Hz) was extracted to characterize the naturalness of the fundamental frequency variation and obtain the fundamental frequency variation data. Abnormal pauses are speech segments with a pause duration greater than 300 milliseconds, and the energy spectrum of the second-order derivative fluctuation of the fundamental frequency is the energy spectrum with a frequency range of 0-20Hz. Based on the self-attention mechanism, the features of the ±20% region of speech rate mutation in the fundamental frequency change data are weightedly fused, and a 128-dimensional prosodic feature vector is generated to detect abnormal behaviors of unnatural speech rate jumps, among which abnormal behaviors include robotic voices.

[0034] In this embodiment, the abnormal behavior of unnatural speech speed jump is detected by using a 128-dimensional prosodic feature vector, which can break through the limitation of traditional methods that rely on a single fundamental frequency mean and effectively improve the accuracy and comprehensiveness of prosodic behavior analysis.

[0035] Understand the semantics of pre-processed speech data, including: Based on the BERT-wwm (BERT-based natural language processing) model, the preprocessed speech data is input into the BERT-wwm model to extract the 768-dimensional [CLS] semantic vector and construct a conversation logic graph. A dual-channel strategy is used to detect sensitive words.

[0036] In this embodiment, MFCC is a feature extraction method widely used in audio signal processing and speech recognition. Specifically, MFCC is used for noise reduction and, combined with first-order and / or second-order dynamic differences, can better describe the time-domain variation characteristics of speech signals (such as changes in pitch or speaking rate), thereby improving the robustness of speech recognition or classification systems. MFCC simulates the way the human ear perceives sound and converts the audio signal into a set of coefficients that can characterize its spectral characteristics. These coefficients can be used to train machine learning models or perform speech analysis.

[0037] Preferably, constructing a dialogue logic map includes: The nodes in the 768-dimensional [CLS] semantic vector are defined as named entities, and the BiLSTM-CRF (bidirectional long short-term memory network and conditional random field) model is used to identify the named entities. Based on the co-occurrence frequency of semantic vectors and the correlation with BERT (pre-trained language model), the edge weight is calculated using a threshold of cosine similarity greater than 0.7 as the strong correlation judgment criterion.

[0038] In this embodiment, by calculating edge weights based on the co-occurrence frequency of semantic vectors and the correlation between them and BERT (pre-trained language model), the semantic coherence and reasoning ability of the graph can be enhanced.

[0039] Preferably, a dual-channel strategy is used to detect sensitive words, including: The first channel is set as an exact matching phase, and in the exact matching phase, the words are strictly compared with the preset word library, wherein the preset word library includes standard sensitive words and variant combinations of standard sensitive words; The second channel is set as the fuzzy matching link, and the Sentence-BERT (an improved model based on BERT) model is used to calculate the semantic similarity between the speech text and the sensitive concepts.

[0040] In this embodiment, precise matching is performed on a preset vocabulary (including variant combinations such as "identity-verification"), and fuzzy matching is performed through semantic similarity (Sentence-BERT) to detect circumventing expressions (such as "account permission confirmation" is mapped to "identity verification"). By setting up a first channel and a second channel, the first channel can strictly compare with the preset vocabulary to ensure efficient interception of direct expressions. At the same time, the second channel can effectively detect circumventing expressions, improving the system's recognition coverage of variants or obscure expressions.

[0041] Preferably, a dynamic adaptation mechanism for the acoustic environment is established simultaneously, including: Set the KL divergence threshold of the spectral feature to ≤ 0.15, compare the KL divergence of the spectral features of the registered voiceprint and the real-time sound scene, train the spectral correlation decision model, and obtain the target spectral correlation decision model; A target spectrum correlation decision model is used to detect the credibility mapping and forgery traces of acoustic features, where the forgery traces include artificial harmonic distortion components.

[0042] Preferably, based on the user's mobile terminal device and preset scoring rules, the user's fingerprint information, geographic location and historical behavior log are collected to identify the user's voice signal recognition risk, including: Set the scoring rules as: When a user uses a device for the first time or has not bound the device, the device unfamiliarity score is increased; When a device logs in from a different location or is triggered by a high-risk IP, the geographic location anomaly score is increased; When high-frequency voice acquisition attempts fail or unusual operations are detected, the behavioral deviation score is increased. High-frequency voice acquisition refers to more than 3 acquisitions per minute, and unusual operations include accessing sensitive functions at 3 a.m. Set the weight of each mode to: The weight of device unfamiliarity is 30%, the weight of geographic location abnormality is 40%, and the weight of behavioral deviation is 30%.

[0043] Preferably, multimodal feature fusion and model training are performed based on noise-resistant MFCC acoustic features, prosodic behavior and semantics, including: The system integrates noise-resistant MFCC acoustic features, prosodic and semantic features, and uses the DTW algorithm in the multimodal Transformer network and the cross-modal attention mechanism to align the acoustic features and prosodic timing of the speech. The attention mechanism is introduced to dynamically assign weights to each modality based on risk assessment results, and a learnable parameter matrix is ​​used to adaptively adjust the weights of each modality to obtain fused features. The fused features are input into SVM (kernel switching support vector machine), and RBF (radial basis function) kernel is used to process nonlinear data; When the kernel processes nonlinear data in a sparse dataset, it switches to a linear kernel to effectively prevent overfitting. A multi-task learning framework is used to integrate the losses of voiceprint recognition, semantic analysis, and behavior prediction, and the F1-score indicator is used to optimize weight distribution to achieve overall performance balance.

[0044] like Figure 3 As shown, this embodiment provides a speech recognition authentication system based on multimodal features and dynamic evaluation, which is applied to the speech recognition authentication method based on multimodal features and dynamic evaluation in the above-mentioned solution. The system includes a data preprocessing module 1, a feature extraction module 2, an adaptation mechanism establishment module 3, a risk scoring module 4, a model training module 5, a scenario authentication module 6, an execution module 7, and an authentication pass module 8; The data preprocessing module 1 is configured to collect the user's original voice signal using a microphone, preprocess the original voice signal, and obtain preprocessed voice data; The feature extraction module 2 is configured to extract feature data of the pre-processed speech data using a multi-dimensional feature hierarchical extraction technique, wherein the feature data includes feature alignment benchmarks, prosodic behavior, and semantics of the pre-processed speech data; Adaptation mechanism establishment module 3 is configured as the initial training phase based on the voiceprint feature model. It injects multi-source noise samples into the acoustic feature space of the voiceprint feature model to construct a noise-resistant hybrid voiceprint spectrum. It deploys the Adv-GAN discriminator and performs feature deconstruction on the deep fake speech generated by WaveGAN in the discriminator based on the spectrum-time domain joint identification mechanism. It also establishes a dynamic adaptation mechanism for the acoustic environment. The risk scoring module 4 is configured to collect the user's fingerprint information, geographic location and historical behavior log based on the user's mobile terminal device and preset scoring rules, identify the user's voice signal recognition risk, and obtain a voice recognition risk score; The model training module 5 is configured to perform multimodal feature fusion and model training based on noise-resistant MFCC acoustic features, prosodic behavior and semantics; The scenario authentication module 6 is configured to, when determining that the authentication scenario is a low-risk scenario, require only voice authentication; when determining that the authentication scenario is a medium-risk scenario, enable two-factor authentication for enhanced verification; when determining that the authentication scenario is a high-risk scenario, forcibly enable the second factor authentication, wherein the second factor authentication is fingerprint or dynamic password authentication; The execution module 7 is configured to distribute the liveness detection, feature comparison, and second factor verification tasks to the multi-core processor for parallel execution in high-risk scenarios; The authentication module 8 is configured to authorize access and generate an encrypted token when authentication is passed, and store the login time, device information, risk score and authentication results in the security log.

[0045] The above are only preferred specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solutions and inventive concepts of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A speech recognition authentication method based on multimodal features and dynamic evaluation, characterized in that: include: Using a microphone to collect the user's original voice signal, preprocessing the original voice signal to obtain preprocessed voice data; Extracting feature data of the pre-processed speech data using a multi-dimensional feature hierarchical extraction technique, wherein the feature data includes feature alignment benchmarks, prosodic behavior, and semantics of the pre-processed speech data; During the initial training phase based on the voiceprint feature model, multi-source noise samples are injected into the acoustic feature space of the voiceprint feature model to construct a noise-resistant hybrid voiceprint map. The Adv-GAN discriminator is deployed to perform feature deconstruction on the deep fake speech generated by WaveGAN in the discriminator based on a spectral-temporal joint identification mechanism, and a dynamic adaptation mechanism for the acoustic environment is simultaneously established. Based on the user's mobile terminal device and preset scoring rules, the user's fingerprint information, geographic location and historical behavior logs are collected, and a risk scoring model is used to identify the user's voice signal recognition risk to obtain a voice recognition risk score; Multimodal feature fusion and model training based on noise-resistant MFCC acoustic features, prosodic behavior and semantics; When the authentication scenario is judged to be a low-risk scenario, only voice authentication is required; when the authentication scenario is judged to be a medium-risk scenario, two-factor authentication is enabled for enhanced verification; when the authentication scenario is judged to be a high-risk scenario, the second factor authentication is forcibly enabled, where the second factor authentication is fingerprint or dynamic password authentication; In high-risk scenarios, liveness detection, feature matching, and second-factor authentication tasks are assigned to multi-core processors for parallel execution. When authentication is successful, access is authorized and an encrypted token is generated, and the login time, device information, risk score, and authentication results are stored in the security log.

2. The method for voice recognition and authentication based on multimodal features and dynamic evaluation according to claim 1, characterized in that: The user's original voice signal is collected by a microphone, and the original voice signal is preprocessed to obtain preprocessed voice data, including: Read the original speech signal frame by frame and remove repeated or invalid segments to obtain valid speech data, store the valid speech data as a standardized data set, convert the standardized data set into a WAV format file and output it; Using the RNNoise deep learning model to eliminate ambient noise in the WAV format file to obtain a de-noised audio signal, filtering the non-speech audio segment of the de-noised audio signal to obtain a frequency domain speech signal, wherein the non-speech audio segment includes mechanical noise and / or sudden noise; The frequency domain speech signal is divided into short-time frame speech signals by using a speech segmentation method, and speech start and end points of the short-time frame speech signals are detected based on short-time energy and zero-crossing rate to obtain framed speech data, wherein the short-time frame speech signals are speech signals with a frame length not exceeding 20 ms. Speech segments with insufficient speech length or too low signal-to-noise ratio in the framed speech data are eliminated to obtain preprocessed speech data, wherein the speech segment with insufficient speech length is a speech segment with a time length of less than 0.5 seconds, and the speech segment with too low signal-to-noise ratio is a speech segment with a signal-to-noise ratio of less than 15 dB.

3. The method for voice recognition and authentication based on multimodal features and dynamic evaluation according to claim 2, characterized in that: The method of extracting feature data of the pre-processed speech data using a multi-dimensional feature hierarchical extraction technology includes: Obtaining a feature alignment benchmark for the preprocessed speech data; analyzing prosodic behavior of the preprocessed speech data; The semantics of the pre-processed speech data is understood.

4. The method for voice recognition and authentication based on multimodal features and dynamic evaluation according to claim 3, characterized in that: Obtain feature alignment datums, including: Extracting short-term energy envelope prosodic behavior markers of the preprocessed speech data using a noise-resistant MFCC and fundamental frequency trajectory fluctuation analysis method including first-order and / or second-order dynamic difference; generating a behavior-enhanced voiceprint of the prosodic behavior marker based on time-frequency domain joint coding to obtain a feature alignment benchmark for the preprocessed speech data, wherein the prosodic behavior marker includes a syllable starting point and / or an energy transition gradient; The analyzing the prosodic behavior of the pre-processed speech data includes: Segmenting the speech stream of the pre-processed speech data into independent syllables based on an energy envelope mutation detection algorithm; Calculating a dynamic coefficient of variation of the speaking rate of the independent syllable, and quantifying pronunciation stability based on the dynamic coefficient of variation of the speaking rate, wherein the dynamic coefficient of variation of the speaking rate includes a standard deviation and / or a mean; An LSTM classifier is used to analyze the distribution of abnormal pauses in the independent syllables, and the energy spectrum of the second-order derivative fluctuation of the fundamental frequency of the independent syllables is extracted to characterize the naturalness of the fundamental frequency change, thereby obtaining fundamental frequency change data. An abnormal pause is a speech segment with a pause duration greater than 300 milliseconds, and the energy spectrum of the second-order derivative fluctuation of the fundamental frequency is an energy spectrum with a frequency range of 0-20 Hz. Based on the self-attention mechanism, weighted fusion is performed on the features of the ±20% region of the speech rate mutation in the fundamental frequency variation data, and a 128-dimensional prosodic feature vector is generated to detect abnormal behaviors of unnatural speech rate jumps, wherein abnormal behaviors include robotic voices; The understanding of the semantics of the pre-processed speech data includes: Based on the BERT-wwm model, the speech text of the preprocessed speech data is input into the BERT-wwm model to extract a 768-dimensional semantic vector and construct a dialogue logic graph; A dual-channel strategy is used to detect sensitive words.

5. The method for voice recognition and authentication based on multimodal features and dynamic evaluation according to claim 4, characterized in that: The construction of the dialogue logic graph includes: The nodes in the 768-dimensional semantic vector are defined as named entities, and the BiLSTM-CRF model is used to identify the named entities. Based on the co-occurrence frequency of the semantic vector and the correlation with BERT, the edge weight is calculated using a threshold of cosine similarity greater than 0.7 as a strong correlation judgment criterion.

6. The method for voice recognition and authentication based on multimodal features and dynamic evaluation according to claim 5, characterized in that: The dual-channel strategy for detecting sensitive words includes: The first channel is set as an exact matching phase, and in the exact matching phase, the words are strictly compared with a preset word library, wherein the preset word library includes standard sensitive words and variant combinations of the standard sensitive words; The second channel is set as the fuzzy matching link, and the Sentence-BERT model is used to calculate the semantic similarity between the speech text and the sensitive concepts.

7. The method for voice recognition and authentication based on multimodal features and dynamic evaluation according to claim 6, characterized in that: The synchronous establishment of a dynamic adaptation mechanism for the acoustic environment includes: Set the KL divergence threshold of the spectral feature to ≤ 0.15, compare the KL divergence of the spectral features of the registered voiceprint and the real-time sound scene, train the spectral correlation decision model, and obtain the target spectral correlation decision model; A target spectrum correlation decision model is used to detect the credibility mapping and forgery traces of acoustic features, where the forgery traces include artificial harmonic distortion components.

8. The method for voice recognition and authentication based on multimodal features and dynamic evaluation according to claim 7, characterized in that: The method, based on the user's mobile terminal device and preset scoring rules, collects the user's fingerprint information, geographic location, and historical behavior logs, and uses a risk scoring model to identify the user's voice signal identification risk, including: Set the scoring rules as: When a user uses a device for the first time or has not bound the device, the device unfamiliarity score is increased; When a device logs in from a different location or is triggered by a high-risk IP, the geographic location anomaly score is increased; When high-frequency voice acquisition attempts fail or unusual operations are detected, the behavioral deviation score is increased. High-frequency voice acquisition refers to more than three acquisitions per minute. Unusual operations include accessing sensitive functions between 1:00 AM and 3:00 AM. Set the weight of each mode to: The weight of device unfamiliarity is 30%, the weight of geographic location anomaly is 40%, and the weight of behavioral deviation is 30%. Using a risk scoring model: , calculate the recognition risk of the speech signal, where, For the comprehensive risk score, Score the device unfamiliarity. is the geographic location anomaly score, is the behavioral deviation score, , , All are weights, 0.3, 0.4, and 0.3 respectively; when When the risk is low, it is judged as a low-risk scenario; when When the risk is too high, it is judged as a medium-risk scenario; when It is determined to be a high-risk scenario.

9. The method for voice recognition and authentication based on multimodal features and dynamic evaluation according to claim 8, characterized in that: The multimodal feature fusion and model training based on noise-resistant MFCC acoustic features, prosodic behavior and semantics include: The system integrates noise-resistant MFCC acoustic features, prosodic and semantic features, and uses the DTW algorithm in the multimodal Transformer network and the cross-modal attention mechanism to align the acoustic features and prosodic timing of the speech. The attention mechanism is introduced to dynamically assign weights to each modality based on risk assessment results, and a learnable parameter matrix is ​​used to adaptively adjust the weights of each modality to obtain fused features. The fusion features are input into SVM, and the RBF kernel is used to process nonlinear data; When the kernel processes nonlinear data in a sparse dataset, switch to a linear kernel. A multi-task learning framework is used to integrate the losses of voiceprint recognition, semantic analysis, and behavior prediction, and the F1-score indicator is used to optimize the weight distribution.

10. A speech recognition and authentication system based on multimodal features and dynamic evaluation, characterized by: Applied to the speech recognition authentication method based on multimodal features and dynamic evaluation according to any one of claims 1 to 9, the system comprises: A data preprocessing module is configured to collect the user's original voice signal using a microphone, preprocess the original voice signal, and obtain preprocessed voice data; a feature extraction module configured to extract feature data of the preprocessed speech data using a multi-dimensional feature hierarchical extraction technique, wherein the feature data includes a feature alignment benchmark, prosodic behavior, and semantics of the preprocessed speech data; An adaptation mechanism establishment module, configured to, based on the initial training phase of the voiceprint feature model, inject multi-source noise samples into the acoustic feature space of the voiceprint feature model to construct a noise-resistant hybrid voiceprint map, deploy an Adv-GAN discriminator, perform feature deconstruction on the deepfake speech generated by WaveGAN in the discriminator based on a joint spectrum-time domain identification mechanism, and simultaneously establish a dynamic adaptation mechanism for the acoustic environment; A risk scoring module is configured to collect the user's fingerprint information, geographic location, and historical behavior logs based on the user's mobile terminal device and preset scoring rules, identify the user's voice signal recognition risk, and obtain a voice recognition risk score; A model training module configured to perform multimodal feature fusion and model training based on noise-resistant MFCC acoustic features, prosodic behavior, and semantics; A scenario authentication module, wherein the scenario authentication module is configured to require only voice authentication when the authentication scenario is determined to be a low-risk scenario; enable two-factor authentication for enhanced verification when the authentication scenario is determined to be a medium-risk scenario; and forcibly enable a second factor authentication when the authentication scenario is determined to be a high-risk scenario, wherein the second factor authentication is fingerprint or dynamic password authentication; an execution module configured to distribute liveness detection, feature comparison, and second factor verification tasks to a multi-core processor for parallel execution in high-risk scenarios; The authentication module is configured to authorize access and generate an encrypted token when the authentication is passed, and store the login time, device information, risk score and authentication results in the security log.

Citation Information

Patent Citations

  • A voiceprint recognition method and a voiceprint recognition device

    CN109920435B

  • Multimodal fusion speech translation method, system and equipment

    CN118692446A

  • Audio and video identification system based on multi-modal analysis

    CN118692662A

  • Speech recognition and natural language processing integration method and system

    CN120220652A

  • Sound standardization evaluation system for multi-dimensional acoustic parameter and emotional scene dynamic analysis

    CN120340540A

Cited By

  • Television terminal education application multi-mode user authentication method and system

    CN120956974A

  • Intelligent glasses voiceprint anti-counterfeiting recognition method and system based on NFC (Near Field Communication)

    CN121054002A

  • Service robot behavior recognition method based on multi-modal perception fusion

    CN121350967A

  • Multi-modal radio frequency authentication method based on multi-scale signal representation

    CN121418820A

  • A multi-modal radio frequency authentication method based on multi-scale signal representation

    CN121418820B