Language processing system for speech recognition combining machine learning and deep learning

By combining machine learning and deep learning technology in the speech recognition system, mainstream tone learning models and dynamic tone compensation mechanisms are built, and misjudgment problems in the noise environment and user tone changes in the existing technology are solved, and the robustness and personalized adaptation of speech recognition are achieved.

CN120220666AActive Publication Date: 2025-06-27YUNNAN BEIFEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510551165.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-06-27
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

The existing Deep Speaker Embeddings technology cannot adapt to the tone line in a noisy environment or when the user's tone changes, resulting in the system misjudging it as a different speaker and rejecting the user's voice input.

Method used

A speech recognition language processing system combining machine learning and deep learning is designed, including an audio acquisition and processing unit, a mainstream tone learning unit, a noise learning storage unit, and a natural speech recognition unit. The mainstream tone learning model is constructed through a variational autoencoder combined with timing dynamic modeling technology, and dynamically determine whether the current tone deviates from the normal tone distribution, and performs real-time tone compensation when deviating.

Benefits of technology

It realizes the robustness of speech recognition in the noise environment and user tone changes, avoids the system's misjudgment as different speakers, and ensures the accuracy and personalized adaptation of user voice input.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220666A_ABST
    Figure CN120220666A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of language processing, in particular to a speech recognition language processing system combining machine learning and deep learning, which comprises an audio acquisition and processing unit, a speech recognition unit, a speech recognition unit and a speech recognition unit, the ontology timbre learning unit continuously learns normal timbre distribution of a user based on a variational auto-encoder in combination with a time sequence dynamic modeling technology, constructs a main stream timbre learning model, constructs a main and auxiliary stream dynamic gating mechanism through a generative adversarial network in combination with a KL divergence detection technology, and judges whether the current timbre deviates from the normal timbre distribution; the noise learning storage unit is used for continuously learning wind noise features and KTV noise features; the natural-color speech recognition unit screens user speech signals based on a mainstream timbre learning model and a noise feature library for speech processing. The speech processing system for speech recognition combining machine learning and deep learning locks and extracts ontology user speech in a high-noise environment based on a mainstream timbre learning model in combination with a noise feature library.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of language processing. Specifically, it relates to a language processing system for speech recognition that combines machine learning and deep learning. Background Art

[0002] The language processing system for speech recognition that combines machine learning and deep learning aims to accurately recognize and continuously adapt to the individual timbre of users and improve the robustness of speech recognition in complex noise environments. Through long-term timbre modeling of the mainstream timbre learning model and the sub-stream timbre compensation mechanism, it controls the short-term timbre changes caused by users' illness and fatigue or the recognition errors caused by environmental noise interference, realizes the dynamic adaptation of users' personalized timbre, and enables the system to accurately lock the voice signal of the original user even in the case of overlapping voices of multiple people, wind noise environment during cycling, noisy environment in KTV, and changes in users' timbre.

[0003] Existing language processing systems for speech recognition usually have difficulty coping with the short-term timbre changes of users, such as the voice changes caused by colds and illnesses. Moreover, when in a noise environment or when the user's timbre changes, the existing Deep Speaker Embeddings technology cannot perform online timbre adaptation, which will lead to the problem that the system misjudges different speakers and rejects the user's voice input. Therefore, a language processing system for speech recognition that combines machine learning and deep learning is designed. Summary of the Invention

[0004] The purpose of the present invention is to provide a language processing system for speech recognition that combines machine learning and deep learning, so as to solve the problem proposed in the above background art that when in a noise environment or when the user's timbre changes, the existing Deep Speaker Embeddings technology cannot perform online timbre adaptation, which will lead to the problem that the system misjudges different speakers and rejects the user's voice input.

[0005] To achieve the above purpose, the present invention aims to provide a language processing system for speech recognition that combines machine learning and deep learning, including:

[0006] An audio acquisition and processing unit, which is used to acquire audio signals, perform preprocessing, and extract audio features;

[0007] It further includes an original timbre learning unit. The original timbre learning unit continuously learns the normal timbre distribution of users and constructs a mainstream timbre learning model based on the variational autoencoder combined with the time series dynamic modeling technology, and constructs a main and sub-stream dynamic gating mechanism through the generative adversarial network combined with the KL divergence detection technology to dynamically determine whether the current timbre deviates from the normal timbre distribution;

[0008] It also includes a noise learning and storage unit, which is used to continuously learn the wind noise characteristics and KTV noise characteristics and store them in the noise feature library;

[0009] It also includes a natural voice recognition unit, which is based on a mainstream timbre learning model and a noise feature library, and is used to screen user voice signals in a noisy environment and convert them into text language using end-to-end voice recognition technology.

[0010] As a further improvement of this technical solution, the audio acquisition and processing unit includes an audio acquisition module and an audio processing module;

[0011] Among them, the audio acquisition module uses sensors to collect user audio signals and environmental audio signals;

[0012] The audio processing module is used to preprocess the user audio signal and the environmental audio signal, and extract the audio features of the user audio signal and the environmental audio signal.

[0013] As a further improvement of this technical solution, the ontology timbre learning unit includes a mainstream timbre learning module and a sub-stream timbre monitoring module;

[0014] Among them, the mainstream timbre learning module constructs a mainstream timbre learning model based on a variational autoencoder combined with a time series dynamic modeling technique, and is used to continuously learn the normal timbre distribution of users;

[0015] The sub-stream timbre monitoring module constructs a main-sub-stream dynamic gating mechanism through a generative adversarial network combined with KL divergence detection technology, dynamically judges whether the user's real-time timbre deviates from the normal timbre distribution, and performs real-time compensation on the abnormal timbre when the user's timbre is abnormal.

[0016] As a further improvement of this technical solution, the mainstream timbre learning module constructs a mainstream timbre learning model based on a variational autoencoder combined with a time series dynamic modeling technique. The specific method steps are as follows:

[0017] S2.1.1. Extract the user audio signal and the user audio feature from the audio acquisition and processing unit;

[0018] The user audio signal includes: a time-domain user audio signal and a frequency-domain user audio signal;

[0019] The user audio features include: user MFCC timbre features, user LPC cepstral coefficients, and user timbre time series features;

[0020] S2.1.2. Construct a user timbre feature vector based on the user audio feature, and reconstruct the user timbre distribution through a variational autoencoder and optimize the variational autoencoder;

[0021] S2.1.3. Track the timbre change and analyze the current timbre state using the time series dynamic modeling technique based on the user's timbre distribution;

[0022] S2.1.4. Construct a mainstream timbre learning model based on the user's timbre distribution and the current timbre state.

[0023] As a further improvement of this technical solution, in S2.1.2, construct a user timbre feature vector based on the user's audio features, and reconstruct the user's timbre distribution and optimize the variational autoencoder through the variational autoencoder:

[0024] Input the user timbre feature vector into the encoder:

[0025] x = [M u , L u , T u ;

[0026]

[0027] where x is the user timbre feature vector; z is the user timbre latent variable; M u is the user's MFCC timbre feature; L u is the user's LPC cepstral coefficient; T u is the user's timbre time series feature; μ u is the timbre mean; is the timbre variance; q(z|x) is the variational posterior distribution; is the Gaussian distribution;

[0028] The user timbre feature reconstructed by the decoder:

[0029]

[0030] where p(x|z) is the reconstructed distribution of the timbre feature; f(z) is the mapping function of the decoder; is the timbre feature variance;

[0031] Calculate the reconstruction error to optimize the variational autoencoder:

[0032]

[0033] where is the total loss function of the variational autoencoder; p(z) is the prior distribution; D KL (q(z|x)||p(z)) is the KL divergence between the variational posterior distribution q(z|x) and the prior distribution p(z); is the reconstruction error.

[0034] As a further improvement of this technical solution, in S2.1.3, based on the user's timbre distribution, a time-series dynamic modeling technique is used to track the timbre change and analyze the current timbre state. The specific method is as follows:

[0035] Use a Kalman filter to track the timbre change:

[0036] X t = AX t-1 + w t ;

[0037] where t is the time step; X t is the timbre state vector at time step t; A is the timbre change transition matrix; X t-1 is the timbre state vector at time step t-1; w t is the Gaussian noise term, representing the uncertainty of the timbre state change at time t, that is, the model error term;

[0038] Use Gated LSTM to analyze the current timbre state:

[0039] h t = σ(W h X t + U h h t-1 + b h );

[0040] where h t is the timbre state at the current time step t; σ(·) is the Sigmoid gating activation function; W h is the timbre input weight matrix; U h is the weight matrix of the previous time step state; h t-1 is the timbre state at time step t-1; b h is the Gated LSTM bias term;

[0041] In S2.1.4, based on the user's timbre distribution and the current timbre state, a mainstream timbre learning model is constructed. The specific method is as follows:

[0042] p(x normal |z t ) = αp(x normal |z t-1 )+(1-α)p(x t |z t );

[0043] α = σ(W α h t + b α );

[0044] where p(x normal |zt ) is the normal tone distribution of the user at time step t; α is the dynamic smoothing factor; p(x normal |z t-1 ) is the normal tone distribution of the user at time step t-1; p(x t |z t ) is the reconstructed distribution of the tone features at time step t; W α is the Gated LSTM weight matrix; b α is the bias term.

[0045] As a further improvement of this technical solution, the secondary flow tone monitoring module constructs a main-secondary flow dynamic gating mechanism through a generative adversarial network combined with KL divergence detection technology, dynamically determines whether the user's real-time tone deviates from the normal tone distribution, and performs real-time compensation on the abnormal tone when the user's tone is abnormal. The specific method steps are as follows:

[0046] Main-secondary flow dynamic gating mechanism:

[0047] S2.2.1. Calculate the KL divergence of the user's tone based on the tone features of the user reconstructed by the decoder and the mainstream tone learning model:

[0048]

[0049] Among them, D KL (p(x|z)||p(x normal |z)) is the KL divergence of the user's tone;

[0050] S2.2.2. Set the user's tone divergence threshold ∈, and judge whether the user's tone is abnormal:

[0051] If D KL (p(x|z)||p(x normal |z)) ≤ ∈, then the user's tone is normal, and the mainstream tone learning model continues to be used;

[0052] If D KL (p(x|z)||p(x normal |z)) > ∈, then tone compensation is triggered, and S2.2.3 is executed to perform tone compensation using a generative adversarial network;

[0053] S2.2.3. Use the generator of the generative adversarial network to generate a corrected tone, and use the discriminator to evaluate whether the corrected tone is reasonable. If it is reasonable, input the corrected tone into the mainstream tone learning model.

[0054] As a further improvement of this technical solution, the noise learning and storage unit includes a noise feature learning module and a noise feature storage module;

[0055] Among them, the noise feature learning module is used to collect environmental noise features in real time and learn wind noise features and KTV background noise features;

[0056] The noise feature storage module is used to store and manage noise feature data to form a noise feature library.

[0057] As a further improvement of this technical solution, the natural voice recognition unit is based on a mainstream timbre learning model and a noise feature library, and is used to screen user voice signals in a noisy environment and convert them into text language using end-to-end speech recognition technology. The specific method steps are as follows:

[0058] S4.1. Obtain a new audio signal and determine whether the user audio signal in the new audio signal conforms to the user's own timbre. The specific method is as follows:

[0059] Obtain a new audio signal from the audio acquisition and processing unit (1), use the KL divergence detection technology to calculate the current timbre offset, and compare it with the user's own timbre in the mainstream timbre learning model for compliance:

[0060] If the current timbre offset is greater than the user timbre divergence threshold ∈, it does not conform to the user's own timbre and the new audio signal is rejected;

[0061] If the current timbre offset is less than or equal to the user timbre divergence threshold ∈, it conforms to the user's own timbre, and the new audio signal is input to the body timbre learning unit (2) for analysis and S4.2 is executed;

[0062] S4.2. Based on the noise feature library, use the Wiener filtering technology to separate the noise signals existing in the noise feature library from the new audio signal, and only retain the pure user audio signal;

[0063] S4.3. Based on the pure user audio signal, use end-to-end speech recognition technology to convert the pure user audio signal into text language.

[0064] As a further improvement of this technical solution, in S4.1, the user audio signal is input to the body timbre learning unit, which is used to trigger the main and auxiliary flow dynamic gating mechanism for timbre compensation when it is detected that the user's timbre deviates greatly from the existing distribution, and is also used for the mainstream timbre learning model to continuously and dynamically learn the user's normal timbre every time the user's voice is recognized.

[0065] Compared with the prior art, the beneficial effects of the present invention:

[0066] 1. In the speech recognition language processing system combining machine learning and deep learning, a mainstream timbre learning model that combines a variational autoencoder with a temporal dynamic modeling technique is constructed to continuously learn and update the long-term timbre features of users, establish a personalized timbre distribution, and enable the system to adapt to the natural timbre changes of users in the long term.

[0067] 2. In the speech recognition language processing system combining machine learning and deep learning, by constructing a noise feature library, in a multi-person speech overlap or high-noise environment, based on the noise feature library and the mainstream timbre learning model, the speech of the original user is locked and extracted. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 is the overall flowchart of the present invention;

[0069] The meanings of the reference numerals in the figure are as follows:

[0070] 1. Audio acquisition and processing unit; 11. Audio acquisition module; 12. Audio processing module; 2. Ontological timbre learning unit; 21. Mainstream timbre learning module; 22. Sub-stream timbre monitoring module; 3. Noise learning and storage unit; 31. Noise feature learning module; 32. Noise feature storage module; 4. Natural voice recognition unit. DETAILED DESCRIPTION OF THE INVENTION

[0071] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0072] Please refer to Figure 1 as shown, a speech recognition language processing system combining machine learning and deep learning is provided, including:

[0073] An audio acquisition and processing unit 1, which is used to acquire audio signals, perform preprocessing, and extract audio features;

[0074] In this embodiment, the audio acquisition and processing unit 1 includes an audio acquisition module 11 and an audio processing module 12;

[0075] Among them, the audio acquisition module 11 uses a sensor to acquire user audio signals and environmental audio signals;

[0076] The audio processing module 12 is used to perform preprocessing on the user audio signals and environmental audio signals, and extract the audio features of the user audio signals and environmental audio signals.

[0077] In this embodiment, a microphone array with a six-microphone circular array collects audio signals from multiple directions; the adaptive beamforming technology is used to locate and enhance the speech of the target user, and the Generalized Sidelobe Canceller algorithm is adopted for directional speech enhancement; the Time Difference of Arrival method is used to calculate the sound source direction, and the SRP-PHAT method is combined to improve the positioning accuracy of multi-source speech, so as to distinguish the user audio near-field audio and the environmental noise far-field audio;

[0078] Preprocessing is performed on the user audio signal and the environmental audio signal, and the specific method is as follows:

[0079] Noise cancellation is performed on the user audio signal and the environmental audio signal based on spectral subtraction and SEGANGAN-based Speech Enhancement; for the user audio signal, Mel Frequency Cepstral Coefficients are extracted as the audio features of the user audio signal, and at the same time, LPC Cepstral Coefficients are extracted to track the short-term dynamic characteristics of the speech, which helps subsequent timbre modeling; for the environmental audio signal, MFCC, spectral entropy, and zero-crossing rate are used to identify the noise type: wind noise usually has low-frequency characteristics; KTV noise has strong mid-high frequency components;

[0080] It also includes an ontological timbre learning unit 2. The ontological timbre learning unit 2 continuously learns the normal timbre distribution of users and constructs a mainstream timbre learning model based on a variational autoencoder combined with a time series dynamic modeling technique. A main and auxiliary flow dynamic gating mechanism is constructed through a generative adversarial network combined with a KL divergence detection technique to dynamically judge whether the current timbre deviates from the normal timbre distribution;

[0081] In this embodiment, the ontological timbre learning unit 2 includes a mainstream timbre learning module 21 and an auxiliary flow timbre monitoring module 22;

[0082] Among them, the mainstream timbre learning module 21 constructs a mainstream timbre learning model based on a variational autoencoder combined with a time series dynamic modeling technique for continuously learning the normal timbre distribution of users;

[0083] In this embodiment, the variational autoencoder is a probabilistic generative model used to learn the latent distribution of data. It can map the timbre features of the input speech to a low-dimensional latent variable space and generate a personalized distribution of the user's timbre. The temporal dynamic modeling technology is used to track the changes in the user's timbre over time to ensure that the system can adapt to short-term and long-term timbre changes. In traditional speech recognition systems, especially those using fixed Deep Speaker Embeddings, changes in the user's timbre, whether long-term or short-term, often lead to incorrect matching by the system and are unable to effectively update the user's timbre model. This system tracks the changes in the user's timbre based on the variational autoencoder combined with the temporal dynamic modeling technology to ensure the stability and accuracy of speech recognition and can dynamically and online learn the user's timbre in real time. Specifically, in the case of short-term timbre changes caused by the user's illness, such as a cold or a sore throat, the variational autoencoder can still recognize the user because it learns the overall distribution of the timbre rather than a fixed embedding.

[0084] The secondary stream timbre monitoring module 22 constructs a primary-secondary flow dynamic gating mechanism through a generative adversarial network combined with KL divergence detection technology to dynamically determine whether the user's real-time timbre deviates from the normal timbre distribution and perform real-time compensation for the abnormal timbre when the user's timbre is abnormal.

[0085] In this embodiment, the generative adversarial network is a deep learning model based on the game theory, consisting of a generator and a discriminator. The generator learns to generate realistic data samples from the input data, that is, map the current abnormal timbre back to the normal timbre, while the discriminator is responsible for distinguishing between the generated samples and the real samples to ensure that the generated timbre conforms to the user's normal timbre. The two are trained against each other, enabling the generator to generate more realistic samples. Generally speaking, the generative adversarial network in this module is used to perform timbre compensation when the timbre is abnormal, so that the user can still maintain normal speech recognition when ill, when the device changes, or when there is environmental noise interference.

[0086] The KL divergence detection technology is a measure of the difference between two probability distributions, representing the information loss between a distribution and a reference distribution, and is used to determine whether the current user's timbre significantly deviates from the normal timbre distribution.

[0087] The primary-secondary flow dynamic gating mechanism is used to coordinate the main stream timbre learning module and the secondary stream timbre monitoring module. When the KL divergence detection does not exceed the threshold and the timbre is normal, only the main stream timbre learning model is used for recognition. When the KL divergence exceeds the threshold and the timbre is abnormal, the generative adversarial network timbre compensation is triggered to return the timbre to the normal distribution.

[0088] Construct a main - secondary flow dynamic gating mechanism by combining generative adversarial networks with KL - divergence detection technology to dynamically determine whether the user's real - time timbre deviates from the normal timbre distribution, and perform real - time compensation on abnormal timbres when the user's timbre is abnormal. For example, a user's normal timbre is a clear, medium - frequency male voice. When the user has a cold and the voice becomes hoarse, the system calculates the current timbre distribution, and the KL - divergence detection finds that the timbre deviation is too large, triggering the timbre compensation mechanism to prevent the system from rejecting the user's recognition due to short - term timbre changes. Even when the user has a cold, the system can still correctly recognize the user's speech. For another example, in an environment where multiple people are talking simultaneously, such as a meeting, a conference call, or a KTV, the system needs to distinguish the original user from other users. The KL - divergence detection can determine whether the current voice belongs to the original user. If the KL - divergence is extremely large, it means that the timbre does not conform to the long - term timbre distribution of the original user and rejects the input.

[0089] In this embodiment, the mainstream timbre learning module 21 constructs a mainstream timbre learning model based on a variational auto - encoder combined with a time - series dynamic modeling technique. The specific method steps are as follows:

[0090] S2.1.1. Extract the user audio signal and the user audio features from the audio acquisition and processing unit 1;

[0091] The user audio signal includes: a time - domain user audio signal and a frequency - domain user audio signal;

[0092] The user audio features include: the user's MFCC timbre feature, the user's LPC cepstral coefficient, and the user's timbre time - series feature;

[0093] S2.1.2. Construct a user timbre feature vector based on the user audio features, and reconstruct the user timbre distribution through a variational auto - encoder and optimize the variational auto - encoder;

[0094] S2.1.3. Based on the user timbre distribution, use the time - series dynamic modeling technique to track timbre changes and analyze the current timbre state;

[0095] S2.1.4. Based on the user timbre distribution and the current timbre state, construct a mainstream timbre learning model.

[0096] In this embodiment S2.1.1, the user audio signal includes: a time-domain user audio signal and a frequency-domain user audio signal, which is the most original sound input of the system and is used to extract user audio features subsequently. The frequency-domain user audio signal is transformed from the time-domain user audio signal using the short-time Fourier transform technique; the user audio features include: user MFCC timbre features, user LPC cepstral coefficients, and user timbre temporal features; among them, the user MFCC timbre feature is the Mel-frequency cepstral coefficient, which is used for speech recognition and timbre analysis, and can reflect the main auditory features of the audio. In this system, the user MFCC timbre feature is mainly used to depict the timbre spectrum distribution of the user; the user LPC cepstral coefficient is used to analyze the excitation source and the vocal tract transfer function of the speech, and can help track the short-time dynamic characteristics of the speech. The user LPC cepstral coefficient, as an auxiliary feature, together with the MFCC timbre feature, constitutes a more complete timbre representation; the user timbre temporal feature is used to represent the time change pattern of the user's timbre; it is obtained by modeling the statistical features of the audio frame (including the changes in energy, pitch, zero-crossing rate, etc. over time). The user timbre temporal feature assists the VAE and subsequent temporal dynamic modeling; the three types of features, namely the user MFCC timbre feature, the user LPC cepstral coefficient, and the user timbre temporal feature, are combined into a user timbre feature vector and input into the variational autoencoder. In this embodiment S2.1.2, a user timbre feature vector is constructed based on the user audio features, and the variational autoencoder is used to reconstruct the user timbre distribution and optimize the variational autoencoder:

[0097] Input the user timbre feature vector into the encoder:

[0098] x = [M u , L u , T u ;

[0099]

[0100] where x is the user timbre feature vector; z is the user timbre latent variable; M u is the user MFCC timbre feature; L u is the user LPC cepstral coefficient; T u is the user timbre temporal feature; μ u is the timbre mean; is the timbre variance; q(z|x) is the variational posterior distribution; is the Gaussian distribution;

[0101] The user timbre feature reconstructed by the decoder:

[0102]

[0103] where p(x|z) is the reconstructed distribution of the timbre feature; f(z) is the mapping function of the decoder; is the variance of the timbre feature;

[0104] Calculate the reconstruction error to optimize the variational autoencoder:

[0105]

[0106] Among them, is the total loss function of the variational autoencoder; p(z) is the prior distribution; D KL (q(z|x)||p(z)) is the KL divergence between the variational posterior distribution q(z|x) and the prior distribution p(z); is the reconstruction error.

[0107] In this embodiment, in S2.1.3, based on the user's timbre distribution, use the time series dynamic modeling technology to track the timbre change and analyze the current timbre state. The specific method is as follows:

[0108] Use the Kalman filter to track the timbre change:

[0109] X = AX t-1 + w t ;

[0110] Among them, t is the time step; X t is the timbre state vector at time step t; A is the timbre change transition matrix, indicating the change relationship of the user's timbre from the previous time step to the current time step; X t-1 is the timbre state vector at time step t-1; w t is the Gaussian noise term, indicating the uncertainty of the timbre state change at time t, that is, the model error term;

[0111] Use Gated LSTM to analyze the current timbre state:

[0112] h t = σ(W h Xt + U h h t-1 + b h );

[0113] Among them, h t is the timbre state at the current time step t; σ(·) is the Sigmoid gating activation function, used to control the state update of the LSTM calculation; W h is the timbre input weight matrix; U h is the weight matrix of the previous time step state; h t-1 is the timbre state at time step t-1; b h is the Gated LSTM bias term;

[0114] In this embodiment, in S2.1.4, based on the user's tone color distribution and the current tone color state, a mainstream tone color learning model is constructed. The specific method is as follows:

[0115] p(x normal |z t ) = αp(x normal |z t-1 )+(1 - α)p(x t |z t );

[0116] α = σ(W α h t +b α );

[0117] Among them, p(x normal |z t ) is the user's normal tone color distribution at time step t; α is the dynamic smoothing factor; p(x normal |z t-1 ) is the user's normal tone color distribution at time step t - 1; p(x t |z t ) is the reconstructed distribution of the tone color features at time step t; W α is the Gated LSTM weight matrix; b α is the bias term.

[0118] In this embodiment, the current tone color state h t calculated by Gated LSTM is used as a dynamic adjustment factor to provide smooth control when updating the long-term tone color distribution of the mainstream tone color learning model; α of the mainstream tone color learning model is not fixed, but a dynamic weight parameter controlled by the current tone color state h t : When the system detects that the tone color change is small, such as the normal voice fluctuation of the user, the small change in the current tone color state h t makes α maintain a large value, and the mainstream tone color model mainly follows the past tone color distribution; when the system detects a large tone color change, such as the long-term tone color change trend, the large change in the current tone color state h t makes α smaller, and more relies on the current tone color distribution to iteratively update the mainstream tone color learning model.

[0119] In this embodiment, the secondary tone color monitoring module 22 constructs a main-secondary flow dynamic gating mechanism through a generative adversarial network combined with KL divergence detection technology, dynamically determines whether the user's real-time tone color deviates from the normal tone color distribution, and compensates for the abnormal tone color in real time when the user's tone color is abnormal. The specific method steps are as follows:

[0120] S2.2.1. Calculate the KL divergence of the user's tone color based on the user's tone color features reconstructed by the decoder and the mainstream tone color learning model:

[0121]

[0122] Among them, D KL (p(x|z)||p(x normal |z)) is the KL divergence of the user's voice timbre;

[0123] S2.2.2. Set the user voice timbre divergence threshold ∈, and determine whether the user voice timbre is abnormal:

[0124] If D KL (p(x|z)||p(x normal |z)) ≤ ∈, then the user voice timbre is normal, and continue to use the mainstream voice timbre learning model;

[0125] If D KL (p(x|z)||p(x normal |z)) > ∈, then trigger timbre compensation, and execute S2.2.3 to use the generative adversarial network for timbre compensation;

[0126] S2.2.3. Use the generator of the generative adversarial network to generate a corrected timbre, and use the discriminator to evaluate whether the corrected timbre is reasonable. If it is reasonable, input the corrected timbre into the mainstream voice timbre learning model.

[0127] In this embodiment, through the main and auxiliary flow dynamic gating mechanism, it is judged in real time whether the user's voice timbre deviates from the normal voice timbre distribution, and real-time compensation is performed when the voice timbre is abnormal:

[0128] Dynamically monitor the deviation of the user's voice timbre to prevent speech recognition failure or misrecognition caused by short-term voice timbre changes due to illness; when the voice timbre deviates from the normal range, trigger the timbre compensation mechanism to ensure the voice recognition system's ability to adapt to the user's personalized voice timbre and is not affected by short-term voice timbre changes; ensure that the mainstream voice timbre learning model can still operate stably during long-term voice timbre changes;

[0129] Using KL divergence detection can measure the information deviation between two probability distributions and is suitable for statistical modeling of voice timbre features; using the generative adversarial network for timbre compensation to learn the timbre mapping relationship can generate features closer to the normal timbre when the voice timbre is abnormal.

[0130] It also includes a noise learning and storage unit 3, and the noise learning and storage unit 3 is used to continuously learn the wind noise characteristics and KTV noise characteristics and store them in the noise feature library;

[0131] In this embodiment, the noise learning and storage unit 3 includes a noise feature learning module 31 and a noise feature storage module 32;

[0132] Among them, the noise feature learning module 31 is used to collect environmental noise features in real time and learn the wind noise characteristics and KTV background noise characteristics;

[0133] The noise feature storage module 32 is used to store and manage noise feature data, forming a noise feature library.

[0134] It further includes a natural voice recognition unit 4. Based on a mainstream timbre learning model and a noise feature library, the natural voice recognition unit 4 is used to screen user voice signals in a noisy environment and convert them into text language using end-to-end voice recognition technology;

[0135] In this embodiment, based on a mainstream timbre learning model and a noise feature library, the natural voice recognition unit 4 is used to screen user voice signals in a noisy environment and convert them into text language. The specific method steps are as follows:

[0136] S4.1. Obtain a new audio signal and determine whether the user audio signal in the new audio signal conforms to the user's own timbre. The specific method is as follows:

[0137] Obtain a new audio signal from the audio acquisition and processing unit 1, use the KL divergence detection technology to calculate the current timbre offset, and perform a conformity comparison with the user's own timbre in the mainstream timbre learning model:

[0138] If the current timbre offset is greater than the user timbre divergence threshold ∈, it does not conform to the user's own timbre and the new audio signal is rejected;

[0139] If the current timbre offset is less than or equal to the user timbre divergence threshold ∈, it conforms to the user's own timbre, and the new audio signal is input into the body timbre learning unit 2 for analysis and S4.2 is executed;

[0140] In this embodiment, the KL divergence detection technology is a detection method based on the Kullback-Leibler divergence to measure the difference between two probability distributions, used to determine whether there is a significant deviation between the current data distribution and the reference distribution. In this system, the KL divergence detection is used to monitor the distance between the user's current new timbre distribution and the user's normal timbre distribution;

[0141] The user timbre divergence threshold ∈ is the user timbre divergence threshold ∈ set in S2.2.2, and is also used to determine whether the new audio signal conforms to the user's own timbre;

[0142] S4.2. Based on the noise feature library, use the Wiener filtering technology to separate the noise signals existing in the noise feature library from the new audio signal, and only retain the pure user audio signal;

[0143] In this embodiment S4.2, the Wiener filtering technology is an optimal linear filtering technology used to estimate the optimal estimated value of the target signal under noise interference. It is widely used in fields such as speech enhancement and signal denoising. In this system, it is used to separate known types of noise from the new audio signal based on the noise feature library, and only retain the pure user audio signal; and to specifically remove the existing noise types based on the noise feature library.

[0144] S4.3. Based on the pure user audio signal, use the end-to-end speech recognition technology to convert the pure user audio signal into text language.

[0145] In this embodiment S4.3, the end-to-end speech recognition technology is a technology that directly maps from the speech waveform or spectral features to the text output, without relying on traditional staged modeling. Common end-to-end speech recognition technologies include: DeepSpeech based on CTC, the Seq2Seq model based on the attention mechanism, and the end-to-end ASR based on Transformer.

[0146] In S4.1, the user audio signal is input into the body timbre learning unit 2, which is used to trigger the main and auxiliary flow dynamic gating mechanism for timbre compensation when it is detected that the user's timbre deviates greatly from the existing distribution, and is also used for the mainstream timbre learning model to continuously and dynamically learn the user's normal timbre every time the user's voice is recognized.

[0147] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. The above embodiments and descriptions in the specification are only preferred examples of the present invention and do not limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.

Claims

1. A language processing system for speech recognition combining machine learning and deep learning, characterized in that: include: An audio acquisition processing unit (1), the audio acquisition processing unit (1) is used to acquire audio signals, perform preprocessing and extract audio features; The ontology timbre learning unit (2) continuously learns the normal timbre distribution of users and builds a mainstream timbre learning model based on a variational autoencoder combined with a time series dynamic modeling technology, and builds a main-side dynamic gating mechanism through a generative adversarial network combined with a KL divergence detection technology to dynamically determine whether the current timbre deviates from the normal timbre distribution; A noise learning storage unit (3), wherein the noise learning storage unit (3) is used to continuously learn wind noise characteristics and KTV noise characteristics, and store them in a noise characteristic library; The native voice recognition unit (4) is based on a mainstream timbre learning model and a noise feature library, and is used to screen user voice signals in a noisy environment and convert them into text language using end-to-end voice recognition technology.

2. The language processing system for speech recognition combining machine learning and deep learning according to claim 1, characterized in that: The audio acquisition and processing unit (1) comprises an audio acquisition module (11) and an audio processing module (12); The audio acquisition module (11) uses a sensor to acquire user audio signals and environmental audio signals; The audio processing module (12) is used to pre-process the user audio signal and the ambient audio signal, and to extract audio features of the user audio signal and the ambient audio signal.

3. The language processing system for speech recognition combining machine learning and deep learning according to claim 2, characterized in that: The main body timbre learning unit (2) comprises a main stream timbre learning module (21) and a secondary stream timbre monitoring module (22); The mainstream timbre learning module (21) constructs a mainstream timbre learning model based on a variational autoencoder combined with a temporal dynamic modeling technology, and is used to continuously learn the normal timbre distribution of users; The secondary stream timbre monitoring module (22) constructs a primary and secondary stream dynamic gating mechanism by generating an adversarial network combined with KL divergence detection technology, dynamically determines whether the user's real-time timbre deviates from the normal timbre distribution, and compensates for the abnormal timbre in real time when the user's timbre is abnormal.

4. The language processing system for speech recognition combining machine learning and deep learning according to claim 3, characterized in that: The mainstream timbre learning module (21) constructs a mainstream timbre learning model based on a variational autoencoder combined with a temporal dynamic modeling technology. The specific method steps are as follows: S2.1.1, extracting a user audio signal and a user audio feature from the audio acquisition processing unit (1); The user audio signal includes: a time domain user audio signal and a frequency domain user audio signal; The user audio features include: user MFCC timbre features, user LPC cepstral coefficients and user timbre time series features; S2.1.

2. Construct a user timbre feature vector based on the user audio features, reconstruct the user timbre distribution through a variational autoencoder, and optimize the variational autoencoder; S2.1.3, based on the user's timbre distribution, use the time series dynamic modeling technology to track the timbre changes and analyze the current timbre status; S2.1.

4. Build a mainstream timbre learning model based on the user's timbre distribution and current timbre status.

5. The language processing system for speech recognition combining machine learning and deep learning according to claim 4, characterized in that: In S2.1.2, a user timbre feature vector is constructed based on the user audio feature, and the user timbre distribution is reconstructed and the variational autoencoder is optimized by a variational autoencoder: Input the user's timbre feature vector into the encoder: ; ; in, is the user's timbre feature vector; is the user's timbre latent variable; It is the user's MFCC timbre feature; is the user's LPC cepstral coefficient; It is the user's timbre timing characteristics; is the mean value of timbre; is the timbre variance; is the variational posterior distribution; is a Gaussian distribution; User timbre features reconstructed by the decoder: ; in, is the reconstructed distribution of timbre features; is the mapping function of the decoder; is the timbre characteristic variance; Compute the reconstruction error to optimize the variational autoencoder: ; in, is the total loss function of the variational autoencoder; is the prior distribution; is the variational posterior distribution and the prior distribution KL divergence of is the reconstruction error.

6. The language processing system for speech recognition combining machine learning and deep learning according to claim 5, characterized in that: In S2.1.3, based on the user's timbre distribution, the time series dynamic modeling technology is used to track the timbre changes and analyze the current timbre status. The specific method is as follows: Use the Kalman filter to track timbre changes: ; in, is the time step; is the time step The timbre state vector of is the timbre change transfer matrix; is the time step The timbre state vector of is a Gaussian noise term, indicating the timbre state at time uncertainty of change, i.e., the model error term; Use Gated LSTM to analyze the current timbre status: ; in, is the current time step Timbre status; is the Sigmoid gate activation function; Enter the weight matrix for the timbre; is the state weight matrix of the previous time step; is the time step Timbre status; is the Gated LSTM bias item; In S2.1.4, based on the user's timbre distribution and the current timbre state, a mainstream timbre learning model is constructed, and the specific method is as follows: ; ; in, is the time step Normal timbre distribution of users; is the dynamic smoothing factor; is the time step Normal timbre distribution of users; is the time step Reconstructed distribution of timbre features; is the Gated LSTM weight matrix; is the bias term.

7. The language processing system for speech recognition combining machine learning and deep learning according to claim 6, characterized in that: The secondary stream timbre monitoring module (22) constructs a primary and secondary stream dynamic gating mechanism by generating an adversarial network combined with KL divergence detection technology, dynamically determines whether the user's real-time timbre deviates from the normal timbre distribution, and compensates for the abnormal timbre in real time when the user's timbre is abnormal. The specific method steps are as follows: S2.2.

1. Based on the user timbre features reconstructed by the decoder and the mainstream timbre learning model, the user timbre KL divergence is calculated: ; in, is the KL divergence of the user's timbre; S2.2.

2. Setting the user's timbre dispersion threshold , determine whether the user's tone is abnormal: like , the user's timbre is normal, and the mainstream timbre learning model continues to be used; like , then the timbre compensation is triggered, and S2.2.3 is executed to use the generative adversarial network to perform timbre compensation; S2.2.

3. Use the generator of the generative adversarial network to generate the corrected timbre, and use the discriminator to evaluate whether the corrected timbre is reasonable. If it is reasonable, the corrected timbre is input into the mainstream timbre learning model.

8. The language processing system for speech recognition combining machine learning and deep learning according to claim 7, characterized in that: The noise learning storage unit (3) comprises a noise feature learning module (31) and a noise feature storage module (32); The noise feature learning module (31) is used to collect environmental noise features in real time, and learn wind noise features and KTV background noise features; The noise feature storage module (32) is used to store and manage noise feature data to form a noise feature library.

9. The language processing system for speech recognition combining machine learning and deep learning according to claim 8, characterized in that: The native voice recognition unit (4) is based on a mainstream timbre learning model and a noise feature library, and is used to screen user voice signals in a noisy environment and convert them into text language using end-to-end voice recognition technology. The specific method steps are as follows: S4.

1. Obtain a new audio signal and determine whether the user audio signal in the new audio signal matches the user's own timbre. The specific method is as follows: The new audio signal is obtained from the audio acquisition processing unit (1), the KL divergence detection technology is used to calculate the current timbre offset, and the consistency is compared with the user's own timbre in the mainstream timbre learning model: If the current timbre deviation is greater than the user timbre dispersion threshold If it does not match the user's own tone, the new audio signal will be rejected; If the current timbre deviation is less than or equal to the user timbre dispersion threshold , then it matches the user's own timbre, and the new audio signal is input into the main body timbre learning unit (2) for analysis and execution of S4.2; S4.

2. Based on the noise feature library, the Wiener filtering technology is used to separate the noise signal in the noise feature library from the new audio signal, and only the pure user audio signal is retained; S4.

3. Based on the clean user audio signal, the clean user audio signal is converted into text language using end-to-end speech recognition technology.

10. The language processing system for speech recognition combining machine learning and deep learning according to claim 9, characterized in that: In the above S4.1, the user audio signal is input into the main body timbre learning unit (2), which is used to trigger the main and sub-stream dynamic gating mechanism to perform timbre compensation when it is detected that the user's timbre deviates greatly from the existing distribution, and is also used for the mainstream timbre learning model to continuously and dynamically learn the user's normal timbre each time the user's voice is recognized.

Citation Information

Patent Citations

  • Data updating method, client and electronic device

    CN109145145A

  • Voice data recognition method and device, chip and electronic equipment

    CN116312503A

  • Self-learning method and system for voiceprint recognition

    CN116597842A

  • Method and system capable of dynamically tracking and identifying long-term progressive change of personal timbre

    CN118762718A

  • Adaptive voice authentication system and method

    US20170140760A1