A method and system for streaming audio language identification

By extracting features through speech activity detection, Fourier transform, and Mel filter bank, an encoder-decoder model is constructed. Combined with window-level and frame-level discrimination methods, the accuracy and real-time performance issues of multilingual mixed speech recognition in streaming scenarios are solved, achieving efficient recognition of multilingual mixed speech.

CN119811383BActive Publication Date: 2025-11-25BEIJING YUNSHANG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411918988.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-11-25
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

In streaming scenarios, existing technologies struggle to achieve the accuracy and real-time performance of multilingual mixed speech recognition, making traditional sentence-level language recognition and monolingual ASR aggregation methods no longer applicable.

Method used

The audio data is preprocessed using a speech activity detection method, and features are extracted using Fourier transform and Mel filter bank. An encoder-decoder model is constructed for language recognition feature training, and language conversion points are determined by window-level and frame-level language discrimination and clustering methods.

Benefits of technology

It has achieved sentence-level, window-level, and frame-level language recognition and prediction functions, accurately located language conversion points, expanded the application scenarios of language recognition, and improved the accuracy and real-time performance of multilingual mixed speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811383B_ABST
    Figure CN119811383B_ABST
Patent Text Reader

Abstract

The application discloses a streaming audio language recognition method and system, and belongs to the technical field of language recognition. The method is implemented as follows: 1. original audio data is preprocessed by using a voice activity detection method to obtain language recognition training data; 2. feature extraction is performed on the language recognition training data; 3. an encoder-decoder model is constructed and language recognition feature training is performed; 4. language recognition test data is input into the trained encoder-decoder model to obtain language recognition audio data, and the language recognition audio data is formed into an audio data stream in a data accumulation manner; 5. audio data stream is detected by using a voice activity detection method; 6. window-level language discrimination is performed on the audio data detected by the activity detection; specifically, the audio data of the current window and the previous window are compared to obtain the timestamp and the language result of the current state; compared with the prior art, the application realizes multi-language mixed voice recognition in a streaming scenario.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a streaming audio language recognition method and system, and belongs to the technical field of language recognition. BACKGROUND

[0002] In recent years, cross-language communication is becoming more and more frequent. In order to break the communication barriers among countries in the world, speech translation technology is attracting more and more attention. Among them, the combination of speech recognition and machine translation is a mainstream scheme of current speech translation technology. In this scheme, the accuracy of speech recognition plays a key role in the overall performance. Mixed language ASR has the ability to recognize multiple languages, but compared with single language ASR, its performance will be greatly lost. Therefore, the combination of language recognition and single language ASR to realize multi-language ASR has become a commonly used scheme in the industry.

[0003] In the speech translation task, in addition to the accuracy of the translation result, the real-time performance of the returned result is also an important indicator to measure the performance of the system. Therefore, the traditional sentence-level language recognition and single language ASR combination mode is no longer applicable.

[0004] Therefore, how to realize multi-language mixed speech recognition in a streaming scenario has become a problem to be solved. SUMMARY

[0005] In view of the technical problem of realizing multi-language mixed speech recognition in a streaming scenario, the purpose of the application is to provide a streaming audio language recognition method and system, so as to combine language recognition and single language ASR.

[0006] The purpose of the application is realized by the following technical scheme:

[0007] The application discloses a streaming audio language recognition method, which comprises the following steps:

[0008] Step 1: using a voice activity detection method to pre-process the original audio data to obtain language recognition training data;

[0009] Step 2: performing feature extraction on the language recognition training data, and further using Fourier transform and a mel filter bank to obtain language recognition features;

[0010] Step 2.1: converting time domain features into frequency domain features by using a Fourier transform method on the language recognition training data;

[0011] Step 2.2: converting the frequency domain features into language recognition features by filtering through a mel filter bank;

[0012] Step 3: constructing an encoder-decoder model and performing language recognition feature training;

[0013] Step 3.1: constructing an encoder-decoder model for language recognition feature training;

[0014] Step 3.2: training language recognition features using the encoder-decoder model;

[0015] Step 3.2.1: in the encoder-decoder model, the encoder parameters are initialized using speech recognition encoder parameters; the decoder parameters are initialized with random variables;

[0016] Step 3.2.2: training the encoder-decoder model in a cyclic iteration manner;

[0017] Step 4: inputting language recognition test data into the trained encoder-decoder model to obtain language recognition audio data, and forming an audio data stream in a data accumulation manner;

[0018] Step 5: using a voice activity detection method to detect the activity of the audio data stream;

[0019] Step 6: performing window-level language discrimination on the audio data detected by the activity;

[0020] Step 6.1: when the current window is the starting window, returning the timestamp and language result of the current state;

[0021] Step 6.2: when the previous window is a silent window, returning the timestamp and language result of the current state;

[0022] Step 6.3: when the previous window is not a silent window and the current window is not a starting window, comparing the current language result with the previous time language result to obtain the timestamp and language result of the current state;

[0023] Step 6.3.1: when the current language result is the same as the previous time language result, returning the timestamp and language result of the current state;

[0024] Step 6.3.2: when the current language result is different from the previous time language result, performing frame-level detection to obtain the timestamp and language result of the current state;

[0025] Step 6.3.2.1: splicing the audio data of the previous window and the current window in a time sequence manner to form spliced audio data;

[0026] Step 6.3.2.2: using a sliding window to perform frame-level language prediction on the spliced audio data, and obtaining a corresponding language conversion point prediction result through clustering, and correcting the language results of the previous window and the current window;

[0027] The application discloses a streaming audio language recognition system for realizing the method.

[0028] The VAD module is used for voice activity detection on speech in audio data, and further screening of training data for language recognition; and the VAD module is used as input of the language recognition module.

[0029] The language recognition module is used for window-level language discrimination, comparison between audio data of a current window and previous window, and further obtaining of a timestamp and a language result of a current state; and the language recognition module is used as input of the language conversion point detection module.

[0030] The language conversion point detection module is used for obtaining a timestamp and a language result of a current state by means of a modified language conversion point when the language of the current window is different from that of the previous window. Advantages

[0031] Compared with the prior art, the application has the following advantages:

[0032] 1. The language recognition model is trained, and sentence-level, window-level and frame-level language recognition and prediction functions can be realized simultaneously.

[0033] 2. After predicting the frame-level language result, the kmeans clustering scheme is used to accurately locate the language conversion point position, so that the streaming language recognition and prediction purpose is realized.

[0034] 3. On the basis of the conventional language recognition, the streaming judgment scheme is creatively proposed, and the application scene of language recognition is expanded. DETAILED DESCRIPTION

[0035] Figure 1 FIG. 1 is a flowchart of a streaming audio language recognition system according to an embodiment of the application;

[0036] Figure 2 FIG. 3 is a structural diagram of a language recognition conversion point detection module according to an embodiment of the application;

[0037] Figure 3 FIG. 4 is a structural diagram of a language recognition decoder according to an embodiment of the application. DETAILED DESCRIPTION

[0038] In order to better illustrate the purposes and advantages of the application, the following further describes the application content in combination with the drawings and examples. It should be noted that the implementation of the application is not limited to the following examples, and any form of modification or change of the application will fall within the protection scope of the application. EMBODIMENT

[0039] As Figure 1 shown, the specific implementation steps of the streaming audio language recognition method of the embodiment are as follows:

[0040] Step 1: Preprocess the original audio data by using a voice activity detection method to obtain language recognition training data;

[0041] In the embodiment, the original audio data is processed by the VAD (voice activity detection) method to remove the silent segments, so as to ensure that the human voice of each training data is long enough, such as greater than 0.5s;

[0042] Step 2: Feature extraction is performed on the language recognition training data, and further Fourier transform and mel filter bank are used to obtain language recognition features;

[0043] Step 2.1: The time domain features are converted into frequency domain features by using the Fourier transform method on the language recognition training data;

[0044] Step 2.2: The frequency domain features are converted into language recognition features by filtering through the mel filter bank;

[0045] In the embodiment, the language recognition training data is pre-emphasized, framed, and windowed, the frame length is set to 25ms, the frame shift is 10ms, the window type is selected as the Hamming window; the time domain features are converted into frequency domain features by using the FFT (Fourier transform), and then the energy spectrum is sent into the mel filter bank to obtain 80-dimensional frequency domain features, and finally the log value is taken to obtain the fbank features. By using this method, a feature representation with lower dimension and more compactness can be obtained, so as to facilitate the training and inference of the model.

[0046] Step 3: Construct an encoder-decoder model and train the language recognition features;

[0047] Step 3.1: Construct an encoder-decoder model for language recognition feature training;

[0048] In the embodiment, a language recognition model with an Encoder-Decoder architecture is built, wherein the Encoder parameters are migrated from the ASR task, so the Encoder structure needs to be consistent with the trained ASR model; the Decoder adopts a 7-layer Tdnn model to realize the language classification and frame-level language information prediction tasks.

[0049] Step 3.2: Train the language recognition features by using the encoder-decoder model;

[0050] Step 3.2.1: In the encoder-decoder model, the encoder parameters are initialized by using the speech recognition encoder parameters; the decoder parameters are initialized by using random variables;

[0051] Step 3.2.2: training the encoder-decoder model in a cyclic iteration manner;

[0052] In the embodiment, during training, only the Decoder parameters are trained with the Encoder parameters fixed, and after the model converges, the full parameters are trained. Specifically, the Encoder parameters of the encoder-decoder model are initialized with the trained ASR Encoder parameters, and the Decoder parameters are randomly initialized. The training is divided into two parts. First, the Decoder parameters are trained with an initial learning rate of 0.002 for 20 iterations, and then the full parameters are trained with an initial learning rate of 0.0001. The total number of iterations is about 30.

[0053] Step 4: input the language recognition test data into the trained encoder-decoder model to obtain language recognition audio data, and form an audio data stream in a data accumulation manner;

[0054] In the embodiment, the data receiving module buffers the test data stream, and the data is accumulated to 1s in length before being sent to the downstream for judgment. For scenarios with higher real-time requirements, the window length can be reduced. Similarly, for scenarios with higher accuracy requirements, the window length can be increased. In order to balance the real-time performance and accuracy of the system, the window length of the present application is set to 1s.

[0055] Step 5: performing activity detection on the audio data stream using a voice activity detection method;

[0056] In the embodiment, VAD is used to judge the human voice. Specifically, the proportion of human voice in 1s audio is judged, and when the human voice is more than 0.5s, it is considered valid and sent to the downstream language recognition module for judgment. Otherwise, it is considered to be silent.

[0057] Step 6: performing window-level language discrimination on the audio data that passes the activity detection;

[0058] Step 6.1: when the current window is the starting window, returning the timestamp and language result of the current state;

[0059] Step 6.2: when the previous window is a silent window, returning the timestamp and language result of the current state;

[0060] In the embodiment, the current window language recognition judgment specifically includes: the audio data with sufficient human voice is sent to the same feature extractor as in the training stage to extract 80-dimensional fbank features, and then sent to the language recognition model to obtain the window-level language discrimination result. Specifically, the language result of the current state is judged. If the current window is the starting window or the previous window is a silent window, the current timestamp and language discrimination result are directly returned.

[0061] Step 6.3: When the previous window is not a mute window and the current window is not a start window, the current language result is compared with the previous time language result, and then the timestamp and language result of the current state are obtained;

[0062] Step 6.3.1: When the current language result is the same as the previous time language result, the timestamp and language result of the current state are returned;

[0063] In the embodiment, when the previous window is not a mute window or the current window is not a start window, the current language result is compared with the previous time language result, and when the results are the same, the current timestamp and language discrimination result are directly returned.

[0064] Step 6.3.2: When the current language result is different from the previous time language result, frame-level detection is adopted to obtain the timestamp and language result of the current state;

[0065] Step 6.3.2.1: The audio data of the previous window and the current window are spliced in a time sequence to form spliced audio data;

[0066] Step 6.3.2.2: The spliced audio data is predicted by using a sliding window at a frame level, and a corresponding language conversion point prediction result is obtained by clustering, and the language results of the previous window and the current window are corrected;

[0067] In the embodiment, as shown in Figure 2 and Figure 3 , specifically, first, the previous window audio and the current window audio are spliced in the time dimension, second, the input audio is extracted for features and then sent to a language recognition model to calculate frame-level language discrimination information, the language recognition model adopts an Encoder-Decoder architecture, wherein the core of the Encoder module is a Confomer component; the first 5 layers of the Decoder module are Tdnn components for obtaining frame-level language discrimination information; then, the frame-level features are processed by sliding window and taking the mean value in the time dimension, considering that language discrimination needs to include semantic information, and the frame-level features include too little semantic information, so the frame-level features are processed by sliding window and taking the mean value in the time dimension, the window length is set to 1s and the window shift is 0.2s, so that 6 pieces of compressed frame-level language recognition information are obtained; finally, the language conversion point prediction result is obtained by kmeans clustering, that is, the 6 pieces of language recognition information are clustered by kmeans with a class specified as 2, and the class conversion point position in the clustering result is the corresponding language conversion point prediction result; the language results of the previous window and the current window are corrected, the previous time prediction language corresponds to the language conversion point from the start point of the previous window to the language conversion point, and the current time prediction language corresponds to the language from the language conversion point to the end point of the current window.

[0068] The embodiment discloses a streaming audio language recognition system for realizing the method.

[0069] The VAD module is used for voice activity detection on the voice in the audio data, and further screening out the training data for language recognition; and the VAD module is used as the input of the language recognition module.

[0070] The language recognition module is used for language discrimination at a window level, and the timestamp and the language result of the current state are obtained by comparing the audio data of the current window with the previous window; and the language recognition module is used as the input of the language conversion point detection module.

[0071] The language conversion point detection module is used for obtaining the timestamp and the language result of the current state by using a modified language conversion point when the language of the current window is different from that of the previous window.

[0072] Results show that the application realizes language recognition in a streaming scenario, and multi-language mixed voice recognition can be realized by combining with a single-language ASR set.

Claims

1. A method for identifying the language of streaming audio, characterized in that: Includes the following steps, Step 1: Preprocess the raw audio data using the speech activity detection method to obtain language recognition training data; Step 2: Extract features from the language recognition training data, and further use Fourier transform and Mel filter bank to obtain language recognition features; Step 3: Construct the encoder-decoder model and train it using language recognition features; Step 4: Input the language recognition test data into the trained encoder-decoder model to obtain language recognition audio data, and form an audio data stream by accumulating the language recognition audio data; Step 5: Perform activity detection on the audio data stream using a speech activity detection method; Step 6: Perform window-level language identification on the audio data that has passed the activity detection; Step 6.1: When the current window is the starting window, return the timestamp and language result of the current state; Step 6.2: If the previous window is a mute window, return the timestamp and language result of the current state; Step 6.3: When the previous window is not a silent window and the current window is not a start window, compare the current language result with the language result of the previous moment to obtain the timestamp and language result of the current state.

2. The streaming audio language recognition method as described in claim 1, characterized in that: Step 2 is implemented as follows: Step 2.1: Use the Fourier transform method to convert the time-domain features into frequency-domain features for the language recognition training data; Step 2.2: Convert the frequency domain features into language recognition features by filtering with a Mel filter bank.

3. The streaming audio language recognition method as described in claim 1, characterized in that: Step 3 is implemented as follows: Step 3.1: Construct an encoder-decoder model for language identification feature training; Step 3.2: Train the language recognition features using an encoder-decoder model.

4. The streaming audio language recognition method as described in claim 3, characterized in that: Step 3.2 is implemented as follows: Step 3.2.1: In the encoder-decoder model, the encoder parameters are initialized using the speech recognition encoder parameters; the decoder parameters are initialized using random variables. Step 3.2.2: Train the encoder-decoder model in a loop-iterative manner.

5. The streaming audio language recognition method as described in claim 1, characterized in that: Step 6.3 is implemented as follows: Step 6.3.1: If the current language result is the same as the language result of the previous time step, return the timestamp and language result of the current state; Step 6.3.2: When the current language result is different from the language result of the previous time step, frame-level detection is performed to obtain the timestamp and language result of the current state.

6. The streaming audio language recognition method as described in claim 5, characterized in that: The implementation method for step 6.3.2 is as follows: Step 6.3.2.1: Concatenate the audio data from the previous window and the current window in a time series manner to form concatenated audio data; Step 6.3.2.2: Use a sliding window to perform frame-level language prediction on the spliced ​​audio data, obtain the corresponding language conversion point prediction results through clustering, and correct the language results of the previous window and the current window.

7. A streaming audio language recognition system implementing the method as described in claim 1, characterized in that: It includes a VAD module, a language recognition module, and a language conversion point detection module; The VAD module is used to perform speech activity detection on the speech in the audio data, and then filter out the training data for language recognition; it will be used as the input of the language recognition module. The language identification module is used for window-level language discrimination. It compares the audio data of the current window with that of the previous window to obtain the timestamp and language result of the current state. This will be used as input for the language conversion point detection module; The language conversion point detection module is used to correct the language conversion point when the language of the current window is different from that of the previous window, and obtain the timestamp and language result of the current state.

Citation Information

Patent Citations

  • Language identification method and system

    CN112530407A

  • End-to-end multi-language continuous speech stream speech content identification method and system

    CN113077785A