Multilingual subtitle generation method based on Whisper and FunASR dual-path speech recognition large model
By introducing Whisper and FunASR dual-channel speech recognition models into the audio recognition system, combining technical means of sliding windows and pinyin similarity, the problem of insufficient recognition accuracy and robustness in the existing technology is solved, and more efficient long audio processing and multilingual subtitle generation are achieved.
Patent Information
- Application Number
- CN202510074500.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-17
AI Technical Summary
The existing audio recognition automatic subtitle generation scheme has problems such as limited recognition accuracy, insufficient robustness, poor long audio processing and limited scalability.
A multilingual subtitle generation method based on Whisper and FunASR dual-channel speech recognition large model is adopted. Audio clips are divided through sliding windows, and the two output results are compared using pinyin similarity to perform splicing and verification of recognition results. Finally, the speech recognition results are input into the translation model to generate multilingual subtitles.
Improves recognition accuracy and robustness, improves the effect and scalability of long audio processing, and ensures reliability and multilingual support for subtitle generation.
Smart Images

Figure CN119517038B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech recognition technology, and in particular to a multilingual subtitle generation method based on a Whisper and FunASR dual-path speech recognition large model. Background Art
[0002] Existing audio recognition and automatic subtitle generation solutions mainly use the following technical means to achieve audio content analysis and text conversion:
[0003] Speech recognition (ASR, Automatic Speech Recognition) technology: Use deep learning models (such as RNN, Transformer, CTC, etc.) to process audio signals and convert speech content into text. Common frameworks include Google Speech-to-Text, DeepSpeech, Wav2Vec 2.0, etc.
[0004] Audio signal processing: Preprocess the audio signal through methods such as short-time Fourier transform (STFT) to extract speech features such as Mel-frequency cepstral coefficients (MFCC) or spectrogram.
[0005] Language model: Combined with the context, the recognition results are optimized through language models (such as GPT and BERT) to improve semantic understanding and text accuracy.
[0006] Timestamp alignment: Through dynamic time warping (DTW) or other time alignment algorithms, the recognized text is accurately aligned with the audio content to generate accurate time-annotated subtitle files.
[0007] Development languages and tools: The implementation of audio recognition technology is usually based on development languages such as Python and C++. Common frameworks include TensorFlow, PyTorch, and open source tools such as Kaldi and OpenAI Whisper.
[0008] However, the above technical solutions have the disadvantages of limited recognition accuracy, insufficient robustness, poor long audio processing effect and limited scalability. Summary of the invention
[0009] In order to help solve the above technical problems, this application provides a multilingual subtitle generation method based on Whisper and FunASR dual-path speech recognition large model, using the following technical solutions:
[0010] A multilingual subtitle generation method based on Whisper and FunASR dual-path speech recognition large model, wherein the method comprises:
[0011] Step S1: inputting the original audio segmentation fragments into the two-way speech recognition large model through a sliding window method;
[0012] Step S2: Based on the pinyin similarity, the first recognition result and the existing result are obtained by comparing the two-way output results of the two-way speech recognition large model;
[0013] Step S3: based on the pinyin similarity, concatenating the first recognition result and the existing result into a second recognition result;
[0014] Step S4: setting a next sliding window at the end position of the second recognition result, and continuing to recognize the next audio segment until all audio segments are recognized to obtain a speech recognition result;
[0015] Step S5: input the speech recognition result into a translation model to generate multilingual subtitle content.
[0016] Preferably, the step S1 comprises:
[0017] Step S11: resample the audio stream to a sampling rate of 16000 Hz and configure a custom hot word list;
[0018] Step S12: determining segments containing speech content through the two-way speech recognition large model, and recording the start timestamps and end timestamps of all the segments;
[0019] Step S13: taking out a segment to be processed from the head of the audio to be processed according to the specified time length, judging whether the segment to be processed contains speech content according to the start timestamp and the end timestamp, if so, inputting the segment to be processed into the two-way speech recognition large model, if not, skipping the segment and selecting the next segment as the segment to be processed;
[0020] Step S14: inputting the segment to be processed into the two-way speech recognition large model.
[0021] Preferably, step S2 comprises:
[0022] Compare the pinyin similarity of the two-way output results of the two-way speech recognition model to determine whether the output result is valid. The pinyin similarity S p The calculation formula is as follows:
[0023] ;
[0024] in, is the modified distance between the two output results based on the pinyin consistency judgment, The number of words in Whisper's recognition results. The number of words recognized by FunASR;
[0025] Preferably, the step S2 further comprises:
[0026] If the modified distance between the two output results based on the pinyin consistency judgment is less than or equal to 1, the two output results are considered to be consistent and are used as the first recognition result and the existing result respectively;
[0027] If the pinyin similarity is greater than a preset comparison threshold, the two output results are considered to be substantially consistent, and the Whisper recognition result is selected as the first recognition result, and the FunASR recognition result is selected as the existing result;
[0028] If the pinyin similarity is less than a preset comparison threshold, the two output results are considered inconsistent, and the FunASR recognition result is selected as the first recognition result, and the Whisper recognition result is selected as the existing result.
[0029] Preferably, step S3 comprises:
[0030] The head part of the first recognition result is compared with the tail part of the existing result based on pinyin similarity, and the part with the maximum length that meets the preset similarity threshold is found. The maximum length part is removed from the first recognition result, and the maximum length part is spliced to the tail of the existing result to obtain the second recognition result.
[0031] Preferably, the maximum length is calculated as follows:
[0032] ;
[0033] in, The maximum length that satisfies the preset similarity threshold, The tail of the existing results Length and header of the new recognition result length of phonetic similarity, is the preset similarity threshold.
[0034] Preferably, step S4 comprises:
[0035] The start timestamp of the second-to-last sentence at the end of the second recognition result is used as the start timestamp of the next sliding window segment.
[0036] In summary, this application has the following advantages:
[0037] 1. Improve recognition accuracy: Use advanced large models to improve the recognition accuracy of the large models themselves;
[0038] 2. Improve robustness: Use dual-channel large models for simultaneous recognition and two-way verification to determine whether the large model results are valid, avoiding the situation where a single-channel large model outputs unusable results. At the same time, each model can also play its advantages and maximize the overall effect;
[0039] 3. Improve the usability of long audio results: Sliding windows and splicing based on pinyin similarity and sentence by sentence have better segmentation and splicing accuracy and stronger robustness compared to sliding windows and splicing at the token level directly performed on long segments within the large model;
[0040] 4. Improve scalability: Develop a scalable multi-channel large-model audio recognition solution, which can subsequently expand the number of large models as needed to further improve overall performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 A schematic block diagram of an embodiment of a multilingual subtitle generation method based on Whisper and FunASR dual-path speech recognition large model of the present application;
[0042] Figure 2 The present invention is a flowchart of an embodiment of a method for generating multilingual subtitles based on a large dual-path speech recognition model of Whisper and FunASR. DETAILED DESCRIPTION
[0043] The present application is further described below in conjunction with the accompanying drawings, and the structure and principle of the present application are very clear to people in the field. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0044] Figure 1 This is a schematic block diagram of an embodiment of a multilingual subtitle generation method based on Whisper and FunASR dual-path speech recognition large model of the present application, Figure 2 The present invention is a flowchart of an embodiment of a method for generating multilingual subtitles based on a large dual-path speech recognition model of Whisper and FunASR.
[0045] Combination Figure 1 and Figure 2 It is understood that the method of the present application may include:
[0046] Step S1: Input the original audio segment into the two-way speech recognition model through a sliding window method. In step S1, it can include: step S11: resample the audio stream to a sampling rate of 16000Hz, and configure a custom hot word list; step S12: determine the segment containing speech content through the two-way speech recognition model, and record the start timestamp and end timestamp of all segments; step S13: take out the segment to be processed from the head of the audio to be processed according to the specified time length, and determine whether the segment to be processed contains speech content according to the start timestamp and the end timestamp. If so, input the segment to be processed into the two-way speech recognition model, if not, skip the segment and select the next segment as the segment to be processed; step S14: input the segment to be processed into the two-way speech recognition model.
[0047] Specifically, step S1 may include audio preprocessing, configuring hot words, judging whether there are content segments, and sliding window input to the big model. The audio stream is resampled to a sampling rate of 16000 Hz, and the required custom hot word list is configured. The voice content detection big model is used to determine whether there are audio content segments, and the start and end timestamps of all segments containing voice content are taken. When the audio is subsequently processed into the big model, if an audio segment has no speaking content, the segment is skipped. Then, according to the specified time length, the segment is taken from the head of the audio to be processed (the segment also needs to be compared with the content segment obtained above to determine whether to skip it), and input into the big models of Whisper and FunASR.
[0048] Step S2: Based on the pinyin similarity, the first recognition result and the existing result are obtained by comparing the two-way output results of the two-way speech recognition large model. In step S2, the pinyin similarity of the two-way output results of the two-way speech recognition large model is compared to determine whether the output result is valid. The pinyin similarity S p The calculation formula is as follows:
[0049] ;
[0050] in, is the modified distance between the two output results based on the pinyin consistency judgment, The number of words in Whisper's recognition results. is the number of characters in the FunASR recognition result. If the modification distance between the two output results is less than or equal to 1, the two output results are considered to be consistent and are respectively used as the first recognition result and the existing result; if the pinyin similarity is greater than the preset comparison threshold, the two output results are considered to be basically consistent, and the Whispe recognition result is selected as the first recognition result, and the FunASR recognition result is selected as the existing result; if the pinyin similarity is less than the preset comparison threshold, the two output results are considered to be inconsistent, and the FunASR recognition result is selected as the first recognition result, and the Whispe recognition result is selected as the existing result.
[0051] Specifically, when judging whether the pinyin is consistent, considering that speech recognition itself has a certain degree of error rate, when comparing the pinyin of the characters of the two recognition results, if the modification distance of the two pinyins is less than or equal to 1, the two pinyins are considered to be consistent, otherwise they are judged to be inconsistent. If the ratio of the pinyin modification distance and the pinyin length of the two segments is greater than the set comparison threshold, it is considered that the output results of the two large models are basically consistent, and the Whisper large model result is selected; if the ratio is less than the set comparison threshold, it is considered that the two large model results are inconsistent, and the FunASR large model result is selected. This selection method is determined based on the characteristics of the two large models verified by experiments: FunASR is more robust and Whisper is more accurate.
[0052] Step S3: Based on the pinyin similarity, the first recognition result and the existing result are spliced into the second recognition result. In step S3, the head part of the first recognition result is compared with the tail part of the existing result based on the pinyin similarity, and the maximum length part that meets the preset similarity threshold is found. The maximum length part is removed from the first recognition result, and the maximum length part is spliced to the tail of the existing result to obtain the second recognition result. The maximum length is calculated as follows:
[0053] ;
[0054] in, To meet the maximum length of the preset similarity threshold, The tail of the existing results Length and header of the new recognition result length of phonetic similarity, is the preset similarity threshold.
[0055] Step S4: Set the next sliding window at the end position of the second recognition result, and continue to recognize the next audio segment until all audio segments are recognized to obtain the speech recognition result. In step S4, the start timestamp of the second-to-last sentence at the end of the second recognition result is used as the start timestamp of the next sliding window segment.
[0056] Step S5: Input the speech recognition result into the translation model to generate multilingual subtitle content.
Claims
1. A multilingual subtitle generation method based on Whisper and FunASR dual-path speech recognition large model, characterized in that: The method comprises: Step S1: inputting the original audio segmentation fragments into the two-way speech recognition large model through a sliding window method; Step S2: Based on the pinyin similarity, a first recognition result and another recognition result are obtained by comparing the two-way output results of the two-way speech recognition large model; Step S3: based on the pinyin similarity, concatenating the first recognition result and the other recognition result into a second recognition result; Step S4: setting a next sliding window at the end position of the second recognition result, and continuing to recognize the next audio segment until all audio segments are recognized to obtain a speech recognition result; Step S5: inputting the speech recognition result into a translation model to generate multilingual subtitle content; The step S2 comprises: The pinyin similarity of the two-way output results of the two-way speech recognition large model is compared to determine whether the output results are valid. The pinyin similarity Sp calculation formula is as follows: ; in, is the modified distance between the two output results based on the pinyin consistency judgment, The number of words in Whisper's recognition results. The number of words recognized by FunASR.
2. The multilingual subtitle generation method based on Whisper and FunASR two-way speech recognition large model according to claim 1, characterized in that: The step S1 comprises: Step S11: resample the audio stream to a sampling rate of 16000 Hz and configure a custom hot word list; Step S12: determining segments containing speech content through the two-way speech recognition large model, and recording the start timestamps and end timestamps of all the segments; Step S13: taking out a segment to be processed from the head of the audio to be processed according to the specified time length, judging whether the segment to be processed contains speech content according to the start timestamp and the end timestamp, if so, inputting the segment to be processed into the two-way speech recognition large model, if not, skipping the segment and selecting the next segment as the segment to be processed; Step S14: inputting the segment to be processed into the two-way speech recognition large model.
3. The multilingual subtitle generation method based on Whisper and FunASR two-way speech recognition large model according to claim 1, characterized in that: The step S2 further comprises: If the modified distance between the two output results based on the pinyin consistency judgment is less than or equal to 1, the two output results are considered to be consistent and are used as the first recognition result and the other recognition result respectively; If the pinyin similarity is greater than a preset comparison threshold, the two output results are considered to be substantially consistent, and the Whisper recognition result is selected as the first recognition result, and the FunASR recognition result is selected as the other recognition result; If the pinyin similarity is less than a preset comparison threshold, the two output results are considered inconsistent, and the FunASR recognition result is selected as the first recognition result, and the Whisper recognition result is selected as the other recognition result.
4. The multilingual subtitle generation method based on Whisper and FunASR two-way speech recognition large model according to claim 1, characterized in that: The step S3 comprises: The head part of the first recognition result is compared with the tail part of the other recognition result based on pinyin similarity, and the part with the maximum length that meets the preset similarity threshold is found. The part is removed from the first recognition result, and the part is spliced to the tail of the other recognition result to obtain the second recognition result.
5. The multilingual subtitle generation method based on Whisper and FunASR two-way speech recognition large model according to claim 4, characterized in that: The maximum length is calculated as follows: ; in, The maximum length that satisfies the preset similarity threshold, The tail of the other recognition result Length and header of the new recognition result length of phonetic similarity, is the preset similarity threshold.
6. The multilingual subtitle generation method based on Whisper and FunASR two-way speech recognition large model according to claim 1, characterized in that: The step S4 comprises: The start timestamp of the second-to-last sentence at the end of the second recognition result is used as the start timestamp of the next sliding window segment.
Citation Information
Patent Citations
Voice recognition method capable of realizing multi-language mixed use
CN105096953A
Voice command word recognition method and device, storage medium and electronic equipment
CN113971953A