A live caption generation method and related device
Through the combination of streaming speech recognition and offline speech recognition technology, the live audio stream data is segmented and corrected by using short-term audio energy and voice activity detection, which solves the problem of difficult to balance real-time and accuracy in live subtitles generation, and realizes real-time and high-precision subtitles generation.
Patent Information
- Application Number
- CN202510357222.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The existing live subtitle generation technology is difficult to balance in real-time and accuracy, and the asynchronous recognition and correction method can easily lead to delay accumulation, affecting the real-time nature of live subtitles.
Streaming voice recognition technology is used to initially identify live audio stream data, and the audio stream data is divided by combining audio short-time energy and voice activity detection confidence, multiple audio blocks are obtained, and each block is identified through offline voice recognition technology, and the timestamp alignment is used to correct it to realize the coordinated output of stream recognition and offline high-precision correction.
It achieves the balance between real-time and accuracy of live subtitles, ensures the consistency between subtitles text and audio content, and improves the real-time response and accuracy of subtitles.
Smart Images

Figure CN119865669B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of live broadcast technology, and more particularly, to a method for generating live subtitles and related devices. Background Art
[0002] With the rapid development of the network and media platforms, video live broadcast has become the main carrier of information transmission, and its influence is increasing day by day. The existing live streaming subtitle technology is not yet mature, and the accuracy of real-time generated live subtitles is not high. Although the asynchronous recognition and correction method can correct the generated live subtitle text asynchronously, it is prone to delay accumulation and affects the real-time performance of live subtitles.
[0003] Therefore, how to provide a method for generating live subtitles to achieve a balance between real-time performance and accuracy has become a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0004] In view of this, the present invention discloses a method for generating live subtitles and related devices to realize the collaborative output of real-time response of streaming speech recognition and offline high-precision correction of speech recognition results, so that the live subtitles can achieve a balance between real-time performance and accuracy.
[0005] A method for generating live subtitles includes:
[0006] Obtaining live audio stream data;
[0007] Performing streaming speech recognition on the live audio stream data to obtain a streaming subtitle text;
[0008] Segmenting the live audio stream data into multiple audio chunks based on audio short-time energy and speech activity detection confidence, and performing offline speech recognition on each audio chunk to obtain an offline subtitle text;
[0009] Aligning the time stamps of the streaming subtitle text and the offline subtitle text to obtain a target streaming subtitle text and a target offline subtitle text;
[0010] Correcting the target streaming subtitle text with the target offline subtitle text to obtain the final live subtitle.
[0011] Optionally, segmenting the live audio stream data into multiple audio chunks based on audio short-time energy and speech activity detection confidence, and performing offline speech recognition on each audio chunk to obtain an offline subtitle text, includes:
[0012] Segment the live audio stream data into multiple audio chunks based on the short-time energy of the audio and the confidence of the voice activity detection. Each audio chunk is either a first audio chunk representing a voice-active region or a second audio chunk obtained by merging multiple consecutive voice-inactive regions.
[0013] Adopt a cache management strategy of the Least Recently Used - K times algorithm to discard K consecutive adjacent second audio chunks among all the second audio chunks, obtaining an audio chunk set composed of the remaining second audio chunks and all the first audio chunks.
[0014] Perform offline speech recognition on each audio chunk in the audio chunk set to obtain the offline subtitle text.
[0015] Optionally, performing offline speech recognition on each audio chunk in the audio chunk set to obtain the offline subtitle text includes:
[0016] Use a voice activity detection processing method for each audio chunk in the audio chunk set to obtain the silence probability. Among them, the audio chunk with a high silence probability is a potential voice tail point.
[0017] Determine the cosine similarity between the Mel Frequency Cepstral Coefficient (MFCC) features of each audio chunk in the audio chunk set. Among them, the audio chunk with a low cosine similarity is a potential voice tail point.
[0018] Perform fundamental frequency continuity analysis on each audio chunk in the audio chunk set to determine whether there is a fundamental frequency mutation. Among them, the audio chunk with a fundamental frequency mutation is a potential voice tail point.
[0019] Determine the multi-modal feature comprehensive score based on the silence probability, the cosine similarity, and whether there is a fundamental frequency mutation corresponding to each audio chunk in the audio chunk set.
[0020] For the first target audio chunk whose multi-modal feature comprehensive score exceeds the feature comprehensive score threshold, trigger the offline speech recognition model to perform offline speech recognition on the first target audio chunk to obtain the corresponding offline subtitle text.
[0021] For n consecutive cumulative second target audio chunks whose multi-modal feature comprehensive score does not exceed the feature comprehensive score threshold, forcibly start the offline speech recognition model to perform offline speech recognition on each second target audio chunk to obtain the corresponding offline subtitle text, where n is a positive integer.
[0022] Optionally, the optimization training process of the offline speech recognition model is:
[0023] A student model is obtained by knowledge distillation of a teacher model. The quantization technique is applied to the student model to convert the weights and activation functions of the student model from high-precision floating-point type to low-precision integer type, and the pruning technique is used to remove the parameters or connections with a contribution degree less than the contribution degree threshold to the model performance, thereby obtaining the offline speech recognition model.
[0024] Optionally, the target streaming caption text is corrected using the target offline caption text to obtain the final live caption, including:
[0025] The target streaming caption text is corrected using the target offline caption text, and text error correction is performed on the corrected target streaming caption text to obtain the final live caption.
[0026] Optionally, it further includes:
[0027] The live caption is output to the front-end display interface in the form of a fade-in animation.
[0028] Optionally, it further includes:
[0029] During the process of performing offline speech recognition on each audio chunk, if there is a third target audio chunk whose offline speech recognition time exceeds the time threshold, the third target audio chunk is discarded and the corresponding log is recorded.
[0030] Optionally, the live audio stream data is segmented into multiple audio chunks based on audio short-time energy and speech activity detection confidence, including:
[0031] According to the monitored central processing unit utilization rate, the chunking strategy for segmenting the live audio stream data into multiple audio chunks based on the audio short-time energy and the speech activity detection confidence and the corresponding thread pool size are dynamically adjusted.
[0032] A live caption generation device includes:
[0033] An audio stream acquisition unit for acquiring live audio stream data;
[0034] A streaming speech recognition unit for performing streaming speech recognition on the live audio stream data to obtain a streaming caption text;
[0035] An offline speech recognition unit for segmenting the live audio stream data into multiple audio chunks based on audio short-time energy and speech activity detection confidence, and performing offline speech recognition on each audio chunk to obtain an offline caption text;
[0036] A timestamp alignment unit for aligning the timestamps of the streaming caption text and the offline caption text to obtain a target streaming caption text and a target offline caption text;
[0037] A subtitle correction unit, configured to correct the target streaming subtitle text by using the target offline subtitle text to obtain the final live subtitle.
[0038] An electronic device, comprising: a memory and a processor;
[0039] The memory is used to store at least one instruction;
[0040] The processor is configured to execute the at least one instruction to implement any live subtitle generation method.
[0041] As can be seen from the above technical solutions, the present invention discloses a live subtitle generation method and related devices. The method includes: acquiring live audio stream data, performing streaming speech recognition on the live audio stream data to obtain a streaming subtitle text, segmenting the live audio stream data based on audio short-time energy and speech activity detection confidence to obtain a plurality of audio chunks, performing offline speech recognition on each audio chunk to obtain an offline subtitle text, aligning the time stamps of the streaming subtitle text and the offline subtitle text to obtain a target streaming subtitle text and a target offline subtitle text, and correcting the target streaming subtitle text by using the target offline subtitle text to obtain the final live subtitle. First, the present application uses streaming speech recognition technology to perform preliminary recognition on the live audio stream data to obtain a streaming subtitle text, and then uses offline speech recognition technology to recognize each audio chunk obtained by segmenting the live audio stream data to obtain an offline subtitle text. After aligning the time stamps of the streaming subtitle text and the offline subtitle text to ensure that the subtitle text is consistent with the audio content, the target offline subtitle text after time stamp alignment is used to correct the target streaming subtitle text, realizing the collaborative output of the real-time response of streaming recognition speech and the offline high-precision correction of the speech recognition result, so that the live subtitle achieves a balance between real-time performance and accuracy. Description of the Drawings
[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the disclosed drawings.
[0043] Figure 1 It is a flowchart of a live subtitle generation method disclosed in an embodiment of the present invention;
[0044] Figure 2 It is a schematic structural diagram of a live subtitle generation device disclosed in an embodiment of the present invention;
[0045] Figure 3Schematic structural diagram of an electronic device disclosed in an embodiment of the present invention. Detailed implementation manners
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0047] An embodiment of the present invention discloses a live caption generation method and related device. First, the streaming speech recognition technology is used to perform preliminary recognition on the live audio stream data to obtain the streaming caption text. Then, the offline speech recognition technology is used to recognize each audio segment obtained by splitting the live audio stream data to obtain the offline caption text. After aligning the timestamps of the streaming caption text and the offline caption text to ensure that the caption text is consistent with the audio content, the target offline caption text after timestamp alignment is used to correct the target streaming caption text, realizing the collaborative output of the real-time response of the streaming recognized speech and the offline high-precision corrected speech recognition result, so that the live caption reaches a balance in terms of real-time performance and accuracy.
[0048] See Figure 1 , a flowchart of a live caption generation method disclosed in an embodiment of the present application. The method includes:
[0049] Step S101, obtain the live audio stream data.
[0050] During the live broadcast process of the user using the live broadcast device, the generated live audio stream data is obtained in real time.
[0051] Step S102, perform streaming speech recognition on the live audio stream data to obtain the streaming caption text.
[0052] Streaming speech recognition, also known as real-time speech recognition, is a technology that can recognize while the user is speaking and return the recognition result in real time. The core of the streaming speech recognition technology lies in its ability to output results while inputting speech, thus greatly reducing the processing time of speech recognition in the human-computer interaction process.
[0053] Specifically, the live audio stream data can be recognized by using a streaming speech recognition model to obtain the streaming caption text.
[0054] In practical applications, compared with the offline speech recognition model, the streaming speech recognition model obtains less context information. Therefore, the recognition accuracy of the streaming speech recognition model is not as good as that of the offline speech recognition model. Based on this, the present application improves the recognition ability of the streaming speech recognition model through knowledge distillation. Knowledge distillation is a technology for model compression and transfer learning, and its core idea is to transfer the knowledge of a large model (teacher model) to a small model (student model) to improve the performance and speed of the small model. When using knowledge distillation to improve the recognition ability of the streaming speech recognition model, the offline model can be used as the teacher and the streaming model as the student, and the context-dependent knowledge is transferred by minimizing the KL (Kullback-Leibler) divergence loss.
[0055] Step S103: Segment the live audio stream data based on the short-time energy of the audio and the confidence of voice activity detection to obtain multiple audio chunks, and perform offline speech recognition on each audio chunk to obtain offline caption text.
[0056] The short-time energy of an audio refers to the energy accumulation of an audio signal within a short-time window, usually expressed in decibels (dB). The short-time energy of an audio frame is specifically: perform an amplitude square operation on the windowed speech signal, and accumulate the results to obtain the energy value of this frame, which represents the amplitude size of this frame signal.
[0057] Voice Activity Detection (VAD) confidence refers to the measure of the reliability of the detection result when performing voice activity detection, reflecting the accuracy of the system's judgment on whether there is voice activity in the current audio signal.
[0058] In the present application, the live audio stream data is segmented based on the short-time energy of the audio and the confidence of voice activity detection to obtain multiple audio chunks, and the duration of each audio chunk can be dynamically adjusted in combination with the short-time energy threshold of the audio and the confidence threshold of voice activity detection.
[0059] Offline speech recognition refers to the recognition and conversion of speech signals locally on the device without relying on an Internet connection or cloud services.
[0060] Step S104: Align the timestamps of the streaming caption text and the offline caption text to obtain the target streaming caption text and the target offline caption text.
[0061] In practical applications, the DTW (Dynamic Time Warping) algorithm can be used to determine the similarity between two time series and minimize the distance between the series. The present application uses DTW to align the timestamps of the streaming caption text and the offline caption text.
[0062] Step S105: Correct the target streaming subtitle text with the target offline subtitle text to obtain the final live subtitle.
[0063] Generally, the accuracy of the recognition result of offline speech recognition is higher than that of streaming speech recognition. Based on this, in this application, the target offline subtitle text obtained by timestamp alignment is used to correct the target streaming subtitle text to obtain the final live subtitle.
[0064] In practical applications, the minimum edit operations (insertion, deletion, replacement) can be generated based on the OT (Operational Transformation) algorithm to correct the target streaming subtitle text to obtain the final live subtitle.
[0065] The following is an example for illustration:
[0066] Example 1
[0067] The target streaming subtitle text is: A B C D;
[0068] The target offline subtitle text is: A X D;
[0069] The final live subtitle is: A X D, and the correction operations include: deleting B and C in the target streaming subtitle text, and inserting X between A and D.
[0070] Example 2
[0071] The target streaming subtitle text is: [[Hello, ], [The weather is clear and sunny today.]];
[0072] The target offline subtitle text is: [[Hello, ], [The weather is sunny and bright today.]];
[0073] The final live subtitle is: [[Hello, ], [The weather is sunny and bright today.]];
[0074] The correction operations include: replacing "beautiful" with "charming".
[0075] Example 3
[0076] The target streaming subtitle text is: [[The weather is clear and sunny today.]];
[0077] The target offline subtitle text is: [[The weather is sunny and bright today.]];
[0078] The final live subtitle is: [[The weather is sunny and bright today.]];
[0079] The correction operations include: replacing "清" with "晴", replacing "美" with "媚", inserting a "天" in "今日氣清晴", and inserting a "啊" after the replaced "媚".
[0080] In summary, the present application discloses a method for generating live subtitles, which obtains live audio stream data, performs streaming speech recognition on the live audio stream data to obtain streaming subtitle text, divides the live audio stream data into multiple audio blocks based on audio short-time energy and voice activity detection confidence, and performs offline speech recognition on each audio block to obtain offline subtitle text, aligns the streaming subtitle text and the offline subtitle text by timestamps, obtains the target streaming subtitle text and the target offline subtitle text, and uses the target offline subtitle text to correct the target streaming subtitle text to obtain the final live subtitle. The present application first uses streaming speech recognition technology to perform preliminary recognition on the live audio stream data to obtain streaming subtitle text, and then uses offline speech recognition technology to recognize each audio block obtained by segmenting the live audio stream data to obtain offline subtitle text, and after aligning the streaming subtitle text and the offline subtitle text by timestamps to ensure that the subtitle text is consistent with the audio content, the target offline subtitle text after timestamp alignment is used to correct the target streaming subtitle text, so as to achieve the coordinated output of the real-time response of the streaming recognition speech and the offline high-precision correction speech recognition result, so that the live subtitles achieve a balance in real-time and accuracy.
[0081] In one embodiment, step S103 may specifically include:
[0082] (1) The live audio stream data is divided into multiple audio blocks based on the audio short-term energy and voice activity detection confidence.
[0083] Each audio block is: a first audio block representing a speech active area or a second audio block representing a combination of multiple continuous speech inactive areas.
[0084] In the present application, the first audio block is a speech active area (high energy area). In practical applications, the duration of the first audio block is relatively short, such as 0.3s, to improve the real-time performance of streaming speech recognition.
[0085] The second audio block is an audio block formed by merging multiple continuous speech inactive regions (low energy regions). The second audio block has a longer duration, such as 1.2s, to reduce the amount of calculation during speech recognition.
[0086] The length of the audio chunk, Chunk size, can be determined using the following formula: (1)
[0087] (1);
[0088] Wherein, H represents the short-time energy of audio, represents the short-time energy threshold of audio, C represents the confidence of voice activity detection, represents the confidence threshold of voice activity detection, represents others, that is the situation of.
[0089] The meaning of the formula is: initially, the duration of each audio segment obtained by splitting the live audio stream data is 0.3s, and the final length of the audio segment is determined according to the short-time energy of audio and the confidence of voice activity detection for each audio segment.
[0090] If the short-time energy H of the audio segment is greater than the short-time energy threshold , and the confidence of voice activity detection is greater than the confidence threshold of voice activity detection , it is determined that the audio segment is a voice active region, and the audio segment is determined as the first audio segment, and the duration of the first audio segment remains 0.3s unchanged.
[0091] If the short-time energy H of the audio segment is not greater than the short-time energy threshold , or, the confidence of voice activity detection is not greater than the confidence threshold of voice activity detection , it is determined that the audio segment is a voice inactive region, and the next audio segment of the audio segment is continuously judged. If four consecutive audio segments are all the situation of, the durations of the four audio segments are merged to obtain the second audio segment, and the duration of the second audio segment is 1.2s.
[0092] It should be noted that the values of 0.3s and 1.2s for the duration Chunk size of the audio segment in formula (1) are both examples, and the values can be adjusted according to needs in actual applications.
[0093] (2) Adopt the cache management strategy of the Least Recently Used-K algorithm, discard K consecutive adjacent second audio segments among all the second audio segments, and obtain an audio segment set composed of the remaining second audio segments and all the first audio segments;
[0094] In actual applications, adopt the LRU-K (Least Recently Used-K) cache management strategy, preferentially cache the audio segments with high VAD confidence (that is, the first audio segments), that is, cache the audio segments that may contain important voice information, and at the same time release the audio segments with low VAD confidence (that is, the second audio segments) to reduce the cache data volume and reduce the silent area.
[0095] (3) Perform offline speech recognition on each audio chunk in the audio chunk set to obtain the offline subtitle text.
[0096] Taking the live audio stream data [0.3, 0.3, 0.3, 0.3, 0.3, 0.3, 0.3, 0.3, 0.3......] as an example, the total audio length of the live audio stream data is 0.3×n, that is, the live audio stream data includes n audio chunks, and the duration of each audio chunk is 0.3s.
[0097] According to formula (1), it can be known that the duration of the audio chunk that satisfies remains unchanged at 0.3s. If it does not satisfy , it is determined that the audio chunk contains less speech information, and subsequent merge operations are performed. Taking the merge of the durations of four consecutive audio chunks that do not satisfy as an example, the merged duration is 1.2s. The live audio stream data that simultaneously includes the first audio chunk and the second audio chunk is obtained as follows:
[0098] Live audio stream data [0.3, 0.3, 0.3, 1.2, 0.3, 0.3, 0.3, 0.3, 0.3, 0.3, 1.2......].
[0099] Then, use the LRU-K cache management strategy to process the live audio stream data after merging the audio chunk durations, discard the middle consecutive K second audio chunks, retain the high-information audio chunks for recognition, reduce the real-time memory occupancy of the stream, and finally obtain the following audio chunk set:
[0100] Audio chunk set [0.3, 0.3, 0.3, 1.2, 0.3, 0.3, 0.3, 0.3, 1.2......].
[0101] It should be particularly noted that in the priority thread pool, each task will be assigned a priority. The order of assigning priorities in this application is: streaming recognition (high priority) > offline correction (medium priority) > cache management (low priority), and the token bucket algorithm (Token Bucket Algorithm, TBA) is used to limit the concurrency of offline tasks to avoid system overload, that is, to avoid CPU (Central Processing Unit) overload.
[0102] The token bucket algorithm (Token Bucket Algorithm, TBA) is a traffic control and rate algorithm widely used in the field of network communication. In this application, streaming recognition, offline recognition, and cache management are synchronized in the thread pool. To avoid system overload, this application sets the task assignment priority to ensure the normal and stable operation of the system.
[0103] In one embodiment, the process of performing offline speech recognition on each audio chunk in the audio chunk set to obtain the offline subtitle text may include:
[0104] (1) Obtain the silence probability for each audio chunk in the audio chunk set by using a voice activity detection processing method.
[0105] In practical applications, when processing each audio chunk by using a voice activity detection processing method, voice activity detection processing can be performed on each audio chunk based on the VAD model obtained by an RNN (Recurrent Neural Network), and the silence probability of each audio chunk can be obtained. Among them, the audio chunk with a high silence probability is a potential speech end point.
[0106] (2) Determine the cosine similarity between the MFCC features of each audio chunk in the audio chunk set.
[0107] MFCC (Mel-Frequency Cepstral Coefficients) is the result of feature extraction of an audio signal based on the Mel scale (a frequency scale that simulates the human ear's perception). It can effectively capture the spectral features of the audio signal, especially the features that are more important for the human auditory system.
[0108] In this application, by determining the cosine similarity between the MFCC features of each audio chunk, it is determined whether the audio chunk is a potential speech end point, and the audio chunk with a low cosine similarity is determined as a potential speech end point.
[0109] (3) Perform fundamental frequency continuity analysis on each audio chunk in the audio chunk set to determine whether there is a fundamental frequency mutation.
[0110] Fundamental frequency mutation refers to a significant change in the fundamental frequency in the sound signal, and based on this, it is determined whether the corresponding audio chunk is a potential speech end point, that is, the audio chunk with a fundamental frequency mutation is a potential speech end point.
[0111] (4) Determine the multi-modal feature comprehensive score based on the silence probability, the cosine similarity, and whether there is a fundamental frequency mutation corresponding to each audio chunk in the audio chunk set.
[0112] After this application obtains the audio chunk set [0.3, 0.3, 0.3, 1.2, 0.3, 0.3, 0.3, 0.3, 1.2......], an important step in speech recognition is to determine the speech end point, which has an important impact on offline speech recognition, sentence segmentation, punctuation determination, etc.
[0113] In practical applications, any single method among the following three methods can be used to identify the end point of speech: using a voice activity detection processing method for audio chunks to obtain the silence probability, determining the cosine similarity between the MFCC features of audio chunks, and performing fundamental frequency continuity analysis on audio chunks to determine whether there is a fundamental frequency mutation. In order to improve the accuracy of determining the end point of speech, this application comprehensively analyzes audio chunks in combination with three parameters to obtain a comprehensive multi-modal feature score. When the comprehensive multi-modal feature score exceeds the feature comprehensive score threshold, it is determined that the corresponding audio chunk contains the end point of speech.
[0114] (5)For the first target audio chunk whose comprehensive multi-modal feature score exceeds the feature comprehensive score threshold, trigger the offline speech recognition model to perform offline speech recognition on the first target audio chunk to obtain the corresponding offline subtitle text.
[0115] In this embodiment, the process of triggering offline speech recognition for the first target audio chunk whose comprehensive multi-modal feature score exceeds the feature comprehensive score threshold is defined as: end point trigger.
[0116] Taking the audio chunk set [0.3, 0.3, 0.3, 1.2, 0.3, 0.3, 0.3, 0.3, 1.2......] as an example, the obtained first target audio chunks are as follows:
[0117] [[0.3, 0.3, 0.3], [1.2], [0.3, 0.3, 0.3, 0.3], [1.2],......].
[0118] The first target audio chunks are respectively: [0.3, 0.3, 0.3], [1.2], [0.3, 0.3, 0.3, 0.3], [1.2], that is, each [] can be understood as the start and end of a sentence.
[0119] (6)For n consecutive and cumulative second target audio chunks whose comprehensive multi-modal feature score does not exceed the feature comprehensive score threshold, force the offline speech recognition model to perform offline speech recognition on each second target audio chunk to obtain the corresponding offline subtitle text, where n is a positive integer.
[0120] In this application, the process of forcing the start of offline speech recognition for n consecutive and cumulative second target audio chunks whose comprehensive multi-modal feature score does not exceed the feature comprehensive score threshold is defined as: forced trigger.
[0121] To avoid causing excessive memory pressure and real-time performance pressure on the system, this application forces the start of the offline speech recognition task when n audio chunks have been accumulated without detecting the end point of speech.
[0122] Taking n = 10 as an example, if the comprehensive scores of the multimodal features of 10 consecutive audio chunks do not exceed the comprehensive score threshold of the features, that is, the voice end point is not detected, forced sentence breaking is triggered at this time to initiate speech recognition.
[0123] Taking the audio chunk set [0.3, 0.3, 0.3, 1.2, 0.3, 0.3, 0.3, 0.3, 1.2......] as an example, the processing result after forced sentence breaking triggers offline speech recognition is:
[0124] [[0.3, 0.3, 0.3, 0.3, 0.3, 0.3, 0.3, 0.3, 0.3, 0.3], [0.3, 0.3, 0.3, 0.3, 0.3, 0.3, 0.3, 0.3]......].
[0125] Among them, each [] is a set of 10 audio chunks for which forced sentence breaking triggers offline speech recognition.
[0126] The optimization training process of the offline speech recognition model in this application is as follows:
[0127] Knowledge distillation is performed on the teacher model to obtain the student model. The quantization technique is applied to the student model to convert the weights and activation functions of the student model from high-precision floating-point type to low-precision integer type, and the pruning technique is used to remove the parameters or connections with a contribution degree to the model performance less than the contribution degree threshold, thereby obtaining the offline speech recognition model.
[0128] Knowledge distillation is a technology for model compression and transfer learning. Its core idea is to transfer the knowledge of a large model (teacher model) to a small model (student model) to improve the performance and speed of the small model.
[0129] The quantization technique reduces the storage space size and computational complexity of the model by converting the weights and activation functions of the model from high-precision floating-point type to low-precision integer type.
[0130] The pruning technique reduces the number of parameters and computational amount of the model and improves the running efficiency of the model by accurately identifying and removing the parameters or connections with a relatively small contribution degree to the model performance.
[0131] This application improves the performance, speed and running efficiency of the offline speech recognition model through knowledge distillation, quantization technique and pruning technique, and reduces the storage space and computational complexity of the offline speech recognition model.
[0132] In practical applications, when adopting quantization technology for the student model, dynamic 8-bit quantization can be used, that is, converting the model weights from FP32 to INT8 to reduce memory occupancy. Utilize OpenMP (Open Multi-Processing) to parallelize matrix multiplication, and optimize the operation method for the SSE (Streaming SIMD Extensions) single instruction multiple data instruction set to optimize the CPU operator.
[0133] In one embodiment, step S105 may specifically include:
[0134] Use the target offline caption text to correct the target streaming caption text, and perform text error correction on the corrected target streaming caption text to obtain the final live caption.
[0135] In practical applications, the corrected target streaming caption text can be subjected to text error correction by superimposing a lightweight NLP (Natural Language Processing) language model, such as a Transformer model with only 4 layers and a hidden speech vector feature space of 128 dimensions, to obtain the final live caption.
[0136] For example, perform text error correction on [[Today's weather is clear and sunny.]] to obtain the following content:
[0137] [[Today's weather is sunny and clear.]]
[0138] In one embodiment, the live caption generation method may further include:
[0139] Output the live caption to the front-end display interface in the form of a fade-in animation.
[0140] In practical applications, a fade-in animation can be added to the live caption at the UI (User Interface) layer, rather than a direct jump, to present a word-by-word fade-in and fade-out effect, improving the user's viewing experience.
[0141] In one embodiment, the live caption generation method may further include:
[0142] During the process of performing offline speech recognition on each audio chunk, if there is a third target audio chunk whose offline speech recognition time exceeds the time threshold, discard the third target audio chunk and record the corresponding log.
[0143] During the process of performing offline speech recognition on each audio chunk, if due to special circumstances that may occur, such as system fluctuations or overly long audio, a third target audio chunk whose offline speech recognition time exceeds the time threshold appears, discard the third target audio chunk (i.e., actively abandon the task) and record the corresponding log to prevent blocking the main thread and ensure the normal progress of subsequent streaming speech recognition and output correction data stream.
[0144] Among them, the value of the time threshold is determined according to actual needs and is not limited in this application.
[0145] In practical applications, during the process of generating live subtitles, the audio chunk strategy and the size of the thread pool can be dynamically adjusted by monitoring the CPU utilization rate in real time.
[0146] Therefore, in one embodiment, the process of splitting the live audio stream data into multiple audio chunks based on audio short-time energy and voice activity detection confidence can include:
[0147] According to the monitored CPU utilization rate, dynamically adjust the chunking strategy for splitting the live audio stream data into multiple audio chunks based on audio short-time energy and voice activity detection confidence and the corresponding size of the thread pool.
[0148] Corresponding to the above method embodiment, this application also discloses a live subtitle generation device.
[0149] See Figure 2 , a schematic structural diagram of a live subtitle generation device disclosed in an embodiment of this application, the device includes:
[0150] An audio stream acquisition unit 201, configured to acquire live audio stream data.
[0151] During the live broadcast by the user using the live broadcast device, acquire the generated live audio stream data in real time.
[0152] A streaming speech recognition unit 202, configured to perform streaming speech recognition on the live audio stream data to obtain a streaming subtitle text.
[0153] Streaming speech recognition, also known as real-time speech recognition, is a technology that can perform recognition while the user is speaking and return the recognition result in real time. The core of streaming speech recognition technology lies in its ability to output results while inputting speech, thereby greatly reducing the processing time of speech recognition in the human-computer interaction process.
[0154] The streaming speech recognition unit 202 can specifically be configured to perform recognition on the live audio stream data using a streaming speech recognition model to obtain a streaming subtitle text.
[0155] In practical applications, compared with offline speech recognition models, streaming speech recognition models obtain less context information. Therefore, the recognition accuracy of streaming speech recognition models is not as good as that of offline speech recognition models. Based on this, the present application improves the recognition ability of streaming speech recognition models through knowledge distillation. Knowledge distillation is a technology for model compression and transfer learning. Its core idea is to transfer the knowledge of a large model (teacher model) to a small model (student model) to improve the performance and speed of the small model. When using knowledge distillation to improve the recognition ability of streaming speech recognition models, the offline model can be used as the teacher and the streaming model as the student, and the context-dependent knowledge can be transferred by minimizing the KL (Kullback-Leibler) divergence loss.
[0156] The offline speech recognition unit 203 is configured to segment the live audio stream data based on the short-time energy of the audio and the confidence of voice activity detection to obtain a plurality of audio chunks, and perform offline speech recognition on each of the audio chunks to obtain offline caption texts.
[0157] The short-time energy of an audio refers to the energy accumulation of an audio signal within a short-time window, usually expressed in decibels (dB). The short-time energy of an audio frame is specifically: performing an amplitude square operation on the windowed speech signal and accumulating the results to obtain the energy value of this frame, which represents the amplitude size of this frame signal.
[0158] The confidence of voice activity detection (VAD) is a measure of the reliability of the system's detection results during voice activity detection, reflecting the accuracy of the system's judgment on whether there is voice activity in the current audio signal.
[0159] In the present application, the live audio stream data is segmented based on the short-time energy of the audio and the confidence of voice activity detection to obtain a plurality of audio chunks, and the duration of each audio chunk can be dynamically adjusted in combination with the short-time energy threshold of the audio and the confidence threshold of voice activity detection.
[0160] Offline speech recognition refers to the recognition and conversion of speech signals locally on a device without relying on an Internet connection or cloud services.
[0161] The timestamp alignment unit 204 is configured to align the timestamps of the streaming caption text and the offline caption text to obtain a target streaming caption text and a target offline caption text.
[0162] In practical applications, the DTW (Dynamic Time Warping) algorithm can be used to determine the similarity between two time series and minimize the distance between the sequences. In this application, DTW is used to align the timestamps of the streaming subtitle text and the offline subtitle text.
[0163] The subtitle correction unit 205 is configured to correct the target streaming subtitle text by using the target offline subtitle text to obtain the final live subtitle.
[0164] Generally, the accuracy of the recognition result of offline speech recognition is higher than that of streaming speech recognition. Based on this, in this application, the target streaming subtitle text is corrected by using the target offline subtitle text obtained by timestamp alignment to obtain the final live subtitle.
[0165] In practical applications, the OT (Operational Transformation) algorithm can be used to generate the minimum edit operations (insertion, deletion, replacement) to correct the target streaming subtitle text to obtain the final live subtitle.
[0166] In summary, this application discloses a live subtitle generation device, which obtains live audio stream data, performs streaming speech recognition on the live audio stream data to obtain streaming subtitle text, segments the live audio stream data based on the short-time energy of the audio and the confidence of voice activity detection to obtain multiple audio chunks, performs offline speech recognition on each audio chunk to obtain offline subtitle text, aligns the timestamps of the streaming subtitle text and the offline subtitle text to obtain the target streaming subtitle text and the target offline subtitle text, and corrects the target streaming subtitle text by using the target offline subtitle text to obtain the final live subtitle. This application first uses the streaming speech recognition technology to perform preliminary recognition on the live audio stream data to obtain the streaming subtitle text, and then uses the offline speech recognition technology to recognize each audio chunk obtained by segmenting the live audio stream data to obtain the offline subtitle text. After aligning the timestamps of the streaming subtitle text and the offline subtitle text to ensure that the subtitle text is consistent with the audio content, the target streaming subtitle text is corrected by using the target offline subtitle text after timestamp alignment, realizing the collaborative output of the real-time response of the streaming recognized speech and the offline high-precision corrected speech recognition result, and making the live subtitle achieve a balance between real-time performance and accuracy.
[0167] The offline speech recognition unit 203 may specifically include:
[0168] An audio chunk splitting sub-unit, configured to split the live audio stream data into a plurality of audio chunks based on the short-time energy of the audio and the confidence of the voice activity detection. Each audio chunk is either a first audio chunk representing a voice active region or a second audio chunk obtained by merging a plurality of consecutive voice inactive regions;
[0169] A cache management sub-unit, configured to adopt a cache management strategy of the least recently used - K times algorithm to discard K consecutive adjacent second audio chunks among all the second audio chunks, so as to obtain an audio chunk set composed of the remaining second audio chunks and all the first audio chunks;
[0170] An offline recognition sub-unit, configured to perform offline speech recognition on each audio chunk in the audio chunk set to obtain the offline subtitle text.
[0171] In one embodiment, the offline recognition sub-unit may specifically be configured to:
[0172] Perform voice activity detection processing method on each audio chunk in the audio chunk set to obtain a silence probability. Among them, the audio chunk with a high silence probability is a potential voice tail point;
[0173] Determine the cosine similarity between the Mel Frequency Cepstral Coefficient (MFCC) features of each audio chunk in the audio chunk set. Among them, the audio chunk with a low cosine similarity is a potential voice tail point;
[0174] Perform fundamental frequency continuity analysis on each audio chunk in the audio chunk set to determine whether there is a fundamental frequency mutation. Among them, the audio chunk with a fundamental frequency mutation is a potential voice tail point;
[0175] Determine a multi-modal feature comprehensive score based on the silence probability, the cosine similarity, and whether there is a fundamental frequency mutation corresponding to each audio chunk in the audio chunk set;
[0176] For a first target audio chunk whose multi-modal feature comprehensive score exceeds the feature comprehensive score threshold, trigger an offline speech recognition model to perform offline speech recognition on the first target audio chunk to obtain the corresponding offline subtitle text;
[0177] For n consecutive cumulative second target audio chunks whose multi-modal feature comprehensive score does not exceed the feature comprehensive score threshold, force-start the offline speech recognition model to perform offline speech recognition on each second target audio chunk to obtain the corresponding offline subtitle text, where n is a positive integer.
[0178] Among them, the optimization training process of the offline speech recognition model is:
[0179] Knowledge distillation is performed on the teacher model to obtain the student model. The quantization technique is applied to the student model to convert the weights and activation functions of the student model from high-precision floating-point type to low-precision integer type, and the pruning technique is used to remove the parameters or connections with a contribution degree less than the contribution degree threshold to the model performance, thereby obtaining the offline speech recognition model.
[0180] In one embodiment, the subtitle correction unit 205 can specifically be used for:
[0181] Using the target offline subtitle text to correct the target streaming subtitle text, and performing text error correction on the corrected target streaming subtitle text to obtain the final live subtitle.
[0182] In one embodiment, the live subtitle generation device may further include:
[0183] An output unit, configured to output the live subtitle to the front-end display interface in the form of a fade-in animation.
[0184] In one embodiment, the live subtitle generation device may further include:
[0185] An audio chunk discarding unit, configured to discard a third target audio chunk with an offline speech recognition time exceeding a time threshold and record the corresponding log during the process of performing offline speech recognition on each audio chunk.
[0186] In one embodiment, the offline speech recognition unit 203 can specifically be used for:
[0187] According to the monitored central processing unit utilization rate, dynamically adjust the chunking strategy for splitting the live audio stream data into multiple audio chunks based on the audio short-time energy and the speech activity detection confidence level, and the corresponding thread pool size.
[0188] It should be noted that for the specific working principles of the components in the device embodiments, please refer to the corresponding parts of the method embodiments, which will not be elaborated here.
[0189] Corresponding to the above embodiments, the present application also discloses a computer storage medium, and the computer storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, the steps shown in the method embodiments of the live subtitle generation method are implemented.
[0190] A computer storage medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. The computer storage medium can be a machine-readable signal medium or a machine-readable storage medium. The computer storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0191] Corresponding to the above embodiments, as Figure 3 shown, the present invention also provides an electronic device, which may include: a processor 1 and a memory 2;
[0192] wherein, the processor 1 and the memory 2 communicate with each other through a communication bus 3;
[0193] The processor 1 is configured to execute at least one instruction;
[0194] The memory 2 is configured to store at least one instruction;
[0195] The processor 1 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0196] The memory 2 may include high-speed RAM memory and may also include non-volatile memory, such as at least one disk memory.
[0197] Wherein, the processor executes at least one instruction to implement the steps shown in the embodiment of the live caption generation method.
[0198] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0199] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. For the same or similar parts among the various embodiments, reference may be made to each other.
[0200] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A live caption generation method, characterized in that, Including: Obtaining live audio stream data; Performing streaming speech recognition on the live audio stream data to obtain streaming subtitle text; Segmenting the live audio stream data into multiple audio chunks based on audio short-time energy and speech activity detection confidence, and performing offline speech recognition on each audio chunk to obtain offline subtitle text; Aligning the time stamps of the streaming subtitle text and the offline subtitle text to obtain target streaming subtitle text and target offline subtitle text; Correcting the target streaming subtitle text using the target offline subtitle text to obtain the final live subtitle; Among them, segmenting the live audio stream data into multiple audio chunks based on audio short-time energy and speech activity detection confidence, and performing offline speech recognition on each audio chunk to obtain offline subtitle text, including: Segmenting the live audio stream data into multiple audio chunks based on the audio short-time energy and the speech activity detection confidence, and each audio chunk is: a first audio chunk representing a speech active region or a second audio chunk obtained by merging multiple consecutive speech inactive regions; Adopting a cache management strategy of the least recently used - K times algorithm, discarding K consecutive second audio chunks among all the second audio chunks to obtain an audio chunk set composed of the remaining second audio chunks and all the first audio chunks; Performing offline speech recognition on each audio chunk in the audio chunk set to obtain the offline subtitle text; The performing offline speech recognition on each audio chunk in the audio chunk set to obtain the offline subtitle text includes: Performing speech activity detection processing method on each audio chunk in the audio chunk set to obtain a silence probability, and among them, an audio chunk with a high silence probability is a potential speech end point; Determining the cosine similarity between the Mel frequency cepstral coefficient MFCC features of each audio chunk in the audio chunk set, and among them, an audio chunk with a low cosine similarity is a potential speech end point; Performing fundamental frequency continuity analysis on each audio chunk in the audio chunk set to determine whether there is a fundamental frequency mutation, and among them, an audio chunk with a fundamental frequency mutation is a potential speech end point; Determining a multi-modal feature comprehensive score based on the silence probability, the cosine similarity, and whether there is a fundamental frequency mutation corresponding to each audio chunk in the audio chunk set; For a first target audio chunk whose multi-modal feature comprehensive score exceeds the feature comprehensive score threshold, triggering an offline speech recognition model to perform offline speech recognition on the first target audio chunk to obtain the corresponding offline subtitle text; For n consecutive cumulative second target audio chunks whose multi-modal feature comprehensive score does not exceed the feature comprehensive score threshold, forcibly starting the offline speech recognition model to perform offline speech recognition on each second target audio chunk to obtain the corresponding offline subtitle text, where n is a positive integer.
2. The live caption generation method according to claim 1, wherein The optimization training process of the offline speech recognition model is: The teacher model is distilled to obtain a student model. The quantization technique is applied to the student model to convert the weights and activation functions of the student model from high-precision floating-point type to low-precision integer type, and the pruning technique is used to remove the parameters or connections with a contribution degree less than the contribution degree threshold to the model performance, thereby obtaining the offline speech recognition model.
3. The live caption generation method according to any one of claims 1 to 2, characterized in that Using the target offline subtitle text to correct the target streaming subtitle text to obtain the final live subtitle, including: Using the target offline subtitle text to correct the target streaming subtitle text, and performing text error correction on the corrected target streaming subtitle text to obtain the final live subtitle.
4. The live caption generation method according to any one of claims 1 to 2, characterized in that It also includes: Outputting the live subtitle to the front-end display interface in the form of a fade-in animation.
5. The live caption generation method according to any one of claims 1 to 2, characterized in that It also includes: During the process of performing offline speech recognition on each audio chunk, if there is a third target audio chunk whose offline speech recognition time exceeds the time threshold, discard the third target audio chunk and record the corresponding log.
6. The live caption generation method according to any one of claims 1 to 2, characterized in that Segmenting the live audio stream data based on audio short-time energy and speech activity detection confidence to obtain multiple audio chunks, including: Dynamically adjusting the chunking strategy for segmenting the live audio stream data based on the audio short-time energy and the speech activity detection confidence and the corresponding thread pool size according to the monitored CPU utilization rate.
7. A live caption generation device, characterized in that, It includes: An audio stream acquisition unit for acquiring live audio stream data; A streaming speech recognition unit for performing streaming speech recognition on the live audio stream data to obtain a streaming subtitle text; An offline speech recognition unit for segmenting the live audio stream data based on audio short-time energy and speech activity detection confidence to obtain multiple audio chunks, and performing offline speech recognition on each audio chunk to obtain an offline subtitle text; A timestamp alignment unit for aligning the timestamps of the streaming subtitle text and the offline subtitle text to obtain a target streaming subtitle text and a target offline subtitle text; A subtitle correction unit for using the target offline subtitle text to correct the target streaming subtitle text to obtain the final live subtitle; Among them, the offline speech recognition unit segments the live audio stream data based on audio short-time energy and speech activity detection confidence to obtain multiple audio chunks, and performs offline speech recognition on each audio chunk to obtain an offline subtitle text, including: Segmenting the live audio stream data into multiple audio chunks based on the audio short-time energy and the speech activity detection confidence, and each audio chunk is: a first audio chunk representing a speech active region or a second audio chunk obtained by merging multiple consecutive speech inactive regions; Adopting a cache management strategy of the least recently used - K times algorithm, discarding K consecutive second audio chunks among all the second audio chunks to obtain an audio chunk set composed of the remaining second audio chunks and all the first audio chunks; Performing offline speech recognition on each audio chunk in the audio chunk set to obtain the offline subtitle text; Performing offline speech recognition on each audio chunk in the audio chunk set to obtain the offline subtitle text includes: Obtaining a silence probability for each audio chunk in the audio chunk set by using a voice activity detection processing method, where an audio chunk with a high silence probability is a potential voice end point; Determining the cosine similarity between the Mel Frequency Cepstral Coefficient (MFCC) features of each audio chunk in the audio chunk set, where an audio chunk with a low cosine similarity is a potential voice end point; Performing fundamental frequency continuity analysis on each audio chunk in the audio chunk set to determine whether there is a fundamental frequency mutation, where an audio chunk with a fundamental frequency mutation is a potential voice end point; Determining a multi-modal feature comprehensive score based on the silence probability, the cosine similarity, and whether there is a fundamental frequency mutation corresponding to each audio chunk in the audio chunk set; For a first target audio chunk whose multi-modal feature comprehensive score exceeds a feature comprehensive score threshold, triggering an offline speech recognition model to perform offline speech recognition on the first target audio chunk to obtain the corresponding offline subtitle text; For n consecutive cumulative second target audio chunks whose multi-modal feature comprehensive score does not exceed the feature comprehensive score threshold, forcibly starting the offline speech recognition model to perform offline speech recognition on each second target audio chunk to obtain the corresponding offline subtitle text, where n is a positive integer.
8. An electronic device, characterized in that, The electronic device includes: a memory and a processor; The memory is used to store at least one instruction; The processor is used to execute the at least one instruction to implement the live subtitle generation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Video searching system based on content analysis
CN101021857A
Voice interaction method and device, voice recognition method and device, equipment and storage medium
CN114446280A