A low-latency real-time speech transcription method and system

By introducing speech activity detection algorithms and error correction mechanisms into the speech transcription system, the problems of high latency and unstable data transmission in the prior art are solved, and low latency and high efficiency real-time speech transcription is achieved, improving the user experience.

CN119811372BActive Publication Date: 2025-05-16SICHUAN YIXUN INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510295239.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-05-16
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

Existing speech transcription methods have problems of high latency and unstable data transmission, which cannot meet the needs of real-time interaction and user experience.

Method used

A low-latency real-time voice transcription method is adopted, and the audio data is encoded and compressed and transmitted to the server through the client, an error correction and packet loss retransmission mechanism is introduced, and a voice activity detection algorithm is used to segment the voice fragments and transcription is performed in real time on the server side.

Benefits of technology

It realizes low-latency and high-efficiency real-time voice transcription, significantly reduces processing delays, improves user experience, and ensures the stability and semantic integrity of audio data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811372B_ABST
    Figure CN119811372B_ABST
Patent Text Reader

Abstract

The present invention provides a low-latency real-time speech transcription method and system thereof, comprising the following steps: A. encoding and compressing original audio data collected by a client, and then transmitting the obtained audio encoding data to a server through a real-time transmission protocol, introducing an error correction and packet loss retransmission mechanism, and increasing the reliability of transmission; B. using a server-side voice activity detection algorithm to detect speech in audio data, dividing the audio data judged as speech into multiple independent parts, and adding a time sequence identifier to each part; C. transcribing the divided audio data in real time in time sequence to generate corresponding text content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of speech recognition, and in particular relates to a low-delay real-time speech transcription method and a system thereof. Background Art

[0002] With the rapid development of information technology, the demand for real-time speech transcription and translation is growing, especially in the fields of cross-language communication, meeting records, online education, etc. However, existing speech processing technologies face several challenges. First, traditional speech transcription methods often rely on high-latency batch processing modes and cannot meet the needs of real-time interaction. Secondly, when processing continuous audio streams, existing technologies often cause data loss or increased latency due to network fluctuations, affecting user experience. Therefore, it is particularly important to develop a real-time speech transcription system that can achieve low latency and high efficiency.

[0003] CN117640976A provides a low-delay speech recognition method in a live broadcast environment, including the following steps: receiving audio data and preprocessing, performing preliminary processing on the unpacked audio data, audio segmentation, speech-to-text conversion, text translation, and text display. The audio segmentation step of the patent is to segment according to the time segment length set between 200ms and 500ms. Although the Hamming window and other methods are used to reduce the distortion caused by the segmentation, this method only reduces the distortion but cannot avoid it, and may segment the sentence, resulting in incomplete speech segments and affecting the transcription effect. Summary of the invention

[0004] In order to solve the delay problem of existing speech transcription methods, the present invention provides a low-delay real-time speech transcription method, which is a reliable and smooth real-time speech transcription and translation solution.

[0005] In order to achieve the purpose of the present invention, the technical solution adopted by the present invention is:

[0006] A low-latency real-time speech transcription method comprises the following steps:

[0007] A. Encode and compress the original audio data collected by the client, and then transmit the obtained audio encoding data to the server through the real-time transport protocol, introduce error correction and packet loss retransmission mechanism to increase the reliability of transmission;

[0008] B. The server uses a voice activity detection algorithm to detect speech in the audio data, splits the audio data judged to be speech into multiple independent parts, and adds a timestamp to each part;

[0009] C. The segmented audio data is transcribed in real time in chronological order to generate corresponding text content.

[0010] The step A of the present invention comprises:

[0011] A1. Create an array A to store the audio to be detected for voice activity;

[0012] A2. Each time, 256ms-512ms of audio data is taken from the collected audio and appended to array A. If the current accumulated audio segment duration L reaches more than 4 times the duration of the single appended audio, voice activity detection is performed, and a timestamp is added to each voice segment. The audio consists of several voice segments containing voice and silent segments without voice. The segment between two adjacent voice segments is a silent segment.

[0013] A3, the difference between the duration L of the audio segment and the end time of the last voice segment is diff,

[0014] If diff<1s, the audio data before the last voice segment is uploaded to the server, and the uploaded audio data in array A is deleted;

[0015] If diff>=1s, all the audio segment data are uploaded to the server and array A is cleared;

[0016] The Opus coding transmission bit rate of the silence segment is lower than that of the speech segment;

[0017] A4. If the client has not finished the collection work, continue to execute the loop from step A2.

[0018] In step A of the present invention, based on the UDP transmission protocol, the UDP protocol data packet is encoded based on the forward error correction mechanism to obtain the encoded data block, and the encoded data block is interleaved to combat burst packet loss.

[0019] Step A of the present invention adopts key frame retransmission, and triple redundancy is implemented for the key frames of the audio during retransmission:

[0020] Original frame F_t;

[0021] Differential frame ΔF_t = F_t ⊕ F_{t-1};

[0022] Check frame C_t = CRC32(F_t) || (F_t >> 8);

[0023] A sliding window mechanism is used with a window size of W∈[2,5] and a redundancy insertion period of T∈[10ms,50ms], satisfying:

[0024]

[0025]

[0026] Where t is the serial number of the audio frame, F_t is the audio data of the current frame, and F_{t-1} is the audio data of the previous frame.

[0027] CRC refers to cyclic redundancy check, ⊕ represents Galois field addition, is the packet loss rate threshold,

[0028] is the weight coefficient.

[0029] In step A of the present invention, spiral check is used to detect residual errors during audio decoding: in audio data transmission, the sending end generates spiral check bytes for the data block and appends them to the end of the data block; the receiving end extracts the data block and recalculates the check bytes to detect whether there is an error; when the receiving end detects a data error through spiral check, it first attempts to locate and correct the error through a spiral check polynomial. If the spiral check cannot correct the error, the receiving end will mark the frame as a residual error and notify the sending end to transmit triple redundant information to recover the lost or damaged data.

[0030] Preferably, a spiral check polynomial is defined:

[0031]

[0032]

[0033] in,

[0034] :Indicates the first A value of bytes, each byte is treated as a Galois field The elements in

[0035] :Position weight of the polynomial, corresponding to the position of the byte in the data stream,

[0036] : A reduced polynomial defined on GF(2) used to construct a 16-dimensional spiral space;

[0037] mod means modular operation;

[0038] is a prime number permutation table, ⊕ represents Galois Field addition, represents Galois Field multiplication,

[0039] Indicates the first j + 1 byte (the byte following the current byte), Indicates circular left shift Bit operations.

[0040] The server side of the present invention configures an array as an audio buffer, and fills the array each time an audio data block is received and placed after the previous data block; when the audio data accumulated in the buffer reaches a preset threshold, voice activity detection is performed on the audio stream to distinguish between voice segments and silent segments, and the audio stream is segmented into independent voice segments based on timestamps, and the segmented voice segments are transcribed in real time in chronological order.

[0041] Preferably, when segmenting, two adjacent voice segments of a silent segment with a duration of less than 1 second are regarded as a continuous voice segment, and the last voice segment after the audio stream is segmented is regarded as incomplete audio, and the voice segment and the subsequent audio data are temporarily retained in the buffer and merged and processed after subsequent data arrives.

[0042] A system for implementing the low-latency real-time speech transcription method of the present invention comprises:

[0043] The client obtains the audio data, encodes and compresses the original audio data, and transmits it to the server in real time through the data transmission module;

[0044] The server receives the audio data, detects the voice in the audio using the voice activity detection module, segments the audio data judged to be voice, and transcribes the segmented audio data in real time in chronological order.

[0045] Preferably, the server side calls the translation model or transmits it to a third-party translation software to translate the text information.

[0046] The beneficial effects of the present invention are:

[0047] 1. The present invention ensures the accuracy of speech transcription and the immediacy of translation through the collaborative work of the client and the server, providing users with an efficient and reliable real-time speech processing solution, significantly reducing processing delays and improving user experience.

[0048] 2. The present invention introduces a voice activity detection (VAD) algorithm to optimize the audio data transmission protocol, uses voice detection to distinguish between voice segments and silent segments in the audio, and dynamically encodes the voice segments and silent segments to reduce the amount of data transmission. Then, by limiting the audio acquisition duration and the detection duration, the audio segment is transmitted every 256ms-512ms. The single transmission amount is small and the transmission is faster. The integrity of the voice is judged according to the duration of the last audio segment during the transmission process. While solving the problems of unstable audio data transmission and high processing delay in the prior art, the audio quality is also guaranteed.

[0049] 3. UDP protocol data packets are encoded based on the forward error correction mechanism, and the data blocks are interleaved. The improved spiral check is used to detect residual errors. The perfect error correction mechanism can detect packet loss or redundancy in time, and then implement triple redundancy for the audio key frames to request the client to retransmit or delete redundancy, thereby ensuring the stability of audio data transmission.

[0050] 4. After the present invention transmits the audio data to the server, the voice activity detection (VAD) algorithm is used to segment the audio again, where the silence begins, and the silence segments less than 1 second are considered as incomplete voice segments, which can ensure that the segmented voice segments are complete sentences, and the semantics will not be destroyed or distorted due to the segmentation of complete sentences. The accuracy and semantic integrity of the transcribed results are guaranteed, which is more friendly to the translation engine. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 The figure is a flow chart of the low-latency real-time speech transcription method of the present invention. DETAILED DESCRIPTION

[0052] In order to more clearly and in detail illustrate the technical solution of the present invention, the present invention is further described below through relevant embodiments. The following embodiments are only for specific explanation of the implementation method of the present invention, and do not limit the protection scope of the present invention.

[0053] Embodiment 1:

[0054] A low-latency real-time speech transcription method comprises the following steps:

[0055] A. Encode and compress the original audio data collected by the client, and then transmit the obtained audio encoding data to the server through the real-time transport protocol, introduce error correction and packet loss retransmission mechanism to increase the reliability of transmission;

[0056] B. The server uses a voice activity detection algorithm to detect speech in the audio data, splits the audio data judged to be speech into multiple independent parts, and adds a timestamp to each part;

[0057] C. The segmented audio data is transcribed in real time in chronological order to generate corresponding text content.

[0058] Embodiment 2:

[0059] This embodiment is based on embodiment 1:

[0060] The step A comprises:

[0061] A1. Create an array A to store the audio to be detected for voice activity;

[0062] A2, each time 256ms of audio data is taken from the collected audio and appended to array A. When the current accumulated audio segment duration L reaches more than 4 times the single appended audio duration, that is, 1024 s, voice activity detection is performed, and a timestamp is added to each voice segment. The audio consists of several voice segments containing voice and silent segments without voice. The segment between two adjacent voice segments is a silent segment.

[0063] A3, the difference between the duration L of the audio segment and the end time of the last voice segment is diff,

[0064] If diff<1s, the audio data before the last voice segment is uploaded to the server, and the uploaded audio data in array A is deleted;

[0065] If diff>=1s, all the audio segment data are uploaded to the server and array A is cleared;

[0066] The Opus coding transmission bit rate of the silence segment is lower than that of the speech segment;

[0067] A4. If the client has not finished the collection work, continue to execute the loop from step A2.

[0068] Voice activity detection considers the last silent segment of the audio segment less than 1s as a continuous voice segment. Therefore, only when the last silent segment is greater than 1s will the entire audio segment be considered complete. The entire audio segment is uploaded to the server and array A is cleared, which is conducive to improving the integrity of the uploaded voice.

[0069] Embodiment 3:

[0070] This embodiment is based on embodiment 1:

[0071] Each time, 512ms of audio data is taken from the collected audio and appended to array A.

[0072] The silence segment is transmitted by Opus encoding at a bit rate of 8Khz; the voice segment is transmitted by Opus encoding at a bit rate of 16Khz.

[0073] Embodiment 4:

[0074] This embodiment is based on embodiment 1:

[0075] In the step A, based on the UDP transmission protocol, the UDP protocol data packet is encoded based on the forward error correction mechanism to obtain the encoded data block, and the encoded data block is interleaved to combat burst packet loss.

[0076] Forward Error Correction (FEC) Matrix Coding

[0077] A hybrid mode of Reed-Solomon (RS) coding and XOR check is used, and the coding parameters are defined as RS(n,k), where:

[0078] k = number of original packets

[0079] n = k + m (m is the number of redundant packets)

[0080] The generator matrix G is a Vandermonde matrix:

[0081]

[0082] Among them, α_i∈GF(2^8) Galois Field elements. The encoding process is:

[0083]

[0084] D is the original data matrix, and C is the set of encoded data packets.

[0085] Dynamic block interleaving: Define the interleaving depth L = 2^t (t∈[3,5]), and the data packets are arranged in the following matrix:

[0086]

[0087] The packets are sent in row order during transmission and reassembled in column order at the receiving end. The relationship between the packet loss rate δ and the recovery probability P is:

[0088]

[0089] Where m is the number of redundant packets.

[0090] Embodiment 5:

[0091] This embodiment is based on embodiment 1:

[0092] The A step described above uses key frame retransmission.

[0093] Implement triple redundancy for audio keyframes during retransmission:

[0094] Original frame F_t;

[0095] Differential frame ΔF_t = F_t ⊕ F_{t-1};

[0096] Check frame C_t = CRC32(F_t) || (F_t >> 8);

[0097] Here, || is a byte-level concatenation operation, CRC32(F_t) || (F_t >> 8) = [CRC32 checksum] + [original frame data after right shifting 8 bits].

[0098] The sliding window mechanism is adopted, the window size W=3, and the redundancy insertion period T=20ms, which satisfies:

[0099]

[0100]

[0101] Where t is the serial number of the audio frame, F_t is the audio data of the current frame, F_{t-1} is the audio data of the previous frame, and each t corresponds to a 20ms audio data block. For example, t=5 represents the audio frames from 100ms to 120ms. W∈[2,5] can be adjusted as needed, and the redundancy insertion period T∈[10ms,50ms].

[0102] CRC refers to cyclic redundancy check, ⊕ represents Galois field addition, is the packet loss rate threshold,

[0103] is the weight coefficient, the empirical value is [0.5, 0.3, 0.2].

[0104] Triple redundancy is implemented in the following situations:

[0105] 1. When the server requests the client to retransmit;

[0106] 2. The client detects that the weighted packet loss rate exceeds the threshold, and monitors the historical packet loss rate through the sliding window (reference formula). When the weighted packet loss rate exceeds the threshold When , redundancy insertion is performed on the current frame and unconfirmed frames in the window.

[0107] The traditional triple redundancy formula only determines whether to trigger redundancy insertion based on the packet loss rate at the current moment, which is too simple. The present invention improves it by using a sliding window and weighted evaluation. The advantages of the improvement are as follows:

[0108] 1. More accurate network status assessment: Through sliding windows and weighted assessment, the network status can be more comprehensively reflected, avoiding misjudgment caused by instantaneous fluctuations.

[0109] 2. Adaptive to complex network environments: When network conditions fluctuate greatly, the improved formula can more stably evaluate the overall network quality and reduce the frequency of redundant insertion.

[0110] 3. Improve system reliability: When redundant insertion is really needed (for example, when there is a continuous high packet loss rate), the redundant mechanism can be triggered in time to improve the reliability of data transmission.

[0111] Embodiment 6:

[0112] This embodiment is based on Embodiment 5:

[0113] In the step A, spiral check is used to detect residual errors during audio decoding: in audio data transmission, the sending end generates a spiral check byte for the data block and appends it to the end of the data block; the receiving end extracts the data block and recalculates the check byte to detect whether there is an error; when the receiving end detects a data error through the spiral check, it first attempts to locate and correct the error through the spiral check polynomial. If the spiral check cannot correct the error, the receiving end will mark the frame as a residual error and notify the sending end to transmit triple redundant information to recover the lost or damaged data and continue normal transmission.

[0114] Define the spiral check polynomial:

[0115]

[0116] in,

[0117] :Indicates the first A value of bytes, each byte is treated as a Galois field The elements in

[0118] :Position weight of the polynomial, corresponding to the position of the byte in the data stream,

[0119] : A reduced polynomial defined on GF(2) used to construct a 16-dimensional spiral space;

[0120] mod means modular operation;

[0121] Check byte generation algorithm:

[0122]

[0123] in,

[0124] is a prime number permutation table, ⊕ represents Galois Field addition, represents Galois Field multiplication,

[0125] Indicates the first j + 1 byte (the byte following the current byte), Indicates circular left shift Bit operations.

[0126] The improved spiral check polynomial of the present invention is improved by adding modular operation , significantly reducing the computational complexity and improving real-time performance. By introducing prime number permutation tables and Galois field multiplication, the randomness and anti-interference ability of the check bytes are improved, and the error correction capability is enhanced.

[0127] The results show that an effective recovery rate of 99.2% can be achieved at a packet loss rate of 15%.

[0128] Embodiment 7:

[0129] This embodiment is based on embodiment 1:

[0130] The server configures an array as an audio buffer. Each time an audio data block is received, the array is filled and placed after the previous data block. When the accumulated audio data in the buffer reaches a preset threshold, voice activity detection is performed on the audio stream to distinguish between voice segments and silent segments, and the audio stream is segmented into independent voice segments based on timestamps. The segmented voice segments are transcribed in real time in chronological order.

[0131] When segmenting, two adjacent voice segments of a silent segment with a duration of less than 1 second are regarded as a continuous voice segment, and the last voice segment after the audio stream is segmented is regarded as incomplete audio. The voice segment and the subsequent audio data are temporarily retained in the buffer and merged and processed after subsequent data arrives.

[0132] The present invention uses the silero_vad library to quickly call the silero_vad speech detection model to detect speech activity on the audio. In the speech transcription and translation process, the Whisper speech transcription model is used. In addition, a hybrid model can also be used to combine traditional statistical models (such as hidden Markov models HMM) and modern deep learning models (such as recurrent neural networks RNN or Transformer models) to improve the accuracy and robustness of transcription and translation.

[0133] Embodiment 8:

[0134] A system for implementing the low-latency real-time speech transcription method of the present invention comprises:

[0135] The client obtains the audio data, encodes and compresses the original audio data, and transmits it to the server in real time through the data transmission module;

[0136] The server receives the audio data, detects the voice in the audio using the voice activity detection module, segments the audio data judged to be voice, and transcribes the segmented audio data in real time in chronological order.

[0137] The server side calls the translation model or transmits it to a third-party translation software to translate the text information.

[0138] The above-mentioned embodiments only express the specific implementation of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.

Claims

1. A low-latency real-time speech transcription method, characterized in that , including the following steps: A. Encode and compress the original audio data collected by the client, and then transmit the obtained audio encoding data to the server through the real-time transport protocol, introduce error correction and packet loss retransmission mechanism to increase the reliability of transmission; B. The server uses a voice activity detection algorithm to detect speech in the audio data, splits the audio data judged to be speech into multiple independent parts, and adds a timestamp to each part; C. Transcribe the segmented audio data in real time in chronological order to generate corresponding text content; The step A comprises: A1. Create an array A to store the audio to be detected for voice activity; A2. Each time, 256ms-512ms of audio data is taken from the collected audio and appended to array A. If the current accumulated audio segment duration L reaches more than 4 times the duration of the single appended audio, voice activity detection is performed, and a timestamp is added to each voice segment. The audio consists of several voice segments containing voice and silent segments without voice. The segment between two adjacent voice segments is a silent segment. A3, the difference between the duration L of the audio segment and the end time of the last voice segment is diff, If diff<1s, the audio data before the last voice segment is uploaded to the server, and the uploaded audio data in array A is deleted; If diff>=1s, all the audio segment data are uploaded to the server and array A is cleared; The Opus coding transmission bit rate of the silence segment is lower than that of the speech segment; A4. If the client has not finished the collection work, continue to execute the loop from step A2.

2. The low-latency real-time speech transcription method according to claim 1, characterized in that In the step A, based on the UDP transmission protocol, the UDP protocol data packet is encoded based on the forward error correction mechanism to obtain the encoded data block, and the encoded data block is interleaved to combat burst packet loss.

3. The low-delay real-time speech transcription method according to claim 1, characterized in that ,The described A step adopts key frame retransmission, and implements triple redundancy for the key frames of the audio during retransmission: Original frame F_t; Differential frame ΔF_t=F_t⊕F_{t-1}; Check frame C_t = CRC32(F_t)||(F_t>>8); A sliding window mechanism is used with a window size of W∈[2,5] and a redundancy insertion period of T∈[10ms,50ms], satisfying: Where t is the serial number of the audio frame, F_t is the audio data of the current frame, and F_{t-1} is the audio data of the previous frame. CRC refers to cyclic redundancy check. represents Galois Field addition, θ is the packet loss rate threshold, W i is the weight coefficient.

4. The low-delay real-time speech transcription method according to claim 3, characterized in that ,Step A described above, uses spiral checksum to detect residual errors during audio decoding: in audio data transmission, the sender generates spiral checksum bytes for the data block and appends them to the end of the data block; the receiver extracts the data block and recalculates the checksum bytes to detect whether there is an error; when the receiver detects a data error through spiral checksum, it first attempts to locate and correct the error through the spiral checksum polynomial. If the spiral checksum cannot correct the error, the receiver will mark the frame as a residual error and notify the sender to transmit triple redundant information to recover the lost or damaged data.

5. The low-delay real-time speech transcription method according to claim 4, characterized in that , define the spiral check polynomial: in, h i ∈GF(2 8 ): represents the value of the i-th byte in the data block, and each byte is regarded as a Galois field GF(2 8 ), x i :Position weight of the polynomial, corresponding to the position of the byte in the data stream, x 16 +1: A reduced polynomial defined on GF(2) to construct a 16-dimensional spiral space. mod means modular operation; Check byte generation algorithm: in, π(i) is the prime number permutation table, represents Galois Field addition, represents Galois Field multiplication, b j+i It represents the j+1th byte in the input data block (the successor byte of the current byte), and <<i represents a circular left shift operation of i bits.

6. The low-latency real-time speech transcription method according to claim 1, characterized in that ,The server configures an array as an audio buffer. Every time an audio data block is received, the array is filled and placed after the previous data block. When the audio data in the buffer reaches the preset threshold, the audio data is ,detected for voice activity to distinguish voice segments from silence segments, and ,segmented into independent voice segments based on timestamps. The segmented voice segments are ,transcribed in real time in chronological order.

7. The low-delay real-time speech transcription method according to claim 6, characterized in that When segmenting, two adjacent voice segments of a silent segment with a duration of less than 1s are regarded as a continuous voice segment, and the last voice segment after the audio data is segmented is regarded as an incomplete audio. The voice segment and the subsequent audio data are temporarily retained in the buffer and merged and processed after the subsequent data arrives.

8. A system for implementing the low-delay real-time speech transcription method according to any one of claims 1 to 7, characterized in that: include: The client obtains the audio data, encodes and compresses the original audio data, and transmits it to the server in real time through the data transmission module; The server receives the audio data, detects the voice in the audio using the voice activity detection module, segments the audio data judged to be voice, and transcribes the segmented audio data in real time in chronological order.

9. The system of the low-latency real-time speech transcription method according to claim 8, characterized in that: The server side calls the translation model or transmits it to a third-party translation software to translate the text information.

Citation Information

Patent Citations

  • Voice wake-up method and device, storage medium and electronic equipment

    CN115831109A

  • Speech recognition method and device, equipment and storage medium

    CN115938397A

  • Low-delay real-time voice-to-text and text-to-voice transmission method

    CN118865942A