Voice stream segmentation method based on double-model dynamic triggering

Through the dual-model dynamically triggered speech flow slicing method, combined with fast segmentation and high-precision segmentation model, the problem of speech flow slicing efficiency and accuracy in high concurrent scenarios is solved, and efficient and accurate speech recognition results are achieved, suitable for scenarios such as intelligent customer service and conference transcription.

CN120260546AActive Publication Date: 2025-07-04THE 28TH RES INST OF CHINA ELECTRONICS TECH GROUP CORP

Patent Information

Application Number
CN202510726884.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-04
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The prior art is difficult to take into account the processing efficiency and accuracy of speech stream segmentation in high concurrency scenarios. The sliding window method has a high misjudgment rate in a low signal-to-noise ratio environment, while the high computational complexity of the end-to-end deep learning model leads to a decrease in the long speech processing efficiency.

Method used

The speech stream slicing method based on dual-model dynamic triggering is adopted. By constructing a data stream buffer management mechanism for multiple voice streams, the fast slicing model is used to filter and output it to a high-precision slicing model, and combined with dynamic adjustment of buffer thresholds, efficient slicing of voice fragments is achieved.

Benefits of technology

Under the guarantee of high-precision speech recognition results, the demand for computing resources is significantly reduced, and is suitable for multi-channel speech stream recognition scenarios, especially intelligent customer service and conference transcription programs that process real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260546A_ABST
    Figure CN120260546A_ABST
Patent Text Reader

Abstract

The invention discloses a voice stream segmentation method based on double-model dynamic triggering, and the method comprises the following steps: 1, constructing a data stream buffer management mechanism of multiple voice streams, building an independent processing channel for each voice stream, and enabling voice data accumulated to a threshold duration to form a to-be-processed voice set; 2, screening, analyzing and processing a to-be-processed voice set through the rapid segmentation model, and selecting voice segments meeting conditions and outputting the selected voice segments to a high-precision segmentation model; step 3, according to a screening result of the rapid segmentation model, splicing data which does not meet conditions with data in the data stream buffer, and adjusting a threshold duration of a buffer area corresponding to the voice segment; 4, processing the voice segments screened by the rapid segmentation model by using a high-precision segmentation model; and step 5, outputting the segmented audio segments to a voice recognition system and other systems according to a processing result, splicing the remaining data with the data in the data stream buffer, and updating the threshold duration of the corresponding buffer area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech processing, and in particular to a speech stream segmentation method based on dual-model dynamic triggering. Background Art

[0002] In the field of speech processing, speech recognition models are divided into two types: streaming and non-streaming. The recognition accuracy of the latter is much higher than that of the former. In scenarios where speech recognition of a speech stream is required, the speech stream can be segmented into multiple speech segments by combining speech stream segmentation technology, so that the non-streaming speech recognition model can be used to achieve accurate speech recognition and provide accurate text information for subsequent processing. As a basic technology in the field of speech processing, speech stream segmentation technology has important application value in scenarios such as intelligent customer service systems, multi-party conference transcription, and real-time speech analysis. With the exponential growth of real-time speech processing requirements, the existing technology faces a key technical bottleneck that it is difficult to balance processing efficiency and segmentation accuracy in high-concurrency scenarios.

[0003] The current mainstream technical solutions have the following technical defects: 1. The speech activity detection scheme based on a sliding window adopts an energy detection method with a fixed threshold: Although this scheme has the advantage of millisecond-level real-time performance, it is highly sensitive to environmental noise, and the misjudgment rate exceeds 35% in low signal-to-noise ratio scenarios; 2. The end-to-end deep learning model solution: Although a segmentation accuracy of more than 90% is achieved through a neural network model, its computational complexity results in a linear increase in model inference time consumption with the speech duration, and the processing efficiency of long speech drops sharply. Summary of the Invention

[0004] Object of the Invention: The technical problem to be solved by the present invention is to provide a speech stream segmentation method based on dual-model dynamic triggering in view of the deficiencies of the prior art.

[0005] To solve the above technical problem, the present invention discloses a speech stream segmentation method based on dual-model dynamic triggering, including the steps of:

[0006] Step 1: Construct a data stream buffer management mechanism for multiple speech streams, establish an independent processing channel for each speech stream, and form a set of speech data to be processed by accumulating speech data up to a threshold duration;

[0007] Step 2: Screen, analyze, and process the set of speech data to be processed through a fast segmentation model, and select eligible speech segments and output them to a high-precision segmentation model;

[0008] Step 3: According to the screening results of the fast segmentation model, splice the ineligible data with the data in the data stream buffer, and adjust the threshold duration of the corresponding buffer of the speech segment;

[0009] Step 4: Process the speech segments screened by the fast segmentation model using the high-precision segmentation model;

[0010] Step 5: Output the segmented audio segments to other systems such as speech recognition based on the processing results, concatenate the remaining data with the data in the data stream buffer, and update the threshold duration of the corresponding buffer.

[0011] The data stream buffer management mechanism for constructing multiple voice streams described in step 1 includes the following steps:

[0012] Step 1-1: Create an independent data buffer structure for each voice stream. The data buffer structure of the kth voice stream includes: the audio data that has been received and resampled to a sampling rate of 16000 Hz , the duration of the audio received (in seconds), threshold duration (in seconds);

[0013] Step 1-2: After receiving the new data, store it in the data buffer structure of the corresponding voice stream, resample the new audio data to a sampling rate of 16000Hz and append it to the audio data After that, update the audio duration ;

[0014] Step 1-3: Compare audio duration and threshold duration ,when , all the voice data in the buffer structure Output to the speech collection to be processed , otherwise execute steps 1-2.

[0015] The fast segmentation model screening described in step 2 includes the following steps:

[0016] Step 2-1: Check the speech collection to be processed Is it empty? If it is empty, wait for the next check, otherwise execute step 2-2;

[0017] Step 2-2: From the speech collection to be processed Select the speech segment of the kth speech stream , and its last last_second( The data of 10 seconds is divided into num_frame frames according to the sliding window with a frame length of 30ms and a frame shift of 10ms. The spectrum features of each frame are extracted and input into the fast segmentation model to obtain the probability list of whether each frame is a silent frame. , statistical probability list middle , the ratio of frames greater than 0.5, when this ratio exceeds 0.4, the speech segment is judged Eligible.

[0018] In step 2-2, the data of its last last_second seconds is divided into num_frame frames according to a sliding window with a frame length of 30 ms and a frame shift of 10 ms, and the spectral features of each frame are extracted as follows:

[0019] Step 2-2-1: Apply a Hamming window to the last_second seconds of speech data for frame segmentation. The frame length is 30 ms and the frame shift is 10 ms, and a total of num_frame ( frames are obtained. The signal amplitude of each frame is , where i represents the i-th frame, , r represents the r-th sampling point, ;

[0020] Step 2-2-2: Perform a fast Fourier transform on each frame of speech data, calculate the energy spectrum, and obtain the speech frequency domain signal and the energy spectrum , where i represents the i-th frame, , k represents the k-th frequency point, , r represents the r-th sampling point, ;

[0021] Step 2-2-3: Pass the energy spectrum through a Mel filter bank to obtain M-dimensional Mel spectral features , where represents the band-pass triangular filter of the Mel filter bank, M represents the number of filters, and m represents the m-th filter;

[0022] Step 2-2-4: Take the logarithm of the Mel spectral features and perform a discrete cosine transform to obtain the Mel frequency cepstral coefficients (MFCC coefficients), , where n is the frequency point after DCT. For simplicity, the value range of n is the same as that of m, that is, the spectral features of the i-th frame.

[0023] The fast segmentation model described in step 2-2 is a lightweight binary classification model based on a one-dimensional convolutional neural network. Its network structure includes:

[0024] Input layer: Receive ×M-dimensional MFCC feature matrix, where M is the number of Mel filters;

[0025] One-dimensional convolutional layer: Use 32 convolutional kernels with a width of 5 and a stride of 1 to perform one-dimensional convolution along the time axis, and the output dimension is ×32;

[0026] Max pooling layer: The pooling window size is 2 and the stride is 2, and the output dimension is ×32;

[0027] Flatten layer: Flatten the feature map into a 1D vector;

[0028] Fully connected layer: Through a fully connected layer with

[0029] Output layer: Output a single-node probability value through the Sigmoid activation function, representing the probability that each frame of the input frame is not a silent frame.

[0030] Step 3 includes the following steps:

[0031] Step 3-1: Statistically calculate the proportion of frames in the probability list whose silent probability is greater than 0.5. When this proportion does not exceed 0.4, select the kth voice stream corresponding to the voice segment and concatenate the voice segment with the latest received data in the corresponding voice stream buffer at the head and tail to form new buffered data ; ;

[0032] Step 3-2: Update the threshold duration = + 0.3 according to the formula to avoid repeatedly triggering the fast segmentation model screening step.

[0033] The network structure of the high-precision segmentation model described in Step 4 is as follows:

[0034] Input layer: Receive the MFCC feature sequence of the entire speech, with a dimension of T×M, where T is the number of frames obtained after the entire speech passes through a sliding window with a frame length of 30 ms and a frame shift of 10 ms, and M is the number of Mel filters;

[0035] Two bidirectional LSTM layers: A bidirectional LSTM layer with 128 hidden units, and the output dimension is T×256;

[0036] Fully connected layer: A fully connected layer with 64 neurons, and the activation function is ReLU;

[0037] Output layer: Output a T-dimensional probability sequence through the Sigmoid activation function, representing the probability that each time frame is not a silent frame;

[0038] Boundary decision module: Determine as a segmentation point when the probability value is greater than 0.7 and is a local maximum.

[0039] The specific steps of Step 4 include the following:

[0040] Step 4-1: Extract complete MFCC features from the input voice segment, with a frame length of 30 ms and a frame shift of 10 ms;

[0041] Step 4-2: Obtain the boundary probability sequence through a high-precision segmentation model;

[0042] Step 4-3: Segment at the probability peak points to obtain a set of audio segments with someone speaking , where k represents the k-th voice stream, L represents the total number of audio segments included in the set of audio segments with someone speaking. If no one speaking is detected in the input voice segment, an empty set is output.

[0043] Step 5 includes the following steps:

[0044] Step 5-1: Check whether the set of audio segments with someone speaking output by the high-precision segmentation model is empty:

[0045] If it is not empty, send each set of audio segments (i = 1, 2,..., L) to the speech recognition system in sequence and execute Step 5-2;

[0046] If it is empty, directly execute Step 5-3;

[0047] Step 5-2: Extract the remaining data in the original input voice segment after , concatenate with the latest received data in the corresponding voice stream buffer at the head and tail to form new buffered data , and update the audio duration ; ;

[0048] Step 5-3: Update the threshold duration :

[0049] When it is not empty, update according to the formula = max(2.0, average_duration + 0.5), where average_duration is the average duration of each audio segment in the set of audio segments ;

[0050] When it is empty, the threshold duration remains unchanged;

[0051] Step 5-4: Reset the audio duration to the actual duration of the current concatenated buffer.

[0052] The fast segmentation model has been trained; the high-precision segmentation model has been trained.

[0053] Beneficial effects:

[0054] 1. The present invention proposes a dual-model dynamic trigger mechanism to achieve a balance between efficiency and accuracy through the dynamic cooperation of a fast detection model and a high-precision model.

[0055] 2. The present invention dynamically adjusts the buffer threshold in combination with the dual-model segmentation results of the audio stream to effectively suppress invalid triggers.

[0056] This model can efficiently segment multiple voice streams while ensuring a high segmentation accuracy; it can provide a method for achieving high-precision speech recognition results in the scenario of recognizing multiple voice streams, cooperating with a non-streaming speech recognition model, and providing auxiliary support for intelligent customer service and conference transcription programs, having certain practical value. Brief description of the drawings

[0057] Figure 1 It is a schematic diagram of the overall process of the present invention.

[0058] Figure 2 It is a schematic diagram of the fast segmentation model.

[0059] Figure 3 It is a schematic diagram of the high-precision segmentation model.

[0060] Figure 4 It is a schematic diagram of the data flow in the actual system of the present invention. Detailed implementation manners

[0061] A method for segmenting a voice stream based on dual-model dynamic trigger (such as Figure 1 ) includes the following steps:

[0062] Step 1: Construct a data flow buffer management mechanism for multiple voice streams, establish an independent processing channel for each voice stream, and form a set of voice data to be processed by accumulating voice data up to the threshold duration;

[0063] Step 2: Screen, analyze, and process the set of voice data to be processed through the fast segmentation model, and select the voice segments that meet the conditions and output them to the high-precision segmentation model;

[0064] Step 3: According to the screening results of the fast segmentation model, splice the data that does not meet the conditions with the data in the data flow buffer, and adjust the threshold duration of the corresponding buffer of the voice segment;

[0065] Step 4: Use the high-precision segmentation model to process the voice segments screened by the fast segmentation model;

[0066] Step 5: Output the segmented audio segments to other systems such as speech recognition according to the processing results, splice the remaining data with the data in the data stream buffer, and update the threshold duration of the corresponding buffer.

[0067] The data stream buffer management mechanism for constructing a multi-channel voice stream described in Step 1 includes the following steps:

[0068] Step 1-1: Establish an independent data buffer structure for each voice stream. The data buffer structure of the k-th voice stream includes: audio data that has been received and resampled to a sampling rate of 16,000 Hz , the received audio duration (in seconds), the threshold duration (in seconds);

[0069] Step 1-2: After receiving new data, store it in the data buffer structure of the corresponding voice stream. Resample the new audio data to a sampling rate of 16,000 Hz and append it to the audio data and update the audio duration ;

[0070] Step 1-3: Compare the audio duration and the threshold duration . When , output all the voice data in the buffer structure to the set of voices to be processed , otherwise execute Step 1-2.

[0071] The fast segmentation model screening described in Step 2 includes the following steps:

[0072] Step 2-1: Check whether the set of voices to be processed is empty. If it is empty, wait for the next check; otherwise, execute Step 2-2;

[0073] Step 2-2: Select the voice segment of the k-th voice stream from the set of voices to be processed . Divide the last last_second( seconds of data into num_frame frames according to a sliding window with a frame length of 30 ms and a frame shift of 10 ms. Extract the spectral features of each frame and input them into the fast segmentation model to obtain a list of probabilities indicating whether each frame is a silent frame . Count the proportion of frames in the probability list where is greater than 0.5. When this proportion exceeds 0.4, it is determined that the voice segment meets the conditions. Meets the conditions.

[0074] In step 2-2, the data in its last last_second seconds is divided into num_frame frames according to a sliding window with a frame length of 30 ms and a frame shift of 10 ms, and the spectral features of each frame are extracted as follows:

[0075] Step 2-2-1: Apply a Hamming window to the last_second seconds of speech data for frame processing, with a frame length of 30 ms and a frame shift of 10 ms, and a total of num_frame ( frames are obtained. The signal amplitude of each frame is , where i represents the i-th frame, , r represents the r-th sampling point, ;

[0076] Step 2-2-2: Perform a fast Fourier transform on each frame of speech data, calculate the energy spectrum, and obtain the speech frequency domain signal and the energy spectrum , where i represents the i-th frame, , k represents the k-th frequency point, , r represents the r-th sampling point, ;

[0077] Step 2-2-3: Pass the energy spectrum through a Mel filter bank to obtain M-dimensional Mel spectral features , where represents the band-pass triangular filter of the Mel filter bank, M represents the number of filters, and m represents the m-th filter;

[0078] Step 2-2-4: Take the logarithm of the Mel spectral features and perform a discrete cosine transform to obtain the Mel frequency cepstral coefficients (MFCC coefficients), , where n is the frequency point after DCT. For simplicity, the value range of n is required to be the same as that of m, that is, the spectral features of the i-th frame.

[0079] The fast segmentation model described in step 2-2 is a lightweight binary classification model based on a one-dimensional convolutional neural network (such as Figure 2 ), and its network structure includes:

[0080] Input layer: Receive ×M-dimensional MFCC feature matrix, where M is the number of Mel filters;

[0081] One-dimensional convolutional layer: Use 32 convolutional kernels with a width of 5 and a stride of 1 to perform one-dimensional convolution along the time axis, and the output dimension is ×32;

[0082] Max pooling layer: The pooling window size is 2 and the stride is 2, and the output dimension is ×32;

[0083] Flatten layer: Flatten the feature map into a 1D vector;

[0084] Fully connected layer: Through a fully connected layer with

[0085] Output layer: Output a single-node probability value through the Sigmoid activation function, representing the probability that each frame of the input frame is not a silent frame.

[0086] Step 3 includes the following steps:

[0087] Step 3-1: Calculate the proportion of frames in the probability list whose silent probability is greater than 0.5 in the probability list. When this proportion does not exceed 0.4, select the k-th speech stream corresponding to the speech segment and concatenate the speech segment with the latest received data in the corresponding speech stream buffer at the head and tail to form new buffered data ; ;

[0088] Step 3-2: Update the threshold duration = + 0.3 according to the formula to avoid repeatedly triggering the fast segmentation model screening step. ;

[0089] The network structure of the high-precision segmentation model described in Step 4 (such as Figure 3 ) is as follows:

[0090] Input layer: Receive the MFCC feature sequence of the entire speech, with a dimension of T×M, where T is the number of frames obtained after the entire speech passes through a sliding window with a frame length of 30 ms and a frame shift of 10 ms, and M is the number of Mel filters;

[0091] Two bidirectional LSTM layers: A bidirectional LSTM layer with 128 hidden units, and the output dimension is T×256;

[0092] Fully connected layer: A fully connected layer with 64 neurons, and the activation function is ReLU;

[0093] Output layer: Output a T-dimensional probability sequence through the Sigmoid activation function, representing the probability that each time frame is not a silent frame;

[0094] Boundary decision module: Determine as a segmentation point when the probability value is greater than 0.7 and is a local maximum.

[0095] The specific steps of Step 4 include the following:

[0096] Step 4-1: Extract complete MFCC features from the input voice segment, with a frame length of 30 ms and a frame shift of 10 ms;

[0097] Step 4-2: Obtain the boundary probability sequence through a high-precision segmentation model;

[0098] Step 4-3: Perform segmentation at the probability peak points to obtain a set of audio segments with people speaking , where k represents the k-th voice stream, L represents the total number of audio segments included in the set of audio segments with people speaking. If no one speaking is detected in the input voice segment, an empty set is output.

[0099] Step 5 includes the following steps:

[0100] Step 5-1: Check whether the set of audio segments with people speaking output by the high-precision segmentation model is empty:

[0101] If it is not empty, send each set of audio segments (i = 1, 2,..., L) to the speech recognition system in sequence (such as Figure 4 ), and execute Step 5-2;

[0102] If it is empty, directly execute Step 5-3;

[0103] Step 5-2: Extract the remaining data in the original input voice segment after , concatenate with the latest received data in the corresponding voice stream buffer at the head and tail to form new buffered data , and update the audio duration ; ;

[0104] Step 5-3: Update the threshold duration :

[0105] When it is not empty, update according to the formula = max(2.0, average_duration + 0.5), where average_duration is the average duration of each audio segment in the set of audio segments ;

[0106] When it is empty, the threshold duration remains unchanged;

[0107] Step 5-4: Reset the audio duration to the actual duration of the current concatenated buffer.

[0108] The fast segmentation model has been trained; the high-precision segmentation model has been trained.

[0109] Embodiment:

[0110] This embodiment takes the air traffic control voice recorder system for land-air communication as an example to illustrate the application scenario of the method described in the present invention in the real-time segmentation of multi-channel voice streams. Combining Figure 1 with the schematic diagram shown, the specific implementation steps are as follows:

[0111] Step 1: Establish a multi-channel voice stream buffer management mechanism:

[0112] The system establishes an independent buffer structure for each radio channel, and defines the structure of the k-th voice stream as:

[0113] {

[0114] 'audio_buffer': np.array([]), / / Audio data resampled to 16000Hz

[0115] 'audio_len': 0.0, / / Cumulative duration (seconds)

[0116] 'threshold': 2.0 / / Initial trigger threshold

[0117] }

[0118] When new audio data arrives, execute: resample to a sampling rate of 16000Hz; append the data to the corresponding audio_buffer; update audio_len = len(audio_buffer) / 16000; when audio_len > threshold, pack the complete buffer data into the set of data to be processed audio_set.

[0119] Step 2: Dynamically trigger the fast segmentation model (corresponding to Figure 2 ):

[0120] a. Extract the audio of the last 1.5 seconds for feature analysis:

[0121] Frame the audio with a frame length of 30ms and a frame shift of 10ms

[0122] Calculate 13-dimensional MFCC features (M = 13)

[0123] b. Perform silence detection through a lightweight model:

[0124] c. Trigger high-precision processing when the proportion of silence frames ≤ 40%

[0125] Step 3: Buffer Dynamic Adjustment:

[0126] Not triggered: Concatenate the current data with the previous buffer and update the threshold formula: threshold_new = current_len + 0.3. Example: When the original buffer duration of 2.1 seconds is not triggered, the new threshold is set to 2.1 + 0.3 = 2.4 seconds

[0127] 4. High-precision Speech Segmentation (corresponding to Figure 3 )

[0128] Process the complete audio segment using a bidirectional LSTM model:

[0129] Step 5: Result Output and System Linkage (corresponding to Figure 4 )

[0130] Push the valid speech segments to the speech recognition engine and write the remaining data back to the buffer.

[0131] In this embodiment, through the collaborative work of a two-stage model (such as the process shown in Figure 1 ), while ensuring high-precision segmentation, the computational resource requirements are significantly reduced compared to traditional single-model solutions, and it is particularly suitable for scenarios such as air traffic control that require real-time processing of multiple voice streams.

[0132] The present invention provides a method for segmenting a voice stream based on dynamic triggering of a dual model. There are many methods and ways to specifically implement this technical solution. The above description is only a preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented using existing technologies.

Claims

1. A speech stream segmentation method based on dual-model dynamic triggering, characterized in that, It includes the following steps: Step 1: Construct a data stream buffer management mechanism for multi-channel voice streams, establish an independent processing channel for each voice stream, and form a set of voice data to be processed by accumulating voice data up to the threshold duration; Step 2: Screen, analyze and process the set of voice data to be processed through a fast segmentation model, and select eligible voice segments and output them to a high-precision segmentation model; Step 3: According to the screening results of the fast segmentation model, splice the ineligible data with the data in the data stream buffer, and adjust the threshold duration of the corresponding buffer of the voice segment; Step 4: Process the voice segments screened by the fast segmentation model using a high-precision segmentation model; Step 5: Output the segmented audio segments to the speech recognition system according to the processing results, splice the remaining data with the data in the data stream buffer, and update the threshold duration of the corresponding buffer.

2. The method for segmenting a speech stream based on dual-model dynamic triggering according to claim 1, wherein The construction of the data stream buffer management mechanism for multi-channel voice streams described in Step 1 includes the following steps: Step 1-1: Establish an independent data buffer structure for each voice stream. The data buffer structure of the k-th voice stream includes: the audio data that has been received and resampled , the received audio duration , the threshold duration ; Step 1-2: After receiving new data, store it in the data buffer structure corresponding to the voice stream, resample the new audio data and append it to the audio data After that, update the audio duration ; Step 1-3: Compare the audio duration with the threshold duration , when , output all the voice data in the buffer structure to the voice set to be processed , otherwise execute Step 1-2.

3. A method for segmenting a speech stream based on dual-model dynamic triggering according to claim 1, characterized in that, The screening of the fast segmentation model described in Step 2 includes the following steps: Step 2-1: Check the set of voices to be processed to see if it is empty. If it is empty, wait for the next check; otherwise, execute Step 2-2; Step 2-2: Select the speech segment of the k-th speech stream from the speech set to be processed and divide the data of its last last_second seconds into num_frame frames through a sliding window, extract the spectral features of each frame, input them into the fast segmentation model, and obtain a probability list of whether each frame is a silent frame Then, count the probability list to determine whether the speech segment meets the conditions in it ​ 4. The voice stream segmentation method based on dual-model dynamic triggering according to claim 3, wherein, Regarding the data in its last last_second seconds in Step 2-2, divide the data into num_frame frames through a sliding window, and extract the spectral features of each frame, specifically as follows: Step 2-2-1: Frame the voice data of the last_second seconds. A total of num_frame frames are obtained, and the signal amplitude of each frame is , where i represents the i-th frame, , and r represents the r-th sampling point; Step 2-2-2: Perform a fast Fourier transform on each frame of speech data, calculate the energy spectrum, and obtain the speech frequency-domain signal of each frame and the energy spectrum , where i represents the i-th frame, , k represents the k-th frequency point, , and r represents the r-th sampling point; Step 2-2-3: Pass the energy spectrum through a Mel filter bank to obtain M-dimensional Mel spectrum features , where represents the band-pass triangular filter of the Mel filter bank, M represents the number of filters, and m represents the m-th filter; Step 2-2-4: Take the logarithm of the Mel-spectrum features and perform discrete cosine transform to obtain Mel-frequency cepstral coefficients , , where n is the frequency point after DCT, and the value range of n is the same as that of m, that is, the spectral features of the i-th frame.

5. The method for segmenting a speech stream based on dual-model dynamic triggering according to claim 3, wherein The fast segmentation model described in Step 2-2 is a lightweight binary classification model based on a one-dimensional convolutional neural network, and its network structure includes: Input layer: Receive ×MFCC feature matrix of M dimensions, where M is the number of Mel filters; 1D Convolutional Layer: Using 32 convolutional kernels with a width of 5 and a stride of 1, perform 1D convolution along the time axis, and the output dimension is ×32; Max pooling layer: the pooling window size is 2, the stride is 2, and the output dimension is ×32; Flattening layer: Flatten the feature map into a 1D vector; Fully connected layer: Through fully connected layer of neurons, with the activation function ReLU; Output layer: The single-node probability value is output through the Sigmoid activation function, representing the probability that each frame of the input frame is not a silent frame.

6. The voice stream segmentation method based on dual-model dynamic triggering according to claim 1, wherein Step 3 includes the following steps: Step 3-1: Statistic probability list probability list The proportion of frames with a silent probability greater than 0.5 in When this proportion does not exceed 0.4, select the kth voice stream corresponding to the voice segment Concatenate the voice segment with the latest received data in the corresponding voice stream buffer at the head and tail to form new buffered data ; Step 3-2: According to the formula = + 0.3 to update the threshold duration .

7. A method for segmenting a speech stream based on dual-model dynamic triggering according to claim 1, characterized in that The network structure of the high-precision segmentation model described in Step 4 is as follows: Input layer: Receive the MFCC feature sequence of the entire voice segment, with a dimension of T×M, where T is the number of frames obtained after the entire voice segment passes through a sliding window with a frame length of 30 ms and a frame shift of 10 ms, and M is the number of Mel filters; Two bidirectional LSTM layers: A bidirectional LSTM layer with 128 hidden units, and the output dimension is T×256; Fully connected layer: A fully connected layer with 64 neurons, and the activation function is ReLU; Output layer: Output a T-dimensional probability sequence through the Sigmoid activation function, indicating the probability that each time frame is not a silent frame; Boundary decision module: Determine as a segmentation point when the probability value is greater than 0.7 and is a local maximum.

8. A method for segmenting speech streams based on dual-model dynamic triggering according to claim 7, characterized in that The specific steps of Step 4 include the following: Step 4-1: Extract complete MFCC features from the input voice segment; Step 4-2: Obtain a boundary probability sequence through a high-precision segmentation model; Step 4-3: Perform segmentation at the probability peak point to obtain a set of audio segments with someone speaking , where k represents the k-th voice stream, L represents the total number of audio segments included in the set of audio segments with someone speaking. If no one speaking is detected in the input audio segment, an empty set is output.

9. A method for speech stream segmentation based on dual-model dynamic triggering according to claim 1, characterized in that Step 5 includes the following steps: Step 5-1: Check whether the set of audio segments with people speaking output by the high-precision segmentation model is empty: If the set of audio segments is non-empty, then send each set of audio segments (i = 1, 2,..., L) to the speech recognition system in sequence and execute step 5-2; If the set of audio segments is empty, directly execute step 5-3; Step 5-2: Extract the original input voice segment from the audio segment set and the remaining data thereafter are concatenated end-to-end with the latest received data in the corresponding voice stream buffer to form new buffered data and the audio duration is updated ; Step 5-3: Update the threshold duration : When the set of audio segments is non-empty, update according to the formula = max(2.0, average_duration + 0.5), where average_duration is the average duration of each audio segment in the set of audio segments ; When the set of audio segments is empty, the threshold duration remains unchanged; Step 5-4: Reset the audio duration is the actual duration of the current buffer after splicing.

10. A method for segmenting a speech stream based on dual-model dynamic triggering according to claim 1, characterized in that, The fast segmentation model has been trained; the high-precision segmentation model has been trained.

Citation Information

Patent Citations

  • Audio analysis system based on content

    CN101021854A

  • Deep learning-based unusual speech distinguishing method

    CN108766419A

  • Synthetic speech detection method based on speech segmentation

    CN113012684A

  • End-to-end speech recognition method based on fusion neural network structure

    CN114187898A

  • Grounding knife switch opening and closing sound recognition method and device

    CN114283792A

Cited By

  • Digital human smooth image generation method and system based on dynamic threshold triggering

    CN121616673A