Audio segmentation method and device, electronic equipment and storage medium
By extracting acoustic features from audio for semantic boundary annotation and segmentation, the semantic integrity problem of existing speech translation systems is solved, achieving more accurate and reliable audio segmentation, which is suitable for a variety of application scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HKUST IFLYTEK (SHANGHAI) TECH CO LTD
- Filing Date
- 2022-12-27
- Publication Date
- 2026-04-17
AI Technical Summary
Existing speech translation systems using cascading and end-to-end methods suffer from poor semantic integrity during audio segmentation, failing to effectively utilize tone and pause information in the audio for semantic sentence segmentation.
By extracting acoustic features from the audio to be segmented, semantic boundary sequence annotation is performed based on the acoustic features. Frames with strong and weak semantic boundaries are identified as candidate frames. Segmentation is then performed according to the segmentation application scenario, and the semantic boundary labels obtained by the boundary annotation model are used for audio segmentation.
It improves the accuracy and reliability of audio segmentation, preserves the complete semantic information of the audio, expands the application scope of audio segmentation, and is suitable for cascaded and end-to-end speech translation systems.
Smart Images

Figure CN116013263B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing technology, and in particular to an audio segmentation method, apparatus, electronic device, and storage medium. Background Technology
[0002] From a technical perspective, speech translation can be divided into two types: cascaded and end-to-end. Cascaded translation refers to audio being processed by speech recognition and machine translation engines separately to obtain the final translation result, while end-to-end translation refers to speech being directly processed by a single speech translation model to obtain the translation result.
[0003] In existing technologies, cascaded speech translation systems typically perform speech recognition first, then re-type punctuation based on the recognized punctuation or a punctuation model for sentence segmentation. This approach suffers from recognition and punctuation errors, hindering effective semantic segmentation. Furthermore, it fails to utilize intonation and pauses in the audio to aid semantic segmentation, thus limiting the semantic segmentation capabilities of cascaded speech translation systems. End-to-end speech translation systems, which generally use VAD (Voice Activity Detection) to identify empty sounds in the audio and segment the audio based on these empty sounds, or segment the audio using a fixed length, result in even poorer semantic integrity. Summary of the Invention
[0004] This invention provides an audio segmentation method, apparatus, electronic device, and storage medium to address the shortcomings of poor semantic integrity in audio segmentation of existing cascaded and end-to-end speech translation systems.
[0005] This invention provides an audio segmentation method, comprising:
[0006] Get the audio to be segmented;
[0007] The acoustic features of each frame in the audio to be segmented are extracted, and based on the acoustic features of each frame, the semantic boundary sequence is labeled to obtain the semantic boundary labeling results of each frame.
[0008] Based on the semantic boundary annotation results of each frame, the audio to be segmented is segmented.
[0009] According to an audio segmentation method provided by the present invention, the segmentation of the audio to be segmented based on the semantic boundary annotation results of each frame includes:
[0010] Based on the semantic boundary annotation results of each frame, candidate frames are determined from each frame;
[0011] Based on the candidate frames, the audio to be segmented is segmented.
[0012] According to an audio segmentation method provided by the present invention, determining candidate frames from the frames based on the semantic boundary annotation results of each frame includes:
[0013] Based on the segmentation application scenario, frames whose speech boundary annotation results are strong semantic boundaries and / or weak semantic boundaries are identified as candidate frames.
[0014] According to an audio segmentation method provided by the present invention, the step of determining frames whose speech boundary annotation results are strong semantic boundaries and / or weak semantic boundaries as candidate frames based on the segmentation application scenario includes:
[0015] In the case where the segmentation application scenario is the first scenario, the frames whose speech boundary annotation results are strong semantic boundaries are determined as candidate frames;
[0016] In the case where the segmentation application scenario is the second scenario, frames with strong semantic boundaries and frames with weak semantic boundaries are identified as candidate frames.
[0017] The semantic integrity requirement of the first scenario is greater than that of the second scenario, while the real-time requirement of the second scenario is greater than that of the first scenario.
[0018] According to an audio segmentation method provided by the present invention, the step of segmenting the audio to be segmented based on the candidate frames includes:
[0019] Obtain the frame number of the candidate audio segment in the audio to be segmented, wherein the candidate audio segment consists of consecutive candidate frames;
[0020] Based on candidate audio segments with a frame count greater than a preset threshold, the audio to be segmented is segmented.
[0021] According to an audio segmentation method provided by the present invention, the step of extracting acoustic features of each frame in the audio to be segmented, and performing semantic boundary sequence annotation on the audio to be segmented based on the acoustic features of each frame to obtain the semantic boundary annotation results of each frame includes:
[0022] Based on the boundary labeling model, the acoustic features of each frame in the audio to be segmented are extracted, and the acoustic features of each frame are applied to perform semantic boundary sequence labeling on the audio to be segmented, so as to obtain the semantic boundary labeling results of each frame.
[0023] The boundary labeling model is trained based on sample audio and semantic boundary labels of each frame in the sample audio. The semantic boundary labels are determined based on the punctuation marks of the text corresponding to the sample audio.
[0024] According to an audio segmentation method provided by the present invention, the semantic boundary label is any one of non-boundary, strong semantic boundary, and weak semantic boundary, wherein the strong semantic boundary corresponds to the sentence-ending symbol in the punctuation marks, and the weak semantic boundary corresponds to the sentence-intermediate symbol in the punctuation marks.
[0025] The present invention also provides an audio segmentation device, comprising:
[0026] The acquisition unit is used to acquire the audio to be segmented;
[0027] The semantic boundary labeling unit is used to extract the acoustic features of each frame in the audio to be segmented, and to perform semantic boundary sequence labeling on the audio to be segmented based on the acoustic features of each frame, so as to obtain the semantic boundary labeling results of each frame.
[0028] The audio segmentation unit is used to segment the audio to be segmented based on the semantic boundary annotation results of each frame.
[0029] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the audio segmentation methods described above.
[0030] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the audio segmentation method as described above.
[0031] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the audio segmentation methods described above.
[0032] The audio segmentation method, apparatus, electronic device, and storage medium provided by this invention extract the acoustic features of each frame in the audio to be segmented, and perform semantic boundary sequence annotation on the audio to be segmented based on the acoustic features of each frame to obtain the semantic boundary annotation results of each frame. Then, based on the semantic boundary annotation results of each frame, the audio to be segmented is segmented. Thus, the tone and pause information in the acoustic features of each frame can be used to assist semantic sentence segmentation, preserving the complete semantic information of the audio and avoiding punctuation recognition errors, thereby improving the accuracy and reliability of audio segmentation. Furthermore, this method can be applied to cascaded and end-to-end speech translation systems, expanding the application scope of audio segmentation. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0034] Figure 1 This is one of the flowcharts illustrating the audio segmentation method provided by the present invention;
[0035] Figure 2 This is the second flowchart illustrating the audio segmentation method provided by the present invention;
[0036] Figure 3 This is a flowchart illustrating step 130 in the audio segmentation method provided by the present invention;
[0037] Figure 4 This is a flowchart illustrating step 210 in the audio segmentation method provided by the present invention;
[0038] Figure 5 This is a flowchart illustrating step 132 in the audio segmentation method provided by the present invention;
[0039] Figure 6 This is a schematic diagram of the audio segmentation device provided by the present invention;
[0040] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0042] The terms "first," "second," etc., used in the specification and claims of this invention are used to distinguish similar objects and are not used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and that the objects distinguished by "first," "second," etc., are generally of the same class.
[0043] In related technologies, speech translation can be divided into two types based on its technical approach: cascaded approach and end-to-end approach. The cascaded approach refers to audio being processed by speech recognition and machine translation engines separately to obtain the final translation result, while the end-to-end approach refers to speech being directly processed by a speech translation model to obtain the translation result.
[0044] The cascaded approach works as follows: First, the complete audio is segmented into audio segments of fixed length or with complete semantics. Then, speech recognition is performed on these segments to obtain the recognition results. Since the punctuation in the recognition results is often poor, a separate punctuation model is needed to re-punctuate the results. Finally, translation is performed based on the re-punctuated recognition results to obtain the final translation. In this cascaded approach, the units fed into the final translation model can be units from the initial audio segmentation or units from the re-punctuated results of the complete speech recognition, followed by sentence segmentation based on the punctuation. The initial audio segmentation needs to meet the requirements of the speech recognition model. Currently, the VAD model is generally used for segmentation. Its main method is to identify empty sounds in the audio and then segment the audio using these empty sounds as segmentation boundaries. Typically, audio segmented by VAD does not have complete semantic units.
[0045] End-to-end methods are simpler, directly feeding the segmented audio into the speech translation model to obtain the translation result. In end-to-end speech translation systems, the units fed into the translation model can be at the granularity of the preceding audio segmentation (VAD granularity in cascaded systems). Some end-to-end systems use a more crude audio segmentation method, directly employing a fixed audio length segmentation approach.
[0046] Some current recognition models can not only predict the recognition result but also simultaneously predict punctuation using audio information. However, it's important to note that these punctuation predictions are often localized and cannot utilize the complete semantic information of the entire audio stream. This is because the recognition model's training only needs to focus on local information, and its training corpus is typically at the VAD (Voice-Audio Digest) level. The recognition model also requires VAD-level audio input for prediction. Since punctuation within a VAD is predicted independently, it still cannot accurately predict complete semantic units.
[0047] To address the above problems, this invention provides an audio segmentation method. Figure 1 This is one of the flowcharts illustrating the audio segmentation method provided by the present invention. Figure 2 This is the second flowchart illustrating the audio segmentation method provided by this invention, as shown below. Figure 1 , Figure 2 As shown, the method includes:
[0048] Step 110: Obtain the audio to be segmented.
[0049] Specifically, the audio to be segmented can be obtained, which is the audio that needs to be segmented later.
[0050] The audio to be segmented can be obtained through a sound pickup device, which can be a smartphone, tablet, or smart appliance such as a speaker, television, or air conditioner. After the sound pickup device obtains the audio to be segmented through a microphone array, it can also amplify and reduce noise in the audio to be segmented. This embodiment of the invention does not specifically limit this.
[0051] Step 120: Extract the acoustic features of each frame in the audio to be segmented, and based on the acoustic features of each frame, perform semantic boundary sequence annotation on the audio to be segmented to obtain the semantic boundary annotation results of each frame.
[0052] Specifically, after obtaining the audio to be segmented, the acoustic features of each frame in the audio to be segmented can be extracted. Here, the acoustic features of each frame in the audio to be segmented can be extracted using an acoustic model. The acoustic model can be a HuBERT (Hidden-Unit Bidirectional Encoder Representations from Transformer) model, or an HMM-GMM (Hidden Markov Model-Gaussian Mixture Model), or a DNN-HMM (Deep Neural Networks-Hidden Markov Model), etc. The embodiments of the present invention do not specifically limit this.
[0053] The acoustic features of each frame here can also be obtained by extracting the acoustic features of each frame through Fast Fourier Transform (FFT) after the audio to be segmented is windowed and segmented. The acoustic features can be Mel Frequency Cepstrum Coefficient (MFCC) features, or Perceptual Linear Predictive (PLP) features, etc. The embodiments of the present invention do not specifically limit this.
[0054] After extracting the acoustic features of each frame in the audio to be segmented, semantic boundary sequence annotation can be performed on the audio to be segmented based on the acoustic features of each frame, and the semantic boundary annotation results of each frame can be obtained.
[0055] Here, semantic boundary sequence labeling of the segmented audio can be performed using a sequence labeling model. This sequence labeling model can be a Transformer model, an HMM (Hidden Markov Model), or a maximum entropy model, etc. The embodiments of this invention do not specifically limit this.
[0056] The semantic boundary annotation results for each frame here reflect the semantic boundary annotation information of each frame. The semantic boundary annotation results for each frame can include non-boundaries, strong semantic boundaries, and weak semantic boundaries. The strong semantic boundaries here correspond to sentence-ending symbols in punctuation, such as periods (.), question marks (?), and exclamation marks (!), etc., but this embodiment of the invention does not specifically limit them.
[0057] The weak semantic boundaries here correspond to punctuation marks in sentences, such as commas (,), colons (:), semicolons (;), and pause marks (、), etc. This embodiment of the invention does not specifically limit these.
[0058] The non-boundary here means that there is no complete semantic boundary, which can be represented by "O".
[0059] Understandably, the semantic boundary annotation results of each frame can assist in semantic sentence segmentation based on the tone and pause information in the acoustic features of each frame, preserving the complete semantic information of the audio and avoiding punctuation recognition errors, thus improving the accuracy and reliability of audio segmentation.
[0060] Step 130: Based on the semantic boundary annotation results of each frame, the audio to be segmented is segmented.
[0061] Specifically, after obtaining the semantic boundary annotation results for each frame, the audio to be segmented can be segmented based on the semantic boundary annotation results for each frame.
[0062] For example, candidate frames can be identified from each frame by referring to the segmentation application scenario and the semantic boundary annotation results of each frame, and the audio to be segmented can be segmented based on the candidate frames.
[0063] The method provided in this invention extracts the acoustic features of each frame in the audio to be segmented, and performs semantic boundary sequence annotation on the audio to be segmented based on the acoustic features of each frame to obtain the semantic boundary annotation results of each frame. Then, based on the semantic boundary annotation results of each frame, the audio to be segmented is segmented. Thus, the tone and pause information in the acoustic features of each frame can be used to assist semantic sentence segmentation, preserving the complete semantic information of the audio and avoiding punctuation recognition errors. This improves the accuracy and reliability of audio segmentation. Furthermore, this method can be applied to cascaded and end-to-end speech translation systems, expanding the application scope of audio segmentation.
[0064] Based on the above embodiments, Figure 3 This is a flowchart illustrating step 130 of the audio segmentation method provided by the present invention, as shown below. Figure 3 As shown, step 130 includes:
[0065] Step 131: Based on the semantic boundary annotation results of each frame, candidate frames are determined from each frame;
[0066] Step 132: Based on the candidate frames, segment the audio to be segmented.
[0067] Specifically, after obtaining the semantic boundary annotation results of each frame, candidate frames can be determined from each frame based on the semantic boundary annotation results of each frame. The candidate frames here are the audio frames used to segment the audio to be segmented.
[0068] The semantic boundary annotation results for each frame here reflect the semantic boundary annotation information of each frame. The semantic boundary annotation results for each frame can include non-boundary, strong semantic boundary and weak semantic boundary.
[0069] Here, the segmentation application scenario is also taken into account when determining candidate frames from each frame.
[0070] After identifying candidate frames from each frame, the audio to be segmented can be segmented based on these candidate frames. For example, candidate frames can be extracted from the audio to be segmented.
[0071] The method provided in this invention determines candidate frames from each frame based on the semantic boundary annotation results of each frame, ensuring the accuracy and reliability of the determined candidate frames. Then, based on the candidate frames, the audio to be segmented is segmented, thereby improving the accuracy and reliability of audio segmentation.
[0072] In related technologies, the specific application scenarios of audio segmentation are not considered when performing audio segmentation, resulting in low accuracy of audio segmentation. Therefore, the embodiments of the present invention take into account the application scenarios of audio segmentation when performing audio segmentation.
[0073] Based on the above embodiments, step 131 includes:
[0074] Step 210: Based on the segmentation application scenario, determine the frames whose speech boundary annotation results are strong semantic boundaries and / or weak semantic boundaries as candidate frames.
[0075] Specifically, based on the segmentation application scenario, frames with strong semantic boundaries and weak semantic boundaries can be identified as candidate frames. Alternatively, frames with strong semantic boundaries can be identified as candidate frames, or frames with weak semantic boundaries can be identified as candidate frames. This embodiment of the invention does not impose specific limitations on these aspects.
[0076] The segmentation application scenario here refers to the application scenario of audio segmentation, which can include the first scenario and the second scenario. The semantic integrity requirement of the first scenario is greater than that of the second scenario, while the real-time requirement of the second scenario is greater than that of the first scenario.
[0077] The strong semantic boundaries here correspond to sentence-ending punctuation marks, such as periods (.), question marks (?), and exclamation marks (!), etc., but this embodiment of the invention does not specifically limit them.
[0078] The weak semantic boundaries here correspond to punctuation marks in sentences, such as commas (,), colons (:), semicolons (;), and pause marks (、), etc. This embodiment of the invention does not specifically limit these.
[0079] It is understandable that frames with strong semantic boundaries in speech boundary annotation have complete semantic information, while frames with weak semantic boundaries in speech boundary annotation do not have complete semantic information, but have real-time semantic information.
[0080] The method provided in this invention, based on the segmentation application scenario, determines frames with strong semantic boundaries and / or weak semantic boundaries as candidate frames, thereby improving the accuracy and reliability of the determined candidate frames.
[0081] Based on the above embodiments, Figure 4 This is a flowchart illustrating step 210 of the audio segmentation method provided by the present invention, as shown below. Figure 4 As shown, step 210 includes:
[0082] Step 211: In the case that the segmentation application scenario is the first scenario, the frames whose speech boundary annotation results are strong semantic boundaries are determined as candidate frames.
[0083] Step 212: In the case that the segmentation application scenario is the second scenario, frames with strong semantic boundaries and frames with weak semantic boundaries are identified as candidate frames.
[0084] The semantic integrity requirement of the first scenario is greater than that of the second scenario, while the real-time requirement of the second scenario is greater than that of the first scenario.
[0085] Specifically, considering the different semantic integrity and real-time requirements of different segmentation application scenarios, in the case of the first segmentation application scenario, frames with strong semantic boundaries in the speech boundary annotation results can be identified as candidate frames. The semantic integrity requirement of the first scenario here is greater than that of the second scenario, and the first scenario here can be a speech translation scenario.
[0086] When the application scenario is segmented as the second scenario, frames with strong semantic boundaries and frames with weak semantic boundaries can be identified as candidate frames.
[0087] The real-time requirements of the second scenario here are greater than those of the first scenario. The second scenario here can be a speech recognition scenario, a simultaneous interpretation scenario, or a combination of speech recognition and simultaneous interpretation scenarios. This embodiment of the invention does not specifically limit this.
[0088] Understandably, the semantic integrity requirement is greater in speech translation than in simultaneous interpretation, while the real-time requirement is greater in simultaneous interpretation than in speech translation.
[0089] Based on the above embodiments, Figure 5 This is a flowchart illustrating step 132 of the audio segmentation method provided by the present invention, as shown below. Figure 5 As shown, step 132 includes:
[0090] Step 510: Obtain the frame number of the candidate audio segment in the audio to be segmented, wherein the candidate audio segment is composed of consecutive candidate frames;
[0091] Step 520: Based on the candidate audio segments with a frame count greater than a preset threshold, the audio to be segmented is segmented.
[0092] Specifically, considering that the duration of an audio frame is relatively short, for example, 10ms, while the duration of frames with strong and weak semantic boundaries, i.e. candidate frames, is generally more than 200ms, there will be 20 consecutive frames that are identified as candidate frames.
[0093] The number of frames of a candidate frame in the candidate audio segment of the audio to be segmented can be obtained. Here, the candidate audio segment is composed of consecutive candidate frames. The number of frames of a candidate frame in the candidate audio segment of the audio to be segmented can be 20, 25, or 30. This embodiment of the invention does not make a specific limitation on this.
[0094] Considering that when the number of candidate frames in the candidate audio segment is small, it may be a noisy frame, therefore, after obtaining the number of candidate frames in the candidate audio segment, the audio to be segmented can be segmented based on the candidate audio segment with a number of frames greater than a preset threshold.
[0095] The preset threshold here can be 20, 25, or 22, etc., and the embodiments of the present invention do not specifically limit it.
[0096] The method provided in this invention segmentes the audio to be segmented based on candidate audio segments with a frame count greater than a preset threshold, which can avoid interference from noise frames and further improve the accuracy and reliability of audio segmentation.
[0097] Based on the above embodiments, step 120 includes:
[0098] Step 121: Based on the boundary labeling model, extract the acoustic features of each frame in the audio to be segmented, and apply the acoustic features of each frame to perform semantic boundary sequence labeling on the audio to be segmented, so as to obtain the semantic boundary labeling results of each frame.
[0099] The boundary labeling model is trained based on sample audio and semantic boundary labels of each frame in the sample audio. The semantic boundary labels are determined based on the punctuation marks of the text corresponding to the sample audio.
[0100] Specifically, in order to obtain the semantic boundary annotation results for each frame, the boundary annotation model needs to be obtained through the following steps before step 121:
[0101] Sample audio data and semantic boundary labels for each frame within the sample audio can be collected in advance. These semantic boundary labels are determined based on the punctuation marks of the corresponding text in the sample audio. An initial boundary labeling model can also be pre-constructed. This model extracts the acoustic features of each frame in the audio to be segmented and applies these features to perform semantic boundary sequence labeling on the sample audio, yielding the semantic boundary labeling results for each frame. The parameters of this initial boundary labeling model can be randomly generated or pre-set.
[0102] The initial boundary labeling model here can be a Transformer model, an HMM model, a maximum entropy model, etc., and this embodiment of the invention does not specifically limit it. The Transformer model here can include a 12-layer Transformer module.
[0103] Subsequently, the sample audio can be input into the initial boundary labeling model, which extracts the acoustic features of each frame in the sample audio and applies the acoustic features of each frame to perform semantic boundary sequence labeling on the sample audio, thus obtaining the semantic boundary labeling results for each frame.
[0104] After obtaining the semantic boundary annotation results of each frame based on the initial boundary annotation model, the semantic boundary annotation results can be compared with the semantic boundary labels of each frame in the pre-collected sample audio. The loss function value is calculated based on the degree of difference between the two, and the parameters of the initial boundary annotation model are iterated based on the loss function value. The initial boundary annotation model after parameter iteration is denoted as the boundary annotation model.
[0105] The loss function here can be either the cross-entropy loss function or the mean squared error loss function; this embodiment of the invention does not specifically limit it.
[0106] Here, the backpropagation algorithm or similar can be used to iterate the parameters of the initial boundary labeling model based on the loss function value, and this embodiment of the invention does not impose specific limitations on this.
[0107] Understandably, the greater the difference between the semantic boundary annotation results and the semantic boundary labels of each frame in the pre-collected sample audio, the larger the loss function value; conversely, the smaller the difference between the semantic boundary annotation results and the semantic boundary labels of each frame in the pre-collected sample audio, the smaller the loss function value.
[0108] During the training process of the boundary labeling model, it learned to extract the acoustic features of each frame in the audio to be segmented, and to apply the acoustic features of each frame to perform semantic boundary sequence labeling on the audio to be segmented, so as to obtain the semantic boundary labeling results of each frame.
[0109] After training the boundary labeling model, the acoustic features of each frame in the audio to be segmented can be extracted based on the boundary labeling model. Then, the acoustic features of each frame can be applied to perform semantic boundary sequence labeling on the audio to be segmented, and the semantic boundary labeling results of each frame can be obtained.
[0110] The method provided in this embodiment of the invention uses a boundary labeling model trained on sample audio and semantic boundary labels of each frame in the sample audio. The semantic boundary labels are determined based on the punctuation marks of the text corresponding to the sample audio, thereby ensuring the accuracy and reliability of the semantic boundary labeling results of each frame.
[0111] Based on the above embodiments, the semantic boundary label is any one of non-boundary, strong semantic boundary, and weak semantic boundary. The strong semantic boundary corresponds to the sentence-ending symbol in the punctuation marks, and the weak semantic boundary corresponds to the sentence-intermediate symbol in the punctuation marks.
[0112] Specifically, the semantic boundary label can be any one of non-boundary, strong semantic boundary, and weak semantic boundary. Here, the strong semantic boundary corresponds to the sentence-ending punctuation marks, such as the period (.), question mark (?), and exclamation mark (!). This embodiment of the invention does not specifically limit this.
[0113] The weak semantic boundaries here correspond to the sentence-level symbols in the punctuation marks, such as commas (,), colons (:), semicolons (;), and pause marks (、), etc. This embodiment of the invention does not specifically limit these.
[0114] The non-boundary here means that there is no complete semantic boundary, which can be represented by "O".
[0115] Based on any of the above embodiments, an audio segmentation method includes the following steps:
[0116] The first step is to obtain the audio to be segmented.
[0117] The second step is to extract the acoustic features of each frame in the audio to be segmented based on the boundary labeling model, and apply the acoustic features of each frame to perform semantic boundary sequence labeling on the audio to be segmented, so as to obtain the semantic boundary labeling results of each frame.
[0118] The boundary labeling model here is trained based on sample audio and the semantic boundary labels of each frame in the sample audio. The semantic boundary labels here are determined based on the punctuation marks of the text corresponding to the sample audio.
[0119] The semantic boundary label here can be any one of non-boundary, strong semantic boundary, and weak semantic boundary. The strong semantic boundary here corresponds to the sentence-ending mark in punctuation, and the weak semantic boundary here corresponds to the sentence-intermediate mark in punctuation.
[0120] The third step is to identify frames with strong semantic boundaries in the first scenario when the application scenario is segmented.
[0121] In the case of segmentation application scenario 2, frames with strong semantic boundaries and frames with weak semantic boundaries are identified as candidate frames.
[0122] The semantic integrity requirement of the first scenario is greater than that of the second scenario, while the real-time requirement of the second scenario is greater than that of the first scenario.
[0123] The fourth step is to obtain the frame number of the candidate audio segment in the audio to be segmented. Here, the candidate audio segment consists of consecutive candidate frames.
[0124] The fifth step is to segment the audio to be segmented based on candidate audio segments with a frame count greater than a preset threshold.
[0125] The audio segmentation device provided by the present invention is described below. The audio segmentation device described below can be referred to in correspondence with the audio segmentation method described above.
[0126] Based on any of the above embodiments, the present invention provides an audio segmentation device. Figure 6 This is a schematic diagram of the audio segmentation device provided by the present invention, as shown below. Figure 6 As shown, the device includes:
[0127] Acquisition unit 610 is used to acquire the audio to be segmented;
[0128] The semantic boundary labeling unit 620 is used to extract the acoustic features of each frame in the audio to be segmented, and to perform semantic boundary sequence labeling on the audio to be segmented based on the acoustic features of each frame, so as to obtain the semantic boundary labeling results of each frame.
[0129] The audio segmentation unit 630 is used to segment the audio to be segmented based on the semantic boundary annotation results of each frame.
[0130] The apparatus provided in this invention extracts acoustic features from each frame of the audio to be segmented, and performs semantic boundary sequence annotation on the audio to be segmented based on the acoustic features of each frame, obtaining semantic boundary annotation results for each frame. Then, based on the semantic boundary annotation results of each frame, the audio to be segmented is segmented. Thus, the tone and pause information in the acoustic features of each frame can be used to assist semantic sentence segmentation, preserving the complete semantic information of the audio and avoiding punctuation recognition errors, thereby improving the accuracy and reliability of audio segmentation. Furthermore, this method can be applied to cascaded and end-to-end speech translation systems, expanding the application scope of audio segmentation.
[0131] Based on any of the above embodiments, the audio segmentation unit is specifically used for:
[0132] The candidate frame unit is used to determine candidate frames from the frames based on the semantic boundary annotation results of each frame;
[0133] The segmentation unit is used to segment the audio to be segmented based on the candidate frames.
[0134] Based on any of the above embodiments, determining the candidate frame unit is specifically used for:
[0135] Candidate frame sub-units are determined to identify frames whose speech boundary annotation results are strong semantic boundaries and / or weak semantic boundaries as candidate frames based on the segmentation application scenario.
[0136] Based on any of the above embodiments, determining the candidate frame subunit is specifically used for:
[0137] In the case where the segmentation application scenario is the first scenario, the frames whose speech boundary annotation results are strong semantic boundaries are determined as candidate frames;
[0138] In the case where the segmentation application scenario is the second scenario, frames with strong semantic boundaries and frames with weak semantic boundaries are identified as candidate frames.
[0139] The semantic integrity requirement of the first scenario is greater than that of the second scenario, while the real-time requirement of the second scenario is greater than that of the first scenario.
[0140] Based on any of the above embodiments, the segmentation unit is specifically used for:
[0141] The acquisition unit is used to acquire the frame number of the candidate audio segment in the audio to be segmented, wherein the candidate audio segment is composed of consecutive candidate frames;
[0142] The segmentation unit is used to segment the audio to be segmented based on candidate audio segments whose frame number is greater than a preset threshold.
[0143] Based on any of the above embodiments, the semantic boundary annotation unit is specifically used for:
[0144] Based on the boundary labeling model, the acoustic features of each frame in the audio to be segmented are extracted, and the acoustic features of each frame are applied to perform semantic boundary sequence labeling on the audio to be segmented, so as to obtain the semantic boundary labeling results of each frame.
[0145] The boundary labeling model is trained based on sample audio and semantic boundary labels of each frame in the sample audio. The semantic boundary labels are determined based on the punctuation marks of the text corresponding to the sample audio.
[0146] Based on any of the above embodiments, the semantic boundary label is any one of non-boundary, strong semantic boundary, and weak semantic boundary, the strong semantic boundary corresponds to the sentence-ending symbol in the punctuation marks, and the weak semantic boundary corresponds to the sentence-intermediate symbol in the punctuation marks.
[0147] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7As shown, the electronic device may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740, wherein the processor 710, the communications interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute an audio segmentation method, which includes: acquiring the audio to be segmented; extracting the acoustic features of each frame in the audio to be segmented, and performing semantic boundary sequence annotation on the audio to be segmented based on the acoustic features of each frame to obtain semantic boundary annotation results for each frame; and segmenting the audio to be segmented based on the semantic boundary annotation results for each frame.
[0148] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0149] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the audio segmentation method provided by the above methods. The method includes: acquiring audio to be segmented; extracting acoustic features of each frame in the audio to be segmented, and performing semantic boundary sequence annotation on the audio to be segmented based on the acoustic features of each frame to obtain semantic boundary annotation results for each frame; and segmenting the audio to be segmented based on the semantic boundary annotation results for each frame.
[0150] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the audio segmentation method provided by the above methods. The method includes: acquiring audio to be segmented; extracting acoustic features of each frame in the audio to be segmented, and performing semantic boundary sequence annotation on the audio to be segmented based on the acoustic features of each frame to obtain semantic boundary annotation results for each frame; and segmenting the audio to be segmented based on the semantic boundary annotation results for each frame.
[0151] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0152] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0153] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An audio segmentation method, characterized by, include: Get the audio to be segmented; The acoustic features of each frame in the audio to be segmented are extracted, and based on the acoustic features of each frame, the semantic boundary sequence is labeled to obtain the semantic boundary labeling results of each frame. In the case where the segmentation application scenario is the first scenario, the frames whose speech boundary annotation results are strong semantic boundaries are determined as candidate frames; In the case where the segmentation application scenario is the second scenario, frames with strong semantic boundaries and frames with weak semantic boundaries are identified as candidate frames. The semantic integrity requirement of the first scenario is greater than that of the second scenario, and the real-time requirement of the second scenario is greater than that of the first scenario. Based on the candidate frames, the audio to be segmented is segmented.
2. The audio segmentation method of claim 1, wherein, The step of segmenting the audio to be segmented based on the candidate frames includes: Obtain the frame number of the candidate audio segment in the audio to be segmented, wherein the candidate audio segment consists of consecutive candidate frames; Based on candidate audio segments with a frame count greater than a preset threshold, the audio to be segmented is segmented.
3. The audio segmentation method according to any one of claims 1-2, wherein, The process of extracting acoustic features from each frame of the audio to be segmented, and performing semantic boundary sequence annotation on the audio to be segmented based on the acoustic features of each frame to obtain the semantic boundary annotation results for each frame, includes: Based on the boundary labeling model, the acoustic features of each frame in the audio to be segmented are extracted, and the acoustic features of each frame are applied to perform semantic boundary sequence labeling on the audio to be segmented, so as to obtain the semantic boundary labeling results of each frame. The boundary labeling model is trained based on sample audio and semantic boundary labels of each frame in the sample audio. The semantic boundary labels are determined based on the punctuation marks of the text corresponding to the sample audio.
4. The audio segmentation method of claim 3, wherein, The semantic boundary label is any one of non-boundary, strong semantic boundary, and weak semantic boundary. The strong semantic boundary corresponds to the sentence-ending symbol in the punctuation marks, and the weak semantic boundary corresponds to the sentence-intermediate symbol in the punctuation marks.
5. An audio segmentation apparatus characterized by comprising: include: The acquisition unit is used to acquire the audio to be segmented; The semantic boundary labeling unit is used to extract the acoustic features of each frame in the audio to be segmented, and to perform semantic boundary sequence labeling on the audio to be segmented based on the acoustic features of each frame, so as to obtain the semantic boundary labeling results of each frame. The first determining unit is used to determine, in the case of segmenting the application scenario as the first scenario, the frames whose speech boundary annotation results are strong semantic boundaries as candidate frames. The second determining unit is used to determine, when the segmentation application scenario is the second scenario, frames with strong semantic boundaries and frames with weak semantic boundaries as candidate frames. The semantic integrity requirement of the first scenario is greater than that of the second scenario, and the real-time requirement of the second scenario is greater than that of the first scenario. The segmentation unit is used to segment the audio to be segmented based on the candidate frames.
6. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the audio segmentation method as described in any one of claims 1 to 4.
7. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by a processor, implements the audio segmentation method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Method and device of audio segmentation
CN107680584A
Speech recognition method and device, storage medium and electronic equipment
CN112634876A