A voice data processing method and device, electronic equipment and storage medium
By identifying the phoneme types in speech data, the boundaries of speech activity can be accurately located, solving the problem of inaccurate boundaries in the speech recognition process in existing technologies and improving the accuracy of speech data processing.
Patent Information
- Application Number
- CN202210878651.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-25
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2042-07-25
AI Technical Summary
In existing technologies, the large granularity of binary classification in speech recognition leads to inaccurate recognition of speech activity boundaries.
By identifying the phonemes of each audio frame in the speech data and using the phoneme type to determine the silence and non-silence types of audio frames, the boundaries of speech activity can be accurately located.
It achieves more accurate speech activity boundary detection and improves the accuracy of speech data processing.
Smart Images

Figure CN115240650B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of speech recognition, and particularly relates to a speech data processing method and device, an electronic device and a storage medium. BACKGROUND
[0002] With the rapid development and wide application of multimedia technology and network technology, speech recognition needs to be performed in many scenarios. However, in the current speech recognition process, a two-classification manner is usually adopted to determine whether an audio frame is a silent frame, so as to determine the speech boundary. Due to the large granularity of two-classification, the accuracy of the recognition result corresponding to the audio frame is low, resulting in inaccurate final speech activity boundary. SUMMARY
[0003] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides a speech data processing method and device, and a model training method.
[0004] According to an aspect of an embodiment of the present disclosure, a speech data processing method is provided, comprising:
[0005] obtaining target speech data to be processed, wherein the target speech data comprises a plurality of audio frames arranged in time sequence;
[0006] detecting the target speech data to obtain a target phoneme corresponding to each of the audio frames;
[0007] determining a target type of the audio frame corresponding to the target phoneme based on a phoneme type corresponding to the target phoneme, wherein the target type comprises a silent type or a non-silent type;
[0008] determining a speech activity boundary in the target speech data based on the target type corresponding to the audio frame, and taking the speech activity boundary as a detection result of the target speech data, wherein the speech activity boundary is used to divide a speech segment and a silent segment in the target speech data.
[0009] According to another aspect of an embodiment of the present disclosure, a speech data processing device is also provided, comprising:
[0010] an obtaining module configured to obtain target speech data to be processed, wherein the target speech data comprises a plurality of audio frames arranged in time sequence;
[0011] a detecting module configured to detect the target speech data to obtain a target phoneme corresponding to each of the audio frames;
[0012] The classification module is configured to determine a target type of the audio frame corresponding to the target phoneme based on a phoneme type corresponding to the target phoneme, wherein the target type comprises a silence type or a non-silence type.
[0013] The processing module is configured to determine a voice activity boundary in the target voice data based on the target type corresponding to the audio frame, and take the voice activity boundary as a detection result of the target voice data, wherein the voice activity boundary is used to divide a voice segment and a silence segment in the target voice data.
[0014] According to another aspect of the embodiments of the present disclosure, a storage medium is also provided, which includes a stored program. The program performs the steps described above when running.
[0015] According to another aspect of the embodiments of the present disclosure, an electronic device is also provided, which includes a memory configured to store at least one computer program, and at least one processor configured to perform the steps of any of the above methods by running the at least one computer program stored in the memory.
[0016] The embodiments of the present disclosure also provide a computer program product containing instructions, which, when running on a computer, cause the computer to perform the steps of the above methods.
[0017] The above technical solutions provided by the embodiments of the present disclosure have the following advantages compared with the prior art: the method provided by the embodiments of the present disclosure identifies phonemes of each audio frame in voice data, and determines silence type audio frames and non-silence type audio frames by using phonemes, which can more accurately detect the type of audio frames and more accurately locate the voice activity boundary in voice data compared with the prior art using a two-classification method. BRIEF DESCRIPTION OF DRAWINGS
[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and serve to explain the principles of the present disclosure, together with the description.
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, those skilled in the art can obtain other drawings from these drawings without creative effort.
[0020] Figure 1 A flowchart of a voice data processing method according to an embodiment of the present disclosure;
[0021] Figure 2 A flowchart of a voice data processing method according to another embodiment of the present disclosure;
[0022] Figure 3 A flowchart of a voice data processing method provided for another embodiment of the present disclosure;
[0023] Figure 4 A block diagram of a voice data processing apparatus provided for an embodiment of the present disclosure;
[0024] Figure 5 A structural schematic diagram of an electronic device provided for an embodiment of the present disclosure. DETAILED DESCRIPTION
[0025] In order to make the objects, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure, and do not constitute an improper limitation on the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present disclosure.
[0026] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0027] For example, in response to receiving an active request of a user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will need to obtain and use the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium, etc. that performs the operation of the technical solution of the present disclosure according to the prompt information.
[0028] As an optional but non-limiting implementation manner, in response to receiving an active request of a user, the manner of sending prompt information to the user may, for example, be a pop-up window manner, and the prompt information may, for example, be presented in the form of text in the pop-up window. In addition, the pop-up window may, for example, also carry a selection control for the user to select “agree” or “disagree” to provide the personal information to the electronic device.
[0029] It can be understood that the above notification and user authorization process is only illustrative, and does not constitute a limitation on the implementation manner of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0030] It can be understood that the data involved in the technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws and regulations and relevant provisions.
[0031] It should be noted that, in this document, relational terms such as "first" and "second", and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a", "comprises", or "comprising" does not, without more constraints, preclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0032] The embodiments of the present disclosure provide a speech data processing method and device, electronic equipment and storage medium. The method provided by the embodiments of the present disclosure can be applied to any required electronic equipment, for example, a server, a terminal, and the like, which are not limited here. For the convenience of description, the electronic equipment is referred to as electronic equipment hereinafter.
[0033] According to an aspect of the embodiments of the present disclosure, a method embodiment of a speech data processing method is provided. Figure 1 A flowchart of a speech data processing method provided by the embodiments of the present disclosure is shown in Figure 1 The method comprises the following steps.
[0034] In step S11, target speech data to be processed is obtained, wherein the target speech data comprises a plurality of audio frames arranged in time sequence.
[0035] In the embodiments of the present disclosure, the target speech data to be processed refers to speech to be subjected to speech recognition. For example, a certain scene recording can be used as the target speech data to be processed, and a certain dialogue / conversation can be used as the target speech data to be processed. The target speech data to be processed can be acquired in real time or pre-stored. For example, speech data input by a user in real time can be acquired through software (social software, live broadcast software, etc.) as the target speech data to be processed. The target speech data to be processed can also be pre-stored in a database, and when speech detection is required, the target speech data to be processed can be acquired from the database.
[0036] In the embodiments of the present disclosure, the target speech data to be processed can be speech data obtained by preprocessing original speech data. Specifically, the original speech data can be subjected to noise reduction processing to obtain the final target speech data, so as to ensure that the obtained speech data is smoother.
[0037] In step S12, the target speech data is detected to obtain a target phoneme corresponding to each audio frame.
[0038] In the embodiments of the present disclosure, the target speech data can be detected by using a pre-trained speech detection model to extract a target phoneme corresponding to each audio frame. The phoneme (phone) can be an element of each speech, which is the smallest language unit divided according to the natural attributes of language. It can be analyzed according to the pronunciation action of syllables, and one action constitutes one phoneme. For Chinese, the phoneme can be divided into vowels and consonants, for example, the Chinese syllable "a" has one phoneme, "ai" has two phonemes, and "dai" has three phonemes.
[0039] In step S13, a target type of the audio frame corresponding to the target phoneme is determined based on a phoneme type corresponding to the target phoneme, wherein the target type includes a silence type or a non-silence type.
[0040] In the embodiments of the present disclosure, in order to more accurately determine whether the audio frame is of the silence type or the non-silence type, first, the target phoneme is compared with a phoneme set to determine a phoneme type corresponding to the target phoneme, which is used to indicate whether the target phoneme is a silence phoneme or a non-silence phoneme.
[0041] As an example, the target speech data is "Hi, you", and then the target speech data is detected based on the speech detection model to obtain a target phoneme corresponding to each audio frame in the target speech data. The target speech data includes 18 audio frames, wherein "h" corresponds to the 1st-2nd audio frame, "i" corresponds to the 3rd audio frame, "sil" corresponds to the 4th-5th audio frame, "n" corresponds to the 6th-8th audio frame, "i" corresponds to the 9th-10th audio frame, "h" corresponds to the 11th-13th audio frame, "a" corresponds to the 14th audio frame, "o" corresponds to the 15th audio frame, and "sil" corresponds to the 16th-18th audio frame. Based on this, it can be determined that the silence phoneme is located in the 6th-8th audio frame and the 11th-13th audio frame.
[0042] In step S14, a speech activity boundary in the target speech data is determined based on the target type corresponding to the audio frame, and the speech activity boundary is taken as a detection result of the target speech data, wherein the speech activity boundary is used to divide a speech segment and a silence segment in the target speech data.
[0043] In the embodiments of the present disclosure, the speech activity boundary in the target speech data is determined based on the target type corresponding to the audio frame, including the following steps A1-A2:
[0044] Step A1, from the plurality of audio frames arranged in time sequence, a first audio frame adjacent to the second audio frame is obtained, wherein the first audio frame is a silent type audio frame, and the second audio frame is a non-silent type audio frame.
[0045] Step A2, the frame boundary between the first audio frame and the second audio frame is determined as the speech activity boundary.
[0046] In the embodiments of the present disclosure, in order to accurately find the speech activity boundary in the target speech data, the first audio frame of the silent type and the second audio frame of the non-silent type are obtained from the plurality of audio frames arranged in time sequence, and then the frame boundary between the first audio frame and the second audio frame is determined as the speech activity boundary. Through the speech activity boundary, the silent segment and the non-silent segment in the target speech data can be accurately divided. Alternatively, in order to more accurately locate the frame boundary between the first audio frame and the second audio frame, the first audio frame and the second audio frame are first processed by a sliding window. Specifically, the audio segment in the audio frame is intercepted by using the sliding window method, and the audio segment is detected by the speech activity, so as to remove the audio segment with noise, thereby improving the positioning accuracy of the frame boundary.
[0047] As an example, "h" corresponds to the 1st-2nd audio frame, "i" corresponds to the 3rd audio frame, "sil" corresponds to the 4th-5th audio frame, "n" corresponds to the 6th-8th audio frame, "i" corresponds to the 9th-10th audio frame, "h" corresponds to the 11th-13th audio frame, "a" corresponds to the 14th audio frame, "o" corresponds to the 15th audio frame, and "sil" corresponds to the 16th-18th audio frame. Based on this, the frame boundary between the 5th audio frame and the 6th audio frame is determined as the speech activity boundary, the frame boundary between the 8th audio frame and the 9th audio frame is determined as the speech activity boundary, and the frame boundary between the 15th audio frame and the 16th audio frame is determined as the speech activity boundary.
[0048] The method provided by the embodiments of the present disclosure can more accurately detect the type of audio frame and more accurately locate the speech activity boundary in the speech data by identifying the phonemes corresponding to each audio frame in the speech data and determining the silent type audio frame and the non-silent type audio frame by using the phonemes.
[0049] Figure 2 A flowchart of a speech data processing method provided by the embodiments of the present disclosure is shown in Figure 2 The method comprises:
[0050] Step S21, obtaining target speech data to be processed, wherein the target speech data comprises a plurality of audio frames arranged in time sequence. For details, please refer to the relevant description of the above-mentioned embodiments, which will not be repeated here.
[0051] Step S22, detecting the target speech data to obtain a target phoneme corresponding to each audio frame.
[0052] In the embodiments of the present disclosure, detecting the target speech data to obtain a target phoneme corresponding to each audio frame comprises the following steps A1-A2:
[0053] Step A1, obtaining a pre-trained speech detection model, wherein the speech detection model comprises an identification network and a classification network.
[0054] Step A2, extracting a target acoustic feature corresponding to each audio frame in the target speech data through the identification network, and determining a target phoneme corresponding to the target acoustic feature based on a preset corresponding relationship between the acoustic feature and the phoneme.
[0055] In the embodiments of the present disclosure, the speech detection model can be a model based on HMM (Hidden Markov Model), such as LSTM-HMM (Long Short-Term Memory-Hidden Markov Model), GMM-HMM (Gaussian Mixture Model-Hidden Markov Model), or DNN-HMM (Deep Neural Network-Hidden Markov Model), etc.
[0056] In the embodiments of the present disclosure, the speech detection model comprises an identification network and a classification network, the identification network can comprise a plurality of convolutional layers, and the classification network comprises a fully connected layer. Specifically, the identification network extracts a target acoustic feature corresponding to each audio frame in the target speech data, and determines a target phoneme corresponding to the target acoustic feature through a pre-learned corresponding relationship between the acoustic feature and the phoneme.
[0057] Step S23, obtaining a speech score corresponding to the target phoneme, and determining a target type corresponding to the audio frame by using the speech score of each target phoneme.
[0058] In the embodiments of the present disclosure, step S23, obtaining a speech score corresponding to the target phoneme, and determining a target type corresponding to the audio frame by using the speech score of each target phoneme, comprises the following steps B1-B2:
[0059] Step B1, calculating the speech score corresponding to the target phoneme through the target node of the classification network, and obtaining the target mapping result corresponding to the speech score.
[0060] In the embodiments of the present disclosure, the speech score corresponding to the target phoneme is calculated through the target node of the classification network, and the target mapping result corresponding to the speech score is obtained, including the following steps B101-B104:
[0061] Step B101, obtaining the first node corresponding to the first target phoneme from at least one node of the classification network, and determining the nodes of the classification network except the first node as target nodes, wherein the number of nodes in the classification network is the same as the number of target phonemes, and the first target phoneme is the first phoneme in the target phonemes.
[0062] Step B102, determining the acoustic score output by the target node in the classification network through which the second target phoneme passes, and the state transition probability between adjacent second target phonemes, wherein the second target phoneme is a phoneme in the target phonemes except the first target phoneme.
[0063] Step B103, obtaining the speech score corresponding to the second target phoneme based on the acoustic score and the state transition probability.
[0064] Step B104, determining the target state node corresponding to the speech score based on the mapping relationship between the speech score and the state node, and taking the target state node as the target mapping result, wherein the target state node is used to represent the type corresponding to the audio frame.
[0065] In the embodiments of the present disclosure, the classification network includes a full connection layer and an output layer, and each node of the target phoneme is input into the full connection layer. First, the first phoneme in the target phonemes is determined as the first target phoneme, and the first node corresponding to the first target phoneme is searched from the nodes of the full connection layer. The acoustic score output by the first node is directly mapped to the silence node of the output layer, so that the first phoneme in the target phonemes is determined as the silence phoneme.
[0066] At the same time, the phonemes in the target phonemes except the first target phoneme are taken as the second target phonemes. First, the target node corresponding to each second target phoneme is searched from the nodes of the full connection layer, and the acoustic score corresponding to the second target phoneme and the state transition probability between adjacent second target phonemes are output through the target node. Second, the speech score corresponding to the second target phoneme is obtained based on the acoustic score corresponding to each second target phoneme and the state transition probability. Finally, the target state node corresponding to the speech score of each second target phoneme is obtained through the mapping relationship between the preset speech score and the state node, and the target state node is taken as the target mapping result.
[0067] It should be noted that in the process of training the classification network, the classification network learns the state transition probability between adjacent phonemes from the speech data sample. Therefore, after the second target phoneme is input into the full connection layer, the full connection layer outputs the acoustic score output by the target node corresponding to the phoneme, and the state transition probability between adjacent second target phonemes.
[0068] Step B2, determine the type indicated by the target mapping result as the target type corresponding to the audio frame.
[0069] The speech score corresponding to the target phoneme can be used to determine whether the target phoneme is a silent phoneme, so as to determine whether the audio frame corresponding to the target phoneme is a silent type audio frame. The silent phoneme includes a silent phoneme (SIL), noise, laughter, and other non-linguistic directly related pseudo phonemes. Specifically, when the speech score corresponding to the target phoneme indicates that the target phoneme belongs to a silent phoneme, it means that the audio frame corresponding to the target phoneme is a silent type audio frame. Or when the speech score corresponding to the target phoneme indicates that the target phoneme belongs to a non-silent phoneme, it means that the audio frame corresponding to the target phoneme is a non-silent type audio frame.
[0070] Step S24, determine the speech activity boundary in the target speech data based on the target type corresponding to the audio frame, and take the speech activity boundary as the detection result of the target speech data. For details, refer to the related description of the above-mentioned embodiments, which will not be repeated here.
[0071] Figure 3 A flowchart of a training method of a speech detection model provided by an embodiment of the present disclosure is shown in Figure 3 The method comprises:
[0072] Step S31, obtaining a speech data sample, wherein the speech data sample comprises at least one audio frame, and the phoneme label corresponding to the acoustic feature of each audio frame.
[0073] In the embodiment of the present disclosure, the speech data in multiple scenes can be obtained from the speech interaction device as the speech data sample. The speech interaction device can be a smart phone, a smart watch, a smart home companion, or other devices with speech collection / recognition. The multiple scenes can be home scenes, conference scenes, sports scenes, and the like.
[0074] In the embodiment of the present disclosure, after obtaining the speech data sample, the acoustic feature corresponding to each audio frame in the speech data sample can be extracted by using the acoustic model, and the acoustic feature is forcibly aligned with the phoneme, that is, the phoneme label corresponding to the acoustic feature of each audio frame is obtained.
[0075] Step S32, extracting the acoustic feature corresponding to each audio frame in the speech data sample, and obtaining the output phoneme based on the acoustic feature.
[0076] In the embodiments of the present disclosure, the acoustic feature corresponding to each audio frame in the speech data sample can be extracted by using the speech detection model to be trained. The speech detection model includes an identification network and a classification network, and the identification network includes a convolutional subnetwork and a prediction subnetwork. The speech detection model can be an LSTM-HMM (Long Short-Term Memory-Hidden Markov Model), a GMM-HMM (Gaussian Mixture Model-Hidden Markov Model), or a DNN-HMM (Deep Neural Network-Hidden Markov Model).
[0077] Specifically, the acoustic feature corresponding to each audio frame in the speech data sample is extracted, and the output phoneme is obtained based on the acoustic feature, including the following steps C1-C2:
[0078] Step C1, the acoustic feature corresponding to each audio frame in the speech data sample is extracted by the identification network.
[0079] Step C2, the output phoneme corresponding to the audio frame is obtained by the prediction subnetwork based on the acoustic feature.
[0080] In the embodiments of the present disclosure, after the acoustic feature corresponding to each audio frame in the speech data sample is extracted and the phoneme corresponding to the acoustic feature is output, the method further includes: calculating the speech score corresponding to the phoneme by the target node of the classification network, obtaining the mapping result corresponding to the speech score, and determining the type indicated by the mapping result as the output type corresponding to the audio frame.
[0081] Step S33, based on the output phoneme and the phoneme label, the parameters of the speech detection model are adjusted.
[0082] In the embodiments of the present disclosure, based on the phoneme and the phoneme label, the parameters of the speech detection model are adjusted, including: obtaining a first loss between the phoneme and the phoneme label, obtaining a type label corresponding to each audio frame, and determining a second loss between the output type and the type label, based on the first loss and the second loss, adjusting the parameters of the speech detection model.
[0083] In this embodiment, the parameter-tuned speech detection model is tested. Specifically, test speech data is acquired, which is speech data without phoneme labels. The test speech data is input into the parameter-tuned speech detection model to obtain test results. If the test results meet the test requirements, the parameter-tuned speech detection model is considered to have completed training. The test results are the output phonemes corresponding to the test speech data. Meeting the test requirements means that the output phonemes match the phonemes corresponding to each audio frame in the test speech data.
[0084] Figure 4 This is a block diagram of a voice data processing apparatus provided in an embodiment of the present disclosure. This apparatus can be implemented as part or all of an electronic device through software, hardware, or a combination of both. Figure 4 As shown, the device includes:
[0085] The acquisition module 41 is used to acquire the target speech data to be processed, wherein the target speech data includes multiple audio frames arranged in chronological order.
[0086] The detection module 42 is used to detect the target speech data and obtain the target phonemes corresponding to each audio frame.
[0087] The classification module 43 is used to obtain the speech score corresponding to the target phoneme and use the speech score of each target phoneme to determine the target type corresponding to the audio frame.
[0088] The processing module 44 is used to determine the speech activity boundary in the target speech data based on the target type corresponding to the audio frame, and use the speech activity boundary as the detection result of the target speech data.
[0089] This disclosure also provides an electronic device, such as... Figure 5 As shown, the electronic device may include: a processor 1501, a communication interface 1502, a memory 1503, and a communication bus 1504, wherein the processor 1501, the communication interface 1502, and the memory 1503 communicate with each other through the communication bus 1504.
[0090] Memory 1503 is used to store computer programs;
[0091] When the processor 1501 executes the computer program stored in the memory 1503, it implements the steps of the above embodiments.
[0092] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.
[0093] The communication interface is used for communication between the terminal and other devices.
[0094] The memory can include a Random Access Memory (RAM) and can also include a non-volatile memory, such as at least one disk memory. Optionally, the memory can also be at least one storage device located away from the aforementioned processor.
[0095] The processor mentioned above can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; can also be a Digital Signal Processing (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0096] In yet another embodiment provided by the present disclosure, a computer readable storage medium is also provided, and the computer readable storage medium stores instructions, when the instructions run on a computer, cause the computer to execute the voice data processing method in any of the above embodiments.
[0097] In yet another embodiment provided by the present disclosure, a computer program product containing instructions is also provided, and when the instructions run on a computer, cause the computer to execute the voice data processing method in any of the above embodiments.
[0098] According to one or more embodiments of the present disclosure, example 1 provides a voice data processing method, comprising:
[0099] obtaining target speech data to be processed, wherein the target speech data comprises a plurality of audio frames arranged in time sequence;
[0100] detecting the target speech data to obtain a target phoneme corresponding to each audio frame;
[0101] determining a target type of the audio frame corresponding to the target phoneme based on a phoneme type corresponding to the target phoneme, wherein the target type comprises a silence type or a non-silence type;
[0102] determining a voice activity boundary in the target speech data based on the target type corresponding to the audio frame, and taking the voice activity boundary as a detection result of the target speech data, wherein the voice activity boundary is used to divide a voice segment and a silence segment in the target speech data.
[0103] Further, the detecting the target speech data to obtain a target phoneme corresponding to each audio frame comprises:
[0104] obtaining a pre-trained speech detection model, wherein the speech detection model comprises an identification network and a classification network;
[0105] extracting a target acoustic feature corresponding to each audio frame in the target speech data through the identification network, and determining a target phoneme corresponding to the target acoustic feature based on a preset corresponding relationship between an acoustic feature and a phoneme.
[0106] Further, the determining a target type of the audio frame corresponding to the target phoneme based on a phoneme type corresponding to the target phoneme comprises:
[0107] calculating a voice score corresponding to the target phoneme through a target node of the classification network, and obtaining a target mapping result corresponding to the voice score, wherein the voice score is used to represent a phoneme type of the target phoneme;
[0108] determining a type indicated by the target mapping result as the target type corresponding to the audio frame.
[0109] Further, the calculating a voice score corresponding to the target phoneme through a target node of the classification network, and obtaining a target mapping result corresponding to the voice score comprises:
[0110] obtaining a first node corresponding to a first target phoneme from at least one node of the classification network, and determining nodes other than the first node in the classification network as the target node, wherein a number of nodes in the classification network is the same as a number of the target phonemes, and the first target phoneme is a first phoneme in the target phonemes;
[0111] determining an acoustic score output by a target node through which a second target phoneme passes in the classification network, and a state transition probability between adjacent second target phonemes, wherein the second target phoneme is a phoneme other than the first target phoneme in the target phonemes.
[0112] obtaining a speech score corresponding to the second target phoneme based on the acoustic score and the state transition probability;
[0113] determining a target state node corresponding to the speech score based on a mapping relationship between the speech score and the state node, and taking the target state node as a target mapping result, wherein the target state node is used to represent a type corresponding to the audio frame.
[0114] Further, the target type includes: a silence type and a non-silence type.
[0115] determining a voice activity boundary in the target speech data based on the target type corresponding to the audio frame, comprising:
[0116] obtaining a first audio frame and a second audio frame adjacent to each other from a plurality of audio frames arranged in time sequence, wherein the first audio frame is a silence type audio frame, and the second audio frame is a non-silence type audio frame;
[0117] determining a frame boundary between the first audio frame and the second audio frame as the voice activity boundary.
[0118] Further, the method further comprises:
[0119] obtaining a speech data sample, wherein the speech data sample includes at least one audio frame, and a phoneme label corresponding to an acoustic feature of each audio frame;
[0120] extracting an acoustic feature corresponding to each audio frame in the speech data sample, and obtaining an output phoneme based on the acoustic feature;
[0121] adjusting parameters of the speech detection model based on the output phoneme and the phoneme label.
[0122] Further, the speech detection model includes: a recognition network and a classification network, wherein the recognition network includes a convolutional subnetwork and a prediction subnetwork.
[0123] extracting an acoustic feature corresponding to each audio frame in the speech data sample, and obtaining an output phoneme based on the acoustic feature, comprising:
[0124] extracting an acoustic feature corresponding to each audio frame in the speech data sample through the recognition network;
[0125] obtaining an output phoneme corresponding to the audio frame through the prediction subnetwork based on the acoustic feature.
[0126] Further, after extracting the acoustic feature corresponding to each audio frame in the speech data sample and outputting the phoneme corresponding to the acoustic feature, the method further comprises:
[0127] The target node of the classification network is used to calculate a speech score corresponding to the phoneme, obtain a mapping result corresponding to the speech score, and determine a type indicated by the mapping result as the output type corresponding to the audio frame.
[0128] Further, based on the phoneme and the phoneme label, the parameters of the speech detection model are adjusted, including:
[0129] A first loss between the phoneme and the phoneme label is obtained.
[0130] A type label corresponding to each audio frame is obtained, and a second loss between the output type and the type label is determined.
[0131] Based on the first loss and the second loss, the parameters of the speech detection model are adjusted.
[0132] According to one or more embodiments of the present disclosure, example 2 provides a processing apparatus of speech data, comprising:
[0133] An obtaining module is configured to obtain target speech data to be processed, wherein the target speech data comprises a plurality of audio frames arranged in time sequence.
[0134] A detecting module is configured to detect the target speech data to obtain a target phoneme corresponding to each audio frame.
[0135] A classification module is configured to determine a target type of an audio frame corresponding to the target phoneme based on a phoneme type of the target phoneme, wherein the target type comprises a silence type or a non-silence type.
[0136] A processing module is configured to determine a speech activity boundary in the target speech data based on the target type of the audio frame, and take the speech activity boundary as a detection result of the target speech data, wherein the speech activity boundary is used to divide a speech segment and a silence segment in the target speech data.
[0137] In the embodiments of the present disclosure, the detecting module is configured to obtain a pre-trained speech detection model, wherein the speech detection model comprises an identification network and a classification network; the target acoustic feature corresponding to each audio frame in the target speech data is extracted through the identification network, and a target phoneme corresponding to the target acoustic feature is determined based on a preset corresponding relationship between the acoustic feature and the phoneme.
[0138] In the embodiments of the present disclosure, the detecting module is configured to calculate a speech score corresponding to the target phoneme through a target node of the classification network, obtain a target mapping result corresponding to the speech score, and determine a type indicated by the target mapping result as a target type corresponding to the audio frame.
[0139] In the embodiment of the present disclosure, the detection module is configured to obtain a first node corresponding to a first target phoneme from at least one node of the classification network, and determine nodes of the classification network other than the first node as target nodes, wherein the number of nodes of the classification network is the same as the number of target phonemes, the first target phoneme is a first phoneme in the target phonemes; determine an acoustic score output by a target node in the classification network through which a second target phoneme passes, and a state transition probability between adjacent second target phonemes, wherein the second target phonemes are phonemes other than the first target phoneme in the target phonemes; obtain a speech score corresponding to the second target phoneme based on the acoustic score and the state transition probability; determine a target state node corresponding to the speech score based on a mapping relationship between the speech score and the state node, and take the target state node as a target mapping result, wherein the target state node is used to represent a type corresponding to the audio frame.
[0140] In the embodiment of the present disclosure, the target type includes a silence type and a non-silence type.
[0141] The processing module is configured to obtain adjacent first and second audio frames from a plurality of audio frames arranged in time sequence, wherein the first audio frame is a silence type audio frame, and the second audio frame is a non-silence type audio frame; and determine a frame boundary between the first and second audio frames as a voice activity boundary.
[0142] In the embodiment of the present disclosure, the speech data processing apparatus further includes a training module configured to obtain a speech data sample, wherein the speech data sample includes at least one audio frame, and a phoneme label corresponding to an acoustic feature of each audio frame; extract the acoustic feature corresponding to each audio frame in the speech data sample, and obtain an output phoneme based on the acoustic feature; and adjust parameters of a speech detection model based on the output phoneme and the phoneme label.
[0143] In the embodiment of the present disclosure, the speech detection model includes an identification network and a classification network, wherein the identification network includes a convolutional subnetwork and a prediction subnetwork.
[0144] In the embodiment of the present disclosure, the training module is configured to extract the acoustic feature corresponding to each audio frame in the speech data sample through the identification network; and obtain the output phoneme corresponding to the audio frame through the prediction subnetwork based on the acoustic feature.
[0145] In the embodiment of the present disclosure, the training module is configured to calculate a speech score corresponding to a phoneme through a target node of the classification network, obtain a mapping result corresponding to the speech score, and determine a type indicated by the mapping result as an output type corresponding to the audio frame.
[0146] In the embodiments of the present disclosure, the training module is configured to obtain a first loss between the phonemes and the phoneme labels; obtain a type label corresponding to each audio frame, and determine a second loss between the output type and the type label; and adjust parameters of the speech detection model based on the first loss and the second loss.
[0147] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center through wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that includes one or more available media sets. The available media can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk), etc.
[0148] The above only describes the preferred embodiments of the present disclosure and is not intended to limit the protection scope of the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
[0149] The above only describes the specific embodiments of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features applied herein.
Claims
1. A method of processing voice data, characterized by, The method comprises: acquiring target speech data to be processed, wherein the target speech data comprises a plurality of audio frames arranged in time sequence; detecting the target speech data to obtain a target phoneme corresponding to each of the audio frames; determining a target type of the audio frame corresponding to the target phoneme based on a phoneme type corresponding to the target phoneme, wherein the target type comprises a silence type or a non-silence type; determining a voice activity boundary in the target speech data based on the target type corresponding to the audio frame, taking the voice activity boundary as a detection result of the target speech data, wherein the voice activity boundary is used to divide a voice segment and a silence segment in the target speech data; the detection of the target speech data to obtain a target phoneme corresponding to each of the audio frames comprises: acquiring a pre-trained speech detection model, wherein the speech detection model comprises an identification network and a classification network; extracting a target acoustic feature corresponding to each audio frame in the target speech data through the identification network, and determining a target phoneme corresponding to the target acoustic feature based on a preset corresponding relationship between an acoustic feature and a phoneme; the determination of the target type of the audio frame corresponding to the target phoneme based on the phoneme type corresponding to the target phoneme comprises: acquiring a first node corresponding to a first target phoneme from at least one node of the classification network, and determining a target node in the classification network except the first node as a target node, wherein the number of nodes in the classification network is the same as the number of target phonemes, and the first target phoneme is the first phoneme in the target phonemes; determining an acoustic score output by a target node passed by a second target phoneme in the classification network and a state transition probability between adjacent second target phonemes, wherein the second target phoneme is a phoneme in the target phonemes except the first target phoneme; obtaining a speech score corresponding to the second target phoneme based on the acoustic score and the state transition probability; determining a target state node corresponding to the speech score based on a mapping relationship between the speech score and the state node, and taking the target state node as a target mapping result, wherein the target state node is used to represent the type corresponding to the audio frame; and determining the type indicated by the target mapping result as the target type corresponding to the audio frame.
2. The method of claim 1, wherein, The target type comprises a silence type and a non-silence type. The determination of the voice activity boundary in the target speech data based on the target type corresponding to the audio frame comprises: acquiring adjacent first and second audio frames from a plurality of audio frames arranged in time sequence, wherein the first audio frame is a silence type audio frame, and the second audio frame is a non-silence type audio frame; determining a frame boundary between the first and second audio frames as the voice activity boundary.
3. The method of claim 1, wherein, The method further comprises: acquiring a speech data sample, wherein the speech data sample comprises at least one audio frame, and an acoustic feature corresponding to a phoneme label of each audio frame; extracting acoustic features corresponding to each audio frame in the speech data sample, and obtaining output phonemes based on the acoustic features; adjusting parameters of the speech detection model based on the output phonemes and the phoneme labels.
4. The method of claim 3, wherein, The speech detection model comprises a recognition network and a classification network, wherein the recognition network comprises a convolutional subnetwork and a prediction subnetwork. The extracting acoustic features corresponding to each audio frame in the speech data sample, and obtaining output phonemes based on the acoustic features comprises: extracting acoustic features corresponding to each audio frame in the speech data sample through the recognition network; obtaining output phonemes corresponding to the audio frame through the prediction subnetwork based on the acoustic features.
5. The method of claim 4, wherein, After extracting acoustic features corresponding to each audio frame in the speech data sample and outputting phonemes corresponding to the acoustic features, the method further comprises: calculating a speech score corresponding to the phoneme through a target node of the classification network, obtaining a mapping result corresponding to the speech score, and determining a type indicated by the mapping result as an output type corresponding to the audio frame.
6. The method of claim 5, wherein, The adjusting parameters of the speech detection model based on the output phonemes and the phoneme labels comprises: obtaining a first loss between the phonemes and the phoneme labels; obtaining a type label corresponding to each audio frame, and determining a second loss between the output type and the type label; adjusting parameters of the speech detection model based on the first loss and the second loss.
7. A voice data processing apparatus characterized by comprising: comprises: an acquisition module configured to acquire target speech data to be processed, wherein the target speech data comprises a plurality of audio frames arranged in time sequence; a detection module configured to detect the target speech data to obtain target phonemes corresponding to each audio frame; a classification module configured to determine a target type of an audio frame corresponding to the target phoneme based on a phoneme type of the target phoneme, wherein the target type comprises a silence type or a non-silence type; a processing module configured to determine a voice activity boundary in the target speech data based on the target type corresponding to the audio frame, and take the voice activity boundary as a detection result of the target speech data, wherein the voice activity boundary is used to divide a voice segment and a silence segment in the target speech data; The detection module is configured to acquire a pre-trained speech detection model, wherein the speech detection model comprises a recognition network and a classification network; the recognition network is configured to extract target acoustic features corresponding to each audio frame in the target speech data, and determine target phonemes corresponding to the target acoustic features based on a correspondence between preset acoustic features and phonemes; The detection module is configured to acquire a first node corresponding to a first target phoneme from at least one node of the classification network, and determine a target node of the classification network as a node other than the first node, wherein the number of nodes in the classification network is the same as the number of target phonemes, and the first target phoneme is the first phoneme in the target phonemes. An acoustic score of a target node output through which a second target phoneme passes in the classification network is determined, as well as a state transition probability between adjacent second target phonemes, wherein the second target phonemes are phonemes other than the first target phoneme among the target phonemes; a speech score corresponding to the second target phoneme is obtained based on the acoustic score and the state transition probability; a target state node corresponding to the speech score is determined based on a mapping relationship between the speech score and a state node, and the target state node is taken as a target mapping result, wherein the target state node is used to represent a type corresponding to the audio frame; and a type indicated by the target mapping result is determined as a target type corresponding to the audio frame.
8. A storage medium, characterized by The storage medium comprises a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 6.
9. An electronic device, comprising: The method comprises: a memory for storing at least one computer program; at least one processor for executing the method of any one of claims 1 to 6 by running the at least one computer program stored on the memory.
Citation Information
Patent Citations
Voice endpoint detection method and device, equipment and storage medium
CN114155839A