A voice data processing method and device, electronic equipment and storage medium
By acquiring the feature sequences and time series of speech data, and utilizing linear computation and neural network models, the problem of the inability to detect fine-grained acoustic features of speakers in existing technologies has been solved, thereby improving the accuracy of fluency detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING YOUZHUJU NETWORK TECH CO LTD
- Filing Date
- 2023-05-19
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies cannot effectively detect the speaker's fine-grained acoustic features, resulting in low accuracy of fluency test results.
By acquiring the feature sequence and time sequence of the target speech data, and using linear computation and neural network models, phoneme description information and time information are extracted to achieve fine-grained fluency detection.
This improves the reliability of fluency test results and provides a reliable basis for testing.
Smart Images

Figure CN116504273B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of speech processing technology, and in particular to a method, apparatus, electronic device and storage medium for processing speech data. Background Technology
[0002] With the development of computer technology, the application of speech recognition technology is increasing. For example, it's used for speech data synthesis and fluency detection. Currently, methods for detecting fluency through language data are relatively limited and cannot effectively detect the speaker's fine-grained acoustic features, resulting in low accuracy in fluency detection results. Summary of the Invention
[0003] In order to solve the above-mentioned technical problems, or at least partially solve the above-mentioned technical problems, this disclosure provides a method, apparatus, electronic device and storage medium for processing voice data.
[0004] According to one aspect of the present disclosure, a method for processing voice data is provided, including:
[0005] Acquire the target speech data to be recognized;
[0006] The target speech data is detected to obtain a target feature sequence and a target time sequence, wherein the target feature sequence includes phoneme description information of each audio frame corresponding to the phoneme in the target speech data, and the target time sequence includes time information corresponding to each phoneme in the target speech data;
[0007] Linear calculations are performed based on the phoneme description information in the target feature sequence and the time information in the target time sequence to obtain the target fluency corresponding to the target speech data.
[0008] According to another aspect of the embodiments of this disclosure, a voice data processing apparatus is also provided, comprising:
[0009] The acquisition module is used to acquire the target speech data to be recognized;
[0010] The detection module is used to detect the target speech data to obtain a target feature sequence and a target time sequence, wherein the target feature sequence includes phoneme description information of each phoneme corresponding to each audio frame in the target speech data, and the target time sequence includes time information corresponding to each phoneme in the target speech data;
[0011] The prediction module is used to perform linear calculations based on the phoneme description information in the target feature sequence and the time information in the target time sequence to obtain the target fluency corresponding to the target speech data.
[0012] According to another aspect of the embodiments of this disclosure, a storage medium is also provided, the storage medium including a stored program that executes the above steps when the program is run.
[0013] According to another aspect of the present disclosure, an electronic device is also provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; wherein: the memory is used to store computer programs; and the processor is used to execute the steps in the above method by running the programs stored in the memory.
[0014] This disclosure also provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the steps in the above-described method.
[0015] The technical solutions provided in this disclosure have the following advantages: The method provided in this disclosure extracts the target feature sequence and target time sequence of speech data. Through the phoneme description information in the target feature sequence and the time information in the target time sequence, it can accurately express fine-grained acoustic features, providing a reliable basis for the fluency detection of speech data and improving the reliability of the fluency detection results. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0017] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 A flowchart of a voice data processing method provided in this embodiment of the disclosure;
[0019] Figure 2 A schematic diagram of the feature sequences and time series provided in the embodiments of this disclosure;
[0020] Figure 3 A flowchart of a voice data processing method provided in this embodiment of the disclosure;
[0021] Figure 4 A schematic diagram of the masked feature sequence and time sequence provided in the embodiments of this disclosure;
[0022] Figure 5 This is a schematic diagram of the structure of a preset neural network model provided in an embodiment of the present disclosure;
[0023] Figure 6 A block diagram of a voice data processing apparatus provided in an embodiment of this disclosure;
[0024] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. The illustrative embodiments and their descriptions are used to explain this disclosure and do not constitute an improper limitation of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0026] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0027] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0028] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0029] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0030] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0031] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another similar entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0032] This disclosure provides a method, apparatus, electronic device, and storage medium for processing voice data. The method provided in this disclosure can be applied to any electronic device as needed, such as a server, terminal, or other electronic device. No specific limitation is made here, and for ease of description, it will be referred to as an electronic device below.
[0033] According to one aspect of the present disclosure, a method embodiment for processing voice data is provided. Figure 1 A flowchart of a voice data processing method provided in this disclosure embodiment is shown below. Figure 1 As shown, the method includes:
[0034] Step S11: Obtain the target speech data to be recognized.
[0035] The method provided in this disclosure is applied to smart devices capable of voice processing, which may include: voice recording, voice recognition, voice synthesis, fluency detection, etc. Smart devices may include: personal computers, mobile devices (phones, tablets, etc.), smart wearable devices (smartwatches, smart bracelets, etc.), etc.
[0036] In this embodiment, the smart device first records the speech of the target subject, thereby obtaining target voice data. The target subject can be a trainee participating in language training or a speaker. Alternatively, the voice data can be obtained through transmission from other devices, such as other smart terminals or storage media like external hard drives. The target voice data includes multiple audio frames.
[0037] Step S12: Detect the target speech data to obtain the target feature sequence and the target time sequence. The target feature sequence includes the phoneme description information of each audio frame in the target speech data, and the target time sequence includes the time information corresponding to each phoneme in the target speech data.
[0038] In this embodiment of the disclosure, detecting target speech data to obtain target feature sequences and target time sequences includes the following steps A1-A4:
[0039] Step A1: Detect the target speech data to obtain the phonemes corresponding to each audio frame and the phoneme description information corresponding to each phoneme.
[0040] In this embodiment, the intelligent device inputs target speech data into a detection model, uses the detection model to detect the target speech data, and obtains the acoustic features of the target speech data. The acoustic features are Mel Frequency Cepstrum Coefficient (MFCC) features. Then, the phonemes corresponding to each audio frame in the target speech data are obtained from the acoustic features. A phoneme can be an element that makes up each speech sound; it is the smallest unit of language defined according to the natural attributes of language. It can be analyzed based on the pronunciation action of a syllable; one action constitutes one phoneme. For Chinese, phonemes can be divided into vowels and consonants. For example, the Chinese syllable 'a' has one phoneme, 'ai' has two phonemes, and 'dai' has three phonemes. For English, the International Phonetic Alphabet (IPA) has 48 phonemes, including 20 vowel phonemes and 28 consonant phonemes. To make the fluency detection results more accurate, this embodiment treats pauses as a silence phoneme.
[0041] For example, if the target speech data includes N audio frames, and if it includes 50 audio frames, then Mel-frequency cepstral coefficient features for the N audio frames will be generated. Then, an acoustic model is used to determine the phonemes corresponding to the Mel-frequency cepstral coefficient features of the N audio frames. This acoustic model can be a deep neural network-Hidden Markov model; specifically, the acoustic model includes the correspondence between the Mel-frequency cepstral coefficient features and the phonemes.
[0042] In this embodiment, the acoustic model includes an input layer, a hidden layer, a bottleneck layer, and an output layer. Specifically, the input layer passes the input Mel-frequency cepstral coefficient features to the hidden layer. The hidden layer determines the phoneme corresponding to each Mel-frequency cepstral coefficient feature and the initial phoneme description information corresponding to each phoneme. The initial phoneme description information can be features of multiple feature dimensions corresponding to the phoneme. The bottleneck layer averages the features of multiple dimensions corresponding to each phoneme, thereby ensuring that the feature dimensions of the obtained phoneme description information are consistent. Finally, the output layer outputs the phoneme description information of each phoneme.
[0043] Step A2: Obtain the target text corresponding to the target speech data.
[0044] In this embodiment of the disclosure, the target text can be the transcribed text corresponding to the target speech data. For example, the target speech data can be recognized and converted into text content to obtain the target text. Alternatively, the target text can also be a pre-uploaded document.
[0045] Step A3: Align the target text with the audio frames in the target speech data to obtain the time information corresponding to each phoneme. The time information includes the phoneme duration and the phoneme identifier.
[0046] In this embodiment of the disclosure, a corresponding text phoneme sequence is extracted from the target text, and the phonemes in the text phoneme sequence are aligned with the phonemes corresponding to the audio frames in the target speech data. This allows it to determine which audio frames in the target speech data each phoneme in the target text corresponds to. For example, in the speech data "Are you OK", the phoneme "A" corresponds to frames 1-n, and "R" corresponds to frames n+1-n+m, where m is greater than 1.
[0047] This method allows for the rapid identification of the audio frames corresponding to phonemes and the acquisition of time information for each phoneme. The time information includes the phoneme duration and its identifier. In essence, by determining the number of audio frames corresponding to each phoneme, the phoneme duration can be identified. This allows the phoneme identifier and duration to be used as time information to construct a time series. The phoneme identifier is a unique identifier for each phoneme; for example, phoneme A is identified as AA, phoneme R as R, and the silence phoneme as SIL.
[0048] Step A4: Construct a target feature sequence based on the phoneme description information corresponding to the phonemes in the target speech data, and construct a target time series based on the phoneme duration and phoneme identifier.
[0049] In this embodiment, the target feature sequence constructed based on phoneme description information and the target time sequence constructed based on phoneme duration and phoneme identifier are arranged to obtain the input feature sequence. Subsequently, the target fluency corresponding to the target speech data can be calculated based on the input feature sequence.
[0050] As an example, such as Figure 2As shown, the target speech data is "Are you OK", and its corresponding target time series includes: phoneme identifiers and phoneme durations. The phoneme identifiers include: "AA", "R", "SIL", "Y", "UW", "OW", "K", and "EY". The phoneme durations are "T1, T2, T3, T4, T5, T6, T7, and T8". The target feature sequences corresponding to the target speech data include: phoneme description information 1, phoneme description information 2, phoneme description information 3, phoneme description information 4, phoneme description information 5, phoneme description information 6, phoneme description information 7, and phoneme description information 8.
[0051] Step S13: Perform linear calculations based on the phoneme description information in the target feature sequence and the time information in the target time sequence to obtain the target fluency corresponding to the target speech data.
[0052] In this embodiment of the disclosure, a target fluency corresponding to the target speech data is obtained by linear calculation based on the phoneme description information in the target feature sequence and the time information in the target time sequence. This includes: obtaining a pre-trained fluency prediction model; inputting the phoneme description information, phoneme duration, and phoneme identifier into the fluency prediction model so that the fluency prediction model can obtain the target fluency by linear calculation based on the phoneme description information, phoneme identifier, and phoneme duration.
[0053] In this embodiment, the fluency prediction model can be based on a Long Short-Term Memory (LSTM) model or any suitable neural network model. Specifically, phoneme description information, phoneme duration, and phoneme identifier are input into the fluency prediction model. The fluency prediction model performs linear calculations based on the correlation between the three features to obtain the final target fluency.
[0054] The method provided in this disclosure extracts the target feature sequence and target time sequence of speech data. By using the phoneme description information in the target feature sequence and the time information in the target time sequence, it can accurately express fine-grained acoustic features, providing a reliable basis for the fluency detection of speech data and improving the reliability of the fluency detection results.
[0055] Figure 3 A flowchart illustrating a training method for a fluency prediction model provided in this disclosure embodiment is shown below. Figure 3 As shown, the method may include the following steps:
[0056] Step S21: Obtain speech data samples and the fluency labels corresponding to the speech data samples.
[0057] In this embodiment of the disclosure, the speech data sample can be obtained by collecting speech data from different objects, such as the speech data of multiple participants in an English speech contest, the speech data of a character in an English movie, or the speech data of students practicing in a training class, etc. Simultaneously, a fluency label corresponding to the speech data sample is obtained. The fluency label can be understood as a fluency score for the speech data sample; for example, the fluency score ranges from 1 to 10, with a higher score indicating higher fluency.
[0058] Step S22: Detect the speech data sample to obtain a feature sequence and a time sequence. The feature sequence includes phoneme description information of each audio frame in the speech data sample, and the target time sequence includes time information corresponding to each phoneme in the speech data sample.
[0059] In this embodiment, the smart device inputs speech data samples into a detection model, uses the detection model to detect the speech data samples, and obtains the acoustic features of the speech data samples, which are Mel-frequency cepstral coefficient features. Then, the phonemes corresponding to each audio frame in the speech data samples are obtained from the acoustic features. Finally, the phonemes are input into the acoustic model, and the acoustic model outputs phoneme description information for each phoneme, thereby constructing a feature sequence of the speech data samples using the phoneme description information.
[0060] In this embodiment of the disclosure, a text sample corresponding to a speech data sample is obtained, a corresponding text phoneme sequence is extracted from the text sample, and the phonemes in the text phoneme sequence are aligned with the phonemes corresponding to the audio frames in the speech data sample, thereby determining which audio frames in the speech data sample each phoneme in the text sample corresponds to. In this way, the phoneme identifier and the phoneme duration corresponding to each phoneme can be obtained, and finally, a time series is constructed based on the phoneme identifier and the phoneme duration as time information.
[0061] Step S23: Use the phoneme description information in the feature sequence, the time information in the time sequence, and the fluency label to train a preset neural network to obtain the predicted fluency.
[0062] In the application embodiment, a preset neural network is trained using phoneme description information in the feature sequence, time information in the time sequence, and fluency labels to obtain predicted fluency, including the following steps B1-B3:
[0063] Step B1: Determine the target phoneme description information to be masked in the feature sequence, and obtain the target time information corresponding to the target phoneme description information from the time series.
[0064] In this embodiment, a random masking method can be used to determine the target phoneme description information to be masked from the feature sequence. Specifically, a certain number of phoneme descriptions are randomly selected from the feature sequence according to a preset ratio as the target phoneme description information to be masked. For example, the feature sequence includes: phoneme description information 1, phoneme description information 2, phoneme description information 3, phoneme description information 4, phoneme description information 5, and phoneme description information 6. The preset ratio is 20%, and based on this, two phoneme descriptions are randomly selected from the above six phoneme descriptions as the target phoneme description information. It should be noted that when multiple target phoneme descriptions exist, these multiple target phoneme descriptions can be adjacent or not adjacent.
[0065] In this embodiment of the disclosure, after determining the target phoneme description information, the target time information corresponding to the target phoneme description information is found from the time series.
[0066] Step B2 involves masking the target phoneme description information in the feature sequence and the target time information in the time sequence to obtain the masked feature sequence and the masked time sequence.
[0067] In the embodiments disclosed herein, such as Figure 4 As shown, the feature sequence includes: phoneme description information 1, phoneme description information 2, phoneme description information 3, phoneme description information 4, phoneme description information 5, phoneme description information 6, phoneme description information 7, and phoneme description information 8. Phoneme description information 3 and phoneme description information 5 are masked to obtain a masked feature sequence. Simultaneously, the time information (phoneme identifier and phoneme duration) corresponding to phoneme description information 3 and phoneme description information 5 are masked to obtain a masked time sequence.
[0068] Step B3: Input the mask feature sequence and the mask time sequence into the preset neural network to obtain the predicted fluency.
[0069] In the embodiments disclosed herein, such as Figure 5 As shown, the preset neural network includes: a first prediction network, a second prediction network, and a linear network. The preset neural network can be a Long Short-Term Memory (LSTM) neural network model.
[0070] In this embodiment of the disclosure, the masked feature sequence and the masked time sequence are input into a preset neural network to obtain the predicted fluency, including the following process: First, the masked feature sequence and the masked time sequence are input into the preset neural network. The first prediction network of the preset neural network predicts the masked phoneme description information in the masked feature sequence to obtain the predicted phoneme description information. Simultaneously, the second prediction network of the preset neural network predicts the masked time information in the masked time sequence to obtain the predicted time information.
[0071] Secondly, determine the first loss value between the predicted phoneme description information and the target phoneme description information, and the second loss value between the predicted time information and the target time information.
[0072] If the first loss value is less than the first threshold and the second loss value is less than the second threshold, the mask feature sequence, predicted phoneme description information, mask time series, and prediction time information are passed to the linear network. Finally, the linear network performs linear calculations based on the mask feature sequence, predicted phoneme description information, mask time series, and prediction time information to obtain the predicted fluency.
[0073] If the first loss value is greater than or equal to the first threshold, and / or the second loss value is greater than or equal to the second threshold, then the parameters of the first prediction network and / or the second prediction network of the preset neural network are adjusted. This continues until the first loss value between the predicted phoneme description information output by the first prediction network and the target phoneme description information is less than the first threshold, and the second loss value between the predicted time information output by the second prediction network and the target time information is less than the second threshold.
[0074] Step S24: Adjust the model parameters of the preset neural network based on the predicted fluency and fluency label.
[0075] In this embodiment of the disclosure, adjusting the model parameters of a preset neural network based on the predicted fluency and the fluency label includes: determining a third loss value between the predicted fluency and the fluency label; and adjusting the parameters of the linear network in the preset neural network based on the third loss value.
[0076] The training method provided in this disclosure masks the target phoneme description information in the feature sequence and the target time information in the time sequence, and uses the obtained masked feature sequence and masked time sequence for model training. This enables the model to accurately predict the missing phoneme description information or time information even when the phoneme description information or time information is incomplete, and outputs the fluency detection result using the missing phoneme description information or time information. This further improves the accuracy and applicability of the model in the fluency detection process.
[0077] Figure 6 This is a block diagram of a voice data processing apparatus provided in an embodiment of the present disclosure. This apparatus can be implemented as part or all of an electronic device through software, hardware, or a combination of both. Figure 6 As shown, the device includes:
[0078] The acquisition module 61 is used to acquire the target speech data to be recognized.
[0079] The detection module 62 is used to detect the target speech data to obtain the target feature sequence and the target time sequence. The target feature sequence includes the phoneme description information of each audio frame in the target speech data, and the target time sequence includes the time information corresponding to each phoneme in the target speech data.
[0080] The prediction module 63 is used to perform linear calculations based on the phoneme description information in the target feature sequence and the time information in the target time sequence to obtain the target fluency corresponding to the target speech data.
[0081] In this embodiment of the disclosure, the target speech data includes multiple audio frames;
[0082] The detection module 62 is used to detect the target speech data to obtain the phonemes corresponding to each audio frame and the phoneme description information corresponding to each phoneme; to obtain the target text corresponding to the target speech data; to align the target text with the audio frames in the target speech data to obtain the time information corresponding to each phoneme, wherein the time information includes the phoneme duration and the phoneme identifier; to construct a target feature sequence based on the phoneme description information corresponding to the phonemes in the target speech data, and to construct a target time series based on the phoneme duration and the phoneme identifier corresponding to the phonemes.
[0083] In this embodiment of the present disclosure, the prediction module 63 is used to obtain a pre-trained fluency prediction model; input phoneme description information, phoneme duration and phoneme identifier into the fluency prediction model so that the fluency prediction model performs linear calculation based on the phoneme description information, phoneme identifier and phoneme duration to obtain the target fluency.
[0084] In this embodiment of the disclosure, the speech data processing apparatus further includes: a training module, which includes:
[0085] The acquisition unit is used to acquire speech data samples and the corresponding fluency labels for the speech data samples.
[0086] The processing unit is used to detect speech data samples to obtain feature sequences and time sequences. The feature sequences include phoneme description information of each audio frame in the speech data samples, and the target time sequence includes time information corresponding to each phoneme in the speech data samples.
[0087] The execution unit is used to train a preset neural network using phoneme description information in the feature sequence, time information in the time sequence, and fluency labels to obtain the predicted fluency.
[0088] The optimization unit is used to adjust the model parameters of the preset neural network based on the predicted fluency and fluency label.
[0089] In this embodiment of the disclosure, the execution unit is used to determine the target phoneme description information to be masked in the feature sequence, and obtain the target time information corresponding to the target phoneme description information from the time sequence; mask the target phoneme description information in the feature sequence and the target time information in the time sequence respectively to obtain the masked feature sequence and the masked time sequence; input the masked feature sequence and the masked time sequence into a preset neural network to obtain the predicted fluency.
[0090] In this embodiment of the disclosure, the preset neural network includes: a first prediction network, a second prediction network, and a linear network;
[0091] In this embodiment of the disclosure, the execution unit is configured to predict the masked phoneme description information in the masked feature sequence through a first prediction network to obtain the predicted phoneme description information; predict the masked time information in the masked time sequence through a second prediction network to obtain the predicted time information; and perform linear calculations based on the masked feature sequence, the predicted phoneme description information, the masked time sequence, and the predicted time information through a linear network to obtain the predicted fluency.
[0092] In this embodiment of the disclosure, the execution unit is configured to determine a first loss value between the predicted phoneme description information and the target phoneme description information, and a second loss value between the predicted time information and the target time information; and when the first loss value is less than a first threshold and the second loss value is less than a second threshold, to transmit the mask feature sequence, the predicted phoneme description information, the mask time sequence and the predicted time information to the linear network.
[0093] In this embodiment of the present disclosure, the optimization unit is used to determine a third loss value between the predicted fluency and the fluency label; and to adjust the parameters of the linear network in the preset neural network based on the third loss value.
[0094] This disclosure also provides an electronic device, such as... Figure 7 As shown, the electronic device may include: a processor 1501, a communication interface 1502, a memory 1503, and a communication bus 1504, wherein the processor 1501, the communication interface 1502, and the memory 1503 communicate with each other through the communication bus 1504.
[0095] Memory 1503 is used to store computer programs;
[0096] When the processor 1501 executes the computer program stored in the memory 1503, it implements the steps of the above embodiments.
[0097] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0098] The communication interface is used for communication between the aforementioned terminal and other devices.
[0099] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0100] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0101] In another embodiment provided in this disclosure, a computer-readable storage medium is also provided, which stores instructions that, when executed on a computer, cause the computer to perform any of the voice data processing methods described in the above embodiments.
[0102] In yet another embodiment provided in this disclosure, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the voice data processing methods described in the above embodiments.
[0103] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this disclosure are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk).
[0104] The above description is merely a preferred embodiment of this disclosure and is not intended to limit the scope of protection of this disclosure. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure are included within the scope of protection of this disclosure.
[0105] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for processing voice data, characterized in that, include: Acquire target speech data to be identified, wherein the target speech data includes multiple audio frames; The target speech data is detected to obtain a target feature sequence and a target time sequence, wherein the target feature sequence includes phoneme description information of each audio frame corresponding to the phoneme in the target speech data, and the target time sequence includes time information corresponding to each phoneme in the target speech data; Linear calculations are performed based on the phoneme description information in the target feature sequence and the time information in the target time sequence to obtain the target fluency corresponding to the target speech data. The step of detecting the target speech data to obtain the target feature sequence and the target time sequence includes: detecting the target speech data to obtain the phonemes corresponding to each audio frame and the phoneme description information corresponding to each phoneme; obtaining the target text corresponding to the target speech data, aligning the target text with the audio frames in the target speech data to obtain the time information corresponding to each phoneme, wherein the time information includes the phoneme duration and the phoneme identifier; constructing the target feature sequence based on the phoneme description information corresponding to the phonemes in the target speech data, and constructing the target time sequence based on the phoneme duration and the phoneme identifier corresponding to the phonemes; The step of linearly calculating the target fluency corresponding to the target speech data based on the phoneme description information in the target feature sequence and the time information in the target time sequence includes: acquiring a pre-trained fluency prediction model; inputting the phoneme description information, the phoneme duration, and the phoneme identifier into the fluency prediction model, so that the fluency prediction model linearly calculates the target fluency based on the phoneme description information, the phoneme identifier, and the phoneme duration.
2. The method according to claim 1, characterized in that, The training method for the fluency prediction model includes: Obtain speech data samples and the fluency labels corresponding to the speech data samples; The speech data sample is detected to obtain a feature sequence and a time sequence, wherein the feature sequence includes phoneme description information of each audio frame in the speech data sample, and the target time sequence includes time information corresponding to each phoneme in the speech data sample. Using the phoneme description information in the feature sequence, the time information in the time sequence, and the fluency label, a preset neural network is trained to obtain the predicted fluency. Based on the predicted fluency and the fluency label, the model parameters of the preset neural network are adjusted.
3. The method according to claim 2, characterized in that, The step of training a preset neural network using phoneme description information in the feature sequence, time information in the time sequence, and fluency labels to obtain predicted fluency includes: Determine the target phoneme description information to be masked in the feature sequence, and obtain the target time information corresponding to the target phoneme description information from the time sequence; The target phoneme description information in the feature sequence and the target time information in the time sequence are masked respectively to obtain the masked feature sequence and the masked time sequence; The mask feature sequence and the mask time sequence are input into the preset neural network to obtain the predicted fluency.
4. The method according to claim 3, characterized in that, The preset neural network includes: a first prediction network, a second prediction network, and a linear network; The step of inputting the mask feature sequence and the mask time sequence into the preset neural network to obtain the predicted fluency includes: The predicted phoneme description information is obtained by predicting the masked phoneme description information in the masked feature sequence through the first prediction network. The second prediction network is used to predict the masked time information in the masked time series to obtain the predicted time information. The predicted fluency is obtained by performing linear calculations using the linear network based on the mask feature sequence, the predicted phoneme description information, the mask time sequence, and the prediction time information.
5. The method according to claim 4, characterized in that, Before obtaining the predicted fluency by performing linear calculations using the linear network based on the mask feature sequence, the predicted phoneme description information, the mask time sequence, and the prediction time information, the method further includes: Determine a first loss value between the predicted phoneme description information and the target phoneme description information, and a second loss value between the predicted time information and the target time information; If the first loss value is less than the first threshold and the second loss value is less than the second threshold, the mask feature sequence, the predicted phoneme description information, the mask time sequence, and the prediction time information are passed to the linear network.
6. The method according to claim 4, characterized in that, The step of adjusting the model parameters of the preset neural network based on the predicted fluency and the fluency label includes: Determine a third loss value between the predicted fluency and the fluency label; The parameters of the linear network in the preset neural network are adjusted based on the third loss value.
7. A voice data processing apparatus, characterized in that, include: The acquisition module is used to acquire target speech data to be identified, wherein the target speech data includes multiple audio frames; The detection module is used to detect the target speech data to obtain a target feature sequence and a target time sequence, wherein the target feature sequence includes phoneme description information of each phoneme corresponding to each audio frame in the target speech data, and the target time sequence includes time information corresponding to each phoneme in the target speech data; The prediction module is used to perform linear calculations based on the phoneme description information in the target feature sequence and the time information in the target time sequence to obtain the target fluency corresponding to the target speech data. The detection module is used to detect the target speech data to obtain the phonemes corresponding to each audio frame and the phoneme description information corresponding to each phoneme; to obtain the target text corresponding to the target speech data, align the target text with the audio frames in the target speech data, and obtain the time information corresponding to each phoneme, wherein the time information includes the phoneme duration and the phoneme identifier; to construct the target feature sequence based on the phoneme description information corresponding to the phonemes in the target speech data, and to construct the target time series based on the phoneme duration and the phoneme identifier corresponding to the phonemes; The prediction module is used to obtain a pre-trained fluency prediction model; input the phoneme description information, the phoneme duration, and the phoneme identifier into the fluency prediction model, so that the fluency prediction model performs linear calculations based on the phoneme description information, the phoneme identifier, and the phoneme duration to obtain the target fluency.
8. A storage medium, characterized in that, The storage medium includes a stored program, wherein the program executes the method according to any one of claims 1 to 6 when it is run.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, communication interface, and memory communicate with each other through the communication bus; wherein: Memory, used to store computer programs; A processor for performing the method of any one of claims 1 to 6 by running a program stored in memory.