Voice data processing method and device, computer device and storage medium

By automating the alignment and annotation of phoneme labels on speech data, the problem of inconsistent annotation standards in speech synthesis systems is solved, the annotation efficiency and accuracy are improved, and more natural speech output is generated.

CN115223536BActive Publication Date: 2026-02-10VOICEAI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210658186.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-10
Publication Date
2026-02-10
Estimated Expiration
2042-06-10

AI Technical Summary

Technical Problem

In existing speech synthesis systems, it is difficult to unify the standards for speech data annotation, resulting in high error rates and low efficiency, which affects system performance.

Method used

Speech recognition is performed by acquiring speech data to be processed, the time interval between adjacent groups of phoneme label data is determined, the phoneme label data is labeled using preset pause labels, target phoneme label data is generated, and this data is used as training data to train the speech generation model.

Benefits of technology

It improves the accuracy and efficiency of speech data annotation, ensuring that the speech generation model can better express the intonation and rhythm of speech and generate natural speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115223536B_ABST
    Figure CN115223536B_ABST
Patent Text Reader

Abstract

The application discloses a speech data processing method and device, computer equipment and a storage medium, applied to the technical field of data processing, and the method comprises the steps of obtaining speech data to be processed for speech recognition, obtaining phoneme label data corresponding to the speech data to be processed, aligning the phoneme label data with the speech data to be processed to obtain a phoneme alignment result, determining a first pause time interval between adjacent phoneme label groups in the phoneme label data according to the phoneme alignment result, labeling the phoneme label data according to the first pause time to obtain target phoneme label data, and training a speech generation model by taking the target phoneme label data as training data. In this way, the generation, labeling and training of pause time between phoneme label data are automatically performed, and the labeling efficiency of speech data is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, and particularly relates to a voice data processing method and device, computer equipment and a storage medium. BACKGROUND

[0002] Speech synthesis technology is one of the key technologies for converting text into sound, which can enable electronic devices such as computers and robots to have speaking abilities similar to humans, and is an important competitive market in today's information industry. At present, people use deep learning algorithms to build a speech synthesis system, and use a large amount of speech training data to train the speech synthesis system, so as to obtain a speech synthesis system that can be put into application.

[0003] At present, professional recording equipment is usually used to record voice data, and the voice data is manually annotated to obtain voice training data. However, due to the subjective influence of the annotators, the annotation standard is difficult to unify, the error rate is high, and the efficiency is low when the voice data is annotated, thereby affecting the performance of the entire speech synthesis system. SUMMARY

[0004] The present application provides a voice data processing method and device, computer equipment and a storage medium, which improves the annotation efficiency of voice data.

[0005] A voice data processing method comprises the following steps:

[0006] Obtaining voice recognition of the to-be-processed voice data to obtain phoneme label data corresponding to the to-be-processed voice data;

[0007] Aligning the phoneme label data with the to-be-processed voice data to obtain a phoneme alignment result;

[0008] According to the phoneme alignment result, determining a first pause time interval between adjacent phoneme label groups in the phoneme label data;

[0009] According to the first pause time, the phoneme label data is annotated to obtain target phoneme label data;

[0010] The target phoneme label data is used as training data to train a voice generation model.

[0011] A voice data processing device comprises the following steps:

[0012] A voice recognition module is configured to obtain voice recognition of the to-be-processed voice data to obtain phoneme label data corresponding to the to-be-processed voice data;

[0013] an alignment module configured to align the phoneme label data with the to-be-processed speech data to obtain a phoneme alignment result;

[0014] a first pause time determination module configured to determine, according to the phoneme alignment result, a first pause time between adjacent phoneme label groups in the phoneme label data;

[0015] a labeling module configured to label the phoneme label data according to the first pause time to obtain target phoneme label data;

[0016] a training module configured to train a speech generation model using the target phoneme label data as training data.

[0017] A computer device includes a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor implements the steps of the speech data processing method when executing the computer program.

[0018] A computer readable storage medium stores a computer program, and the computer program implements the steps of the speech data processing method when executed by a processor.

[0019] The speech data processing method, device, computer device, and storage medium provided in the present application obtain to-be-processed speech data for speech recognition to obtain phoneme label data corresponding to the to-be-processed speech data, align the phoneme label data with the to-be-processed speech data to obtain a phoneme alignment result, determine a first pause time between adjacent phoneme label groups in the phoneme label data according to the phoneme alignment result, label the phoneme label data according to the first pause time to obtain target phoneme label data, and train a speech generation model using the target phoneme label data as training data. In the present application, the phoneme label data obtained by performing speech recognition on the to-be-processed speech data is aligned with the to-be-processed speech data, the length of time of the first pause time between adjacent phoneme label groups in the phoneme label data is accurately determined, the phoneme label data is quickly labeled, and the labeling accuracy and efficiency of the speech data are improved. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0021] Figure 1This is a schematic diagram of an application environment for a voice data processing method according to an embodiment of this application;

[0022] Figure 2 This is a flowchart of a voice data processing method according to an embodiment of this application;

[0023] Figure 3 This is a schematic diagram of the structure of a voice data processing device in another embodiment of this application;

[0024] Figure 4 This is a schematic diagram of the structure of a voice data processing device in another embodiment of this application;

[0025] Figure 5 This is a schematic diagram of the structure of a voice data processing device in another embodiment of this application;

[0026] Figure 6 This is a schematic diagram of the structure of a voice data processing device in one embodiment of this application;

[0027] Figure 7 This is a schematic diagram of a computer device according to one embodiment of this application. Detailed Implementation

[0028] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0029] The voice data processing method provided in this application embodiment can be applied to, for example, Figure 1 In this application environment, computer devices and terminal devices communicate with the server via a network. These computer devices and terminal devices can be, but are not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be a standalone server or a server cluster consisting of multiple servers.

[0030] System framework 100 may include terminal devices, a network, and a server. Network 104 is used as a medium to provide a communication link between the terminal devices and the server. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0031] Users can use terminal devices to interact with the server over the network to receive or send messages, etc.

[0032] Terminal devices 101, 102, and 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Eperts Group Audio Layer III), MP4 players (Moving Picture Eperts Group Audio Layer IV), laptops, and desktop computers, etc.

[0033] Server 105 can be a server that provides various services, such as a backend server that supports the pages displayed on terminal devices 101, 102, and 103.

[0034] It should be noted that the voice data processing method provided in this application embodiment is executed by the server, and correspondingly, the voice data processing device is set in the server.

[0035] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown in this embodiment is merely illustrative. Depending on the implementation requirements, there can be any number of terminal devices, networks, and servers. The terminal devices in this embodiment can specifically correspond to application systems in actual production.

[0036] In one embodiment, such as Figure 2 As shown, a voice data processing method is provided, which is applied to... Figure 1 Taking the server in the example, the following steps S201 to S205 are explained:

[0037] Step S201: Obtain the speech data to be processed and perform speech recognition to obtain the phoneme tag data corresponding to the speech data to be processed.

[0038] The speech data to be processed can be speech data composed of at least one audio signal. An audio signal represents a mechanical wave and is the information carrier of changes in the wavelength and intensity of the mechanical wave; it can be an analog signal or a digital signal. Phoneme tag data can be data composed of phoneme tags, with each phoneme tag corresponding to one phoneme. A phoneme is the smallest unit or the smallest speech segment that constitutes a syllable; it is the smallest linear unit of speech defined from the perspective of sound quality. It can be a syllable in Chinese (e.g., ā, à, d), a syllable in Japanese (e.g., ku, ā, ah), or a phonetic symbol in English (e.g., dr, s, ...). :), Phoneme labels are determined based on the actual application scenario, and no specific restrictions are made here.

[0039] For example, the phoneme tag data could be “nin hao qing nin gei wo ji fen zhong deshi jian ting wo jie shao xia zhe xiang zeng zhi ye wu”, and the corresponding text could be “Hello, please give me a minute to listen to my introduction of this value-added service”.

[0040] In this application, a recording device can be used to collect voice data of specific personnel in a specific scenario (such as a call center scenario), and this voice data can be used as voice data to be processed. The specific personnel can be customer service representatives, hosts, radio anchors, etc., and the selection of specific personnel can be based on the actual application scenario; no specific limitation is made here.

[0041] Optionally, a pre-trained speech recognition system can be used to perform speech recognition on the speech data to be processed, obtaining phoneme label data corresponding to the speech data to be processed. The speech recognition system includes an encoding layer, a neural network layer, and a decoding layer. The encoding layer encodes the speech data to be processed to obtain encoded speech data. Then, the neural network layer extracts features from the encoded speech data to obtain speech feature data. Finally, the decoding layer identifies phoneme labels from the speech feature data to obtain phoneme label data. The speech feature data can be a feature representation of the speech data, such as timbre, pitch, speech rate, etc. Optionally, the pre-trained speech recognition model can be implemented by training the neural network.

[0042] Step S202: Align the phoneme tag data with the speech data to be processed to obtain the phoneme alignment result.

[0043] The phoneme alignment result includes the alignment result information of the phoneme tag group (e.g., wo) corresponding to the word at the i-th time (e.g., from 23 minutes 01 seconds to 23 minutes 02 seconds) in the phoneme tag data and the audio signal segment at the i-th time (e.g., from 23 minutes 01 seconds to 23 minutes 02 seconds) in the speech data to be processed. The alignment result information may include the time length occupied by each phoneme tag group in the speech data to be processed and the pause time length between adjacent phoneme tag groups.

[0044] Optionally, the phoneme label data and the speech data to be processed are input into a pre-trained acoustic model for prediction to obtain the phoneme alignment result. The pre-trained acoustic model can be implemented by training a neural network.

[0045] Optionally, the phoneme alignment results can be displayed on the terminal page in a visual format such as charts.

[0046] Step S203: Determine the first pause time between adjacent phoneme label groups in the phoneme label data according to the phoneme alignment result.

[0047] Each phoneme label group corresponds to a character (such as "wo" corresponding to "我") or a word (such as corresponding to "word"). The first pause time can be the pause time length between adjacent phoneme label groups in the speech data to be processed. For a better understanding of the first pause time between adjacent phoneme label groups, exemplarily, assume that the characters are "我" and "和", and the phoneme label groups are "wo" and "he". The first pause time is the pause time length between the phoneme label "o" and the phoneme label "h".

[0048] Step S204: Annotate the phoneme label data according to the first pause time to obtain the target phoneme label data.

[0049] Specifically, the preset pause label can be obtained according to the first pause time, and the phoneme label data can be annotated with the pause label to obtain the target phoneme label data.

[0050] Assume that the preset pause labels can be 1, 2, 3, specifically long, medium, and short pauses. The pause label between each phoneme label data group can be matched according to the first pause time, and the phoneme label data can be annotated between every two phoneme label data groups with the pause label to obtain the target phoneme label data. For example, when the phoneme label data is "您~~~好~~,我~是~客~服~人~员", where "~" represents the first pause time length, the pause label between each phoneme label data group can be matched according to the first pause time, that is, the long pause "~~~" matches the pause label "1", the medium pause "~~" matches the pause label "2", and the short pause "~" matches the pause label "3", and the phoneme label data is annotated between every two phoneme label data groups with the pause label. The obtained target phoneme label data can be expressed as "nin1hao2, wo3shi3ke3fu3ren3yuan".

[0051] Optionally, the first pause time and the phoneme label data can be input into a pre-trained annotation model for annotation to obtain the target phoneme label data. Among them, the pre-trained annotation model can be a conditional random field model, and the pause labels for annotation are preset in the pre-trained annotation model.

[0052] Step S205: Use the target phoneme label data as training data to train the speech generation model.

[0053] The speech generation model is a model used to convert text information into speech.

[0054] Understandably, speech generation models can be implemented based on deep neural network modeling. The speech generation model is trained using target phoneme label data until it converges and training is complete. Since there are pause labels between adjacent phoneme label groups in the target phoneme label data, and each pause label represents the corresponding pause time, the speech generation model can learn the pause rules between different phoneme label groups (i.e., words). Subsequently, after inputting the phoneme label data of a text into the trained speech generation model, it can automatically generate speech containing pauses based on context (i.e., intonation), which can better express the effect of speech playback and improve the expressive effect of speech playback.

[0055] In this embodiment, the speech data to be processed is acquired and speech recognition is performed to obtain the phoneme label data corresponding to the speech data to be processed. The phoneme label data is aligned with the speech data to be processed, which can quickly and accurately determine the length of the first pause time between adjacent phoneme label groups in the phoneme label data, so as to facilitate the annotation of the phoneme label data. In this way, by automatically generating, annotating and training the pause time between phoneme label data, the efficiency of speech data annotation is greatly improved.

[0056] In some optional implementations of this embodiment, such as Figure 3 As shown, step S201 involves acquiring the speech data to be processed and performing speech recognition to obtain the phoneme tag data corresponding to the speech data to be processed, including the following steps S2010 to S2011:

[0057] Step S2010: Perform speech recognition on the speech data to be processed according to the preset frame length to obtain the phoneme tag sequence corresponding to each preset frame length.

[0058] The preset frame length can be determined based on the length of the audio signal in the speech data to be processed. For example, if the audio signal is short, the preset frame length can be 10ms. The preset frame length can be adjusted according to the actual situation to reduce the influence of non-stationary and time-varying audio signals in the speech data to be processed, thereby improving the accuracy of speech recognition and ensuring the accuracy of the obtained phoneme tag sequence.

[0059] Step S2011: Merge the phoneme tag sequences corresponding to each preset frame length in chronological order to obtain the phoneme tag data corresponding to the speech data to be processed.

[0060] The time order refers to the order of the times corresponding to each audio signal in the speech data to be processed.

[0061] In this embodiment, performing speech recognition on the speech data to be processed according to a preset normality can reduce the non-stationary and time-varying audio signals in the speech data to be processed, improve the accuracy of speech recognition, thereby obtaining an accurate phoneme label sequence, and then merging the phoneme label data to obtain phoneme label data, which is beneficial to improving the labeling accuracy and labeling efficiency of speech data.

[0062] In some optional implementations of this embodiment, step S202, aligning the phoneme tag data with the speech data to be processed to obtain the phoneme alignment result, includes:

[0063] Extract the temporal and frequency features of the speech data to be processed.

[0064] Specifically, temporal and frequency features can be extracted from the speech data to be processed based on neural networks.

[0065] Based on temporal and frequency characteristics, the phoneme distribution locations corresponding to the speech data to be processed are obtained.

[0066] The phoneme distribution location can be the time interval of each phoneme tag group in the audio signal of the speech data to be processed. For example, the time interval of the phoneme tag group "wo" in the audio signal of the speech data to be processed is from 23 minutes 40 seconds to 23 minutes 42 seconds. Specifically, the phoneme tag group is determined according to the frequency characteristics, and the time interval of the phoneme tag group in the audio signal of the speech data to be processed is determined according to the temporal characteristics, thereby determining the phoneme distribution location.

[0067] Based on the phoneme distribution location, the phoneme label data and the speech data to be processed are phoneme aligned to obtain the phoneme alignment result.

[0068] Specifically, based on the location of phoneme distribution, the phoneme tag groups in the phoneme tag data are aligned with the speech data to be processed to obtain the phoneme alignment result.

[0069] In this embodiment, the phoneme distribution position of the speech data to be processed is determined by the temporal and frequency features of the speech data to be processed. The phoneme distribution position makes it easy to determine the position of the phoneme tag data in the speech data to be processed. The phoneme alignment result is determined quickly and automatically, thereby improving the annotation efficiency of speech data.

[0070] In some optional implementations of this embodiment, such as Figure 4 As shown, step S204 involves labeling the phoneme tags according to the first pause time to obtain the target phoneme tag data, including the following steps S2040 to S2041:

[0071] Step S2040: Based on the first pause time, retrieve the corresponding pause label from the preset label database.

[0072] The preset label database includes multiple pause labels, each with a corresponding index value. Specifically, the index value in the preset label database can be determined by the first pause time. The corresponding pause label is then retrieved from the vehicle label database based on the index value. For example, the pause label can be 1, 2, or 3, where 1 represents a short pause, 2 represents a medium pause, and 3 represents a long pause. A mapping table between the first pause time and the index value can be preset, and the index value in the preset label database is determined using this mapping table.

[0073] For example, assuming the first pause time is 3ms, the index value corresponding to 3ms in the relation mapping table is A, and the pause label corresponding to index value A in the preset label database is 3 (which can represent a short pause label), then 3 can be obtained from the preset label database in 3ms.

[0074] Step S2041: Use pause tags to mark the interval positions of the phoneme tag data to obtain the target phoneme tag data.

[0075] The interval positions are the positions between every two phoneme tag groups in the phoneme tag data, and the positions after the last phoneme tag group in the phoneme tag data.

[0076] In this embodiment, a preset tag database is used, and pause tags are obtained from the preset tag database according to the first pause time. The interval positions of the phoneme tag data are automatically marked, which has a unified marking standard and quantitative processing, which is conducive to improving the marking efficiency of speech data.

[0077] In some optional implementations of this embodiment, the preset tag database includes at least two types of pause tags, each type of pause tag carrying a different second pause time. Step S2040 involves retrieving the corresponding pause tag from the preset tag database based on the first pause time, including:

[0078] The first pause time and the second pause time are matched to obtain the matching result.

[0079] The second pause time is the length of pause time between phoneme tag groups in the historical speech data to be processed, obtained after analyzing historical experience data.

[0080] Optionally, a similarity algorithm can be used to calculate the similarity between the first pause time and the second pause time to obtain a similarity value. If the similarity value is greater than or equal to a preset threshold (the preset threshold can be obtained based on historical experience), then the matching result is determined to be a match between the first pause time and the second pause time. The similarity algorithm can be cosine similarity, Euclidean distance, etc., and the algorithm can be selected according to actual needs; no limitation is made here.

[0081] Based on the matching results, retrieve the pause label of the corresponding type for each first pause time from the preset label database.

[0082] Specifically, if the matching result is that the first pause time and the second pause time match, then the pause tag corresponding to the second pause time is obtained from the preset tag database and used as the pause tag of the type corresponding to the first pause time.

[0083] In this embodiment, by setting different types of pause labels in a preset database, and each type of pause label carries a corresponding second pause time, the pause label can be quickly determined by matching the first pause time and the second pause time, which is beneficial to improving the labeling accuracy and labeling efficiency of speech data.

[0084] In some optional implementations of this embodiment, before step S205, which uses the target phoneme label data as training data to train the speech generation model, the speech data processing method further includes:

[0085] Speech recognition is performed on the speech data to be processed to obtain the corresponding text data.

[0086] Specifically, speech recognition technology can be used to perform speech recognition on the speech data to be processed, thereby obtaining the corresponding text data. Among them, semantic recognition technology, also known as Automatic Speech Recognition (ASR), aims to convert the lexical content of human speech into computer-readable input, such as keystrokes, binary codes, or character sequences. Speech recognition technology can be based on randomized model methods or neural network methods.

[0087] Text data and speech data to be processed are input into a pre-trained emotion recognition model to perform emotion recognition and determine the emotion category of the speech data to be processed.

[0088] Among them, the pre-trained emotion recognition model can be implemented based on neural network modeling, and the emotion categories can be happy, angry, upset, sad, etc.

[0089] Based on the emotion category, obtain the emotion label, and use the emotion label to mark the target phoneme label data with emotion.

[0090] Among them, emotion tags correspond one-to-one with emotion categories. Emotion tags can be preset according to emotion categories, and emotion categories and emotion tags can be set according to actual application scenarios. No specific restrictions are made here.

[0091] In this embodiment, by identifying the emotion category of the speech data to be processed and obtaining the corresponding emotion label based on the emotion category to mark the target phoneme label data, the speech generated by the speech generation model trained with the target phoneme label data can be made more natural.

[0092] In some optional implementations of this embodiment, such as Figure 5 As shown, before using the target phoneme label data as training data to train the speech generation model in step S205, the speech data processing method further includes the following steps a to d:

[0093] Step a: Perform speech recognition on the speech data to be processed to obtain the text data corresponding to each sentence and the time length of each sentence.

[0094] Specifically, speech recognition technology can be used to perform speech recognition on the speech data to be processed, thereby obtaining the text data corresponding to each sentence of the speech data to be processed and the time length of each sentence.

[0095] Step b: Count the number of words in the text data of each sentence to obtain the word count of each sentence.

[0096] Step c: Determine the speech rate category of the speech data to be processed based on the number of words in each sentence and the duration of each sentence.

[0097] The speech rate categories include fast, slow, and normal.

[0098] Step d: Obtain speech rate labels according to speech rate category, and use speech rate labels to mark the speech rate of target phoneme label data.

[0099] Specifically, speech rate tags correspond one-to-one with speech rate categories. Speech rate tags can be preset according to speech rate categories, and speech rate categories can be set according to actual application scenarios. No specific limitations are made here.

[0100] In this embodiment, by identifying the speech rate category of the speech data to be processed and obtaining the corresponding speech rate label according to the speech rate category to mark the target phoneme label data, the speech generated by the speech generation model trained with the target phoneme label data can be made more natural.

[0101] In some optional implementations of this embodiment, step c, determining the speech rate category of the speech data to be processed based on the number of words and the duration of each sentence, includes:

[0102] If at least two sentences have the same number of characters, then obtain the time length corresponding to at least two sentences.

[0103] Calculate the average duration of the time lengths corresponding to at least two sentences.

[0104] The speech rate category of the speech data to be processed is determined based on the average duration.

[0105] Specifically, a time range is preset for each speech rate category. The speech rate category of the speech data to be processed is determined based on the average duration and the time range. The time range for each speech rate category can be obtained by analyzing historical speech data of the person whose speech data is being processed in the same application scenario. For example, when the speech rate category is fast, the corresponding time range is 4 to 7 seconds; when the speech rate category is normal, the corresponding time range is 8 to 10 seconds; and when the speech rate category is slow, the corresponding time range is greater than or equal to 11 seconds. When the average duration is 8.5 seconds, the speech rate category of the speech data to be processed is determined to be normal.

[0106] In this embodiment, by using the average duration of sentences with the same number of words in the speech data to be processed, and based on the relationship between the number of words and the duration, the accuracy of determining the speech rate category of the speech data to be processed is improved. This helps to make the speech generated by the speech generation model trained with target phoneme label data more natural.

[0107] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0108] In one embodiment, a voice data processing apparatus is provided, which corresponds one-to-one with the voice data processing methods described in the above embodiments. For example... Figure 6 As shown, the speech data processing device includes a speech recognition module 30, an alignment module 31, a first pause time determination module 32, an annotation module 33, and a training module 34. Detailed descriptions of each functional module are as follows:

[0109] The speech recognition module 30 is used to acquire the speech data to be processed and perform speech recognition to obtain the phoneme tag data corresponding to the speech data to be processed.

[0110] Alignment module 31 is used to align phoneme tag data with speech data to be processed to obtain phoneme alignment results.

[0111] The first pause time determination module 32 is used to determine the first pause time between adjacent phoneme tag groups in the phoneme tag data based on the phoneme alignment result.

[0112] The annotation module 33 is used to annotate the phoneme label data according to the first pause time to obtain the target phoneme label data.

[0113] Training module 34 is used to train the speech generation model using the target phoneme label data as training data.

[0114] Optionally, the speech recognition module 30 includes:

[0115] The speech recognition submodule is used to perform speech recognition on the speech data to be processed according to a preset frame length, and obtain the phoneme tag sequence corresponding to each preset frame length.

[0116] The merging submodule is used to merge the phoneme tag sequences corresponding to each preset frame length in chronological order to obtain the phoneme tag data corresponding to the speech data to be processed.

[0117] Optionally, alignment module 31 includes:

[0118] The feature extraction submodule is used to extract the temporal and frequency features of the speech data to be processed.

[0119] The phoneme distribution location acquisition submodule is used to obtain the phoneme distribution location corresponding to the speech data to be processed based on temporal and frequency features.

[0120] The alignment submodule is used to perform phoneme alignment processing on the phoneme label data and the speech data to be processed according to the phoneme distribution position, so as to obtain the phoneme alignment result.

[0121] Optionally, standard module 33 includes:

[0122] The pause tag acquisition submodule is used to retrieve the corresponding pause tag from the preset tag database based on the first pause time.

[0123] The annotation submodule is used to annotate the interval positions of the phoneme tag data using pause tags to obtain the target phoneme tag data.

[0124] Optionally, the preset tag database includes at least two types of pause tags, each type of pause tag carrying a different second pause time. The pause tag acquisition submodule includes:

[0125] The matching result acquisition unit is used to match the first pause time and the second pause time to obtain the matching result.

[0126] The pause label acquisition unit is used to retrieve the pause label of the corresponding type for each first pause time from the preset label database based on the matching results.

[0127] Optionally, the voice processing device also includes:

[0128] The text data acquisition module is used to perform speech recognition on the speech data to be processed and obtain the corresponding text data.

[0129] The emotion recognition module is used to input text data and speech data to be processed into a pre-trained emotion recognition model to identify the emotion category of the speech data to be processed.

[0130] The emotion tagging module is used to obtain emotion tags based on emotion categories and to tag the target phoneme tag data with emotion tags.

[0131] Optionally, the voice processing device also includes:

[0132] The time length acquisition module is used to perform speech recognition on the speech data to be processed, and to obtain the text data corresponding to each sentence of the speech data to be processed, as well as the time length of each sentence.

[0133] The word count module is used to count the number of words in the text data of each sentence and obtain the word count of each sentence.

[0134] The speech rate category determination module is used to determine the speech rate category of the speech data to be processed based on the number of words in each sentence and the duration of each sentence.

[0135] The speech rate tagging module is used to obtain speech rate tags according to speech rate category and to use speech rate tags to tag the target phoneme tag data.

[0136] Optional, the speech rate category determination module includes:

[0137] The judgment submodule is used to obtain the time length of at least two sentences if there are at least two sentences with the same number of characters.

[0138] The calculation submodule is used to calculate the average duration of the time lengths corresponding to at least two sentences.

[0139] The speech rate category determination submodule is used to determine the speech rate category of the speech data to be processed based on the average duration.

[0140] The terms "first" and "second" in the above-mentioned modules / units are only used to distinguish different modules / units and are not intended to specify which module / unit has a higher priority or any other limiting meaning. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those steps or modules explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or devices. The module divisions appearing in this application are merely logical divisions; in actual applications, different division methods may be used.

[0141] Specific limitations regarding the voice data processing device can be found in the limitations of the voice data processing method described above, and will not be repeated here. Each module in the aforementioned voice data processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0142] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data involved in the voice data processing method. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements a voice data processing method.

[0143] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the voice data processing method described in the above embodiments, for example... Figure 2 The steps 201 to 205 shown, as well as other extensions and related steps of the method, are examples. Alternatively, when the processor executes a computer program, it implements the functions of each module / unit of the voice data processing device in the above embodiments, for example... Figure 6 The functions of modules 30 to 34 are shown. To avoid repetition, they will not be described again here.

[0144] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the computer device, connecting various parts of the computer device via various interfaces and lines.

[0145] The memory can be used to store the computer programs and / or modules. The processor implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory and by calling data stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (such as audio data, video data, etc.).

[0146] The memory can be integrated into the processor or it can be set up separately from the processor.

[0147] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the steps of the voice data processing method described in the above embodiments, for example... Figure 2 The steps 201 to 205 shown, as well as other extensions and related steps of the method, are examples. Alternatively, when a computer program is executed by a processor, it implements the functions of each module / unit of the voice data processing device in the above embodiments, for example... Figure 6 The functions of modules 30 and 31 are shown. To avoid repetition, they will not be described again here.

[0148] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0149] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0150] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A voice data processing method, characterized in that, include: Acquire the speech data to be processed and perform speech recognition to obtain the phoneme tag data corresponding to the speech data to be processed; Align the phoneme tag data with the speech data to be processed to obtain the phoneme alignment result; extract the temporal and frequency features of the speech data to be processed. Based on the frequency characteristics, phoneme tag groups are determined, and based on the temporal characteristics, the time interval of the phoneme tag groups on the audio signal in the speech data to be processed is determined to obtain the phoneme distribution position. The phoneme distribution position is used to represent the time interval of each phoneme tag group on the audio signal. The phoneme tag data and the speech data to be processed are phoneme aligned to obtain the phoneme alignment result. Based on the phoneme alignment result, determine the first pause time between adjacent phoneme tag groups in the phoneme tag data; According to the first pause time, the phoneme tag data is labeled so that a pause tag corresponding to the first pause time is added between any two adjacent phoneme tag groups to obtain target phoneme tag data. The pause tag corresponding to the first pause time is one of at least two types of pause tags. Different types of pause tags correspond to different pause times. Any adjacent phoneme tag group corresponds to a character or word in the text obtained by converting the speech data to be processed. The target phoneme label data is used as training data to train the speech generation model.

2. The voice data processing method according to claim 1, characterized in that, The process of acquiring the speech data to be processed and performing speech recognition to obtain the phoneme tag data corresponding to the speech data to be processed includes: Speech recognition is performed on the speech data to be processed according to the preset frame length to obtain the phoneme tag sequence corresponding to each preset frame length; In chronological order, the phoneme tag sequences corresponding to each preset frame length are merged to obtain the phoneme tag data corresponding to the speech data to be processed.

3. The voice data processing method according to claim 1, characterized in that, The step of labeling the phoneme tags according to the first pause time to obtain target phoneme tag data includes: Based on the first pause time, the corresponding pause tag is retrieved from the preset tag database; The pause tags are used to mark the interval positions of the phoneme tag data to obtain the target phoneme tag data.

4. The voice data processing method according to claim 3, characterized in that, The preset tag database includes at least two types of pause tags, each type of pause tag carrying a different second pause time. The step of retrieving the corresponding pause tag from the preset tag database based on the first pause time includes: The first pause time and the second pause time are matched to obtain the matching result; Based on the matching results, retrieve the pause tag of the type corresponding to each first pause time from the preset tag database.

5. The voice data processing method according to any one of claims 1 to 4, characterized in that, Before using the target phoneme label data as training data to train the speech generation model, the method further includes: Perform speech recognition on the speech data to be processed to obtain the text data corresponding to the speech data to be processed. The text data and the speech data to be processed are input into a pre-trained emotion recognition model to perform emotion recognition and determine the emotion category of the speech data to be processed. Based on the emotion category, an emotion tag is obtained, and the emotion tag is used to mark the target phoneme tag data with emotion.

6. The voice data processing method according to any one of claims 1-4, characterized in that, Before using the target phoneme label data as training data to train the speech generation model, the method further includes: Speech recognition is performed on the speech data to be processed to obtain the text data corresponding to each sentence of the speech data to be processed and the time length of each sentence; The word count of each sentence is calculated by performing a word count on the text data of each sentence. The speech rate category of the speech data to be processed is determined based on the number of words in each sentence and the duration of each sentence; Based on the speech rate category, obtain speech rate tags, and use the speech rate tags to mark the speech rate of the target phoneme tag data.

7. The voice data processing method according to claim 6, characterized in that, Determining the speech rate category of the speech data to be processed based on the number of words and the duration of each sentence includes: If at least two sentences have the same number of characters, then obtain the time length corresponding to the at least two sentences; Calculate the average duration of the time lengths corresponding to the at least two sentences; The speech rate category of the speech data to be processed is determined based on the average duration.

8. A voice data processing device, characterized in that, The device includes: The speech recognition module is used to acquire speech data to be processed and perform speech recognition to obtain phoneme tag data corresponding to the speech data to be processed. An alignment module is used to align the phoneme tag data with the speech data to be processed to obtain a phoneme alignment result: extracting the temporal and frequency features of the speech data to be processed; determining phoneme tag groups based on the frequency features; determining the time interval of the phoneme tag groups on the audio signal in the speech data to be processed based on the temporal features to obtain the phoneme distribution position, wherein the phoneme distribution position is used to represent the time interval of each phoneme tag group on the audio signal; and performing phoneme alignment processing on the phoneme tag data and the speech data to be processed to obtain a phoneme alignment result. The first pause time determination module is used to determine the first pause time between adjacent phoneme tag groups in the phoneme tag data based on the phoneme alignment result. The annotation module is used to annotate the phoneme tag data according to the first pause time, so as to add a pause tag corresponding to the first pause time between any two adjacent phoneme tag groups to obtain target phoneme tag data. The pause tag corresponding to the first pause time is one of at least two types of pause tags. Different types of pause tags correspond to different pause times. Any adjacent phoneme tag group corresponds to a character or word in the text obtained by converting the speech data to be processed. The training module is used to train the speech generation model using the target phoneme label data as training data.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the voice data processing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the voice data processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Voice synthesis database pause information automatic marking method and system

    CN105632484A

  • Speech synthesis method and device, equipment and storage medium

    CN113327576A