Data Processing Method, Apparatus and Medium

By determining the acoustic features that are suitable for character labels for text data and performing speech synthesis, the problem of low speech nature in the prior art is solved, and a more natural and diverse speech data generation is achieved.

CN114093341BActive Publication Date: 2025-06-24BEIJING SOGOU TECHNOLOGY DEVELOPMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111167846.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-29
Publication Date
2025-06-24
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

When the prior art converts text data into speech data, the speech is not natural, resulting in obvious modifications in the sound of speech.

Method used

By obtaining the pending text data and its corresponding role tags, acoustic features that are suitable for the character tags are determined for the text sequence using a pre-trained acoustic model, and speech synthesis is performed through a vocoder to generate speech data corresponding to the text data.

Benefits of technology

It improves the naturalness of voice, avoids the same sound quality of voice data, and enhances the accuracy of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114093341B_ABST
    Figure CN114093341B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a data processing method, apparatus, and medium, relating to the fields of computer and artificial intelligence technologies. The method includes: obtaining text data to be processed, where the text data to be processed includes at least one text sequence; obtaining role tags corresponding to each text sequence in the text data to be processed; determining acoustic features adapted to the corresponding role tags for the text sequences through a pre-trained acoustic model; and performing speech synthesis on the acoustic features of each text sequence through a vocoder to obtain speech data corresponding to the text data to be processed. The technical solution of the embodiments of the present application can improve the accuracy of data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer and artificial intelligence technologies. Specifically, it relates to a data processing method, apparatus, and medium. Background Art

[0002] In the existing scenario of converting text data into speech data, such as in the scenario of converting a novel text into audiobook speech, usually after directly converting the text data into speech data, post-processing is performed on the speech data, such as changing the pitch, speech rate, and energy, etc. However, this processing method will result in obvious modification components in the listening sense of the speech, and the naturalness of the speech is not high.

[0003] Based on this, how to improve the accuracy of data processing, especially the accuracy of converting text data into speech data, is a technical problem to be solved urgently. Summary of the Invention

[0004] Embodiments of this application provide a data processing method, apparatus, computer program product or computer program, and computer-readable medium, which can, to at least a certain extent, improve the accuracy of data processing.

[0005] Other features and advantages of this application will become apparent through the following detailed description, or be learned in part through the practice of this application.

[0006] According to one aspect of the embodiments of this application, a data processing method is provided, including: obtaining text data to be processed, where the text data to be processed includes at least one text sequence; obtaining role tags corresponding to each text sequence in the text data to be processed; determining, by a pre-trained acoustic model, acoustic features adapted to the corresponding role tags for the text sequence; and performing speech synthesis on the acoustic features of each text sequence through a vocoder to obtain speech data corresponding to the text data to be processed.

[0007] According to one aspect of the embodiments of this application, a data processing apparatus is provided, including: a first obtaining unit configured to obtain text data to be processed, where the text data to be processed includes at least one text sequence; a second obtaining unit configured to obtain role tags corresponding to each text sequence in the text data to be processed; a determining unit configured to determine, by a pre-trained acoustic model, acoustic features adapted to the corresponding role tags for the text sequence; and a synthesizing unit configured to perform speech synthesis on the acoustic features of each text sequence through a vocoder to obtain speech data corresponding to the text data to be processed.

[0008] In some embodiments of the present application, based on the foregoing solution, the determining unit is configured to: determine prosodic features matching the corresponding role tags for the text sequence through a pre-trained acoustic model; and determine acoustic features for the text sequence through the acoustic model according to the prosodic features.

[0009] In some embodiments of the present application, based on the foregoing solution, the apparatus further includes: a third obtaining unit, configured to obtain training text sequences corresponding to multiple role tags and obtain matching text voices matching each training text sequence before determining prosodic features matching the corresponding role tags for the text sequence through a pre-trained acoustic model; and a training unit, configured to train a to-be-trained acoustic model through the training text sequences and the matching text voices to obtain the acoustic model.

[0010] In some embodiments of the present application, based on the foregoing solution, the training unit is configured to: extract actual acoustic features for the training text sequence from the matching text voice; predict predicted acoustic features for the training text sequence through the to-be-trained acoustic model; and correct model parameters in the to-be-trained acoustic model through gradient backpropagation based on the error between the predicted acoustic features and the actual acoustic features to obtain the acoustic model.

[0011] In some embodiments of the present application, based on the foregoing solution, the training unit is configured to: extract prosodic features for the training text sequence from the matching text voice through the to-be-trained acoustic model as matching prosodic features matching the corresponding role tags; and predict predicted acoustic features for the training text sequence through the to-be-trained acoustic model according to the matching prosodic features.

[0012] In some embodiments of the present application, based on the foregoing solution, the to-be-trained acoustic model includes an encoder model and a prosody model, and the training unit is configured to: encode the training text sequence through the encoder model to obtain text hidden layer features corresponding to the training text sequence; and extract prosodic features for the training text sequence from the matching text voice through the prosody model based on the text hidden layer features.

[0013] In some embodiments of the present application, based on the foregoing solution, the training unit is configured to: determine the distribution characteristics of each character in the training text sequence in time through the prosody model according to the matching text voice; and extract prosodic features for the training text sequence based on the distribution characteristics.

[0014] In some embodiments of the present application, based on the foregoing solution, the acoustic model to be trained further includes a decoder model, and the training unit is configured to: extract the timbre features for the training text sequence from the matched text speech; based on the timbre features and the matched prosody features, decode the text hidden layer features through the decoder model to obtain the predicted acoustic features for the training text sequence.

[0015] In some embodiments of the present application, based on the foregoing solution, the device further includes: a fourth acquisition unit, configured to acquire the model parameters in the acoustic model after correcting the model parameters in the acoustic model to be trained by gradient backpropagation to obtain the acoustic model; an update unit, configured to perform fixed-point processing on the model parameters to generate fixed-point model parameters, and update the fixed-point model parameters to the acoustic model.

[0016] In some embodiments of the present application, based on the foregoing solution, the determination unit is configured to: acquire the sentiment labels corresponding to each text sequence in the text data to be processed; determine, through a pre-trained acoustic model, the prosody features that match the corresponding role labels and the corresponding sentiment labels for the text sequence.

[0017] In some embodiments of the present application, based on the foregoing solution, the determination unit is configured to: identify the semantic information corresponding to each text sequence in the text data to be processed; determine, through the semantic information, the sentiment labels corresponding to each text sequence in the text data to be processed.

[0018] In some embodiments of the present application, based on the foregoing solution, the determination unit is configured to: acquire the timbre features corresponding to each text sequence in the text data to be processed; determine, through the acoustic model, the acoustic features for the text sequence based on the timbre features and the prosody features.

[0019] In some embodiments of the present application, based on the foregoing solution, the text data to be processed includes novel text data, and the role labels include a narrator role label and at least one dialogue role label.

[0020] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the data processing method as described in the foregoing embodiments.

[0021] According to one aspect of the embodiments of the present application, there is also provided a data processing device, which is characterized by including a memory and more than one program. The more than one program is stored in the memory and is configured to be executed by more than one processors. The more than one program includes instructions for performing the data processing method as described in the above embodiments.

[0022] According to one aspect of the embodiments of the present application, there is provided a computer-readable storage medium. At least one program code is stored in the computer-readable storage medium, and the at least one program code is loaded and executed by a processor to implement the operations performed by the data processing method as described in the above embodiments.

[0023] In the technical solutions provided by some embodiments of the present application, by obtaining role tags corresponding to each text sequence in the text data to be processed, and using a pre-trained acoustic model to determine acoustic features adapted to the corresponding role tags for the text sequence, a vocoder can perform voice synthesis on the acoustic features of each text sequence to obtain voice data corresponding to the text data to be processed. Since acoustic features adapted to the corresponding role tags are determined for each text sequence, and the vocoder performs voice synthesis on the acoustic features of each text sequence, voice data with different auditory sensations can be generated for different text sequences, avoiding the uniformity of the voice quality of the voice data, improving the naturalness of the voice, and further improving the accuracy of data processing.

[0024] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Description of the Drawings

[0025] The drawings here are incorporated into the specification and constitute a part of this specification, showing the embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0026] Figure 1 A schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of the present application can be applied is shown;

[0027] Figure 2 A flowchart of a data processing method according to an embodiment of the present application is shown;

[0028] Figure 3 A detailed flowchart of determining acoustic features adapted to the corresponding role tags for the text sequence by a pre-trained acoustic model according to an embodiment of the present application is shown;

[0029] Figure 4 Shows a flowchart of a method before determining prosodic features matching corresponding role tags for the text sequence through a pre-trained acoustic model according to an embodiment of the present application;

[0030] Figure 5 Shows a detailed flowchart of training a to-be-trained acoustic model through the training text sequence and the matching text speech according to an embodiment of the present application;

[0031] Figure 6 Shows a detailed flowchart of predicting predicted acoustic features for the training text sequence through the to-be-trained acoustic model according to an embodiment of the present application;

[0032] Figure 7 Shows a flowchart of a method after obtaining the acoustic model according to an embodiment of the present application;

[0033] Figure 8 Shows a detailed flowchart of determining prosodic features matching corresponding role tags for the text sequence through a pre-trained acoustic model according to an embodiment of the present application;

[0034] Figure 9 Shows a detailed flowchart of determining acoustic features for the text sequence through the acoustic model according to the prosodic features according to an embodiment of the present application;

[0035] Figure 10 Shows a schematic framework diagram of training an acoustic model according to an embodiment of the present application;

[0036] Figure 11 Shows a schematic framework diagram of applying an acoustic model according to an embodiment of the present application;

[0037] Figure 12 Shows a block diagram of a data processing device according to an embodiment of the present application;

[0038] Figure 13 Shows a block diagram of a data processing device according to an embodiment of the present application. Detailed Embodiments

[0039] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.

[0040] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be employed. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0041] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0042] The flowcharts shown in the drawings are only illustrative and not necessarily include all the content and operations / steps, nor are they necessarily executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0043] It should be noted that: "a plurality of" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0044] It should be noted that the terms "first", "second", etc. in the specification, claims, and above-mentioned drawings of the present application are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the objects so used can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described.

[0045] The embodiments in this application involve technologies related to artificial intelligence, that is, through artificial intelligence, the full automation of data (such as text data or voice data) processing is achieved. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.

[0046] Figure 1 The figure shows a schematic diagram of an exemplary system architecture to which the technical solutions of the embodiments of this application can be applied.

[0047] As Figure 1 shown, the system architecture may include terminal devices (such as Figure 1 one or more of the smart phone 101, tablet computer 102, and portable computer 103 shown in the figure. Of course, it can also be a desktop computer, etc., but it is not limited thereto, and this application does not make any restrictions here), network 104, and server 105. Network 104 is used to provide a medium for the communication link between the terminal device and server 105. Network 104 may include various connection types, such as wired communication links, wireless communication links, etc.

[0048] In an embodiment of this application, when a user needs to convert a piece of text data into voice data, the terminal device can send the text data to be processed including at least one text sequence and the role tags corresponding to each text sequence to server 105. After obtaining the text data to be processed and the role tags corresponding to each text sequence, server 105 determines the acoustic features adapted to the corresponding role tags for the text sequence through a pre-trained acoustic model, and performs voice synthesis on the acoustic features of each text sequence through a vocoder to obtain the voice data corresponding to the text data to be processed.

[0049] In this embodiment, by determining the acoustic features adapted to the corresponding role tags for each text sequence and performing voice synthesis on the acoustic features of each text sequence through a vocoder, voice data with different listening sensations can be generated for different text sequences, avoiding the monotony of the voice quality of the voice data, improving the naturalness of the voice, and thus this embodiment can improve the accuracy of data processing.

[0050] It should be noted that the data processing method provided by the embodiments of the present application can be executed by the server 105. Correspondingly, the data processing device is generally set in the server 105. However, in other embodiments of the present application, the terminal device can also have a similar function to the server, so as to execute the data processing solution provided by the embodiments of the present application.

[0051] It should also be noted that Figure 1 the numbers of the terminal devices, networks, and servers in

[0052] are merely illustrative. According to the implementation requirements, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0053] The implementation details of the technical solutions of the embodiments of the present application are elaborated in detail below:

[0054] Figure 2 shows a flowchart of a data processing method according to an embodiment of the present application. The data processing method can be executed by a device with computing and processing capabilities, such as Figure 1 the server 105 shown in Figure 1 or can be executed by the terminal device shown in Figure 2 As shown, the data processing method includes at least steps 210 to 270, which are introduced in detail as follows:

[0055] In step 210, the text data to be processed is obtained, and at least one text sequence is included in the text data to be processed.

[0056] First of all, it should be noted that the data processing method proposed in this application can be applied to application scenarios where text data is converted into speech data. For example, in the scenario of converting news text data into news broadcast data (i.e., speech data), and also, in the scenario of converting novel text data into audiobook data (i.e., speech data).

[0057] It can be understood that the text to be processed can be a piece of news text or a piece of novel text.

[0058] Furthermore, a text sequence in the text data to be processed can be a text part with the same attribute content. For example, in news text data, a text sequence can refer to the text part used to describe a news event. In novel text data, a text sequence can refer to the text part used to describe the psychological activities of a novel character, or the text part of a novel character's dialogue.

[0059] Continue to refer to Figure 2 In step 230, obtain the role tags corresponding to each text sequence in the text data to be processed.

[0060] In this application, the role tag can be used to identify the attribute of a text sequence. For example, in novel text data, the text part used to describe the psychological activities of a novel character can be identified with the role tag "narrator", and the text part of the dialogue of the novel character Zhang San can be identified with the role tag "Zhang San".

[0061] Continue to refer to Figure 2 In step 250, determine the acoustic features adapted to the corresponding role tags for the text sequence through a pre-trained acoustic model.

[0062] In this application, the text sequence can be preprocessed and converted into a text feature vector that can be recognized by the acoustic model, and then the text feature vector is input into the acoustic model, and the acoustic model outputs the acoustic features adapted to the corresponding role tags.

[0063] In this application, the acoustic feature refers to the physical quantity representing the acoustic characteristics of speech, and is also the general term for the acoustic manifestations of various elements of sound. Such as the energy concentration area, formant frequency, formant intensity and bandwidth representing timbre, as well as the duration, fundamental frequency, average speech power, etc. representing the prosodic characteristics of speech.

[0064] In this application, the acoustic feature can specifically be the Mel-spectrum feature in the frequency domain.

[0065] Continue to refer to Figure 2 In step 270, perform speech synthesis on the acoustic features of each text sequence through a vocoder to obtain the speech data corresponding to the text data to be processed.

[0066] In this application, a vocoder is a voice decoder that is responsible for synthesizing voice signals in this application and generating a voice waveform corresponding to the acoustic features based on the acoustic features obtained from the acoustic model.

[0067] In this application, a vocoder can be constructed by a high-quality neural network with ultra-low complexity, so that the sound quality effect of voice data can be greatly improved.

[0068] Next, embodiments of each step as Figure 2 shown will be elaborated in detail:

[0069] In an embodiment of step 250 as Figure 2 shown, an acoustic feature adapted to the corresponding role label is determined for the text sequence through a pre-trained acoustic model, and it can be executed according to the steps as Figure 3 shown.

[0070] Referring to Figure 3 , a detailed flowchart of determining an acoustic feature adapted to the corresponding role label for the text sequence through a pre-trained acoustic model according to an embodiment of the present application is shown. It specifically includes steps 251 to 252:

[0071] Step 251, determining a prosodic feature matching the corresponding role label for the text sequence through a pre-trained acoustic model.

[0072] Step 252, determining an acoustic feature for the text sequence through the acoustic model according to the prosodic feature.

[0073] In this application, the prosodic feature can also be called "suprasegmental feature" or "suprasegmental feature" in the technical field of this technology. It is a phonological structure of language and is closely related to other linguistic structures such as syntactic and discourse structures, and information structures. The prosodic feature can be divided into three main aspects: intonation, temporal distribution, and stress, which are realized through suprasegmental features. The suprasegmental features include pitch, intensity, and temporal characteristics, which are carried by phonemes or groups of phonemes. Prosody is a typical feature of human natural language and has many cross-linguistic common features. For example, downdrift of pitch, stress, pause, etc. are all common in different languages. The prosodic feature is also one of the important forms of language and emotional expression.

[0074] In this application, the acoustic model can determine an acoustic feature for the text sequence through a prosodic feature matching the corresponding role label, so that the prosody corresponding to the prosodic feature is included in the voice synthesized through the acoustic feature.

[0075] In the present application, according to the prosodic features determined for the text sequence and matching the corresponding role tags, acoustic features are determined for the text sequence. The advantage is that the prosody matching the corresponding role tags can be made to exist in the speech data corresponding to the text sequence obtained subsequently, so that speech data with different auditory perceptions can be generated for different text sequences, improving the naturalness of the speech.

[0076] In an embodiment before step 251 as shown Figure 3 That is, before determining the prosodic features matching the corresponding role tags for the text sequence through a pre-trained acoustic model, the steps as shown Figure 4 can also be executed.

[0077] Refer to Figure 4 , which shows a flowchart of a method before determining the prosodic features matching the corresponding role tags for the text sequence according to an embodiment of the present application. Specifically, it includes steps 241 to 242:

[0078] Step 241, obtain training text sequences corresponding to multiple role tags, and obtain matching text voices matching each training text sequence.

[0079] Step 242, train the acoustic model to be trained through the training text sequences and the matching text voices to obtain the acoustic model.

[0080] In this embodiment, taking the scenario of converting novel text data into audiobook data (i.e., speech data) as an example, the multiple role tags may include a narrator role tag and at least one dialogue role tag. Correspondingly, the training text sequences may include a text sequence of "narrator" and at least one text sequence of "dialogue". Further, the matching text voices may include a text voice of "narrator" and at least one text voice of "dialogue".

[0081] In an embodiment of step 242 as shown Figure 4 , training the acoustic model to be trained through the training text sequences and the matching text voices to obtain the acoustic model can be executed according to the steps as shown Figure 5 .

[0082] Refer to Figure 5 , which shows a detailed flowchart of training the acoustic model to be trained through the training text sequences and the matching text voices according to an embodiment of the present application. Specifically, it includes steps 2421 to 2423:

[0083] Step 2421, extract the actual acoustic features for the training text sequences from the matching text voices.

[0084] Step 2422: Predict the predicted acoustic features for the training text sequence through the acoustic model to be trained.

[0085] Step 2423: Based on the error between the predicted acoustic features and the actual acoustic features, correct the model parameters in the acoustic model to be trained through gradient backpropagation to obtain the acoustic model.

[0086] In Figure 5 In one embodiment of step 2422 as shown, predicting the predicted acoustic features for the training text sequence through the acoustic model to be trained can be performed according to the steps as shown in Figure 6 shown.

[0087] See Figure 6 , which shows a detailed flowchart of predicting the predicted acoustic features for the training text sequence through the acoustic model to be trained according to an embodiment of the present application. Specifically, it includes steps 24221 to 24222:

[0088] Step 24221: Extract the prosodic features for the training text sequence from the matched text speech through the acoustic model to be trained as the matched prosodic features matching the corresponding role labels.

[0089] Step 24222: Predict the predicted acoustic features for the training text sequence through the acoustic model to be trained according to the matched prosodic features.

[0090] In the present application, the acoustic model to be trained may include an encoder model and a prosody model. Among them, the encoder model is mainly used to encode the text sequence to obtain the text hidden layer features of the text sequence. The prosody model is mainly used to extract prosodic features from speech data and can also be used to predict the prosodic features matching the corresponding role labels for the text sequence.

[0091] In Figure 6 In step 24221 as shown, extracting the prosodic features for the training text sequence from the matched text speech can be performed according to the following steps 610 to 620.

[0092] Step 610: Encode the training text sequence through the encoder model to obtain the text hidden layer features corresponding to the training text sequence.

[0093] Step 620: Based on the text hidden layer features, extract the prosodic features for the training text sequence from the matched text speech through the prosody model.

[0094] In this application, extracting prosodic features for the training text sequence from the matched text speech through the prosodic model can be performed according to the following steps 621 to step 622:

[0095] Step 621, according to the matched text speech, determine the distribution characteristics of each character in the training text sequence in terms of time through the prosodic model.

[0096] Step 622, based on the distribution characteristics, extract the prosodic features for the training text sequence.

[0097] It should be noted that in this application, the distribution characteristics of each character in terms of time can also be understood as the distribution characteristics of the phonemes of each character in terms of time. Among them, a phoneme is the smallest speech unit divided according to the natural attributes of speech. Further, the distribution characteristics of phonemes in terms of time can be the duration information corresponding to the phonemes. For example, the phoneme sequence corresponding to the character sequence "I am very happy" is [wo hao kai xin], which can be used to represent the corresponding pronunciation. If the speech speed is normal, its duration information can be D1 = [1, 1, 1, 1]. If the speech speed is slower and the pronunciation time of the character "hao" is longer, its duration information can be D1 = [2, 3, 2, 2]. Correspondingly, the phoneme sequence corresponding to "I am very happy" can be expanded to [wo wo hao hao hao kai kai xin xin]. If the speech speed is faster and the pronunciation time of the character "hao" is normal, its duration information can be D1 = [0.5, 1, 0.5, 0.5].

[0098] In Figure 6 the step 24221 as shown, after determining the matched prosodic features that match the corresponding role label, the prosodic model in the acoustic model to be trained can learn and remember the matched prosodic features that match the role label. In the follow-up, the corresponding prosodic features can be directly predicted according to the role label.

[0099] In this application, the acoustic model to be trained may include a decoder model, and the decoder model is mainly used to decode the training text hidden layer features corresponding to the training text sequence to obtain acoustic features.

[0100] In Figure 6 the step 24222 as shown, predicting the predicted acoustic features for the training text sequence through the acoustic model to be trained according to the matched prosodic features can be performed according to the following steps 630 to step 640.

[0101] Step 630, extract the timbre features for the training text sequence from the matched text speech.

[0102] Step 640: Based on the timbre feature and the matching prosody feature, decode the text hidden layer feature through the decoder model to obtain the predicted acoustic feature for the training text sequence.

[0103] In step 2423 as Figure 4 shown, based on the error between the predicted acoustic feature and the actual acoustic feature, the model parameters in the acoustic model to be trained are corrected through gradient backpropagation. Actually, through the predicted acoustic feature and the actual acoustic feature, supervised adversarial training is performed on the acoustic model to be trained.

[0104] During the training process, a loss function can be generated according to the predicted acoustic feature and the actual acoustic feature, and the model parameters in the acoustic model to be trained are corrected through the loss function to obtain an acoustic model, which is used to generate the actual acoustic feature matching the text sequence. Specifically, a loss function (such as a mean squared error loss function) can be generated according to the predicted acoustic feature and the actual acoustic feature corresponding to the training text sequence, which is used to represent the gap between the predicted acoustic feature and the true actual acoustic feature. Furthermore, the model parameters in the acoustic model to be trained can be corrected through this loss function to obtain a trained acoustic model. Among them, the acoustic model is used to generate the actual acoustic feature matching the text sequence, and this acoustic model can include a trained encoder model, a prosody model, and a decoder model.

[0105] In an embodiment after step 2423 as Figure 4 shown, that is, after correcting the model parameters in the acoustic model to be trained through gradient backpropagation to obtain the acoustic model, the steps as Figure 7 shown can also be executed.

[0106] Refer to Figure 7 , which shows a flowchart of the method after obtaining the acoustic model according to an embodiment of the present application. Specifically, it includes step 2424 to step 2425:

[0107] Step 2424: Obtain the model parameters in the acoustic model.

[0108] Step 2425: Perform fixed-point processing on the model parameters to generate fixed-point model parameters, and update the fixed-point model parameters to the acoustic model.

[0109] In the present application, fixed-point processing can modify the floating-point parameters in the acoustic model into fixed-point parameters. The decimal places of the floating-point parameters can vary randomly, and the decimal range they can represent is wider than that of the fixed-point parameters. Correspondingly, the computational amount of the floating-point parameters is also very large. The fixed-point parameters refer to the parameters in which the integer part and the decimal part are fixed in a number.

[0110] In the present application, performing fixed-point processing on the model parameters has the advantage of significantly reducing the time complexity and space complexity of the model, thereby reducing the storage space of the acoustic model without sacrificing the model performance.

[0111] In Figure 3 an embodiment of step 251 as shown, determining prosody features matching the corresponding role tags for the text sequences through a pre-trained acoustic model can be performed according to the steps as Figure 8 shown.

[0112] Referring to Figure 8 , a detailed flowchart showing the determination of prosody features matching the corresponding role tags for the text sequences through a pre-trained acoustic model according to an embodiment of the present application is shown. Specifically, it includes steps 2511 to 2512:

[0113] Step 2511: Obtain sentiment tags corresponding to each text sequence in the text data to be processed.

[0114] Step 2512: Determine prosody features matching the corresponding role tags and the corresponding sentiment tags for the text sequences through a pre-trained acoustic model.

[0115] In the present application, each text sequence can correspond to a sentiment tag, and this sentiment tag can be used to identify the type of sentiment expressed by the sequence text. Specifically, for example, when the sentiment tag is "happy", it indicates that the corresponding text sequence corresponds to the expression of a happy sentiment, and when the sentiment tag is "sad", it indicates that the corresponding text sequence corresponds to the expression of a sad sentiment.

[0116] In the present application, determining prosody features matching the corresponding role tags and the corresponding sentiment tags for the text sequences through a pre-trained acoustic model has the advantage that the speech data corresponding to the text sequences obtained subsequently can not only have prosody matching the corresponding role tags but also have prosody matching the corresponding sentiment tags, so that different auditory feelings can be generated for different text sequences, and more accurate speech data can be obtained, improving the naturalness of the speech.

[0117] In Figure 8 an embodiment of step 2511 as shown, obtaining sentiment tags corresponding to each text sequence in the text data to be processed can perform the following steps 25111 to 25112:

[0118] Step 25111: Identify semantic information corresponding to each text sequence in the text data to be processed.

[0119] Step 25112, determine the sentiment labels corresponding to each text sequence in the text data to be processed based on the semantic information.

[0120] In an embodiment of step 252 as shown Figure 3 , according to the prosody features, determine the acoustic features for the text sequence through the acoustic model, which can be executed according to the steps as shown Figure 9 .

[0121] Refer to Figure 9 , which shows the detailed flowchart of determining the acoustic features for the text sequence through the acoustic model according to the prosody features in an embodiment of the present application. Specifically, it includes steps 2521 to 2522:

[0122] Step 2521, obtain the timbre features corresponding to each text sequence in the text data to be processed.

[0123] Step 2522, based on the timbre features and the prosody features, determine the acoustic features for the text sequence through the acoustic model.

[0124] In the present application, the timbre features corresponding to each text sequence can be the same timbre feature or different timbre features.

[0125] For example, in the scenario of converting novel text data into audiobook data (i.e., voice data), in one case, the timbre feature can be just one timbre feature. For example, it can be just the timbre feature corresponding to the narrator role label. That is to say, the effect brought by this is that in the voice data finally determined according to the acoustic features, the voice prosodies corresponding to each text sequence are different, but the voice timbres corresponding to each text sequence are exactly the same. In one case, the timbre features can be different timbre features corresponding to each role label. The effect brought by this is that in the voice data finally determined according to the acoustic features, the voice prosodies corresponding to each text sequence are different, and the voice timbres corresponding to each text sequence are also different.

[0126] It can be understood that in the present application, the proposed acoustic model can at least include an encoder model, a prosody model, and a decoder model. Among them, the acoustic model composed of the encoder model, the prosody model, and the decoder model can be embedded in offline devices, such as most devices like in-vehicle, mobile phones, TVs, speakers, headphones, DSPs, etc. In this way, even in a network-free environment, the offline device can directly convert different text sequences into voice data that matches the corresponding role labels and / or sentiment labels through this acoustic model, enabling the offline device to play this voice data with a hierarchical prosody, thereby improving the user experience.

[0127] To enable those skilled in the art to better understand the present application, the following will take the scenario of converting novel text data into audiobook data (i.e., voice data) as an example, and in combination with Figure 10 and Figure 11 , the embodiments proposed in the present application will be described from the perspectives of training an acoustic model and applying the acoustic model respectively.

[0128] Refer to Figure 10 , which shows a schematic framework diagram of training an acoustic model according to an embodiment of the present application.

[0129] As Figure 10 described, the acoustic model 1001 to be trained includes an encoder model, a prosody model, and a decoder model. First, on the one hand, the training text sequence is input into the encoder model in the acoustic model 1001. The encoder model outputs the text hidden layer features for the training text sequence, and the text hidden layer features are input into the prosody model and the decoder model. The narrator role label (or dialogue role label) corresponding to the training text sequence is also input into the decoder model and the prosody model, so that the decoder model and the prosody model learn and remember the narrator role label (or dialogue role label). On the other hand, the prosody model extracts the prosody features corresponding to the training text sequence from the actual acoustic features, inputs the prosody features into the decoder model, and further learns the matching relationship between the narrator role label (or dialogue role label) and the prosody features. Then, the decoder model predicts the acoustic features of the training text sequence based on the input text hidden layer features and prosody features, and obtains the predicted acoustic features. Finally, based on the error between the predicted acoustic features and the actual acoustic features, the model parameters of the encoder model, the prosody model, and the decoder model in the acoustic model 1001 are corrected through gradient backpropagation, and the trained acoustic model 1001 is obtained.

[0130] Refer to Figure 11 , which shows a schematic framework diagram of applying an acoustic model according to an embodiment of the present application.

[0131] As Figure 11 described, the trained acoustic model 1002 includes an encoder model, a prosody model, and a decoder model.

[0132] First, input the text sequence into the encoder model in the acoustic model 1002. The encoder model outputs the text hidden layer features for the text sequence, and the text hidden layer features are input into the prosody model and the decoder model. At the same time, input the narrator role label (or dialogue role label) corresponding to the text sequence into the prosody model, and input the timbre label (reflecting a timbre feature) into the decoder model. Then, the prosody model predicts the prosody features for the text sequence according to the input narrator role label (or dialogue role label), and inputs the prosody features into the decoder model. Finally, the decoder model determines the acoustic features of the text sequence based on the input text hidden layer features, prosody features, and timbre label.

[0133] In this application, by obtaining the role labels corresponding to each text sequence in the to-be-processed text data, and using a pre-trained acoustic model to determine the acoustic features adapted to the corresponding role labels for the text sequence, the vocoder can perform speech synthesis on the acoustic features of each text sequence to obtain the speech data corresponding to the to-be-processed text data. Since the acoustic features adapted to the corresponding role labels are determined for each text sequence, and the acoustic features of each text sequence are subjected to speech synthesis by the vocoder, speech data with different auditory sensations can be generated for different text sequences, avoiding the monotony of the speech data in terms of sound quality, improving the naturalness of the speech, and further improving the accuracy of data processing.

[0134] The following introduces the apparatus embodiments of this application, which can be used to execute the data processing method in the above embodiments of this application. For the details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the above data processing method of this application.

[0135] Figure 12 The block diagram of a data processing apparatus according to an embodiment of this application is shown.

[0136] Refer to Figure 12 As shown, a data processing apparatus 1200 according to an embodiment of this application includes: a first acquisition unit 1201, a second acquisition unit 1202, a determination unit 1203, and a synthesis unit 1204.

[0137] Among them, the first acquisition unit 1201 is configured to acquire to-be-processed text data, where the to-be-processed text data includes at least one text sequence; the second acquisition unit 1202 is configured to acquire the role labels corresponding to each text sequence in the to-be-processed text data; the determination unit 1203 is configured to determine, by using a pre-trained acoustic model, the acoustic features adapted to the corresponding role labels for the text sequence; and the synthesis unit 1204 is configured to perform speech synthesis on the acoustic features of each text sequence through a vocoder to obtain the speech data corresponding to the to-be-processed text data.

[0138] In some embodiments of the present application, based on the foregoing solution, the determining unit 1203 is configured to: determine, for the text sequence, prosodic features that match the corresponding role tags through a pre-trained acoustic model; and determine, according to the prosodic features, acoustic features for the text sequence through the acoustic model.

[0139] In some embodiments of the present application, based on the foregoing solution, the apparatus further includes: a third obtaining unit, configured to obtain training text sequences corresponding to a plurality of role tags and obtain matching text voices corresponding to each of the training text sequences before determining, for the text sequence, prosodic features that match the corresponding role tags through a pre-trained acoustic model; and a training unit, configured to train a to-be-trained acoustic model through the training text sequences and the matching text voices to obtain the acoustic model.

[0140] In some embodiments of the present application, based on the foregoing solution, the training unit is configured to: extract actual acoustic features for the training text sequence from the matching text voice; predict predicted acoustic features for the training text sequence through the to-be-trained acoustic model; and correct model parameters in the to-be-trained acoustic model through gradient backpropagation based on the error between the predicted acoustic features and the actual acoustic features to obtain the acoustic model.

[0141] In some embodiments of the present application, based on the foregoing solution, the training unit is configured to: extract, through the to-be-trained acoustic model, prosodic features for the training text sequence from the matching text voice as matching prosodic features that match the corresponding role tags; and predict, according to the matching prosodic features, predicted acoustic features for the training text sequence through the to-be-trained acoustic model.

[0142] In some embodiments of the present application, based on the foregoing solution, the to-be-trained acoustic model includes an encoder model and a prosody model, and the training unit is configured to: encode the training text sequence through the encoder model to obtain text hidden layer features corresponding to the training text sequence; and extract, based on the text hidden layer features, prosodic features for the training text sequence from the matching text voice through the prosody model.

[0143] In some embodiments of the present application, based on the foregoing solution, the training unit is configured to: determine, according to the matching text voice, distribution features of each character in the training text sequence in terms of time through the prosody model; and extract, based on the distribution features, prosodic features for the training text sequence.

[0144] In some embodiments of the present application, based on the foregoing solution, the acoustic model to be trained further includes a decoder model, and the training unit is configured to: extract the timbre features for the training text sequence from the matched text speech; based on the timbre features and the matched prosody features, decode the text hidden layer features through the decoder model to obtain the predicted acoustic features for the training text sequence.

[0145] In some embodiments of the present application, based on the foregoing solution, the apparatus further includes: a fourth acquisition unit, configured to, after correcting the model parameters in the acoustic model to be trained through gradient backpropagation to obtain the acoustic model, acquire the model parameters in the acoustic model; an update unit, configured to perform fixed-point processing on the model parameters to generate fixed-point model parameters, and update the fixed-point model parameters to the acoustic model.

[0146] In some embodiments of the present application, based on the foregoing solution, the determination unit 1203 is configured to: acquire the emotion labels corresponding to each text sequence in the text data to be processed; determine, through a pre-trained acoustic model, the prosody features that match the corresponding role labels and match the corresponding emotion labels for the text sequence.

[0147] In some embodiments of the present application, based on the foregoing solution, the determination unit 1203 is configured to: identify the semantic information corresponding to each text sequence in the text data to be processed; determine, through the semantic information, the emotion labels corresponding to each text sequence in the text data to be processed.

[0148] In some embodiments of the present application, based on the foregoing solution, the determination unit 1203 is configured to: acquire the timbre features corresponding to each text sequence in the text data to be processed; determine, through the acoustic model, the acoustic features for the text sequence based on the timbre features and the prosody features.

[0149] As another aspect, an embodiment of the present application further provides another data processing apparatus, including a memory, and more than one program, where the more than one program is stored in the memory and is configured to be executed by more than one processors, and the more than one program includes instructions for performing the data processing method as described in the foregoing embodiments.

[0150] Figure 13 The block diagram of a data processing apparatus according to an embodiment of the present application is shown. For example, the apparatus 1300 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0151] Refer to Figure 13, device 1300 may include one or more of the following components: a processing component 1302, a memory 1304, a power component 1306, a multimedia component 1308, an audio component 1310, an input / output (I / O) interface 1313, a sensor component 1314, and a communication component 1316.

[0152] The processing component 1302 generally controls the overall operation of the device 1300, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing element 1302 may include one or more processors 1320 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 1302 may include one or more modules to facilitate the interaction between the processing component 1302 and other components. For example, the processing component 1302 may include a multimedia module to facilitate the interaction between the multimedia component 1308 and the processing component 1302.

[0153] The memory 1304 is configured to store various types of data to support the operation of the device 1300. Examples of such data include instructions for any application or method operating on the device 1300, contact data, phone book data, messages, pictures, videos, etc. The memory 1304 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0154] The power component 1306 provides power to the various components of the device 1300. The power component 1306 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device 1300.

[0155] The multimedia component 1308 includes a screen that provides an output interface between the device 1300 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of a touch or swipe action but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 1308 includes a front camera and / or a rear camera. When the device 1300 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0156] The audio component 1310 is configured to output and / or input audio signals. For example, the audio component 1310 includes a microphone (MIC) that is configured to receive external audio signals when the device 1300 is in an operating mode, such as a call mode, a recording mode, and a voice message processing mode. The received audio signals can be further stored in the memory 1304 or transmitted via the communication component 1316. In some embodiments, the audio component 1310 further includes a speaker for outputting audio signals.

[0157] The I / O interface 1313 provides an interface between the processing component 1302 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.

[0158] The sensor component 1314 includes one or more sensors for providing an assessment of the various aspects of the state of the device 1300. For example, the sensor component 1314 can detect the on / off state of the device 1300, the relative positioning of components, such as the display and the keypad of the device 1300. The sensor component 1314 can also search for results to show a change in the position of the device 1300 or a component of the device 1300, the presence or absence of user contact with the device 1300, the orientation or acceleration / deceleration of the device 1300, and the temperature change of the device 1300. The sensor component 1314 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 1314 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 1314 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0159] The communication component 1316 is configured to facilitate communication between the device 1300 and other devices in a wired or wireless manner. The device 1300 can access a communication standard-based wireless network, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 1316 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1316 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0160] In an exemplary embodiment, the device 1300 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0161] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 1304 including instructions, and the above instructions can be executed by the processor 1320 of the device 1300 to complete the above data processing method. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0162] As another aspect, the present application also provides a computer program product or a computer program, and the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the data processing method described in the above embodiments.

[0163] As another aspect, the present application also provides a computer-readable storage medium, which can be included in the device described in the above embodiments; or can exist alone without being assembled into the device. At least one program code is stored in the computer-readable storage medium, and the at least one program code is loaded and executed by the processor of the device to implement the operations performed by the data processing method described in the above embodiments.

[0164] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0165] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0166] After considering the specification and practicing the embodiments disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.

[0167] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A data processing method, characterized in that, The method includes: Obtaining text data to be processed, where the text data to be processed includes at least one text sequence; Obtaining role tags corresponding to each text sequence in the text data to be processed; Determining prosody features matching the corresponding role tags for the text sequences through a pre-trained acoustic model, including: obtaining emotion tags corresponding to each text sequence in the text data to be processed; determining prosody features matching the corresponding role tags and matching the corresponding emotion tags for the text sequences through the pre-trained acoustic model; Determining acoustic features for the text sequences through the acoustic model according to the prosody features; Performing speech synthesis on the acoustic features of each text sequence through a vocoder to obtain speech data corresponding to the text data to be processed.

2. The method according to claim 1, characterized in that, Before determining prosody features matching the corresponding role tags for the text sequences through the pre-trained acoustic model, the method further includes: Obtaining training text sequences corresponding to multiple role tags, and obtaining matching text voices matching each training text sequence; Training a to-be-trained acoustic model through the training text sequences and the matching text voices to obtain the acoustic model.

3. The method according to claim 2, wherein The training the to-be-trained acoustic model through the training text sequences and the matching text voices to obtain the acoustic model includes: Extracting actual acoustic features for the training text sequences from the matching text voices; Predicting predicted acoustic features for the training text sequences through the to-be-trained acoustic model; Based on the error between the predicted acoustic features and the actual acoustic features, correcting the model parameters in the to-be-trained acoustic model through gradient backpropagation to obtain the acoustic model.

4. The method according to claim 3, wherein The predicting predicted acoustic features for the training text sequences through the to-be-trained acoustic model includes: Extracting prosody features for the training text sequences from the matching text voices through the to-be-trained acoustic model as matching prosody features matching the corresponding role tags; Predicting predicted acoustic features for the training text sequences through the to-be-trained acoustic model according to the matching prosody features.

5. The method according to claim 4, wherein The to-be-trained acoustic model includes an encoder model and a prosody model. The extracting prosody features for the training text sequences from the matching text voices includes: Encoding the training text sequences through the encoder model to obtain text hidden layer features corresponding to the training text sequences; Based on the text hidden layer features, extracting prosody features for the training text sequences from the matching text voices through the prosody model.

6. The method according to claim 5, wherein The extracting prosody features for the training text sequences from the matching text voices through the prosody model includes: Determining the distribution features of each character in the training text sequences in time through the prosody model according to the matching text voices; Extracting prosody features for the training text sequences based on the distribution features.

7. The method according to claim 5, characterized in that, The to-be-trained acoustic model further includes a decoder model. The predicting, according to the matched prosody features, of the predicted acoustic features for the training text sequence by the to-be-trained acoustic model includes: extracting timbre features for the training text sequence from the matched text speech; decoding the text hidden layer features through the decoder model based on the timbre features and the matched prosody features to obtain the predicted acoustic features for the training text sequence.

8. The method according to claim 3, characterized in that, After correcting the model parameters in the to-be-trained acoustic model through gradient backpropagation to obtain the acoustic model, the method further includes: obtaining the model parameters in the acoustic model; performing fixed-point processing on the model parameters to generate fixed-point model parameters, and updating the fixed-point model parameters to the acoustic model.

9. The method according to claim 1, wherein The obtaining of the emotion labels corresponding to each text sequence in the to-be-processed text data includes: identifying semantic information corresponding to each text sequence in the to-be-processed text data; determining, through the semantic information, the emotion labels corresponding to each text sequence in the to-be-processed text data.

10. The method according to claim 1, characterized in that, The determining, according to the prosody features, of the acoustic features for the text sequence by the acoustic model includes: obtaining timbre features corresponding to each text sequence in the to-be-processed text data; determining the acoustic features for the text sequence by the acoustic model based on the timbre features and the prosody features.

11. A data processing device, characterized in that, The apparatus includes: a first obtaining unit, configured to obtain to-be-processed text data, where the to-be-processed text data includes at least one text sequence; a second obtaining unit, configured to obtain role labels corresponding to each text sequence in the to-be-processed text data; a determining unit, configured to determine, for the text sequence, prosody features matching the corresponding role labels by a pre-trained acoustic model, including: obtaining emotion labels corresponding to each text sequence in the to-be-processed text data; determining, for the text sequence, prosody features matching the corresponding role labels and matching the corresponding emotion labels by the pre-trained acoustic model; determining the acoustic features for the text sequence according to the prosody features by the acoustic model; a synthesizing unit, configured to perform speech synthesis on the acoustic features of each text sequence through a vocoder to obtain speech data corresponding to the to-be-processed text data.

12. The device according to claim 11, characterized in that, The apparatus further includes: a third obtaining unit, configured to obtain training text sequences corresponding to multiple role labels and obtain matched text speech corresponding to each training text sequence before determining, for the text sequence, prosody features matching the corresponding role labels by the pre-trained acoustic model; a training unit, configured to train the to-be-trained acoustic model through the training text sequences and the matched text speech to obtain the acoustic model.

13. The device according to claim 12, characterized in that, The training unit is specifically configured to: extract actual acoustic features for the training text sequence from the matched text speech; predict the predicted acoustic features for the training text sequence through the to-be-trained acoustic model; Based on the error between the predicted acoustic features and the actual acoustic features, the model parameters in the acoustic model to be trained are corrected through backpropagation of gradients to obtain the acoustic model.

14. The device according to claim 13, characterized in that, The training unit is specifically configured to: Extract, from the matched text speech, the prosodic features for the training text sequence through the acoustic model to be trained, as the matched prosodic features that match the corresponding character labels; Predict the predicted acoustic features for the training text sequence through the acoustic model to be trained according to the matched prosodic features.

15. The device according to claim 14, characterized in that, The acoustic model to be trained includes an encoder model and a prosody model. The training unit is specifically configured to: Encode the training text sequence through the encoder model to obtain the text hidden layer features corresponding to the training text sequence; Based on the text hidden layer features, extract, from the matched text speech, the prosodic features for the training text sequence through the prosody model.

16. The device according to claim 15, characterized in that, The training unit is specifically configured to: Determine, according to the matched text speech, the distribution features of each character in the training text sequence in time through the prosody model; Extract the prosodic features for the training text sequence based on the distribution features.

17. The device according to claim 15, characterized in that, The acoustic model to be trained further includes a decoder model. The training unit is specifically configured to: Extract the timbre features for the training text sequence from the matched text speech; Decode the text hidden layer features through the decoder model based on the timbre features and the matched prosodic features to obtain the predicted acoustic features for the training text sequence.

18. The device according to claim 13, characterized in that, The apparatus further includes: A fourth acquisition unit, configured to acquire the model parameters in the acoustic model after correcting the model parameters in the acoustic model to be trained through backpropagation of gradients to obtain the acoustic model; An update unit, configured to perform fixed-point processing on the model parameters to generate fixed-point model parameters and update the fixed-point model parameters to the acoustic model.

19. The device according to claim 11, characterized in that, The determination unit is specifically configured to: Identify the semantic information corresponding to each text sequence in the text data to be processed; Determine the emotion labels corresponding to each text sequence in the text data to be processed through the semantic information.

20. The device according to claim 11, characterized in that, The determination unit is specifically configured to: Acquire the timbre features corresponding to each text sequence in the text data to be processed; Determine the acoustic features for the text sequence through the acoustic model based on the timbre features and the prosodic features.

21. A data processing device, characterized in that, It includes a memory and more than one program, where the more than one program is stored in the memory and is configured to be executed by more than one processor. The more than one program includes instructions for performing the data processing method as described in any one of claims 1 to 10.

22. A computer-readable storage medium, characterized in that, At least one program code is stored in the computer-readable storage medium, and the at least one program code is loaded and executed by the processor to implement the operations performed by the data processing method as described in any one of claims 1 to 10.

23. A computer program product, characterized in that, The computer program product includes computer instructions, and a processor of a computer device executes the computer instructions, so that the computer device performs the operations executed by implementing the data processing method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Speech synthesis method and device, electronic equipment and readable storage medium

    CN112270920A

  • Speech synthesis method and device and device for speech synthesis

    CN113409764A