Speech processing and training methods and electronic devices

By using timbre conversion technology, single-play audiobooks can be converted into multi-play audiobooks, solving the problem of low timbre differentiation when a single person uses pseudo-voice to portray multiple characters, thus improving the expressiveness of the audiobooks and the user experience.

CN115881145BActive Publication Date: 2026-01-30HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111158143.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-30
Publication Date
2026-01-30
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

When a single person uses pseudo-voice to perform in an audiobook with multiple characters, the character timbre is not well distinguished, resulting in insufficient performance of the scene and making it difficult for users to understand the plot.

Method used

By acquiring the audio data of unicast audiobooks, matching the target timbre and expressive features, performing timbre conversion, and generating multicast audiobooks, the audiobooks are generated, ensuring timbre differentiation and emotional consistency.

Benefits of technology

It improves the timbre differentiation and scene performance of multi-play audiobooks, simplifies the production process, and reduces production costs and time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115881145B_ABST
    Figure CN115881145B_ABST
Patent Text Reader

Abstract

This application provides a speech processing and training method and an electronic device. The speech processing method includes: acquiring unicast audio data of a unicast audiobook, the unicast audio data including N source audio data; then, determining N sets of target timbre feature information matching the N source audio data from M sets of reference timbre feature information, wherein the timbre discrimination between any two sets of the M sets of reference timbre feature information is greater than a discrimination threshold; then, acquiring N sets of expressiveness feature information corresponding to the N sets of source audio data; subsequently, based on the N sets of target timbre feature information and the N sets of expressiveness feature information, performing timbre conversion on the N source audio data respectively to generate multicast audio data. In this way, the timbre of the source audio data can be converted while maintaining expressiveness, realizing the conversion of unicast audiobooks into multicast audiobooks, improving the timbre discrimination of characters in the audiobook, thereby improving the expressiveness of the scene portrayal and facilitating user comprehension of the plot.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the field of data processing, and in particular to a voice processing and training method and an electronic device. BACKGROUND

[0002] At present, the mainstream content of audiobooks is novels, such as romance novels, suspense novels, science fiction novels, and martial arts novels. Since the novel content has multiple characters, when recording the audiobook, a single person usually performs multi-character performance through pseudo-voice to realize the recording of the multi-character voice in the novel.

[0003] However, the voice that a single person can fake is limited. When the number of characters in the novel is too large, even if a single voice actor performs multi-character performance through pseudo-voice, it is still impossible to perform all the characters, resulting in low voice differentiation of each character, thereby causing a lack of performance in scene performance, which is not conducive to user understanding of the plot. SUMMARY

[0004] To solve the above technical problems, the present application provides a voice processing and training method and an electronic device. In the voice processing method, the unicast audio data can be converted into multicast audio data under the premise of ensuring the performance of the unicast audiobook, and the unicast audiobook can be converted into a multicast audiobook.

[0005] In a first aspect, the embodiments of the present application provide a voice processing method, comprising: obtaining unicast audio data of a unicast audiobook, the unicast audio data comprising N pieces of source audio data; then, determining N sets of target voice characteristics information matched with the N pieces of source audio data from M sets of reference voice characteristics information, the M sets of reference voice characteristics information corresponding to M kinds of voice, and the voice differentiation degree of any two sets of the M sets of reference voice characteristics information being greater than a differentiation threshold; and obtaining N sets of performance characteristics information corresponding to the N sets of source audio data; then, based on the N sets of target voice characteristics information and the N sets of performance characteristics information, performing voice conversion on the N pieces of source audio data respectively to generate multicast audio data. In this way, the voice of each piece of source audio data can be converted under the premise of ensuring the performance of the unicast audiobook, the unicast audiobook can be converted into a multicast audiobook, the voice differentiation degree of the characters in the audiobook is improved, thereby improving the scene performance degree of the audiobook, and facilitating user understanding of the plot. In addition,

[0006] On the one hand, after the user performs an audiobook reading operation on the unicast audiobook, the present application can automatically match the voice of each source audio data, without the need for the user to manually select the matching voice for each source audio data, thereby improving the efficiency of converting the unicast audio data into multicast audio data, simplifying the user operation, and improving the user experience.

[0007] On the other hand, in the process of making a multicast audiobook by an audiobook producer, the application only needs a single voice actor to record a single audiobook by using pseudo voices to perform multiple roles, and then convert the single audiobook into a multicast audiobook, so as to realize the making of the multicast audiobook, without the need to use multiple voice actors to record according to the roles to realize the making of the multicast audiobook, thereby reducing the making time and cost of the multicast audiobook and improving the making efficiency of the multicast audiobook.

[0008] On the other hand, in the process of making a multicast audiobook by an audiobook producer, the application only needs a single voice actor to record a single audiobook by using pseudo voices to perform multiple roles, and then convert the single audiobook into a multicast audiobook, so as to realize the making of the multicast audiobook, without the need to use multiple voice actors to record according to the roles to realize the making of the multicast audiobook, thereby reducing the making time and cost of the multicast audiobook and improving the making efficiency of the multicast audiobook.

[0009] Exemplarily, N is a positive integer, and M is a positive integer.

[0010] Exemplarily, the timbre feature information can be used to represent information related to timbre, which can include but is not limited to personality characteristics, gender characteristics, age characteristics, and voice region (such as high voice region, medium voice region, and low voice region) characteristics, etc., and can be represented by a vector or a sequence table, and the application does not limit this.

[0011] According to the first aspect, from the M groups of reference timbre feature information, N groups of target timbre feature information matched with the N pieces of source audio data are determined, including: performing audio sentence division on the single audio data to obtain the N pieces of source audio data, the N pieces of source audio data corresponding to the audio sentences one by one; obtaining N groups of source timbre feature information corresponding to the N pieces of source audio data; and determining N groups of target timbre feature information from the M groups of reference timbre feature information based on the N groups of source timbre feature information.

[0012] According to the first aspect, or any one of the implementation manners of the first aspect, the N groups of target timbre feature information matched with the N pieces of source audio data are determined from the M groups of reference timbre feature information based on the N groups of source timbre feature information, including: for the i th piece of source audio data in the N pieces of source audio data: respectively determining the similarity of the M groups of reference timbre feature information and the i th piece of source audio data; and determining the reference timbre feature information with the highest similarity to the i th piece of source audio data as the target timbre feature information matched with the i th piece of source audio data, where i is a positive integer and the value range is between 1 and N. In this way, the timbre with high similarity to the timbre of the role corresponding to the source audio data can be determined as the timbre matched with the source audio data, so that the timbre after conversion conforms to the timbre of the role corresponding to the source audio data.

[0013] Exemplarily, the description between 1 and N includes 1 and N, that is, i can be 1 or N.

[0014] According to the first aspect, or any one of the implementations of the first aspect, the N sets of expressiveness feature information corresponding to the N sets of source audio data are obtained, including: obtaining N sets of prosody feature information corresponding to the N sets of source audio data; obtaining N sets of emotion feature information corresponding to the N sets of source audio data; and generating N sets of expressiveness feature information corresponding to the N sets of source audio data based on the N sets of prosody feature information and the N sets of emotion feature information.

[0015] For example, the prosody feature information can be used to represent voice emotion skill information such as tone, speed, and sound, and can be represented by a vector or a sequence.

[0016] For example, the emotion feature information can be used to represent emotion types (such as happy, sad, high-pitched, and low-pitched) and attitude dimensions (such as affirmative, negative, praise, and irony), and can be represented by a vector or a sequence.

[0017] According to the first aspect, or any one of the implementations of the first aspect, the N sets of expressiveness feature information corresponding to the N sets of source audio data are obtained, including: obtaining N sets of prosody feature information corresponding to the N sets of source audio data; obtaining N sets of emotion feature information corresponding to the N sets of source audio data; and generating N sets of expressiveness feature information corresponding to the N sets of source audio data based on the N sets of prosody feature information and the N sets of emotion feature information.

[0018] According to the first aspect, or any one of the implementations of the first aspect, the N sets of expressiveness feature information corresponding to the N sets of source audio data are obtained, including: obtaining N sets of prosody feature information corresponding to the N sets of source audio data; obtaining N sets of emotion feature information corresponding to the N sets of source audio data; and generating N sets of expressiveness feature information corresponding to the N sets of source audio data based on the N sets of prosody feature information and the N sets of emotion feature information.

[0019] According to the first aspect, or any one of the implementations of the first aspect, the N sets of expressiveness feature information corresponding to the N sets of source audio data are obtained, including: obtaining N sets of prosody feature information corresponding to the N sets of source audio data; obtaining N sets of emotion feature information corresponding to the N sets of source audio data; and generating N sets of expressiveness feature information corresponding to the N sets of source audio data based on the N sets of prosody feature information and the N sets of emotion feature information.

[0020] According to a first aspect, or any possible implementation mode of the first aspect, the N sets of source timbre feature information corresponding to the N pieces of source audio data are obtained, including: determining a role corresponding to each piece of source audio data in the N pieces of source audio data; obtaining N sets of initial timbre feature information corresponding to the N pieces of source audio data; for X pieces of source audio data corresponding to a first role in the N pieces of source audio data: performing weighted calculation based on the X sets of initial timbre feature information corresponding to the X pieces of source audio data, and determining a result of the weighted calculation as source timbre feature information corresponding to each piece of source audio data in the X pieces of source audio data, wherein X is a positive integer. In this way, the source timbre feature information of the source audio data of the same role in the reading material can be made the same, and the uniformity of the timbre of the same role in the reading material is ensured.

[0021] According to the first aspect, or any possible implementation mode of the first aspect, the N pieces of source audio data include P1 pieces of source audio data of a determined role and P2 pieces of source audio data of an undetermined role, N = P1 + P2, P1 and P2 are both positive integers, and the method further includes: for the jth piece of source audio data in the P2 pieces of source audio data: respectively calculating the similarity of the initial timbre feature information corresponding to the P1 pieces of source audio data and the initial timbre feature information corresponding to the jth piece of source audio data; determining the role corresponding to the source audio data of the P1 pieces of source audio data with the highest similarity of the initial timbre feature information corresponding to the jth piece of source audio data as the role of the jth piece of source audio data; wherein j is a positive integer, and the value range is between 1 and P2. In this way, the role of the source audio data of an undetermined role can be accurately determined, and the error rate of reference disambiguation is reduced.

[0022] For example, the description of 1 to P2 includes 1 and P2, that is, j can be equal to 1 or equal to P2.

[0023] In a second aspect, the embodiments of the present application provide a training method, which comprises the following steps. First, training data is collected, the training data comprising training audio data and a reference role label corresponding to the training audio data, the expressiveness feature information of the training audio data satisfying an expressiveness condition, the training audio data comprising audio data recorded by a plurality of users using their own voice timbres and / or audio data recorded by a plurality of users using pseudo voices, the voice timbres of different pseudo voices used by the same user having a voice timbre distinction degree greater than a threshold. Then, the training audio data is input into an emotion extractor, a content extractor and a prosody extractor for calculation to obtain prosody feature information output by the emotion extractor, content feature information output by the content extractor and prosody feature information output by the prosody extractor; and the training audio data and the reference role label are input into a voice timbre extractor for calculation to obtain voice timbre feature information output by the voice timbre extractor. Next, the prosody feature information, the content feature information, the prosody feature information and the voice timbre feature information are input into a spectrum reconstruction model for spectrum reconstruction to obtain audio spectrum feature information, and the audio spectrum feature information is subjected to frequency-time conversion to obtain reconstructed audio data. Subsequently, a first loss function value is calculated based on the reconstructed audio data and the training audio data, and the model parameters of the emotion extractor, the content extractor, the prosody extractor, the voice timbre extractor and the spectrum reconstruction model are jointly adjusted with the aim of minimizing the first loss function value. In this way, the voice timbre extractor capable of extracting accurate voice timbre feature information, the prosody extractor capable of extracting accurate prosody feature information, the emotion extractor capable of extracting accurate emotion feature information, the content extractor capable of extracting accurate content feature information and the spectrum reconstruction model capable of reconstructing audio spectrum feature information can be trained.

[0024] According to the second aspect, the method further comprises the following steps. On the one hand, the voice timbre feature information is input into a first classifier for calculation to obtain a first role label, and a second loss function value is calculated based on the first role label and the reference role label. On the other hand, the emotion feature information is input into a second classifier for calculation to obtain a second role label, and a second loss function value is calculated based on the second role label and the reference role label. Then, the model parameters of the voice timbre extractor are adjusted with the aim of minimizing the second loss function value and the mutual information of the voice timbre feature information and the emotion feature information, and the model parameters of the emotion extractor are adjusted with the aim of maximizing the third loss function value and minimizing the mutual information. In this way, the voice timbre feature information and the emotion feature information can be decoupled by reducing the overlap of the voice timbre feature information and the emotion feature information.

[0025] In a third aspect, the embodiments of the present application provide an electronic device, which comprises a memory and a processor, the memory being coupled to the processor; the memory stores program instructions which, when executed by the processor, cause the electronic device to perform the voice processing method in the first aspect or any possible implementation manner of the first aspect.

[0026] The third aspect and any kind of implementation manner of the third aspect correspond to the first aspect and any kind of implementation manner of the first aspect respectively. The technical effects corresponding to the third aspect and any kind of implementation manner of the third aspect can be referred to the technical effects corresponding to the first aspect and any kind of implementation manner of the first aspect, which will not be described here.

[0027] In a fourth aspect, an embodiment of the present application provides an electronic device, comprising a memory and a processor, the memory being coupled with the processor; the memory stores program instructions, when the program instructions are executed by the processor, causing the electronic device to execute the training method in the second aspect or any possible implementation manner of the second aspect.

[0028] The fourth aspect and any kind of implementation manner of the fourth aspect correspond to the second aspect and any kind of implementation manner of the second aspect respectively. The technical effects corresponding to the fourth aspect and any kind of implementation manner of the fourth aspect can be referred to the technical effects corresponding to the second aspect and any kind of implementation manner of the second aspect, which will not be described here.

[0029] In a fifth aspect, an embodiment of the present application provides a chip, comprising one or more interface circuits and one or more processors; the interface circuit is used to receive a signal from a memory of an electronic device and send a signal to the processor, the signal comprising computer instructions stored in the memory; when the processor executes the computer instructions, causing the electronic device to execute the voice processing method in the first aspect or any possible implementation manner of the first aspect.

[0030] The fifth aspect and any kind of implementation manner of the fifth aspect correspond to the first aspect and any kind of implementation manner of the first aspect respectively. The technical effects corresponding to the fifth aspect and any kind of implementation manner of the fifth aspect can be referred to the technical effects corresponding to the first aspect and any kind of implementation manner of the first aspect, which will not be described here.

[0031] In a sixth aspect, an embodiment of the present application provides a chip, comprising one or more interface circuits and one or more processors; the interface circuit is used to receive a signal from a memory of an electronic device and send a signal to the processor, the signal comprising computer instructions stored in the memory; when the processor executes the computer instructions, causing the electronic device to execute the training method in the second aspect or any possible implementation manner of the second aspect.

[0032] The sixth aspect and any kind of implementation manner of the sixth aspect correspond to the second aspect and any kind of implementation manner of the second aspect respectively. The technical effects corresponding to the sixth aspect and any kind of implementation manner of the sixth aspect can be referred to the technical effects corresponding to the second aspect and any kind of implementation manner of the second aspect, which will not be described here.

[0033] In a seventh aspect, an embodiment of the present application provides a computer storage medium, which stores a computer program. When the computer program is executed on a computer or a processor, the computer or the processor performs the speech processing method in the first aspect or any possible implementation manner of the first aspect.

[0034] The seventh aspect and the any possible implementation manner of the seventh aspect correspond to the first aspect and the any possible implementation manner of the first aspect respectively. The technical effects of the seventh aspect and the any possible implementation manner of the seventh aspect can refer to the technical effects of the first aspect and the any possible implementation manner of the first aspect, which will not be repeated here.

[0035] In an eighth aspect, an embodiment of the present application provides a computer storage medium, which stores a computer program. When the computer program is executed on a computer or a processor, the computer or the processor performs the training method in the second aspect or any possible implementation manner of the second aspect.

[0036] The eighth aspect and the any possible implementation manner of the eighth aspect correspond to the second aspect and the any possible implementation manner of the second aspect respectively. The technical effects of the eighth aspect and the any possible implementation manner of the eighth aspect can refer to the technical effects of the second aspect and the any possible implementation manner of the second aspect, which will not be repeated here.

[0037] In a ninth aspect, an embodiment of the present application provides a computer program product, which contains a software program. When the software program is executed on a computer or a processor, the steps of the speech processing method in the first aspect or any possible implementation manner of the first aspect are executed.

[0038] The ninth aspect and the any possible implementation manner of the ninth aspect correspond to the first aspect and the any possible implementation manner of the first aspect respectively. The technical effects of the ninth aspect and the any possible implementation manner of the ninth aspect can refer to the technical effects of the first aspect and the any possible implementation manner of the first aspect, which will not be repeated here.

[0039] In a tenth aspect, an embodiment of the present application provides a computer program product, which contains a software program. When the software program is executed on a computer or a processor, the steps of the training method in the second aspect or any possible implementation manner of the second aspect are executed.

[0040] The tenth aspect and any implementation form of the tenth aspect correspond to the second aspect and any implementation form of the second aspect, respectively. The technical effects of the tenth aspect and any implementation form of the tenth aspect can be referred to the technical effects of the second aspect and any implementation form of the second aspect, which will not be described here again.

[0041] In a eleventh aspect, an embodiment of the present application provides a voice processing apparatus, comprising:

[0042] a data obtaining module, configured to obtain unicast audio data of a unicast audiobook, the unicast audio data comprising N pieces of source audio data, N being a positive integer;

[0043] a character voice timbre analysis module, configured to determine, from M sets of reference voice timbre feature information, N sets of target voice timbre feature information matched with the N pieces of source audio data, the M sets of reference voice timbre feature information corresponding to M kinds of voice timbres, a voice timbre distinction degree between any two sets of the M sets of reference voice timbre feature information being greater than a distinction degree threshold, M being a positive integer;

[0044] a character voice timbre conversion module, configured to obtain N sets of expressiveness feature information corresponding to the N pieces of source audio data, and perform voice timbre conversion on the N pieces of source audio data respectively based on the N sets of target voice timbre feature information and the N sets of expressiveness feature information, to generate multicast audio data.

[0045] According to the eleventh aspect, the character voice timbre analysis module comprises:

[0046] an audio sentence division module, configured to perform audio sentence division on the unicast audio data to obtain the N pieces of source audio data, the N pieces of source audio data corresponding to audio sentences one by one;

[0047] a voice timbre feature extraction module, configured to extract, by using a voice timbre extractor, N sets of source voice timbre feature information corresponding to the N pieces of source audio data;

[0048] a voice timbre feature matching module, configured to determine, based on the N sets of source voice timbre feature information, N sets of target voice timbre feature information from the M sets of reference voice timbre feature information.

[0049] According to the eleventh aspect, or any implementation form of the eleventh aspect, the voice timbre feature matching module is configured to, for the i th piece of source audio data in the N pieces of source audio data: respectively determine a similarity of the M sets of reference voice timbre feature information with the i th piece of source audio data; and determine, as the target voice timbre feature information matched with the i th piece of source audio data, the reference voice timbre feature information with the highest similarity with the i th piece of source audio data, where i is a positive integer and takes a value in a range between 1 and N.

[0050] According to the eleventh aspect, or any implementation form of the eleventh aspect, the character voice timbre conversion module comprises:

[0051] a prosody feature extraction module configured to extract, by using a prosody extractor, N sets of prosody feature information corresponding to the N sets of source audio data;

[0052] an emotion feature extraction module configured to extract, by using an emotion extractor, N sets of emotion feature information corresponding to the N sets of source audio data;

[0053] a expressiveness feature generation module configured to generate, based on the N sets of prosody feature information and the N sets of emotion feature information, N sets of expressiveness feature information corresponding to the N sets of source audio data.

[0054] According to the eleventh aspect, or any possible implementation mode of the eleventh aspect, the role voice timbre conversion module comprises:

[0055] a content feature extraction module configured to extract, by using a content extractor, N sets of content feature information corresponding to the N sets of source audio data;

[0056] a feature spectrum reconstruction module configured to perform spectrum reconstruction based on the N sets of target voice timbre feature information, the N sets of expressiveness feature information and the N sets of content feature information by using a spectrum reconstruction model, to obtain N sets of audio spectrum feature information;

[0057] a frequency-time conversion module configured to perform frequency-time conversion on the N sets of audio spectrum feature information respectively, to obtain N sets of target audio data;

[0058] a splicing module configured to splice the N sets of target audio data, to obtain the multicast audio data.

[0059] According to the eleventh aspect, or any possible implementation mode of the eleventh aspect, the audio sentence division module is configured to divide the unicast audio data into the N sets of source audio data by performing VAD (Voice Activity Detection) detection on the unicast audio data.

[0060] According to the eleventh aspect, or any possible implementation mode of the eleventh aspect, the audio sentence division module is configured to obtain a reading text of the unicast audiobook, divide the reading text into N text sentences, align the unicast audio data and the reading text to determine N audio time intervals in the unicast audio data corresponding to the N text sentences, and divide the unicast audio data into the N sets of source audio data based on the N audio time intervals.

[0061] According to a twelfth aspect, or any possible implementation mode of the twelfth aspect, the timbre feature extraction module is configured to determine a role corresponding to each piece of source audio data in the N pieces of source audio data; obtain N sets of initial timbre feature information corresponding to the N pieces of source audio data; and for X pieces of source audio data corresponding to a first role in the N pieces of source audio data, perform weighted calculation based on the X sets of initial timbre feature information corresponding to the X pieces of source audio data, and determine a result of the weighted calculation as source timbre feature information corresponding to each piece of source audio data in the X pieces of source audio data, where X is a positive integer.

[0062] According to the twelfth aspect, or any possible implementation mode of the twelfth aspect, the N pieces of source audio data include P1 pieces of source audio data of a determined role and P2 pieces of source audio data of an undetermined role, N = P1 + P2, P1 and P2 are positive integers, and the device further includes:

[0063] a role determination module configured to, for the jth piece of source audio data in the P2 pieces of source audio data, calculate a similarity between initial timbre feature information corresponding to each piece of source audio data in the P1 pieces of source audio data and initial timbre feature information corresponding to the jth piece of source audio data, respectively; and determine a role corresponding to source audio data in the P1 pieces of source audio data with the highest similarity between initial timbre feature information corresponding to the jth piece of source audio data as a role of the jth piece of source audio data, where j is a positive integer and ranges from 1 to P2.

[0064] The twelfth aspect and any possible implementation mode of the twelfth aspect correspond to the first aspect and any possible implementation mode of the first aspect, respectively. The technical effects of the twelfth aspect and any possible implementation mode of the twelfth aspect can refer to the technical effects of the first aspect and any possible implementation mode of the first aspect, which will not be described here.

[0065] According to a twelfth aspect, or any possible implementation mode of the twelfth aspect, the timbre feature extraction module is configured to determine a role corresponding to each piece of source audio data in the N pieces of source audio data; obtain N sets of initial timbre feature information corresponding to the N pieces of source audio data; and for X pieces of source audio data corresponding to a first role in the N pieces of source audio data, perform weighted calculation based on the X sets of initial timbre feature information corresponding to the X pieces of source audio data, and determine a result of the weighted calculation as source timbre feature information corresponding to each piece of source audio data in the X pieces of source audio data, where X is a positive integer.

[0066] a collection module configured to collect training data, the training data including training audio data and a reference role label corresponding to the training audio data, performance feature information of the training audio data satisfying a performance condition, the training audio data including audio data recorded by a plurality of users using their own timbres and / or audio data recorded by the plurality of users using pseudo-voices, and a timbre distinction degree of different pseudo-voices used by a same user being greater than a distinction degree threshold.

[0067] The feature information extraction module is configured to input the training audio data into the emotion extractor, the content extractor, and the prosody extractor respectively for calculation, to obtain prosody feature information output by the emotion extractor, content feature information output by the content extractor, and prosody feature information output by the prosody extractor; and input the training audio data and the reference role label into the timbre extractor for calculation, to obtain timbre feature information output by the timbre extractor.

[0068] The audio data reconstruction module is configured to input the prosody feature information, the content feature information, the prosody feature information, and the timbre feature information into the spectrum reconstruction model for spectrum reconstruction, to obtain audio spectrum feature information, and perform frequency-time conversion on the audio spectrum feature information, to obtain reconstructed audio data.

[0069] The back propagation module is configured to calculate a first loss function value based on the reconstructed audio data and the training audio data, to jointly adjust model parameters of the emotion extractor, the content extractor, the prosody extractor, the timbre extractor, and the spectrum reconstruction model, with a target of minimizing the first loss function value.

[0070] According to the twelfth aspect, the device further includes:

[0071] The loss function value calculation module is configured to input the timbre feature information into the first classifier for calculation, to obtain a first role label, and calculate a second loss function value based on the first role label and the reference role label; input the emotion feature information into the second classifier for calculation, to obtain a second role label, and calculate a second loss function value based on the second role label and the reference role label.

[0072] The timbre extractor training module is configured to adjust model parameters of the timbre extractor, with a target of minimizing the second loss function value and mutual information of the timbre feature information and the emotion feature information.

[0073] The emotion extractor training module is configured to adjust model parameters of the emotion extractor, with a target of maximizing the third loss function value and minimizing the mutual information.

[0074] The twelfth aspect and any one of the implementation manners of the twelfth aspect correspond to the second aspect and any one of the implementation manners of the second aspect respectively. The technical effects corresponding to the twelfth aspect and any one of the implementation manners of the twelfth aspect can be found in the technical effects corresponding to the second aspect and any one of the implementation manners of the second aspect, which will not be described herein again. BRIEF DESCRIPTION OF DRAWINGS

[0075] Figure 1 An application scenario schematic diagram shown for example;

[0076] Figure 2 An application scenario schematic diagram shown for example;

[0077] Figure 3 schematic diagram of a processing procedure shown for illustration;

[0078] Figure 4 schematic diagram of a processing procedure shown for illustration;

[0079] Figure 5 schematic diagram of a model shown for illustration;

[0080] Figure 6 schematic diagram of a training procedure shown for illustration;

[0081] Figure 7 schematic diagram of a training procedure shown for illustration;

[0082] Figure 8 schematic diagram of an information extraction procedure shown for illustration;

[0083] Figure 9a schematic diagram of an information extraction procedure shown for illustration;

[0084] Figure 9b schematic diagram of an information matching procedure shown for illustration;

[0085] Figure 10 schematic diagram of a timbre conversion shown for illustration;

[0086] Figure 11 schematic diagram of a processing procedure shown for illustration;

[0087] Figure 12 schematic diagram of an information extraction procedure shown for illustration;

[0088] Figure 13a schematic diagram of a speech processing apparatus shown for illustration;

[0089] Figure 13b schematic diagram of a speech processing apparatus shown for illustration;

[0090] Figure 14 schematic diagram of a training apparatus shown for illustration;

[0091] Figure 15 schematic diagram of an apparatus shown for illustration. DETAILED DESCRIPTION

[0092] With reference to the drawings and the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts are within the scope of the present application.

[0093] The term "and / or" used herein is only used to describe an association relationship of associated objects, and means that three relationships can exist, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone.

[0094] The terms "first" and "second" and the like in the description and claims of the embodiments of the present application are used to distinguish different objects, and are not used to describe a specific order of the objects. For example, the first target object and the second target object are used to distinguish different target objects, and are not used to describe a specific order of the target objects.

[0095] In the embodiments of the present application, the words "exemplary" or "for example" are used to mean serving as an example, instance, or illustration. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of the words "exemplary" or "for example" is intended to present relevant concepts in a specific way.

[0096] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more. For example, a plurality of processing units means two or more processing units; a plurality of systems means two or more systems.

[0097] Figure 1 The application scenario shown is exemplary.

[0098] Referring to Figure 1 For example, one possible scenario is a scenario in which a user plays an audiobook.

[0099] Referring to Figure 1For example, the user can start an audio book application in the mobile phone, enter an audio book application main interface 101, and the audio book application main interface 101 can include one or more controls, including but not limited to: a search box, a search option, a book list, a text reading option, an audio reading option 102, and the like. For example, the user can enter a query word in the search box and click the search option to query a book to be read. For example, the user can perform a page turning operation or a sliding operation in the book list to find a book to be read. For example, after the user finds a book to be read in the book list or searches for a book to be read, the user can click the text reading option to enter a text reading interface for text reading. For example, after the user finds a book to be read in the book list or searches for a book to be read, the user can click the audio reading option 102, and the mobile phone can play audio data corresponding to the book to be read by the user in response to the operation behavior of the user.

[0100] For example, the audio book application includes unicast audio books and multicast audio books. The unicast audio book refers to an audio book recorded by a single voice actor through pseudo voice for multi-role performance, and the timbre distinction degree of each role is low and the scene performance degree is low. The multicast audio book refers to an audio book recorded by multiple voice actors for roles, and the timbre distinction degree of each role is high and the scene performance degree is high. If the timbre distinction degrees of two timbres are greater than or equal to a distinction degree threshold, it can be determined that the timbre distinction degrees of the two timbres are high. If the timbre distinction degrees of two timbres are less than the distinction degree threshold, it can be determined that the timbre distinction degrees of the two timbres are low. The distinction degree threshold can be set as required, and the present application does not limit this. For example, if the user clicks the audio reading option 102 for a unicast audio book, in the present application, after receiving the user's click on the audio reading option 102, the mobile phone can convert the unicast audio book into a multicast audio book in response to the operation behavior of the user, and then play the converted multicast audio book, so as to improve the scene performance degree of the audio book and facilitate the user to quickly and fully understand the plot.

[0101] Figure 2 For example, the application scenario is schematically shown.

[0102] Referring to Figure 2 For example, one possible scenario is a scenario in which an audio book maker makes an audio book.

[0103] Referring to Figure 2For example, an audiobook producer can start the audiobook production platform, enter the audiobook production platform main interface 201, and produce an audiobook. For example, the audiobook production platform main interface 201 can include one or more controls, including but not limited to: a produced audiobook option, an audiobook production option, and the like. For example, after the user clicks the produced audiobook option, the electronic device can enter the editing interface of the produced audiobook in response to the user's behavior operation, and edit the produced audiobook, such as changing the name, cover, and the like. For example, the user can click the audiobook production option 202, and the electronic device can enter the audiobook production interface in response to the user's behavior operation, and the user can produce an audiobook in the audiobook production interface (such as producing a single-cast audiobook or a multi-cast audiobook).

[0104] For example, since the prior art requires multiple voice actors to be assigned to record according to the characters in the book to produce a multi-cast audiobook, the production time is long and the efficiency is low, and a single-cast audiobook only requires a single voice actor to perform multiple characters, so in this application, the audiobook producer can first record a single-cast audiobook during the production of a multi-cast audiobook, and then convert the recorded single-cast audiobook into a multi-cast audiobook. In this way, the production time and cost of the multi-cast audiobook can be reduced, and the production efficiency of the multi-cast audiobook can be improved.

[0105] For example, in this application, for a single-cast audiobook that has been produced in the audiobook production platform, the audiobook producer can perform a conversion operation to convert the produced single-cast audiobook into a multi-cast audiobook. Subsequently, after the user enters the audiobook application main interface and clicks the audiobook reading option 102 of a certain multi-cast book, the mobile phone can quickly play the audio data of the multi-cast audiobook in response to the user's operation behavior.

[0106] It should be noted that the book in this application can include a book with multiple characters, such as a novel, a story, a sketch, and the like.

[0107] Now how to convert a single-cast audiobook into a multi-cast audiobook will be described by way of example.

[0108] For example, a single-cast audiobook can include single-cast audio data and book text. For example, the single-cast audio data can include multiple source audio data, each piece of source audio data including multiple frames of audio data, and each piece of source audio data corresponding to an audio sentence. Correspondingly, the book text includes multiple text sentences.

[0109] Figure 3 The processing process is shown by way of example.

[0110] Referring to Figure 3For example, the electronic device can include a character voice timbre analysis module and a character voice timbre conversion module. It should be understood that Figure 3 The electronic device shown is only an example of an electronic device, and the electronic device can have more or fewer modules than those shown in the figure, which is not limited in the present application.

[0111] For example, the character voice timbre analysis module is used to analyze target voice timbre feature information matched with each piece of source audio data in the unicast audio data.

[0112] For example, the character voice timbre conversion module is used to convert the voice timbre of each piece of source audio data in the unicast audio data.

[0113] For example, the present application can convert unicast audiobooks into multicast audiobooks through the character voice timbre analysis module and the character voice timbre conversion module in the electronic device, and the process can be as follows:

[0114] S301, input the unicast audio data of the unicast audiobook into the character voice timbre analysis module.

[0115] In one possible scenario, when the user clicks on the audiobook reading option 102 in the application interface 100, the unicast audio data of the unicast audiobook corresponding to the audiobook reading option 102 can be obtained in response to the user operation behavior, and then the obtained unicast audio data is input into the character voice timbre analysis module. Figure 1 In one possible scenario, when the user clicks on the audiobook reading option 102 in the application interface 100, the unicast audio data of the unicast audiobook corresponding to the audiobook reading option 102 can be obtained in response to the user operation behavior, and then the obtained unicast audio data is input into the character voice timbre analysis module.

[0116] Figure 2 In one possible scenario, when the user clicks on the audiobook reading option 102 in the application interface 100, the unicast audio data of the unicast audiobook corresponding to the audiobook reading option 102 can be obtained in response to the user operation behavior, and then the obtained unicast audio data is input into the character voice timbre analysis module.

[0117] In one possible scenario, when the user clicks on the audiobook reading option 102 in the application interface 100, the unicast audio data of the unicast audiobook corresponding to the audiobook reading option 102 can be obtained in response to the user operation behavior, and then the obtained unicast audio data is input into the character voice timbre analysis module.

[0118] S302, the character voice timbre analysis module outputs N pieces of source audio data in the unicast audio data and the corresponding matched N sets of target voice timbre feature information.

[0119] Wherein, N is a positive integer.

[0120] Figure 4 The processing process is schematically shown.

[0121] Referring to Figure 4 ​Exemplarily, after receiving the unicast audio data, the role timbre analysis module can perform audio sentence segmentation on the unicast audio data to obtain N pieces of source audio data. Then, the role timbre analysis module can perform timbre feature information extraction on the N pieces of source audio data respectively to obtain N groups of source timbre feature information corresponding to the N pieces of source audio data, wherein one group of source timbre feature information corresponds to one piece of source audio data. Then, the role timbre analysis module can perform matching based on the N groups of source timbre feature information corresponding to the N pieces of source audio data respectively to determine N groups of target timbre feature information corresponding to the N pieces of source audio data matched, wherein one group of target timbre feature information corresponds to one piece of source audio data.

[0122] Exemplarily, the timbre feature information can be used to represent information related to timbre, which can include but is not limited to personality characteristics, gender characteristics, age characteristics, and voice region (such as high voice region, medium voice region, and low voice region) characteristics, etc., which are not limited in the present application.

[0123] Audio sentence segmentation

[0124] Exemplarily, the reading material can include multiple chapters. The unicast audio data can be first divided into W (W is a positive integer) pieces of chapter audio data, each piece of chapter audio data corresponding to a chapter. Then, each piece of chapter audio data in the W pieces of chapter audio data can be divided into R (R is a positive integer) pieces of source audio data by audio sentence segmentation, so as to improve the efficiency of audio sentence segmentation. Wherein, N=W*R.

[0125] Exemplarily, each frame of audio data in the unicast audio data has a chapter identifier, which is used to uniquely identify a chapter. Further, in one possible manner, the role timbre analysis module can determine the frames of audio data belonging to the same chapter through the chapter identifiers of the frames of audio data in the unicast audio data, so as to divide the unicast audio data into W pieces of chapter audio data.

[0126] Exemplarily, there is a pause between the adjacent two chapters of the reading material, and correspondingly, there is also a pause between the adjacent two pieces of chapter audio data of the unicast audio data. Wherein, the pause between the adjacent two pieces of chapter audio data of the unicast audio data can be a first preset time length, which can be set according to requirements, which are not limited in the present application. Further, in one possible manner, the role timbre analysis module can use VAD (Voice Activity Detection) detection to divide the unicast audio data into W pieces of chapter audio data.

[0127] Exemplarily, the role voice tone analysis module can adopt VAD to detect whether a time interval between two adjacent frames of audio data is greater than or equal to a first preset time length. If the time interval between the two adjacent frames of audio data is greater than or equal to the first preset time length, the chapter division can be performed between the two adjacent frames of audio data, the previous frame of audio data is divided into the end of the chapter audio data of the previous chapter, and the next frame of audio data is divided into the beginning of the chapter audio data of the next chapter.

[0128] Exemplarily, text recognition can be performed on the unicast audio data to obtain a corresponding text recognition result. Then, text analysis can be performed based on the text recognition result, the audio chapter division is performed on the unicast audio data to obtain W pieces of chapter audio data. Exemplarily, when a chapter distinguishing text (such as “chapter”, “chapter”, etc.) is detected, an audio time interval corresponding to the chapter distinguishing text in the unicast audio data can be determined, and the audio chapter division is performed based on the endpoints of the audio time interval corresponding to the chapter distinguishing text.

[0129] Exemplarily, the unicast audio data can be stored in chapters, at this time, a piece of chapter audio data corresponding to each chapter can be directly obtained each time, and then audio sentence division is performed on the chapter audio data to obtain R pieces of source audio data.

[0130] It should be noted that other ways can also be used to perform audio chapter division on the unicast audio text, and the present application does not limit this.

[0131] Exemplarily, there is a pause between two adjacent sentences in the reading material, and correspondingly, there is also a pause between two adjacent pieces of source audio data in each piece of chapter audio data. The pause between the two adjacent pieces of source audio data is greater than a second preset time length and less than a first preset time length, the second preset time length is less than the first preset time length, and the second preset time length can be set according to requirements, and the present application does not limit this. Further, in a possible way, for the kth (k is a positive integer, and the value range is between 1 and W; this description of 1 to W includes 1 and W, that is, k can be 1 or W) piece of chapter audio data in the W pieces of chapter audio data, the role voice tone analysis module can adopt VAD detection to divide the kth piece of chapter audio data into R pieces of source audio data.

[0132] Exemplarily, for the kth chapter audio data, the character timbre analysis module can detect, through the VAD, whether a time interval between two adjacent frames of audio data in the kth chapter audio data is greater than or equal to a second preset time length. If the time interval between the two adjacent frames of audio data is greater than or equal to the second preset time length, the audio sentence division can be performed between the two adjacent frames of audio data, the previous frame of audio data is divided into the end of the source audio data of the previous sentence, and the next frame of audio data is divided into the beginning of the source audio data of the next sentence.

[0133] It should be noted that the audio sentence division can be performed directly without performing the audio chapter division on the unicast audio data, and the present application does not limit this.

[0134] Further, the unicast audio data is divided into audio sentences in the above manner, and N pieces of source audio data can be obtained.

[0135] Timbre feature information extraction

[0136] Exemplarily, the timbre extractor can be pre-trained, and then the trained timbre extractor is used to extract the timbre feature information, so as to obtain N groups of source timbre feature information corresponding to the N pieces of source audio data.

[0137] Figure 5 The structure of the model is exemplarily shown.

[0138] Reference Figure 5 Exemplarily, the conversion model can include a timbre extractor, an expressiveness extraction module, a content extractor, and a spectrum reconstruction model. Exemplarily, the expressiveness extraction module includes but is not limited to a prosody extractor and an emotion extractor. It should be understood that, Figure 5 The conversion model shown in the figure is only an example of the conversion model, and the conversion model can have more or fewer modules than those shown in the figure, and the present application does not limit this.

[0139] Exemplarily, the timbre extractor, the prosody extractor, the emotion extractor, the content extractor, and the spectrum reconstruction model in the conversion model can be jointly trained.

[0140] Exemplarily, a plurality of users can record a plurality of pieces of training audio data with high expressiveness using their own timbres; wherein each user records at least one piece of training audio data. Exemplarily, a plurality of users can record a plurality of pieces of training audio data with high expressiveness using a plurality of timbres with high timbre differentiation; wherein each user uses at least one pseudo-tone for recording, and each user records at least one piece of training audio data using one pseudo-tone. Exemplarily, each piece of training audio data can include at least one training audio data, and each training audio data corresponds to one sentence.

[0141] Exemplarily, the high expressiveness can refer to that the expressiveness characteristic information satisfies an expressiveness condition. The expressiveness characteristic information can include prosody characteristic information and emotion characteristic information, and the expressiveness condition includes a prosody condition and an emotion condition. The high expressiveness can refer to that the prosody characteristic information satisfies the prosody condition and the emotion characteristic information satisfies the emotion condition. The expressiveness condition, the prosody condition and the emotion condition can be set according to requirements, and the present application does not limit this.

[0142] Exemplarily, for each piece of collected training audio data, a corresponding reference timbre label can be added to the piece of training audio data based on the role information of the timbre used by the user corresponding to the piece of training audio data. The role information includes but is not limited to gender, age, personality, sound area, etc. Exemplarily, the role information can be encoded (such as one-hot (one-bit effective encoding) and the like) to obtain the reference timbre label.

[0143] Exemplarily, a piece of training audio data and the reference timbre label corresponding to the piece of training audio data can be taken as a set of training data, and then a plurality of sets of training data can be obtained. The conversion model can be trained by using the plurality of sets of training data. The present application exemplarily illustrates the training of the conversion model by taking one set of training data as an example.

[0144] Figure 6 The training process is exemplarily illustrated by the schematic diagram.

[0145] Exemplarily, the training audio data in the training data can be respectively input into the prosody extractor, the emotion extractor and the content extractor.

[0146] Exemplarily, the training audio data in the training data and the corresponding reference timbre label can be input into the timbre extractor.

[0147] Exemplarily, after the timbre extractor receives the training audio data and the corresponding reference timbre label, the training audio data and the reference timbre label can be forward calculated, and the timbre characteristic information can be output to the spectrum reconstruction model. Exemplarily, the timbre characteristic information can be represented by a vector or a sequence table.

[0148] Exemplarily, after the prosody extractor receives the training audio data, the training audio data can be forward calculated, and the prosody characteristic information can be output to the spectrum reconstruction model. Exemplarily, the prosody characteristic information can be used to represent the voice sound emotion skill information such as lightness, rapidity, virtuality and reality, and can be represented by a vector or a sequence table.

[0149] Exemplarily, after receiving the training audio data, the emotion extractor can perform forward calculation on the training audio data, and output emotion feature information to the spectrum reconstruction model. Exemplarily, the emotion feature information can be used to represent emotion types (such as happy, sad, high-pitched, low-pitched, etc.), and attitude dimensions (such as positive, negative, praise, sarcasm, etc.), and can be represented by a vector or a sequence.

[0150] Exemplarily, after receiving the training audio data, the content extractor can perform forward calculation on the training audio data, and output content feature information to the spectrum reconstruction model. Exemplarily, the content feature information can be used to represent speech content. Exemplarily, the content feature information can be a phoneme feature, which can be represented by a vector or a sequence. Exemplarily, the content feature information can be a phoneme posterior probability (that is, a probability distribution of a phoneme), which can be represented by a matrix.

[0151] In one possible manner, after receiving the timbre feature information, the prosody feature information, the emotion feature information, and the content feature information, the spectrum reconstruction model can perform spectrum reconstruction based on the timbre feature information, the prosody feature information, the emotion feature information, and the content feature information, and output audio spectrum feature information. Exemplarily, the audio spectrum feature information output by the spectrum reconstruction model can be converted into the time domain to obtain reconstructed audio data of the audio spectrum feature information in the time domain. That is, the spectrum reconstruction model only performs spectrum reconstruction, and the time domain conversion is performed by other modules.

[0152] In one possible manner, after receiving the timbre feature information, the prosody feature information, the emotion feature information, and the content feature information, the spectrum reconstruction model can first perform spectrum reconstruction based on the timbre feature information, the prosody feature information, the emotion feature information, and the content feature information to obtain audio spectrum feature information, and then convert the audio spectrum feature information into the time domain to obtain reconstructed audio data of the audio spectrum feature information in the time domain and output. That is, the spectrum reconstruction and the frequency-time conversion are performed by the spectrum reconstruction model.

[0153] It should be noted that the present application does not limit whether the spectrum reconstruction model only performs spectrum reconstruction or performs spectrum reconstruction and frequency-time conversion.

[0154] Exemplarily, the reconstructed audio data of the audio spectrum feature information in the time domain can be compared with the training audio data in the training data to calculate a corresponding first loss function value. Then, the model parameters of the timbre extractor, the prosody extractor, the emotion extractor, the content extractor, and the spectrum reconstruction model in the conversion model are adjusted to minimize the first loss function value.

[0155] Exemplarily, the conversion model can be trained by using each group of training data in the above manner until the first loss function value meets the first loss condition, or the training frequency of each module in the conversion model meets the corresponding training frequency condition, or the performance of each module in the conversion model meets the corresponding performance condition. Exemplarily, the first loss condition, the training frequency condition and the performance condition can be set according to requirements, and the present application does not make any limitation in this regard. The training frequency conditions of different modules in the conversion model can be different or the same, and the performance conditions of different modules can be different, and the present application does not make any limitation in this regard.

[0156] It should be noted that the content extractor can also be trained independently of the timbre extractor, the prosody extractor, the emotion extractor and the spectrum reconstruction model, and the present application does not make any limitation in this regard.

[0157] Exemplarily, since most of the pseudo-voice is realized by adjusting the speech rate, the fundamental frequency, the formant and other shallow features, and the emotion is often expressed through these skills, which makes the timbre feature information overlap with the emotion feature information. Therefore, on the basis of training the timbre extractor and the emotion extractor as described above, the timbre extractor and the emotion extractor are trained by using the following method to reduce the overlap of the timbre feature information and the emotion feature information, so as to decouple the timbre feature information output by the timbre extractor and the emotion feature information output by the emotion extractor.

[0158] Figure 7 The training process is exemplarily shown in the schematic diagram.

[0159] Exemplarily, after the timbre extractor performs forward calculation based on the training audio data and the reference timbre label in the training data to obtain the timbre feature information, the timbre feature information can be output to the first classifier and the mutual information module respectively.

[0160] Exemplarily, after the emotion extractor performs forward calculation based on the training audio data in the training data to obtain the emotion feature information, the emotion feature information can be output to the second classifier and the mutual information module respectively.

[0161] Exemplarily, the first classifier can perform calculation based on the timbre feature information to output the first timbre label.

[0162] Exemplarily, the second classifier can perform calculation based on the emotion feature information to output the second timbre label.

[0163] Exemplarily, the mutual information module can calculate the mutual information between the timbre feature information and the emotion feature information based on the timbre feature information and the emotion feature information. The mutual information can be an amount of information about one variable contained in another variable, and the mutual information between the timbre feature information and the emotion feature information refers to an amount of information that the timbre feature information contains the emotion feature information, or an amount of information that the emotion feature information contains the timbre feature information.

[0164] Exemplarily, the second loss function value can be calculated based on the first timbre label and the reference timbre label in the training data, and the third loss function value can be calculated based on the second timbre label feature and the reference timbre label in the training data. Exemplarily, the model parameters of the timbre extractor are adjusted to minimize the second loss function value and the mutual information. Exemplarily, the model parameters of the emotion extractor are adjusted to minimize the mutual information and maximize the third loss function value. In this way, the difference between the emotion feature information extracted by the emotion extractor and the timbre feature information extracted by the timbre extractor can be increased, so that the emotion feature information extracted by the emotion extractor and the timbre feature information extracted by the timbre extractor are decoupled.

[0165] Exemplarily, the timbre extractor and the emotion extractor can be trained by using each group of training data in the above manner until the second loss function value meets the second loss condition, or the training frequency of the timbre extractor meets the training frequency condition of the timbre extractor, or the performance of the timbre extractor meets the performance condition of the timbre extractor, and the training of the timbre extractor is stopped. And until the third loss function value meets the third loss condition, or the training frequency of the emotion extractor meets the training frequency condition of the emotion extractor, or the performance of the emotion extractor meets the performance condition of the emotion extractor, and the training of the emotion extractor is stopped. Exemplarily, the second loss condition and the third loss condition can be set as required, and the embodiments of the present application do not limit this.

[0166] After each module in the to-be-converted model completes the training, N pieces of source audio data in the unicast audio data can be input to the trained timbre extractor in sequence, and the timbre extractor extracts N groups of source timbre feature information corresponding to the N pieces of source audio data.

[0167] Figure 8 An information extraction process diagram is exemplarily shown.

[0168] Referring to Figure 8 Exemplarily, Figure 8Five pieces of source audio data are shown in the figure: source audio data 1, source audio data 2, source audio data 3, source audio data 4, and source audio data 5. Source audio data 1 is input to the trained timbre extractor, and timbre feature information A can be output, which is the source timbre feature information corresponding to source audio data 1. Source audio data 2 is input to the trained timbre extractor, and timbre feature information B can be output, which is the source timbre feature information corresponding to source audio data 2. Source audio data 3 is input to the trained timbre extractor, and timbre feature information A can be output, which is the source timbre feature information corresponding to source audio data 3. Source audio data 4 is input to the trained timbre extractor, and timbre feature information B can be output, which is the source timbre feature information corresponding to source audio data 4. Source audio data 5 is input to the trained timbre extractor, and timbre feature information C can be output, which is the source timbre feature information corresponding to source audio data 5. For example, the timbre feature information of source audio data 1 and source audio data 3 is the same, that is, source audio data 1 and source audio data 3 are audio data of the same role. For example, the timbre feature information of source audio data 2 and source audio data 4 is the same, that is, source audio data 2 and source audio data 4 are audio data of the same role.

[0169] For example, the training audio data includes training audio data recorded by different users, and the timbres of different users are highly distinguishable, and training audio data recorded by the same user using multiple highly distinguishable pseudo voices; that is, the timbres of each training audio data are highly distinguishable. Therefore, in order to improve the timbre distinguishability of different roles in the unicast audio data, the timbre of each source audio data in the unicast audio data can be converted to the timbre that matches the timbre of the source audio data among the training audio data corresponding timbres.

[0170] Figure 9a The information extraction process is schematically shown.

[0171] Referring to Figure 9a For example, the training audio data in each set of training data can be input to the trained timbre extractor to output the corresponding timbre feature information. In order to facilitate distinction, the timbre feature information extracted by the timbre extractor for the training audio data can be referred to as reference timbre feature information.

[0172] Exemplarily, for each set of training data, the training audio data in the set of training data can be divided into multiple pieces of training audio data in the above manner. Then the trained voice timbre extractor is used to extract the reference voice timbre feature information of each piece of training audio data. Exemplarily, the training audio data includes audio data recorded by using M (M is a positive integer) voice timbres (including the user's own voice timbre and pseudo voice timbres), and multiple pieces of training audio data are recorded for each voice timbre. Exemplarily, for the rth (r is a positive integer, and the value range is between 1 and M; this description between 1 and M includes 1 and M, that is, r can be 1 or M) voice timbre in the M voice timbres, the r reference voice timbre feature information corresponding to the multiple pieces of training audio data recorded by using the rth voice timbre can be weighted and calculated to obtain the reference voice timbre feature information corresponding to the rth voice timbre. Wherein, the reference voice timbre feature information corresponding to the rth voice timbre can be referred to as the rth set of reference voice timbre feature information. Optionally, the weighted calculation can be to calculate the average. Further, the above manner can be used to obtain M sets of reference voice timbre feature information.

[0173] Exemplarily, the reference voice timbre feature information matched with the source voice timbre feature information of each source audio data can be found from the multiple sets of reference voice timbre feature information, respectively, to find the voice timbre matched with the voice timbre of each source audio data in the unicast audio data from the voice timbre of the training audio data. Wherein, in order to facilitate the distinction and explanation, the reference voice timbre feature information matched with the source voice timbre feature information of the source audio data can be referred to as the target voice timbre feature information.

[0174] Exemplarily, taking the ith (i is a positive integer, and the value range is between 1 and N) source audio data in the N source audio data as an example for exemplarily description. Exemplarily, the similarity between the M sets of reference voice timbre feature information and the source voice timbre feature information of the ith source audio data can be calculated respectively, and the reference voice timbre feature information with the highest similarity to the source voice timbre feature information of the ith source audio data is taken as the target voice timbre feature information. In this way, the N sets of target voice timbre feature information matched with the N source audio data can be found from the M sets of reference voice timbre feature information. Exemplarily, the target voice timbre feature information matched with different source audio data can be the same or different.

[0175] Exemplarily, the distance information between the M sets of reference voice timbre feature information and the source voice timbre feature information of the ith source audio data can be calculated respectively, and the distance information is taken as the similarity between the source voice timbre feature information of the ith source audio data and the reference voice timbre feature information. In a possible manner, the distance information is inversely proportional to the similarity, that is, the greater the distance information, the lower the similarity; the smaller the distance information, the higher the similarity.

[0176] Exemplarily, the distance information between the source timbre feature information of the source audio data and each reference timbre feature information can be determined by calculating the Euclidean distance, cosine similarity, Minkowski distance, etc. between the source timbre feature information of the source audio data and each reference timbre feature information. The present application does not limit this.

[0177] Figure 9b An information matching process diagram is exemplarily shown.

[0178] Referring to Figure 9b Exemplarily, the source timbre feature information of the source audio data 3 is taken as an example to search for the matching target timbre feature information from the M groups of reference timbre feature information.

[0179] Continuing to refer to Figure 9b Exemplarily, the source timbre feature information of the source audio data 3 is timbre feature information A. The distance information 1 between the timbre feature information A and the reference timbre feature information 1 can be calculated, the distance information 2 between the timbre feature information A and the reference timbre feature information 2 can be calculated,..., and the distance information M between the timbre feature information A and the reference timbre feature information M can be calculated. Then, the sizes of the distance information 1, the distance information 2,..., and the distance information M can be compared to determine the smallest distance information. Assuming that the distance information 2 is the smallest distance information, it can be determined that the reference timbre feature information 2 matches the timbre feature information A, that is, the reference timbre feature information 2 is the target timbre feature information matching the source timbre feature information of the source audio data.

[0180] S303, the character timbre conversion module outputs the multicast audio data of the multicast audiobook.

[0181] Exemplarily, the character timbre conversion module can use the trained prosody extractor, emotion extractor, content extractor, and spectrum reconstruction model to convert the unicast audiobook into a multicast audiobook based on the target timbre feature information of each source audio data, that is, to convert the unicast audio data into multicast audio data.

[0182] Exemplarily, each of the N pieces of source audio data of the unicast audio data can be sequentially input to the trained prosody extractor, and the trained prosody extractor can extract N groups of prosody feature information corresponding to the N pieces of source audio data.

[0183] Exemplarily, each of the N pieces of source audio data of the unicast audio data can be sequentially input to the trained emotion extractor, and the trained emotion extractor can extract N groups of emotion feature information corresponding to the N pieces of source audio data.

[0184] Exemplarily, each of the N pieces of source audio data of the unicast audio data can be sequentially input to the trained content extractor, and N sets of content feature information corresponding to the N pieces of source audio data can be extracted by the trained content extractor.

[0185] The following exemplarily takes the i-th piece of source audio data in the N pieces of source audio data as an example to illustrate how to convert the role timbre of the source audio data.

[0186] Exemplarily, the emotion feature information extracted by the trained emotion extractor for the i-th piece of source audio data, the prosody feature information extracted by the trained prosody extractor for the i-th piece of source audio data, the content feature information extracted by the trained content extractor for the i-th piece of source audio data, and the matching target timbre feature information corresponding to the i-th piece of source audio data can be input to the trained spectrum reconstruction model.

[0187] Exemplarily, the trained spectrum reconstruction model can perform spectrum reconstruction based on the emotion feature information, the prosody feature information, the content feature information, and the target timbre feature information of the i-th piece of source audio data, to obtain and output the i-th set of audio spectrum feature information. Then, the i-th set of audio spectrum feature information can be converted in the time domain to obtain the i-th piece of target audio data after timbre conversion.

[0188] Figure 10 The timbre conversion schematic diagram exemplarily shown.

[0189] Referring to Figure 10 Exemplarily, the source audio data 3 can be input to the prosody extractor, the emotion extractor, and the content extractor respectively, to obtain the prosody feature information 3 output by the prosody extractor, the emotion feature information 3 output by the emotion extractor, and the content feature information 3 output by the content extractor. Exemplarily, the prosody feature information 3, the emotion feature information 3, the content feature information 3, and the target timbre feature information corresponding to the source audio data, i.e., the reference timbre feature information 2, can be input to the spectrum reconstruction model. The spectrum reconstruction model can perform spectrum reconstruction based on the prosody feature information 3, the emotion feature information 3, the content feature information 3, and the reference timbre feature information 2, to output the audio spectrum feature information 3. Then, the audio spectrum feature information 3 can be converted in the frequency-time domain to obtain the target audio data 3. The target audio data 3 is the audio data after timbre conversion of the source audio data 3.

[0190] Exemplarily, the trained spectrum reconstruction model can perform spectrum reconstruction based on the emotion feature information, the prosody feature information, the content feature information, and the target timbre feature information of the i-th piece of source audio data, to obtain the i-th set of audio spectrum feature information. Then, the i-th set of audio spectrum feature information can be converted in the time domain to obtain and output the i-th piece of target audio data after timbre conversion.

[0191] Exemplarily, in the above manner, N pieces of target audio data can be obtained, and then the N pieces of target audio data can be spliced to obtain the multicast audio data, that is, the audio data of the multicast audiobook.

[0192] Since each piece of source audio data in the unicast audio data has high expressiveness (including prosody expressiveness and emotion expressiveness), the spectrum reconstruction is performed by extracting the emotion feature information and the prosody feature information of the source audio data, and combining the corresponding matched target timbre feature information of the source audio data, to generate the multicast audio data; the role timbre of the source audio data can be converted under the premise of guaranteeing the emotion expressiveness of the source audio data in the unicast audio data, so as to realize the conversion of the unicast audio data into the multicast audio data. Further, the timbre distinguishability of the roles in the audiobook can be improved, so as to improve the scene performance degree of the audiobook, and facilitate the user to understand the plot. In addition,

[0193] On the one hand, from the perspective of the user, compared with the prior art in which the user needs to manually select the matched timbre for each source audio data, the present application can automatically match the timbre of each source audio data, can improve the efficiency of converting the unicast audio data into the multicast audio data, simplify the user operation, and improve the user experience.

[0194] On the other hand, from the perspective of the audiobook producer, compared with the prior art in which multiple voice actors need to be assigned to record according to the roles to realize the production of the multicast audiobook, the present application only needs a single voice actor to record the unicast audiobook through pseudo-voice multi-role performance, and then converts the unicast audiobook into the multicast audiobook to realize the production of the multicast audiobook, which can reduce the production time and cost of the multicast audiobook, and improve the production efficiency of the multicast audiobook.

[0195] On the other hand, from the perspective of the audiobook producer, compared with the prior art in which multiple voice actors need to be assigned to record according to the roles to realize the production of the multicast audiobook, the present application only needs a single voice actor to record the unicast audiobook through pseudo-voice multi-role performance, and then converts the unicast audiobook into the multicast audiobook to realize the production of the multicast audiobook, which can reduce the production time and cost of the multicast audiobook, and improve the production efficiency of the multicast audiobook.

[0196] Exemplarily, the unicast audio data can be divided into sentences in combination with the audiobook text of the unicast audiobook, so as to improve the accuracy of sentence division, and further improve the accuracy of role analysis of the source audio data in the unicast audio data, thereby improving the accuracy of converting the unicast audiobook into the multicast audiobook.

[0197] Figure 11 The processing process is exemplarily shown in the schematic diagram.

[0198] Reference is made to Figure 11For example, the electronic device can include a character voice timbre analysis module and a character voice timbre conversion module. It should be understood that Figure 11 The electronic device shown is only an example of an electronic device, and the electronic device can have more or fewer modules than those shown in the figure, which is not limited in the present application.

[0199] For example, the functions of the character voice timbre analysis module and the character voice timbre conversion module can be referred to the description above, which will not be repeated here.

[0200] The process of converting a unicast audiobook into a multicast audiobook can be as follows:

[0201] S1101, input the unicast audio data and the book text of the unicast audiobook into the character voice timbre analysis module.

[0202] In one possible scenario, when the user clicks on the audiobook reading option 102 in the Figure 1 , the unicast audio data and the book text of the unicast audiobook corresponding to the audiobook reading option 102 can be obtained in response to the user operation behavior, and then the obtained unicast audio data and book text are input into the character voice timbre analysis module.

[0203] In one possible scenario, when the user clicks on the audiobook reading option 102 in the Figure 2 , the unicast audio data and the book text of the unicast audiobook corresponding to the audiobook reading option 102 can be obtained in response to the user operation behavior, and then the obtained unicast audio data and book text are input into the character voice timbre analysis module.

[0204] In one possible scenario, when the user clicks on the audiobook reading option 102 in the , the unicast audio data and the book text of the unicast audiobook corresponding to the audiobook reading option 102 can be obtained in response to the user operation behavior, and then the obtained unicast audio data and book text are input into the character voice timbre analysis module.

[0205] S1102, the character voice timbre analysis module outputs N pieces of source audio data in the unicast audio data and the corresponding matched N sets of target voice timbre feature information.

[0206] Figure 12 For example, the information extraction process is shown in the figure.

[0207] Referring to Figure 12 , for example, the book text can be divided into N pieces of text sentences first; then the unifrequency audio data is divided into N pieces of source audio data in combination with the N pieces of text sentences.

[0208] For example, to improve the efficiency of segmenting text into sentences, the text can first be divided into chapters, resulting in W segments of chapter text. Alternatively, text analysis can be performed on the text, dividing it into W segments based on chapter distinctions (such as "Chapter *", "chapter*", etc.). For instance, text segmentation can be performed before or after chapter distinction.

[0209] For example, after obtaining W sections of chapter text, character name recognition can be performed on each section to identify the characters contained within that section. Then, character dialogue segmentation is performed on that section to obtain multiple dialogue texts, and the characters corresponding to each dialogue are analyzed. Next, each dialogue text is divided into sentences to obtain R text sentences, and the character of the dialogue text to which each text sentence belongs is identified as the character corresponding to that text sentence.

[0210] For example, the text of the reading material can be stored in chapters. In this case, each time a section of text corresponding to a chapter can be directly obtained, and then the text of that chapter can be divided into text statements to obtain R text statements.

[0211] For example, text statements with a known role can be identified using the role's name. Text statements with an unknown role can be identified using an unknown role tag (such as "Unknown").

[0212] For example, a chapter text is:

[0213] As usual, Zhang asked, "Are you Wang?"

[0214] The plate reluctantly and slowly moved to the side of "Yes," then rotated the pointer on the plate to point at that word.

[0215] "age?"

[0216] "Twenty-three."

[0217] The text of this section is divided into sentences, and the resulting sentences and their corresponding character names are shown in Table 1:

[0218] Table 1

[0219]

[0220] Referring to Table 1, for example, the above chapter text is divided into 5 text statements, and each text statement is labeled with the corresponding character's name. Among them, "Unknown" is the tag for unknown characters.

[0221] For example, the text statement "Zhang** asked as usual:" corresponds to the role of "Narrator". The text statement "Are you Wang**?" corresponds to the role of "Zhang**". The text statement "The plate reluctantly moved to the side of 'Yes', rotated the pointer on the plate, and pointed it at that word." corresponds to an unknown role. The text statement "Age?" corresponds to an unknown role. The text statement "Twenty-three." corresponds to the role of "Wang**".

[0222] For example, following the above method, the text of section W is divided into sections, thereby dividing the reading material into N text statements.

[0223] For example, the unicast audio data can first be divided into audio chapters to obtain multiple chapter audio data segments, as described above, and will not be repeated here. Then, for each chapter, the chapter audio data can be divided into audio sentences based on the W text sentences corresponding to that chapter. This application uses a single chapter as an example to illustrate how to divide chapter audio data into multiple source audio data segments.

[0224] For example, text recognition can be performed on the chapter audio data to obtain the text recognition result. Then, based on the text recognition result, the chapter audio data and chapter text are aligned on the timeline, and then the W audio time intervals corresponding to the W text statements in the chapter audio data are determined. Based on the W audio time intervals, the unicast audio data is divided into W source audio data. And based on the role of the text statement corresponding to each source audio data, the role of each source audio data is determined. For example, see Table 2:

[0225] Table 2

[0226]

[0227] For example, in Table 2, the source audio data corresponding to the text statement "Zhang** asked as usual" is the audio data within the time period from 0 to 0:02.582s. The audio data within the time period from 0 to 0:02.582s can be divided into one source audio data and the corresponding role can be determined as "narrator".

[0228] For example, in Table 2, the source audio data corresponding to the text statement "Are you Wang**?" is the audio data within the time period from 0:02.582 to 0:02.048s. The audio data within the time period from 0:02.582 to 0:02.048s can be divided into one source audio data and the corresponding role can be determined as "Zhang**".

[0229] For example, in Table 2, the source audio data corresponding to the text statement "The disc reluctantly moved to the side of 'Yes', rotated the pointer on the disc, and aimed at that word" is the audio data within the time period from 0:02.048 to 0:13.969s. The audio data within the time period from 0:02.048 to 0:13.969s can be divided into one source audio data and the corresponding role can be determined as an unknown role.

[0230] For example, in Table 2, the source audio data corresponding to the text statement "age" is the audio data within the time period from 0:13.969 to 0:14.818s. The audio data within the time period from 0:13.969 to 0:14.818s can be divided into one source audio data and the corresponding role can be determined as an unknown role.

[0231] For example, in Table 2, the source audio data corresponding to the text statement "twenty-three" is the audio data within the time period from 0:14.818 to 0:16.217s. The audio data within the time period from 0:14.818 to 0:16.217s can be divided into one source audio data and the corresponding role can be identified as "Wang**".

[0232] For example, following the above method, the audio data of section W is divided into audio sentences, thereby dividing the unicast audio data into N source audio data. For example, the N source audio data include P1 source audio data with a determined role (P1 is a positive integer) and P2 source audio data with an undetermined role (P2 is a positive integer), where N = P1 + P2.

[0233] One possible approach is to find X (where X is a positive integer) source audio data corresponding to the first character from P1 source audio data with identified characters (i.e., find multiple source audio data with the same character). Then, using a trained timbre extractor, X sets of initial timbre feature information corresponding to the X source audio data are extracted. A weighted calculation is then performed based on these X sets of initial timbre feature information to obtain the weighted calculation result. This weighted calculation result can then be used as the source timbre feature information corresponding to each of the X source audio data.

[0234] One possible approach is to determine the role corresponding to the P2 source audio data points for which the role is currently undetermined. For example, a trained timbre extractor can be used to extract N sets of initial timbre feature information corresponding to the N source audio data points. For the j-th (j is a positive integer) source audio data point in the P2 source audio data points, the similarity between the initial timbre feature information corresponding to the P1 source audio data point and the initial timbre feature information corresponding to the j-th source audio data point can be calculated. Then, the role of the source audio data point in the P1 source audio data point with the highest similarity to the initial timbre feature information corresponding to the j-th source audio data point is determined as the role of the j-th source audio data point. For example, after determining the role corresponding to the P2 source audio data points, X source audio data points corresponding to the first role can be found from the N source audio data points with determined roles. Then, a trained timbre extractor is used to extract X sets of initial timbre feature information corresponding to the X source audio data points, and a weighted calculation is performed based on the X sets of initial timbre feature information to obtain the weighted calculation result. The weighted calculation result is then used to determine the source timbre feature information corresponding to each source audio data in X source audio data.

[0235] For example, referring to Table 3, based on the roles of the source audio data of known roles, the roles of the source audio data of unknown roles are determined as follows:

[0236] Table 3

[0237]

[0238] For example, in Table 3, the timbre characteristics of the source audio data from 0:02.048 to 0:13.969s (corresponding to the text statement "The disc reluctantly moved to the side of 'Yes', rotated the pointer on the disc, and aimed at that word") have the highest similarity to the timbre characteristics of the source audio data from 0 to 0:02.582s (corresponding to the text statement "Zhang** asked as usual"). Therefore, it can be determined that the role of the source audio data from 0:02.048 to 0:13.969s is narration. The timbre characteristics of the source audio data from 0:13.969 to 0:14.818s (corresponding to the text phrase "age") have the highest similarity to the timbre characteristics of the source audio data from 0:02.582 to 0:02.048s (corresponding to the text phrase "Are you Wang**?"). Therefore, the character in the source audio data from 0:13.969 to 0:14.818s can be identified as Zhang**.

[0239] It should be noted that, as described above, after performing VAD detection on the unicast audio data and dividing it into N source audio data, speech recognition can be performed on each of the N source audio data to obtain the corresponding text recognition results. Then, the above method can be used to determine the role corresponding to the N source audio data based on the text recognition results. Using the above method again, first determine the X source audio data corresponding to the first role. By weighting the initial timbre feature information of the X source audio data, the source timbre feature information of each of the X source audio data is determined; this will not be elaborated further here.

[0240] For example, after determining the N sets of source timbre feature information corresponding to the N source audio data, the target timbre feature information that matches the N source audio data can be found from the M sets of reference timbre feature information based on the N sets of source timbre feature information. This can be referred to the description above and will not be repeated here.

[0241] It should be noted that, alternatively, it is not necessary to extract text chapter information from the reading material. Instead, the reading material can be directly divided into N text sentences, and then the unicast audio data and the reading material can be aligned to determine the N audio time intervals corresponding to the N text sentences in the unicast audio data. Based on the N audio time intervals, the unicast audio data can then be divided into N source audio data.

[0242] S1103, the character voice conversion module outputs multicast audio data for multicast audiobooks.

[0243] For example, S1103 can be referred to the description of S303 above, and will not be repeated here.

[0244] Figure 13a This is a schematic diagram of the structure of an exemplary voice processing device.

[0245] Reference Figure 13a For example, the voice processing device 1300 includes:

[0246] Data acquisition module 1301 is used to acquire unicast audio data of unicast audiobooks, wherein the unicast audio data includes N source audio data, where N is a positive integer;

[0247] The character timbre analysis module 1302 is used to determine N sets of target timbre feature information that match the N source audio data from M sets of reference timbre feature information. The M sets of reference timbre feature information correspond to M types of timbre. The timbre discrimination between any two sets of the M sets of reference timbre feature information is greater than the discrimination threshold. M is a positive integer.

[0248] The character timbre conversion module 1303 is used to acquire N sets of expressive feature information corresponding to the N sets of source audio data; and to perform timbre conversion on the N sets of source audio data based on the N sets of target timbre feature information and the N sets of expressive feature information to generate multicast audio data.

[0249] For example, the data acquisition module 1301 acquires unicast audio data of a unicast audiobook and then inputs the unicast audio data into the character timbre analysis module 1302. The character timbre analysis module 1302 can determine N sets of target timbre feature information that match the N source audio data from M sets of reference timbre feature information, and then input the N sets of target timbre feature information into the character timbre conversion module 1303. The character timbre conversion module 1303 can acquire N sets of expressive feature information corresponding to the N sets of source audio data; then, based on the N sets of target timbre feature information and the N sets of expressive feature information, it performs timbre conversion on the N source audio data respectively to generate multicast audio data. Wherein, the M sets of reference timbre feature information correspond to M types of timbres, and the timbre differentiation between any two sets of the M sets of reference timbre feature information is greater than a differentiation threshold. In this way, while ensuring the expressiveness of the unicast audiobook, the timbre of each source audio data can be converted, realizing the conversion of unicast audiobooks into multicast audiobooks, improving the timbre differentiation of characters in the audiobook, thereby enhancing the expressiveness of the audiobook's scene portrayal and making it easier for users to understand the plot.

[0250] Figure 13b This is a schematic diagram of the structure of an exemplary voice processing device.

[0251] Reference Figure 13b For example, the character voice analysis module 1302 includes:

[0252] The audio statement segmentation module 13021 is used to segment unicast audio data into audio statements to obtain N source audio data and N source audio data and audio statements;

[0253] The timbre feature extraction module 13022 is used to extract N sets of source timbre feature information corresponding to N source audio data using a timbre extractor;

[0254] The timbre feature matching module 13023 is used to determine the target timbre feature information from M sets of reference timbre feature information based on N sets of source timbre feature information.

[0255] For example, the timbre feature matching module 13023 is used to: determine the similarity between M sets of reference timbre feature information and the i-th source audio data in N source audio data; and determine the reference timbre feature information with the highest similarity to the i-th source audio data as the target timbre feature information to match the i-th source audio data, where i is a positive integer and its value ranges from 1 to N.

[0256] Reference Figure 13b For example, the character voice conversion module 1303 includes:

[0257] The prosodic feature extraction module 13031 is used to extract N sets of prosodic feature information corresponding to N sets of source audio data using a prosodic extractor.

[0258] The emotion feature extraction module 13032 is used to extract N sets of emotion feature information corresponding to N sets of source audio data using an emotion extractor;

[0259] The expressive feature generation module 13033 is used to generate N sets of expressive feature information corresponding to N source audio data based on N sets of prosodic feature information and N sets of emotional feature information.

[0260] Reference Figure 13b For example, the character voice conversion module 1303 includes:

[0261] The content feature extraction module 13034 is used to extract N sets of content feature information corresponding to N source audio data using a content extractor.

[0262] The feature spectrum reconstruction module 13035 is used to perform spectrum reconstruction based on N sets of target timbre feature information, N sets of expressive feature information and N sets of content feature information using a spectrum reconstruction model to obtain N sets of audio spectrum feature information.

[0263] The frequency-time transformation module 13036 is used to perform frequency-time transformation on N sets of audio spectrum feature information respectively to obtain N sets of target audio data;

[0264] The splicing module 13037 is used to splice N sets of target audio data to obtain multicast audio data.

[0265] For example, the audio sentence segmentation module 13021 is used to segment unicast audio data into N source audio data by performing voice activity detection (VAD) on the unicast audio data.

[0266] For example, the audio sentence segmentation module 13021 is used to obtain the reading text of a unicast audiobook, divide the reading text into N text sentences; align the unicast audio data and the reading text to determine the N audio time intervals in the unicast audio data corresponding to the N text sentences; and divide the unicast audio data into N source audio data based on the N audio time intervals.

[0267] For example, the timbre feature extraction module 13022 is used to determine the role corresponding to each source audio data in N source audio data; obtain N sets of initial timbre feature information corresponding to the N source audio data; for X source audio data corresponding to the first role in the N source audio data: extract X sets of initial timbre feature information corresponding to the X source audio data; perform weighted calculation based on the X sets of initial timbre feature information corresponding to the X source audio data, and determine the result of the weighted calculation as the source timbre feature information corresponding to each source audio data in the X source audio data, where X is a positive integer.

[0268] Reference Figure 13b For example, the voice processing device 1300 also includes:

[0269] The role determination module 1304 is used to calculate the similarity between the timbre feature information corresponding to P1 source audio data and the initial timbre feature information corresponding to the j-th source audio data; and to determine the role corresponding to the source audio data with the highest similarity to the initial timbre feature information of the j-th source audio data in P1 source audio data as the role of the j-th source audio data; where j is a positive integer, and the value range is between 1 and P2.

[0270] Figure 14 This is a schematic diagram of the structure of a training device as an example.

[0271] Reference Figure 14 For example, the training device 1400 includes:

[0272] The collection module 1401 is used to collect training data. The training data includes training audio data and reference character labels corresponding to the training audio data. The expressive feature information of the training audio data meets the expressive conditions. The training audio data includes audio data recorded by multiple users using their own timbre, and / or audio data recorded by multiple users using pseudo-voices. The timbre differentiation of different pseudo-voices used by the same user is greater than the differentiation threshold.

[0273] The feature information extraction module 1402 is used to input the training audio data into the emotion extractor, content extractor and prosody extractor respectively for calculation, so as to obtain the prosodic feature information output by the emotion extractor, the content feature information output by the content extractor and the prosodic feature information output by the prosody extractor; and to input the training audio data and reference character labels into the timbre extractor for calculation, so as to obtain the timbre feature information output by the timbre extractor.

[0274] The audio data reconstruction module 1403 is used to input prosodic feature information, content feature information, prosodic feature information and timbre feature information into the spectrum reconstruction model to perform spectrum reconstruction, so as to obtain audio spectrum feature information, and to perform frequency-time transformation on the audio spectrum feature information to obtain reconstructed audio data.

[0275] The backpropagation module 1404 is used to calculate the first loss function value based on the reconstructed audio data and the training audio data, with the goal of minimizing the first loss function value, and jointly adjust the model parameters of the emotion extractor, content extractor, prosody extractor, timbre extractor and spectrum reconstruction model.

[0276] For example, the device also includes:

[0277] The loss function value calculation module 1405 is used to input timbre feature information into the first classifier for calculation to obtain the first role label, and to calculate the second loss function value based on the first role label and the reference role label; input emotion feature information into the second classifier for calculation to obtain the second role label, and to calculate the second loss function value based on the second role label and the reference role label;

[0278] The timbre extractor training module 1406 is used to adjust the model parameters of the timbre extractor with the goal of minimizing the value of the second loss function and the mutual information between timbre feature information and emotional feature information;

[0279] The sentiment extractor training module 1407 is used to adjust the model parameters of the sentiment extractor with the goal of maximizing the value of the third loss function and minimizing mutual information.

[0280] In one example, Figure 15 A schematic block diagram illustrating an embodiment of the present application shows an apparatus 1500. The apparatus 1500 may include a processor 1501 and a transceiver / transceiver pin 1502, and optionally, a memory 1503.

[0281] The various components of device 1500 are coupled together via bus 1504, which includes a data bus, a power bus, a control bus, and a status signal bus. However, for clarity, all buses are referred to as bus 1504 in the figure.

[0282] Optionally, the memory 1503 can be used for the instructions in the foregoing method embodiments. The processor 1501 can be used to execute the instructions in the memory 1503, control the receive pin to receive signals, and control the transmit pin to transmit signals.

[0283] Device 1500 may be an electronic device or a chip of an electronic device in the above method embodiments.

[0284] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0285] This embodiment also provides a computer storage medium storing computer instructions. When the computer instructions are executed on an electronic device, the electronic device performs the aforementioned method steps to implement the speech processing and training methods in the above embodiments.

[0286] This embodiment also provides a computer program product that, when run on a computer, causes the computer to perform the aforementioned steps to implement the speech processing and training methods described in the above embodiments.

[0287] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component or module. The apparatus may include a connected processor and a memory. The memory is used to store computer execution instructions. When the apparatus is running, the processor can execute the computer execution instructions stored in the memory to cause the chip to execute the speech processing and training methods in the above-described method embodiments.

[0288] In this embodiment, the electronic device, computer storage medium, computer program product or chip are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding method provided above, and will not be repeated here.

[0289] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0290] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0291] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0292] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0293] Any content in the various embodiments of this application, as well as any content in the same embodiment, can be freely combined. Any combination of the above content is within the scope of this application.

[0294] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0295] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

[0296] The steps of the methods or algorithms described in conjunction with the embodiments of this application can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can reside in an ASIC.

[0297] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0298] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A voice processing method, characterized by, The method comprises: obtaining unicast audio data of unicast audiobook, the unicast audio data comprising N pieces of source audio data, N being a positive integer; determining, from M sets of reference timbre feature information, N sets of target timbre feature information matched with the N pieces of source audio data, the M sets of reference timbre feature information corresponding to M kinds of timbres, a timbre distinction degree between any two of the M sets of reference timbre feature information being greater than a distinction degree threshold, M being a positive integer; obtaining N sets of expressiveness feature information corresponding to the N pieces of source audio data; performing, based on the N sets of target timbre feature information and the N sets of expressiveness feature information, timbre conversion on the N pieces of source audio data respectively to generate multicast audio data.

2. The method of claim 1, wherein, The method comprises: performing audio sentence division on the unicast audio data to obtain the N pieces of source audio data, the N pieces of source audio data corresponding to audio sentences one by one; obtaining N sets of source timbre feature information corresponding to the N pieces of source audio data; determining, based on the N sets of source timbre feature information, the N sets of target timbre feature information from the M sets of reference timbre feature information.

3. The method of claim 2, wherein, The method comprises: for the i-th piece of source audio data in the N pieces of source audio data, i being a positive integer and taking a value in a range between 1 and N: determining similarities of the M sets of reference timbre feature information with the i-th piece of source audio data respectively; determining, as target timbre feature information matched with the i-th piece of source audio data, reference timbre feature information having the highest similarity with the i-th piece of source audio data.

4. The method according to any one of claims 1 to 3, characterized in that, The method comprises: obtaining N sets of prosody feature information corresponding to the N pieces of source audio data; obtaining N sets of emotion feature information corresponding to the N pieces of source audio data; generating, based on the N sets of prosody feature information and the N sets of emotion feature information, the N sets of expressiveness feature information corresponding to the N pieces of source audio data.

5. The method according to any one of claims 1 to 3, characterized in that, The method comprises: obtaining N sets of content feature information corresponding to the N pieces of source audio data; performing, based on the N sets of target timbre feature information, the N sets of expressiveness feature information and the N sets of content feature information, spectrum reconstruction to obtain N sets of audio spectrum feature information; performing frequency-time conversion on the N sets of audio spectrum feature information respectively to obtain N sets of target audio data; splicing the N sets of target audio data to obtain the multicast audio data.

6. The method of claim 2, wherein, The method comprises: dividing, by voice activity detection (VAD) on the unicast audio data, the unicast audio data into the N pieces of source audio data.

7. The method of claim 2, wherein, The audio sentence division is performed on the unicast audio data to obtain N pieces of source audio data, including: obtaining a reading text of the unicast audiobook, and dividing the reading text into N pieces of text sentences; aligning the unicast audio data and the reading text to determine N audio time intervals in the unicast audio data corresponding to the N pieces of text sentences; based on the N audio time intervals, dividing the unicast audio data into the N pieces of source audio data.

8. The method according to any one of claims 2, 3, 6, 7, characterized in that, The N sets of source timbre feature information corresponding to the N pieces of source audio data includes: determining the roles corresponding to each piece of source audio data in the N pieces of source audio data; obtaining N sets of initial timbre feature information corresponding to the N pieces of source audio data; for X pieces of source audio data corresponding to the first role in the N pieces of source audio data: based on the X sets of initial timbre feature information corresponding to the X pieces of source audio data, performing weighted calculation, and determining the result of the weighted calculation as the source timbre feature information corresponding to each piece of source audio data in the X pieces of source audio data, wherein X is a positive integer.

9. The method of claim 8, wherein, The N pieces of source audio data include P1 pieces of source audio data of the determined role and P2 pieces of source audio data of the undetermined role, N=P1+P2, P1 and P2 are positive integers, and the method further includes: for the jth piece of source audio data in the P2 pieces of source audio data: respectively calculate the similarity of the initial timbre feature information corresponding to P1 pieces of source audio data and the initial timbre feature information corresponding to the jth piece of source audio data; determining the role corresponding to the source audio data with the highest similarity of initial timbre feature information corresponding to the jth piece of source audio data among P1 pieces of source audio data as the role of the jth piece of source audio data; wherein j is a positive integer, and the value range is between 1 and P2.

10. A training method characterized by comprising: including: collecting training data, the training data including training audio data and reference role labels corresponding to the training audio data, the expressiveness feature information of the training audio data satisfying an expressiveness condition, the training audio data including audio data recorded by multiple users using their own timbre and / or audio data recorded by multiple users using pseudo-voice, and the timbre distinction degree of different pseudo-voices used by the same user is greater than a distinction degree threshold; inputting the training audio data into the emotion extractor, the content extractor and the prosody extractor for calculation to obtain the emotion feature information output by the emotion extractor, the content feature information output by the content extractor and the prosody feature information output by the prosody extractor; and inputting the training audio data and the reference role label into the timbre extractor for calculation to obtain the timbre feature information output by the timbre extractor; inputting the prosody feature information, the content feature information, the prosody feature information and the timbre feature information into a spectrum reconstruction model for spectrum reconstruction to obtain audio spectrum feature information, and performing frequency-time conversion on the audio spectrum feature information to obtain reconstructed audio data; A first loss function value is calculated based on the reconstructed audio data and the training audio data, and model parameters of the emotion extractor, the content extractor, the prosody extractor, the timbre extractor, and the spectral reconstruction model are jointly adjusted to minimize the first loss function value.

11. The method of claim 10, wherein, The method further includes: The timbre feature information is input to a first classifier for calculation to obtain a first role label, and a second loss function value is calculated based on the first role label and the reference role label; The emotion feature information is input to a second classifier for calculation to obtain a second role label, and a third loss function value is calculated based on the second role label and the reference role label; The model parameters of the timbre extractor are adjusted to minimize the second loss function value and mutual information of the timbre feature information and the emotion feature information. The model parameters of the emotion extractor are adjusted to maximize the third loss function value and minimize the mutual information.

12. An electronic device, comprising: Comprise: a memory and a processor, the memory being coupled to the processor; The memory stores program instructions, when the program instructions are executed by the processor, the electronic device executes the speech processing method in any one of claims 1-9.

13. An electronic device, comprising: Comprise: a memory and a processor, the memory being coupled to the processor; The memory stores program instructions, when the program instructions are executed by the processor, the electronic device executes the training method in any one of claims 10-11.

14. A chip, characterized by Comprise one or more interface circuits and one or more processors; the interface circuit is used to receive signals from the memory of the electronic device, and send the signals to the processor, the signals include computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device executes the speech processing method in any one of claims 1-9.

15. A chip, characterized by Comprise one or more interface circuits and one or more processors; the interface circuit is used to receive signals from the memory of the electronic device, and send the signals to the processor, the signals include computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device executes the training method in any one of claims 10-11.

16. A computer storage medium, comprising, The computer readable storage medium stores a computer program, when the computer program runs on a computer or a processor, the computer or the processor executes the method as claimed in any one of claims 1-11.

17. A computer program product, characterised in that, The computer program product contains a software program, when the software program is executed by a computer or a processor, the steps of the method as claimed in any one of claims 1-11 are executed.

Citation Information

Patent Citations

  • Speech synthesis method and related equipment

    CN108962217A

  • Speech synthesis method and device of text, electronic equipment and storage medium

    CN112908292A