A sound processing method and apparatus
By acquiring user input information and voice parameters, generating implicit feature vectors through feature mapping, and combining them with acoustic models to output acoustic features, the problem of poor voice customization for virtual avatars is solved. This achieves multi-dimensional voice avatar synthesis, enhancing the diversity and fun of virtual avatars.
Patent Information
- Application Number
- CN202111433826.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-29
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2041-11-29
AI Technical Summary
In existing technologies, the voice customization of virtual avatars is limited to the appearance image of the character and cannot customize the user's voice, resulting in poor sound processing effects and failing to meet the user's personalized needs.
By acquiring user input information, implicit sound parameters, and explicit sound parameters, feature mapping is performed to generate implicit feature vectors. Combined with an acoustic model, acoustic features are output to achieve multi-dimensional sound image synthesis.
It enriches the sound processing effects, meets users' needs for precise and expressive virtual voice avatars, and enhances the diversity and fun of virtual avatars.
Smart Images

Figure CN116189660B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio technology, and more particularly to a sound processing method and apparatus. Background Technology
[0002] With the development of computer and network technologies and the growing awareness of user avatar creation, more and more users are showcasing and expressing themselves online. For example, in role-playing games, users are keen to create their own virtual appearance by adjusting facial features, skin tone, hairstyle, clothing, etc., and then use these customized virtual avatars to communicate with other users in the game. However, this kind of personalized virtual avatar is limited to the character's appearance and cannot be customized in terms of the user's voice.
[0003] There is currently a sound processing method where users need to upload a small amount of corpus, which is then used to train a model. Once the model is trained, it can be used to output sound effects corresponding to the corpus uploaded by the user.
[0004] The above solution suffers from poor sound processing quality. Summary of the Invention
[0005] This application provides a sound processing method and apparatus to improve the effect of sound processing.
[0006] To address the aforementioned technical problems, this application provides the following technical solutions:
[0007] In a first aspect, embodiments of this application provide a sound processing method, including:
[0008] Acquire sound parameter information, which includes: user input information, implicit sound parameters, and explicit sound parameters; wherein, the implicit sound parameters include at least one of the following: sound parameters of timbre dimension, sound parameters of emotion dimension, and sound parameters of style dimension; the explicit sound parameters include at least one of the following: sound parameters of speech rate dimension, sound parameters of energy dimension, and sound parameters of pitch dimension.
[0009] The implicit sound parameters are feature-mapped to obtain implicit feature vectors; wherein the implicit feature vectors include at least one of the following: timbre implicit feature vector, emotion implicit feature vector, and style implicit feature vector;
[0010] The user input information, the implicit feature vector, and the explicit sound parameters are input into the acoustic model, and the acoustic features are output through the acoustic model.
[0011] In the above-described scheme, the embodiments of this application can obtain user input information, explicit sound parameters, and implicit sound parameters. Since implicit sound parameters cannot be directly recognized by the acoustic model, they need to be mapped to implicit feature vectors. The acoustic model is input with user input information, implicit feature vectors, and explicit sound parameters. Therefore, the acoustic model can output acoustic features through sound parameters of multiple dimensions, making the sound processing effect richer.
[0012] In the above scheme, the implicit sound parameters included in the sound parameter information can include at least one of the following: sound parameters in the timbre dimension, sound parameters in the emotion dimension, and sound parameters in the style dimension. These multiple dimensions of implicit sound parameters can use default configurations, or users can set one or more of these dimensions. Users can adjust sound parameters in multiple dimensions such as emotion, style, and timbre to synthesize a virtual voice avatar, breaking down the voice avatar into quantifiable parameter expressions, thus meeting users' needs for precisely controllable and expressive virtual voice avatar synthesis.
[0013] In the above scheme, the implicit sound parameters included in the sound parameter information can include at least one of the following: sound parameters in the timbre dimension, sound parameters in the emotion dimension, and sound parameters in the style dimension. These multiple dimensions of implicit sound parameters can use default configurations, or users can set one or more of these dimensions. Users can adjust sound parameters in multiple dimensions such as emotion, style, and timbre to synthesize a virtual voice avatar, breaking down the voice avatar into quantifiable parameter expressions, thus meeting users' needs for precisely controllable and expressive virtual voice avatar synthesis.
[0014] In one possible implementation, before acquiring the sound parameter information, the method further includes:
[0015] Obtain a first training text and a first corpus, the first corpus including: first speech data corresponding to the first training text, wherein the first speech data is pre-configured with a corresponding first explicit feature vector and a first acoustic feature;
[0016] The first speech data is input into an initial acoustic model, and the first implicit feature vector is output through the initial acoustic model.
[0017] Obtain the implicit feature vector of the text corresponding to the first training text;
[0018] The text implicit feature vector and the first implicit feature vector are combined into a second implicit feature vector using the initial acoustic model.
[0019] The second implicit feature vector is predicted using the initial acoustic model to obtain the second explicit feature vector.
[0020] Align the second implicit feature vector with the second explicit feature vector to obtain the third implicit feature vector.
[0021] The second acoustic feature is obtained by predicting the second explicit feature vector and the third implicit feature vector using the initial acoustic model.
[0022] The loss is calculated based on the second explicit feature vector and the first explicit feature vector to obtain the first loss calculation result;
[0023] Loss calculation is performed based on the second acoustic feature and the first acoustic feature to obtain the second loss calculation result;
[0024] Based on the first loss calculation result and the second loss calculation result, determine whether to end the training of the initial acoustic model.
[0025] In the above scheme, the sound processing device can use an acoustic model for feature prediction. This acoustic model can be trained by the sound processing device using a machine learning algorithm. The training process of the acoustic model will be illustrated below. The sound processing device can provide training text and a corpus in advance, and then the sound processing device can train the initial acoustic model based on the training text and the corpus.
[0026] In one possible implementation, before acquiring the sound parameter information, the method further includes:
[0027] Obtain a second corpus, which includes: second speech data, which is pre-configured with a corresponding fourth explicit feature vector and a third acoustic feature;
[0028] The second speech data is input into the initial acoustic model, and the fourth implicit feature vector is output through the initial acoustic model.
[0029] Content features are extracted from the second speech data to obtain implicit content feature vectors;
[0030] The fourth implicit feature vector and the content implicit feature vector are combined into a fifth implicit feature vector using the initial acoustic model.
[0031] The fourth acoustic feature is obtained by predicting the fourth explicit feature vector and the fifth implicit feature vector using the initial acoustic model.
[0032] Loss calculation is performed based on the third acoustic feature and the fourth acoustic feature to obtain the third loss calculation result;
[0033] The training of the initial acoustic model is terminated based on the result of the third loss calculation.
[0034] In the above scheme, the sound processing device can use an acoustic model for feature prediction. This acoustic model can be trained by the sound processing device using a machine learning algorithm. The training process of the acoustic model will be illustrated below. The sound processing device can provide a corpus in advance, and then train the initial acoustic model based on the corpus.
[0035] In one possible implementation, obtaining the sound parameter information includes:
[0036] Obtain the implicit and explicit sound parameters selected by the user; or,
[0037] Obtain a sound parameter template, which includes preset implicit sound parameters and preset explicit sound parameters; obtain the implicit sound parameters and the explicit sound parameters according to the sound parameter template.
[0038] In the above scheme, both explicit and implicit sound parameters can be flexibly configured according to the user's needs, with the user selecting the specific explicit and implicit sound parameters. Not limited to this, in this embodiment, the user can also select implicit sound parameters, while explicit sound parameters are obtained through user input or through default configuration. Different sound parameters are decoupled from each other, allowing control over the specific sound image generated according to the user's needs, thus enabling the user to flexibly customize their own sound image according to their personal preferences. Alternatively, multiple sound parameter templates can be pre-set, each with one implicit and one explicit sound parameter. The user can select a sound parameter template according to their needs, and then obtain the corresponding implicit and explicit sound parameters for that template. This allows control over the specific sound image generated according to the user's needs, enabling the user to flexibly customize their own sound image according to their personal preferences.
[0039] In one possible implementation, obtaining the sound parameter information further includes:
[0040] The user input information is obtained by using the text data input by the user; or,
[0041] The user input information is obtained by using the user's voice data.
[0042] In the above scheme, user input information refers to the sound content input by the user into the sound processing device. The user can input text data into the sound processing device, and the sound processing device acquires this text data as user input information, which includes the text content input by the user. Alternatively, the user can also input speech data into the sound processing device, and the sound processing device acquires this speech data as user input information, which includes the speech content emitted by the user collected by the sound processing device. In this embodiment, the user can use either text input or speech input to control the specific voice image to be generated according to the user's needs, thereby allowing the user to flexibly customize their own voice image according to their personal preferences.
[0043] In one possible implementation, the implicit sound parameters include: sound parameters of a first dimension, wherein the first dimension includes one of the following: timbre dimension, emotion dimension, and style dimension;
[0044] The acquisition of sound parameter information further includes: when the sound parameters of the first dimension include multiple directions and preset amplitudes corresponding to the multiple directions respectively, acquiring a first direction selected by the user from the multiple directions and a first scale selected by the user; adjusting the preset amplitude corresponding to the first direction according to the first scale to obtain the adjusted amplitude corresponding to the first direction;
[0045] The method further includes:
[0046] The acoustic feature is amplitude transformed according to the adjusted amplitude corresponding to the first direction.
[0047] In the above scheme, the acoustic processing device also obtains the adjusted amplitude corresponding to the first direction. Then, it is necessary to perform amplitude transformation on the acoustic features according to the adjusted amplitude corresponding to the first direction, so as to meet the user's need to flexibly customize the sound image according to their own needs. By transforming the amplitude of the acoustic features, the user's need for precise, controllable and expressive virtual sound image synthesis is further met.
[0048] In one possible implementation, the implicit sound parameters include: sound parameters in the timbre dimension, and / or sound parameters in the style dimension;
[0049] The user input information includes: text information;
[0050] The step of performing feature mapping on the implicit sound parameters to obtain implicit feature vectors includes:
[0051] Feature mapping is performed on the sound parameters of the timbre dimension to obtain an implicit timbre feature vector, and / or feature mapping is performed on the sound parameters of the style dimension to obtain an implicit style feature vector; and,
[0052] Obtain the emoji icon corresponding to the text information, and obtain the implicit emotional feature vector corresponding to the emoji icon.
[0053] In the above scheme, if the timbre dimension sound parameters are obtained, feature mapping is required to obtain the timbre implicit feature vector. If the style dimension sound parameters are obtained, feature mapping is required to obtain the style implicit feature vector. If both dimensions of sound parameters are obtained simultaneously, implicit feature vectors of both dimensions can be output. When the user inputs text information, the emotion dimension sound parameters are not required. By analyzing the text information, the corresponding emoticon icon can be obtained, and then the emotion implicit feature vector corresponding to the emoticon icon can be obtained. In this embodiment, the emotion implicit feature vector can be mapped from the user-input text information, thereby determining the emotion the user needs to express, and meeting the user's need for precise, controllable, and expressive virtual voice image synthesis.
[0054] In one possible implementation, the method further includes:
[0055] The acoustic features are then restored to obtain the restored speech data.
[0056] In the above scheme, the sound processing device can reconstruct acoustic features to obtain reconstructed speech data. For example, the sound processing device can use a vocoder to reconstruct acoustic features. In the embodiments of this application, the reconstructed speech data can be used to output a virtual voice avatar for the user, meeting the user's need for precise, controllable, and expressive virtual voice avatar synthesis.
[0057] In one possible implementation, the method further includes:
[0058] The restored voice data is added to one or more audio tracks;
[0059] The restored voice data is played through the audio track.
[0060] In the above scheme, the sound processing device can reconstruct acoustic features to obtain reconstructed speech data. As explained above regarding explicit and implicit sound parameters, the acoustic model can output acoustic features using sound parameters of multiple dimensions, resulting in richer sound processing effects. In this embodiment, through different user selections, speech data corresponding to various sound images can be output. The sound processing device can also add this speech data to an audio track, and play the reconstructed speech data through the audio track to complete the speech playback.
[0061] In one possible implementation, the method further includes:
[0062] The restored voice data is stored in the audio library.
[0063] In the above scheme, the sound processing device can reconstruct acoustic features to obtain reconstructed speech data. As explained above regarding explicit and implicit sound parameters, the acoustic model can output acoustic features using sound parameters of multiple dimensions, resulting in richer sound processing effects. In this embodiment, based on different user selections, speech data corresponding to various sound images can be output. The sound processing device can also store this speech data in a sound library, which is a database storing speech data. Through the sound library, speech data for various sound images can be output, thus enabling rapid output of speech data for multiple sound images according to user needs.
[0064] Secondly, embodiments of this application also provide a sound processing apparatus, including:
[0065] The acquisition module is used to acquire sound parameter information, which includes: user input information, implicit sound parameters, and explicit sound parameters; wherein, the implicit sound parameters include at least one of the following: sound parameters of timbre dimension, sound parameters of emotion dimension, and sound parameters of style dimension; the explicit sound parameters include at least one of the following: sound parameters of speech rate dimension, sound parameters of energy dimension, and sound parameters of pitch dimension.
[0066] The feature mapping module is used to perform feature mapping on the implicit sound parameters to obtain implicit feature vectors; wherein, the implicit feature vectors include at least one of the following: timbre implicit feature vector, emotion implicit feature vector, and style implicit feature vector;
[0067] The sound processing module is used to input the user input information, the implicit feature vector and the explicit sound parameters into the acoustic model, and output acoustic features through the acoustic model.
[0068] In some embodiments of this application, the implicit sound parameters include at least one of the following: sound parameters of timbre dimension, sound parameters of emotion dimension, and sound parameters of style dimension;
[0069] The implicit feature vector includes at least one of the following: timbre implicit feature vector, emotion implicit feature vector, and style implicit feature vector.
[0070] In some embodiments of this application, the explicit sound parameters include at least one of the following: sound parameters in the speech rate dimension, sound parameters in the energy dimension, and sound parameters in the pitch dimension.
[0071] In some embodiments of this application, the sound processing device further includes: a model training module, wherein,
[0072] The model training module is used to acquire a first training text and a first corpus before the acquisition module acquires the sound parameter information. The first corpus includes: first speech data corresponding to the first training text, wherein the first speech data is pre-configured with corresponding first explicit feature vectors and first acoustic features; inputting the first speech data into an initial acoustic model, and outputting a first implicit feature vector through the initial acoustic model; acquiring the text implicit feature vector corresponding to the first training text; combining the text implicit feature vector and the first implicit feature vector into a second implicit feature vector through the initial acoustic model; and applying the second implicit feature vector to the initial acoustic model. The initial acoustic model is used to predict the second explicit feature vector to obtain a second explicit feature vector; the second implicit feature vector is aligned with the second explicit feature vector to obtain a third implicit feature vector; the initial acoustic model is used to predict the second explicit feature vector and the third implicit feature vector to obtain a second acoustic feature; a loss is calculated between the second explicit feature vector and the first explicit feature vector to obtain a first loss calculation result; a loss is calculated between the second acoustic feature and the first acoustic feature to obtain a second loss calculation result; and the training of the initial acoustic model is terminated based on the first loss calculation result and the second loss calculation result.
[0073] In some embodiments of this application, the sound processing device further includes: a model training module, wherein,
[0074] The model training module is used to acquire a second corpus before the acquisition module acquires sound parameter information. The second corpus includes: second speech data, which is pre-configured with corresponding fourth explicit feature vectors and third acoustic features; inputting the second speech data into an initial acoustic model, and outputting a fourth implicit feature vector through the initial acoustic model; extracting content features from the second speech data to obtain a content implicit feature vector; combining the fourth implicit feature vector and the content implicit feature vector into a fifth implicit feature vector through the initial acoustic model; predicting the fourth explicit feature vector and the fifth implicit feature vector through the initial acoustic model to obtain a fourth acoustic feature; calculating a loss based on the third acoustic feature and the fourth acoustic feature to obtain a third loss calculation result; and determining whether to end the training of the initial acoustic model based on the third loss calculation result.
[0075] In some embodiments of this application, the acquisition module is specifically used to acquire the implicit sound parameters and the explicit sound parameters selected by the user; or, to acquire a sound parameter template, the sound parameter template including preset implicit sound parameters and preset explicit sound parameters; and to acquire the implicit sound parameters and the explicit sound parameters according to the sound parameter template.
[0076] In some embodiments of this application, the acquisition module is further configured to acquire the user input information through text data input by the user; or, acquire the user input information through voice data input by the user.
[0077] In some embodiments of this application, the implicit sound parameters include: sound parameters of a first dimension, wherein the first dimension includes one of the following: timbre dimension, emotion dimension, and style dimension;
[0078] The acquisition module is further configured to, when the sound parameters of the first dimension include multiple directions and preset amplitudes corresponding to the multiple directions respectively, acquire a first direction selected by the user from the multiple directions and a first scale selected by the user; adjust the preset amplitude corresponding to the first direction according to the first scale to obtain the adjusted amplitude corresponding to the first direction;
[0079] The sound processing module is further configured to perform amplitude transformation on the acoustic feature according to the adjusted amplitude corresponding to the first direction.
[0080] In some embodiments of this application, the implicit sound parameters include: sound parameters in the timbre dimension, and / or sound parameters in the style dimension;
[0081] The user input information includes: text information;
[0082] The feature mapping module is specifically used to perform feature mapping on the sound parameters of the timbre dimension to obtain a timbre implicit feature vector, and / or to perform feature mapping on the sound parameters of the style dimension to obtain a style implicit feature vector; and to obtain the emoticon icon corresponding to the text information and obtain the emotion implicit feature vector corresponding to the emoticon icon.
[0083] In some embodiments of this application, the sound processing module is further configured to restore the acoustic features to obtain restored speech data.
[0084] In some embodiments of this application, the sound processing module is further configured to add the restored speech data to one or more audio tracks; and play the restored speech data through the audio tracks.
[0085] In some embodiments of this application, the sound processing module is further configured to store the restored speech data in a sound library.
[0086] In the second aspect of this application, the constituent modules of the sound processing device may also perform the steps described in the first aspect and various possible implementations, as detailed in the foregoing description of the first aspect and various possible implementations.
[0087] Thirdly, embodiments of this application provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect above.
[0088] Fourthly, embodiments of this application provide a computer program product containing instructions that, when run on a computer, cause the computer to perform the method described in the first aspect above.
[0089] Fifthly, embodiments of this application provide a communication device, which may include entities such as terminal devices or chips. The communication device includes: a processor and a memory; the memory is used to store instructions; the processor is used to execute the instructions in the memory, causing the communication device to perform the method as described in any one of the first or second aspects above.
[0090] Sixthly, this application provides a chip system including a processor for supporting a sound processing device in implementing the functions involved in the foregoing aspects, such as transmitting or processing data and / or information involved in the foregoing methods. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for the sound processing device. This chip system may be composed of chips or may include chips and other discrete devices.
[0091] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:
[0092] In this embodiment, sound parameter information is first obtained, including user input information, implicit sound parameters, and explicit sound parameters. The implicit sound parameters are then feature-mapped to obtain implicit feature vectors. Finally, the user input information, implicit feature vectors, and explicit sound parameters are input into an acoustic model, which outputs acoustic features. This embodiment can obtain user input information, explicit sound parameters, and implicit sound parameters. Since implicit sound parameters cannot be directly recognized by the acoustic model, they need to be mapped to implicit feature vectors. The acoustic model is input with user input information, implicit feature vectors, and explicit sound parameters. Therefore, the acoustic model can output acoustic features through sound parameters of multiple dimensions, resulting in richer sound processing effects. Attached Figure Description
[0093] Figure 1 A flowchart illustrating a sound processing method provided in an embodiment of this application;
[0094] Figure 2 A schematic diagram of input sound parameters in an acoustic model provided in an embodiment of this application;
[0095] Figure 3 A schematic diagram of input sound parameters in an acoustic model provided in an embodiment of this application;
[0096] Figure 4 A schematic diagram of input sound parameters in an acoustic model provided in an embodiment of this application;
[0097] Figure 5 A schematic diagram of a system architecture for an application of a sound processing method provided in this application embodiment;
[0098] Figure 6 A schematic diagram of a system architecture for an application of a sound processing method provided in this application embodiment;
[0099] Figure 7 A schematic diagram illustrating the training process of an acoustic model provided in an embodiment of this application;
[0100] Figure 8 A schematic diagram of a user parameter selection interface provided in an embodiment of this application;
[0101] Figure 9 A vector space diagram illustrating an implicit feature vector of emotion provided in an embodiment of this application;
[0102] Figure 10A schematic diagram of the prediction process of an acoustic model provided in an embodiment of this application;
[0103] Figure 11 A schematic diagram of an audio comment interface provided in an embodiment of this application;
[0104] Figure 12 A schematic diagram illustrating the training process of an acoustic model provided in an embodiment of this application;
[0105] Figure 13 A schematic diagram of a user parameter selection interface provided in an embodiment of this application;
[0106] Figure 14 A schematic diagram of a voice-over tool interface provided in an embodiment of this application;
[0107] Figure 15 This is a schematic diagram of the composition structure of a sound processing device provided in an embodiment of this application;
[0108] Figure 16 This is a schematic diagram of the composition structure of a sound processing device provided in an embodiment of this application. Detailed Implementation
[0109] This application provides a sound processing method and apparatus to improve the effect of sound processing.
[0110] The embodiments of this application will now be described with reference to the accompanying drawings.
[0111] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0112] Currently, the customization of virtual voice avatars cannot achieve personalized customization for users with multiple different voices. Furthermore, current customization methods for virtual voice avatars do not allow for control and adjustment of multiple voice dimensions, resulting in poor sound processing effects and failing to meet users' needs for personalized virtual voice avatar customization.
[0113] In response to the growing demand from users to enhance the display and expression of their virtual avatars online, this application proposes a sound processing method that enables customized interaction of virtual voice avatars. By adjusting various sound parameters, a virtual voice avatar that the user expects can be synthesized, enriching the user's virtual avatar, enhancing self-expression in social networks and game systems, and improving the diversity and fun of the virtual avatar.
[0114] The sound processing method in this application embodiment can be implemented by a sound processing device. This sound processing device can be applied to various terminal devices that require sound processing, wireless devices that require transcoding, and core network devices. For example, the sound processing device can be a sound processing device for the aforementioned terminal devices, wireless devices, or core network devices. For instance, the sound processing device may include a wireless access network, a media gateway of the core network, a transcoding device, a media resource server, a mobile terminal, a fixed-line terminal, etc. The sound processing device can also be a sound processing device applied to virtual reality (VR) streaming services. This sound processing device can also be a device that provides users with customized virtual voice avatars. For example, the sound processing device provided in this application embodiment can provide users with an interactive voice avatar customization system that allows them to clearly and explicitly customize their voice avatar according to their personal preferences from multiple dimensions.
[0115] Next, we will first describe a sound processing method provided in an embodiment of this application. This method can be executed by a terminal device, such as a sound processing apparatus, which may be an audio encoding terminal. Figure 1 As shown, the main sound processing methods include the following:
[0116] 101. Obtain sound parameter information, which includes: user input information, implicit sound parameters, and explicit sound parameters.
[0117] In order to process the user's voice, the sound processing device first acquires sound parameter information, which consists of the sound parameters required for sound processing. This sound parameter information can include user input information, implicit sound parameters, and explicit sound parameters. Specifically, user input information refers to the sound content input by the user into the sound processing device. For example, the user input information may include text input by the user, or it may include speech content emitted by the user that the sound processing device has collected.
[0118] Furthermore, in the process of customizing a virtual voice avatar, the embodiments of this application can further refine the voice parameters. For example, voice parameters can be divided into implicit voice parameters and explicit voice parameters. Explicit voice parameters refer to parameters used to describe the basic characteristics of sound. The dimensions of explicit voice parameters describing sound are clearly known. For example, explicit voice parameters can describe sound from dimensions such as speech rate, energy, and pitch. Implicit voice parameters, in contrast to explicit voice parameters, refer to voice parameters other than those describing the basic characteristics of sound. Implicit voice parameters can be flexibly configured according to user needs, and the specific implementation method of implicit voice parameters is not limited. For example, implicit voice parameters can have various specific implementation methods; for instance, implicit voice parameters can describe sound from dimensions such as timbre, emotion, and style.
[0119] In some embodiments of this application, implicit sound parameters can also be selected by the user. In these embodiments, the sound can be adjusted according to the user's needs, thereby creating a sound image that perfectly matches the user's requirements. This solves the current problem of not being able to intuitively adjust sound parameters.
[0120] In some embodiments of this application, at least one of the explicit and implicit sound parameters can be flexibly configured according to the user's needs. Different sound parameters are decoupled from each other, and the specific sound image to be generated can be controlled according to the user's needs, so that the user can flexibly customize their own sound image according to their personal preferences.
[0121] In some embodiments of this application, implicit sound parameters include at least one of the following: sound parameters of timbre dimension, sound parameters of emotion dimension, and sound parameters of style dimension.
[0122] Implicit sound parameters can be refined into at least one, two, or three dimensions: timbre dimension, emotion dimension, and style dimension.
[0123] The timbre dimension of sound parameters can include parameters such as user gender, region, and dialect scale. Users can set specific values for these parameters according to their needs. In this embodiment, adjustments can be made at higher sound levels than explicit parameters, such as adjusting gender, region, and dialect scale, thereby customizing the timbre dimension sound parameters according to user requirements. For example, adjusting the gender dimension of a timbre refers to adjusting the user's voice to match the gender (e.g., male or female). Region can represent a dialect style; adjusting the region dimension of a timbre refers to adjusting its dialect style, such as changing Mandarin to Northeastern Mandarin. Dialect scale refers to the level of detail in the dialect style; for example, the dialect scale ranges from 0 to 1, where 0 represents Mandarin and 1 represents entirely dialect. Users can flexibly set this dialect scale to configure their desired voice image according to their individual needs.
[0124] The emotional dimension sound parameters can include various emotional dimensions such as anger, excitement, calmness, sadness, and joy. Users can set the specific values of the emotional dimension sound parameters according to their needs. In this embodiment, adjustments can be made on emotional dimensions other than explicit sound parameters. For example, anger, excitement, calmness, sadness, and joy can be adjusted to customize the emotional dimension sound parameters according to the user's needs. For example, if a user needs to express anger, the direction of anger is set in the emotional dimension, thereby reflecting the user's anger in the sound image. In this embodiment, users can flexibly set the emotional dimension sound parameters to configure their desired sound image according to their individual needs.
[0125] The voice parameters for the style dimension can include a variety of styles, such as handsome idol, gentle goddess, professional customer service, famous storyteller, classical poetry, cute child's voice, charming, and anime character imitation. Users can set the specific values of the voice parameters for the style dimension according to their needs. In this embodiment, adjustments can be made to style dimensions other than explicit voice parameters. For example, the style of the character the user wants to display can be adjusted, thus customizing the voice parameters for the style dimension according to the user's needs. For example, if a user wants to display a cute voice image, the "cute child's voice" direction in the style dimension can be set, so that the user's voice image is reflected as a cute child's voice in the voice image. In this embodiment, users can flexibly set the voice parameters for the style dimension to configure the voice image they want according to their individual needs.
[0126] In this embodiment, the implicit sound parameters included in the sound parameter information may include at least one of the following: sound parameters of timbre dimension, sound parameters of emotion dimension, and sound parameters of style dimension. These implicit sound parameters of various dimensions can use default configurations, or the user can set one or more of these dimensions. Users can adjust sound parameters of multiple dimensions such as emotion, style, and timbre to synthesize a virtual voice avatar, breaking down the voice avatar into quantifiable parameter expressions, thus meeting the user's need for precise, controllable, and expressive virtual voice avatar synthesis.
[0127] In some embodiments of this application, explicit sound parameters include at least one of the following: sound parameters in the speech rate dimension, sound parameters in the energy dimension, and sound parameters in the pitch dimension.
[0128] Among them, the speech rate dimension sound parameter can also be called the duration dimension sound parameter. The speech rate dimension sound parameter is used to adjust the speed of speech, the energy dimension sound parameter is used to adjust the energy level of the voice, and the pitch dimension sound parameter is used to adjust the pitch of the voice. The above-mentioned explicit sound parameters of various dimensions can use the default configuration, or the user can set one or more or all of the explicit sound parameters of the above dimensions. Users can adjust multiple dimensions of sound parameters such as speech rate, energy, and pitch to synthesize virtual voice avatars, breaking down the voice avatar into quantifiable parameter expressions, thus meeting the user's needs for precise, controllable, and expressive virtual voice avatar synthesis.
[0129] In this embodiment, the implicit sound parameters included in the sound parameter information may include at least one of the following: sound parameters of timbre dimension, sound parameters of emotion dimension, and sound parameters of style dimension. These implicit sound parameters of various dimensions can use default configurations, or the user can set one or more of these dimensions. Users can adjust sound parameters of multiple dimensions such as emotion, style, and timbre to synthesize a virtual voice avatar, breaking down the voice avatar into quantifiable parameter expressions, thus meeting the user's need for precise, controllable, and expressive virtual voice avatar synthesis.
[0130] In some embodiments of this application, step 101, obtaining sound parameter information, includes:
[0131] A1. Obtain the implicit and explicit sound parameters selected by the user; or,
[0132] A2. Obtain the sound parameter template, which includes preset implicit sound parameters and preset explicit sound parameters; obtain the implicit and explicit sound parameters based on the sound parameter template.
[0133] In step A1, both explicit and implicit sound parameters can be flexibly configured according to the user's needs, allowing the user to select specific explicit and implicit sound parameters. Not limited to this, in this embodiment, the user can also select implicit sound parameters, while explicit sound parameters are obtained through user input or through default configuration. The different sound parameters are decoupled from each other, allowing the user to control the specific sound image generated according to their needs, thus enabling the user to flexibly customize their own sound image according to their personal preferences.
[0134] In step A2, multiple sound parameter templates can be preset. Each sound parameter template can have one implicit sound parameter and one explicit sound parameter. Users can select a sound parameter template according to their needs, and then obtain the implicit and explicit sound parameters corresponding to the selected sound parameter template. Users can control the specific sound image to be generated according to their needs, so that users can flexibly customize their own sound image according to their personal preferences.
[0135] Furthermore, in some embodiments of this application, such as Figure 2 and Figure 3 As shown, in addition to A1 and A2, step 101 of obtaining sound parameter information also includes:
[0136] B1. Obtain user input information through text data input by the user; or,
[0137] B2. Obtain user input information through user-inputted voice data.
[0138] In this context, user input information refers to the audio content input by the user into the audio processing device. The user can input text data into the audio processing device, which then acquires this text data as user input information. This user input information includes the text content input by the user. Alternatively, the user can input speech data into the audio processing device, which then acquires this speech data as user input information. This user input information includes the speech content emitted by the user, collected by the audio processing device. In this embodiment, the user can use either text input or speech input to control the specific voice image generated according to their needs, allowing the user to flexibly customize their voice image according to their personal preferences.
[0139] 102. Perform feature mapping on the implicit sound parameters to obtain implicit feature vectors.
[0140] In this embodiment, implicit sound parameters cannot be directly recognized by the acoustic model. Therefore, feature mapping is required for these implicit sound parameters. Through feature mapping, the implicit sound parameters can be mapped into implicit feature vectors. This feature mapping enables the acoustic model to recognize these implicit feature vectors. It should be noted that the aforementioned feature mapping process refers to the dimensionality transformation of the implicit sound parameters, which can be mapped into implicit feature vectors. Furthermore, implicit feature vectors can also be called feature implicit vectors.
[0141] In some embodiments of this application, the implicit feature vector includes at least one of the following: timbre implicit feature vector, emotion implicit feature vector, and style implicit feature vector.
[0142] The implicit feature vector corresponds to the aforementioned implicit sound parameters. The implicit sound parameters include at least one of the following: sound parameters in the timbre dimension, sound parameters in the emotion dimension, and sound parameters in the style dimension. The timbre dimension sound parameters can be used to obtain the timbre implicit feature vector through feature mapping. The emotion dimension sound parameters can be used to obtain the emotion implicit feature vector through feature mapping. The style dimension sound parameters can be used to obtain the style implicit feature vector through feature mapping.
[0143] In this embodiment, the implicit feature vector includes at least one of the following: timbre implicit feature vector, emotion implicit feature vector, and style implicit feature vector. The above-mentioned implicit feature vectors of multiple dimensions are obtained by feature mapping of implicit sound parameters of multiple dimensions. Users can adjust implicit sound parameters of multiple dimensions such as emotion, style, and timbre, so as to determine the feature vectors of multiple dimensions such as emotion, style, and timbre according to the user's settings. These implicit feature vectors of multiple dimensions are used to synthesize virtual voice images, decompose the voice image into quantifiable parameter expressions, and meet the user's needs for precise, controllable, and expressive virtual voice image synthesis.
[0144] In some embodiments of this application, implicit sound parameters include: sound parameters in the timbre dimension and / or sound parameters in the style dimension; user input information includes: text information. The implicit sound parameters may not include sound parameters in the emotion dimension; the user inputs text information, through which the emotion the user wishes to express can be obtained.
[0145] Specifically, such as Figure 4 As shown, step 102 performs feature mapping on the implicit sound parameters to obtain implicit feature vectors, including:
[0146] C1. Perform feature mapping on the sound parameters of the timbre dimension to obtain an implicit timbre feature vector, and / or, perform feature mapping on the sound parameters of the style dimension to obtain an implicit style feature vector; and,
[0147] C2. Obtain the emoji icon corresponding to the text information, and obtain the implicit emotional feature vector corresponding to the emoji icon.
[0148] In C1, if the timbre dimension sound parameters are obtained, feature mapping is needed to obtain the timbre implicit feature vector. If the style dimension sound parameters are obtained, feature mapping is needed to obtain the style implicit feature vector. If both dimensions of sound parameters are obtained simultaneously, implicit feature vectors of both dimensions can be output. When the user inputs text information, the emotion dimension sound parameters are not required. By analyzing the text information, the corresponding emoticon icon can be obtained, and then the emotion implicit feature vector corresponding to the emoticon icon can be obtained. In this embodiment, the emotion implicit feature vector can be mapped from the user-input text information, thereby determining the emotion the user needs to express, and meeting the user's need for precise, controllable, and expressive virtual voice image synthesis.
[0149] For example, most scenarios in current social networks involve text communication, combined with various emoticons to express emotions. In this case, the implicit feature vector of emotions is mapped through emoticons, allowing users to make corresponding emotional audio comments, achieving the effect of "Why does your comment have sound?"
[0150] 103. Input user input information, implicit feature vectors, and explicit sound parameters into the acoustic model, and output acoustic features through the acoustic model.
[0151] In this process, an acoustic model is pre-trained to predict corresponding acoustic features. Specifically, the acoustic model can be obtained by training an initial model using a machine learning algorithm, such as a neural network model.
[0152] The sound processing device acquires user input information, implicit sound parameters, and explicit sound parameters, and inputs these into an acoustic model. Furthermore, following the aforementioned steps, the implicit sound parameters are converted into implicit feature vectors, which are then input into the acoustic model. The acoustic model can then make predictions based on the input information, and its output is an acoustic feature, also known as a speech feature. In this embodiment, the acoustic model is input with user input information, implicit feature vectors, and explicit sound parameters. Therefore, the acoustic model can output acoustic features using sound parameters from multiple dimensions, resulting in a richer sound processing effect.
[0153] In some embodiments of this application, step 103 inputs user input information, implicit feature vectors, and explicit sound parameters into the acoustic model. The method provided in the embodiments of this application further includes:
[0154] D1. Reconstruct the acoustic features to obtain the reconstructed speech data.
[0155] The sound processing device can reconstruct acoustic features to obtain reconstructed speech data. For example, the sound processing device can use a vocoder to reconstruct acoustic features. In this embodiment, the reconstructed speech data can be used to output a virtual voice avatar for the user, meeting the user's need for precise, controllable, and expressive virtual voice avatar synthesis.
[0156] In some embodiments of this application, the methods provided in the foregoing steps further include:
[0157] E1. Add the restored voice data to one or more audio tracks;
[0158] E2. Play the restored voice data through the audio track.
[0159] The sound processing device can reconstruct acoustic features to obtain reconstructed speech data. As explained above regarding explicit and implicit sound parameters, the acoustic model can output acoustic features using sound parameters of multiple dimensions, resulting in richer sound processing effects. In this embodiment, based on different user selections, speech data corresponding to various sound images can be output. The sound processing device can also add this speech data to an audio track, and play the reconstructed speech data through the audio track to complete the speech playback.
[0160] In practical applications, the sound processing device can be set to single-track or multi-track. The speech data of a specified voice image synthesized in the speech synthesis window is dragged into the audio track, and the preview window allows for a preview of the video dubbing effect. In this embodiment, the time and economic costs of finding voice actors are eliminated, and costs can be saved through virtual voice images. Because this embodiment allows users to freely adjust various dimensions of sound parameters such as timbre, emotion, and style according to their needs, users have many choices, improving the synthesized speech effect. This embodiment can synthesize speech that closely resembles real people and fits the video's plot.
[0161] In some embodiments of this application, the foregoing steps and methods further include:
[0162] F1 stores the restored voice data in the voice library.
[0163] The sound processing device can reconstruct acoustic features to obtain reconstructed speech data. As explained above regarding explicit and implicit sound parameters, the acoustic model can output acoustic features using sound parameters of multiple dimensions, resulting in richer sound processing effects. In this embodiment, based on different user selections, speech data corresponding to various sound images can be output. The sound processing device can also store this speech data in a sound library, which is a database storing speech data. Through the sound library, speech data for various sound images can be output, thus enabling rapid output of speech data for multiple sound images according to user needs.
[0164] In practical applications, the demand for voice libraries is increasing across various fields of speech synthesis technology. These libraries require diverse timbres, styles, and emotional nuances. However, recording human voice libraries is time-consuming, costly, and cumbersome. This application's embodiments utilize quantization functions for dialect scale, emotional scale, energy scale, and pitch scale to construct multi-scale voice libraries. Based on the requirements of the voice library, such as developing a voice library for a specific emotional scale for a particular voice image, it can be used to output emotion-related speech data.
[0165] In some embodiments of this application, implicit sound parameters include: sound parameters of a first dimension, which includes one of the following: timbre dimension, emotion dimension, and style dimension. The first dimension specifically refers to one of the timbre dimension, emotion dimension, and style dimension. The specific sound dimension indicated by the first dimension is not limited here.
[0166] Furthermore, the aforementioned step 101, obtaining sound parameter information, also includes:
[0167] G1. When the sound parameters of the first dimension include multiple directions and preset amplitudes corresponding to each direction, obtain the first direction selected by the user from the multiple directions and the first scale selected by the user; adjust the preset amplitude corresponding to the first direction according to the first scale to obtain the adjusted amplitude corresponding to the first direction.
[0168] In this system, the sound parameters of the first dimension are represented by direction and amplitude. Direction indicates the type of sound parameter corresponding to the first dimension, and amplitude indicates the value of the sound parameter corresponding to the first dimension. Users can select a specific direction and scale for the sound parameters of the first dimension. For example, if the first direction and first scale represent the sound parameters selected by the user, then the preset amplitude corresponding to the first direction needs to be adjusted to obtain the adjusted amplitude for the first direction. For instance, if the first dimension is the emotion dimension, and the user selects the joy direction as the first direction from multiple directions within the emotion dimension, with a preset amplitude of 1 for the joy direction, and if the first scale corresponding to the joy direction selected by the user is 0.5, then the current value of the preset amplitude corresponding to the joy direction needs to be adjusted from 1 to 0.5.
[0169] Furthermore, in performing the aforementioned step G1, the method provided in this application embodiment further includes:
[0170] G2. Perform amplitude transformation on the acoustic features according to the adjusted amplitude corresponding to the first direction.
[0171] In this process, the acoustic processing device performs the aforementioned step 103. The acoustic processing device also obtains the adjusted amplitude corresponding to the first direction. It is then necessary to perform amplitude transformation on the acoustic features according to the adjusted amplitude corresponding to the first direction, so as to meet the user's need to flexibly customize the sound image according to their own needs. By transforming the amplitude of the acoustic features, the user's need for precise, controllable, and expressive virtual sound image synthesis is further met.
[0172] In some embodiments of this application, the sound processing device can use an acoustic model for feature prediction. This acoustic model can be trained by the sound processing device using a machine learning algorithm. The training process of the acoustic model will be illustrated below. The sound processing device can pre-provide training text and a corpus, and then train the initial acoustic model based on the training text and the corpus. Specifically, before obtaining the sound parameter information in step 101, the method provided in this application embodiment further includes:
[0173] H1. Obtain the first training text and the first corpus. The first corpus includes: the first speech data corresponding to the first training text, wherein the first speech data is pre-configured with the corresponding first explicit feature vector and first acoustic feature.
[0174] Users can provide initial training text as training data and prepare a first corpus containing diverse vocal qualities, styles, and emotions. The initial training text corresponds to initial speech data in the first corpus, which includes explicit feature vectors and acoustic features. For example, the initial explicit feature vectors and acoustic features of the initial speech data can be pre-configured. The first acoustic features can be used to determine whether the training process of the acoustic model has converged. For example, acoustic features can be extracted from the initial speech data, including features such as duration, pitch, and energy. For example, the initial explicit feature vector can include explicit vectors of speech rate, energy, and pitch.
[0175] H2. Input the first speech data into the initial acoustic model, and output the first implicit feature vector through the initial acoustic model.
[0176] The initial acoustic model is pre-configured, and the first speech data in the first corpus is input into the acoustic model. The corresponding first implicit feature vector is extracted through the acoustic model. The first implicit feature vector may include: timbre implicit feature vector, emotion implicit feature vector and style implicit feature vector.
[0177] H3. Obtain the implicit feature vector of the text corresponding to the first training text.
[0178] The first training text is preprocessed and converted into a text embedding layer representation, and then converted into a text implicit feature vector by an encoder.
[0179] H4. The text implicit feature vector and the first implicit feature vector are combined into a second implicit feature vector through the initial acoustic model.
[0180] The second implicit feature vector carries information from both the text implicit feature vector and the first implicit feature vector.
[0181] H5. The second implicit feature vector is predicted using the initial acoustic model to obtain the second explicit feature vector.
[0182] After obtaining the second implicit feature vector, it is still necessary to make predictions through the acoustic model in order to obtain the second explicit feature vector that the model can recognize.
[0183] H6. Align the second implicit feature vector with the second explicit feature vector to obtain the third implicit feature vector.
[0184] Specifically, the second implicit feature vector is aligned with the second explicit feature vector so that the second implicit feature vector can be converted into the third implicit feature vector.
[0185] H7. The second acoustic feature is obtained by predicting the second explicit feature vector and the third implicit feature vector using the initial acoustic model.
[0186] The acoustic model takes a second explicit feature vector and a third implicit feature vector as input, makes predictions, and can output a second acoustic feature.
[0187] H8. Calculate the loss based on the second explicit feature vector and the first explicit feature vector to obtain the first loss calculation result.
[0188] Loss can be calculated by comparing the second explicit feature vector with the first explicit feature vector.
[0189] H9. Calculate the loss based on the second acoustic feature and the first acoustic feature to obtain the second loss calculation result.
[0190] Loss can be calculated by comparing the second acoustic feature with the first acoustic feature.
[0191] H10. Determine whether to end the training of the initial acoustic model based on the results of the first loss calculation and the second loss calculation.
[0192] The loss calculations in steps H8 and H9 yield a first loss calculation result and a second loss calculation result, which are then used to train the initial acoustic model.
[0193] It should be noted that if the initial acoustic model does not converge, steps H1 to H10 need to be executed repeatedly to complete the training of the acoustic model and obtain the acoustic model usable in step 103 of this application embodiment.
[0194] In some embodiments of this application, the sound processing device can use an acoustic model for feature prediction. This acoustic model can be trained by the sound processing device using a machine learning algorithm. The training process of the acoustic model will be illustrated below. The sound processing device can provide a second corpus in advance, and then train the initial acoustic model based on this corpus. Specifically, before obtaining the sound parameter information in step 101, the method provided in this application embodiment further includes:
[0195] I1. Obtain the second corpus, which includes: second speech data, which is pre-configured with corresponding fourth explicit feature vectors and third acoustic features.
[0196] This requires preparing a second corpus containing diverse vocal qualities, styles, and emotions. This second corpus includes second speech data, which is pre-configured with corresponding explicit feature vectors and acoustic features. For example, the second speech data may be pre-configured with corresponding fourth explicit feature vectors and third acoustic features. The third acoustic feature can be used to determine whether the acoustic model's training process has converged. For instance, extracting the third acoustic feature from the second speech data can include features such as duration, pitch, and energy. The fourth explicit feature vector can include explicit vectors for speech rate, energy, and pitch.
[0197] I2. Input the second speech data into the initial acoustic model, and output the fourth implicit feature vector through the initial acoustic model.
[0198] The initial acoustic model is pre-configured, and the second speech data from the second corpus is input into the acoustic model. The corresponding fourth implicit feature vector is extracted through the acoustic model. The fourth implicit feature vector may include: timbre implicit feature vector, emotion implicit feature vector, and style implicit feature vector.
[0199] I3. Extract content features from the second speech data to obtain implicit content feature vectors.
[0200] Unlike embodiments H2 and H3, there is no text input here. Instead, content feature extraction is performed on the second speech data to obtain implicit content feature vectors. For example, the second speech data is input into a content feature extraction model to extract frame-level implicit content feature vectors.
[0201] I4. The fourth implicit feature vector and the content implicit feature vector are combined into the fifth implicit feature vector using the initial acoustic model.
[0202] The fifth implicit feature vector carries information from the content implicit feature vector and information from the fourth implicit feature vector.
[0203] I5. The fourth acoustic feature is obtained by predicting the fourth explicit feature vector and the fifth implicit feature vector using the initial acoustic model.
[0204] The acoustic model takes a fourth explicit feature vector and a fifth implicit feature vector as input, makes predictions, and can output the fourth acoustic feature.
[0205] It should be noted that I5 differs from H5 to H7 in the aforementioned embodiments in that the content features here are at the frame level rather than the phoneme level. Therefore, there is no need for duration prediction and extended alignment. The fourth explicit feature vector and the fifth implicit feature vector are directly input into the acoustic model, and prediction can be performed through the acoustic model to obtain the fourth acoustic feature.
[0206] I6. Calculate the loss based on the third and fourth acoustic features to obtain the third loss calculation result.
[0207] Loss can be calculated by comparing the third and fourth acoustic features.
[0208] I7. Determine whether to end the training of the initial acoustic model based on the results of the third loss calculation.
[0209] The loss calculation in step I6 yields the third loss calculation result, which is then used to train the initial acoustic model.
[0210] It should be noted that if the initial acoustic model does not converge, steps I1 to I7 need to be executed multiple times to complete the training of the acoustic model and obtain the acoustic model usable in step 103 of this application embodiment.
[0211] As illustrated by the foregoing embodiments, the process begins by acquiring sound parameter information, including user input information, implicit sound parameters, and explicit sound parameters. The implicit sound parameters are then feature-mapped to obtain implicit feature vectors. Finally, the user input information, implicit feature vectors, and explicit sound parameters are input into the acoustic model, which then outputs acoustic features. In this embodiment, user input information, explicit sound parameters, and implicit sound parameters are acquired. Since implicit sound parameters cannot be directly recognized by the acoustic model, they need to be mapped to implicit feature vectors. The acoustic model is input with user input information, implicit feature vectors, and explicit sound parameters. Therefore, the acoustic model can output acoustic features through sound parameters of multiple dimensions, resulting in richer sound processing effects.
[0212] To facilitate a better understanding and implementation of the above-described solutions in the embodiments of this application, specific examples of corresponding application scenarios are provided below.
[0213] In this embodiment, we take implicit voice parameters including three dimensions and explicit voice parameters including three dimensions as examples. Specifically, implicit voice parameters include voice parameters of timbre dimension, voice parameters of emotion dimension, and voice parameters of style dimension; explicit voice parameters include voice parameters of speech rate dimension, voice parameters of energy dimension, and voice parameters of pitch dimension.
[0214] This application provides a way to customize and display virtual voice avatars for users online, with a wide range of applications. For example, it can be integrated with social platforms such as forums and blogs, allowing users to communicate online using customized virtual voice avatars; it can also be combined with role-playing games to enrich the user's character portrayals; and it can be provided to video production users to create short, episodic dramas. It can be applied to mobile phones, tablets, and other terminal devices, as well as cloud services.
[0215] The system architecture of this application embodiment is as follows: Figure 5 As shown, the user selects relevant parameters for the desired synthesized voice image through the user parameter selection module. Some of these selected parameters are then mapped to timbre, emotion, and style features by the feature mapping module and input into the speech synthesis module. Other parameters do not require mapping and are directly input as explicit features into the speech synthesis module. Combined with the text obtained from the text input module, acoustic features are input through the acoustic model in the speech synthesis module. For example... Figure 6 As shown, the acoustic model outputs acoustic features, and the vocoder in the speech synthesis module restores the acoustic features into speech data, which is then played through the speech playback module, thereby synthesizing a specified virtual voice image according to the user's individual needs.
[0216] In the embodiments of this application, Figure 6 The functions of each module in the system architecture shown are described below:
[0217] User parameter selection module: This module provides users with parameter selection in four directions (timbre, emotion, style, and others). Timbre, emotion, and style-related parameters constitute implicit voice parameters, while other related parameters constitute explicit voice parameters. Timbre includes gender, region, and dialect scale; emotion includes various emotions such as anger, sadness, and joy; style includes diverse styles such as customer service, storytelling, and alluring; and other directions include parameters such as pitch, energy, and speech rate adjustment. Users select the parameters corresponding to their desired synthesized voice image based on the given parameter selection options.
[0218] Text input module: can obtain the specified text entered by the user.
[0219] Feature mapping module: User-provided parameters such as timbre, emotion, and style cannot be directly used as input to the speech synthesis module. They need to be converted into implicit feature vectors that the speech synthesis module can recognize through the feature mapping module. The feature mapping module can be specifically divided into a timbre feature mapping module, an emotion feature mapping module, and a style feature mapping module.
[0220] Speech Synthesis Module: This module mainly includes an acoustic model and a vocoder. The acoustic model takes the user-provided text input and combines it with the implicit feature vector transformed by the feature mapping module to convert it into acoustic features; the vocoder restores the acoustic features output by the acoustic model into speech data.
[0221] Voice playback module: This module refers to the output of the system in the embodiments of this application, which is used to play the audio stream output by the voice synthesis module.
[0222] Based on the above system architecture, the sound processing method provided in this application includes the following process:
[0223] Step 1: The user parameter selection module accepts user input information and related parameters in four directions (timbre, emotion, style, and others) selected by the user. The user input information can be text data and voice data. The parameter selection other than text data and voice data supports real-time input and adjustment by the user, and also supports the user to import the adjusted parameter template file.
[0224] Step 2: The feature mapping module receives the parameters provided by the user and maps some parameters that cannot be directly recognized by the speech synthesis module into input implicit feature vectors that can be recognized by the speech synthesis module.
[0225] Step 3: The speech synthesis module receives text input from the user parameter selection module, other directional parameters, and identifiable implicit timbre feature vectors, style feature vectors, and emotion feature vectors from the feature mapping module. It then synthesizes the corresponding acoustic features through the acoustic model and uses a vocoder to synthesize the acoustic features into speech data.
[0226] Step 4: The voice playback module receives the audio stream output by the voice synthesis module and plays it through the speaker.
[0227] This application embodiment presents a virtual voice avatar interactive synthesis system based on controllable parameters across three dimensions: timbre, emotion, and style. Users can synthesize a desired virtual voice avatar using given parameters in these three dimensions. Timbre, emotion, and style are decoupled from each other; adjusting one direction will not affect the others. Input emotion parameters can be replaced with emoticons. The system supports text input for synthesizing a specified virtual voice avatar, as well as voice input for voice-changing to a specified virtual voice avatar. Users can control the scale of the implicit feature vector representing the voice avatar. User-selected parameters can be exported as parameter templates.
[0228] The application scenarios of this application will be described below with several examples.
[0229] Example 1 of this application
[0230] This application takes the end-to-end text-to-speech (TTS) virtual voice image generation process as an example, and mainly includes the following steps:
[0231] First, the training process of the acoustic model in the speech synthesis module will be introduced, such as... Figure 7 As shown:
[0232] Input: Text, and voice data paired with the text.
[0233] Output: Acoustic features.
[0234] Step S11: Prepare a corpus containing a variety of vocal images with different timbres, styles, and emotions;
[0235] Step S12: Input text. The text is preprocessed and converted into a text embedding layer representation, then encoded into a text implicit feature vector. Input speech data paired with the text from the corpus. The speech data is processed by a timbre encoder, an emotion encoder, and a style encoder to obtain timbre implicit feature vectors, emotion implicit feature vectors, and style implicit feature vectors that match the input speech data. The timbre encoder, emotion encoder, and style encoder are used to train the acoustic model. The timbre encoder takes speech data as input and outputs timbre implicit feature vectors, which characterize timbre features; different timbres have different timbre features. The output of the timbre encoder is the implicit feature vector related only to timbre, retained after filtering redundant information such as text, emotion, and style from the input speech data. The implementation of the emotion encoder and style encoder is similar to that of the timbre encoder and will not be described further here.
[0236] Explicit features such as duration, pitch, and energy can be extracted from the input speech data and directly input into the speech synthesis module.
[0237] In step S13, the text implicit feature vector and the implicit feature vectors from step S12 are input into the encoder. The timbre, emotion, and style encoders extract the corresponding implicit feature vectors. The encoder input is the text implicit feature vector and the implicit feature vectors proposed by the previous three encoders, used to combine text features, timbre features, emotion features, and style features. The encoder outputs the implicit feature vector as input to the duration, energy, and pitch prediction alignment module, which can output features such as duration, pitch, and energy. Loss calculation is performed between these features and the duration, pitch, and energy features of the input speech data. Explicit modeling of relevant features facilitates controllable quantitative adjustment during inference. Explicit modeling refers to using duration, pitch, and energy as part of the intermediate output of the acoustic model, explicitly modeling features such as duration, pitch, and energy, so that some intermediate outputs of the acoustic model have clear physical meaning, facilitating guidance of the model's training convergence direction from the intermediate outputs.
[0238] Step S14: The implicit feature vector after the alignment module is decoded to predict the corresponding speech features, and loss is calculated by comparing it with the speech features extracted by the input speech module.
[0239] like Figures 8 to 10As shown, the following describes the reasoning process based on the acoustic model, which mainly includes the following steps:
[0240] Step S15, the user passes through Figure 8 The user parameter selection interface shown allows users to choose and adjust parameters in four areas: timbre, emotion, style, and others. It also supports exporting and importing adjusted parameters.
[0241] For example, in the timbre category, users can select gender, age, and region. Gender includes male and female, age includes teenager, youth, adult, middle-aged, and elderly, and users can select a region and its concentration. In the style category, users can choose from options such as alluring, thriller, customer service, martial arts, humorous, and fairy tale. In the emotion category, users can choose from joy, anger, sadness, calmness, and excitement, and can also select the intensity of the emotion. Other categories can include speech rate, pitch, and energy.
[0242] The user can input the text: "The price of the shirt is nine pounds and fifteen pence." After selecting sounds from multiple directions, the speech synthesis module outputs speech data, and the user can click the play button to preview the speech effect.
[0243] Step S16: Using the timbre encoder, emotion encoder, and style encoder trained during the training process, extract the average implicit feature vectors of specified timbre, emotion, and style from the speech in the corpus. Taking the extraction of the implicit feature vector of joy emotion as an example, input all speech features labeled "joy" in the corpus into the emotion encoder to obtain the corresponding implicit feature vector of joy emotion. Averaging all implicit feature vectors yields the average implicit feature vector of "joy" emotion. The processing of other timbre, emotion, and style average implicit feature vectors is similar and will not be elaborated here.
[0244] Step S17: The feature mapping module maps the parameters selected by the user in step S15 to the average implicit feature vectors of timbre, emotion, and style extracted in step S16. For example, if the user selects "joy" as the emotion, the implicit feature vector of emotion input in the speech synthesis module will be the average implicit feature vector of "joy" extracted in step S16.
[0245] like Figure 9As shown, the user parameter selection module also provides a scale selection option. After mapping the parameters to the corresponding implicit feature vectors, it is assumed that the implicit feature vectors represent different directions and amplitudes in the vector space. For example, the implicit feature vector of emotion can be divided into different directions and amplitudes in the vector space. The scale control of the corresponding feature is achieved by maintaining the direction of the implicit feature vector and adjusting its amplitude. For example, if the user selects a scale of 0.8 for the emotion "joy," the average implicit feature vector of "joy" is multiplied by the scale of 0.8 for linear amplitude scaling.
[0246] Step S18: The speech synthesis module performs inference, replacing the original timbre, style, and emotion encoders with the specified timbre, style, and emotion implicit feature vectors processed in step S17. Here, the original timbre, style, and emotion encoders refer to the implicit feature vectors extracted from the timbre, style, and emotion encoders during the training phase using speech input. During the inference phase, when there is no speech data input, only text input and user-specified parameters are available. However, the speech synthesis module still requires timbre, style, and emotion features as input. Therefore, the user-specified parameters are used to obtain timbre, style, and emotion features through the feature mapping module, replacing the features extracted by the timbre, style, and emotion encoders for prediction and reconstruction in the speech synthesis module.
[0247] The implicit feature vectors of timbre, style, and emotion, combined with text features, are used as input to the encoder. In the duration, energy, and pitch prediction alignment module, which is a component of the speech synthesis module, the encoder's input is at the character or phoneme level, and needs to be converted to the frame level by the alignment module (predicting the duration of each character / phoneme). The duration, energy, and pitch feature scales selected in step S15 are input, and the predicted duration, energy, and pitch are scaled. Duration corresponds to speech rate; a faster speech rate results in a shorter phoneme duration, and a slower speech rate results in a longer phoneme duration. For example, selecting a speech rate scale of 2 in step S15 multiplies the predicted duration by a coefficient of 1 / 2, achieving a 2x speech rate effect; selecting a speech rate scale of 1 / 3 multiplies the predicted duration by a coefficient of 3, achieving a slower speech rate effect; selecting an energy scale of 1.2 in step S15 multiplies the predicted energy by a coefficient of 1.2, achieving a synthesized speech energy amplification effect; the same applies to pitch.
[0248] In this embodiment, the richness and fun of parameters in the voice image synthesis process are increased. The voice image is decomposed into quantifiable parameter expressions, and parameters related to the voice image are opened to users from multiple directions and dimensions. Users can adjust the parameters according to their personal preferences to synthesize their desired virtual voice image, which increases the interactivity and fun of the speech synthesis system.
[0249] In this embodiment, no target speech is required as guidance, and the synthesis of sound imagery can be achieved "out of nowhere". No user input speech is required as a reference or guidance, and it is not a modification or post-processing of existing speech.
[0250] The decoupling and quantization of the sound image modules involves breaking down the sound image into different modules and then using the implicit feature vectors of each module to represent the characteristics of the sound image in that module. The implicit feature vectors of each module are independent of each other, and the scale quantization of the module features can be achieved by controlling the amplitude of the implicit feature vectors.
[0251] Through Embodiment 1 of this application, the virtual voice image synthesis system based on end-to-end TTS allows users to adjust controllable parameters in four directions—emotion, style, timbre, and others (duration, pitch, energy)—to synthesize virtual voice images. This decomposes the voice image into quantifiable parameter expressions, meeting users' needs for precise, controllable, and expressive virtual voice image synthesis.
[0252] Technical solution of embodiment two of this application
[0253] Example 2 is an extension and supplement to the application scenario of Example 1. In current social networks, most scenarios involve text communication, which uses various emoticons to express emotions. In this case, the virtual voice image synthesis system based on end-to-end TTS described in Example 1 can be used to map the implicit feature vector of emotions through emoticons, allowing users to make corresponding emotional voice comments and achieve the effect of "Why does your comment have sound?"
[0254] The training process of the acoustic model in Example 2 is the same as that in Example 1, and will not be repeated here.
[0255] Figure 11 As shown, the following describes the reasoning process based on the acoustic model, which mainly includes the following steps:
[0256] S21. The online blog post sending and comment replying system can be equipped with the end-to-end TTS-based virtual voice image synthesis system described in Embodiment 1. As shown in Figure 11, the user opens the user parameter selection module through the voice option. The text content entered by the user is: "How could there be such shameless people in the world?" In the emotion feature mapping module, the user can choose to analyze the emotion through the emoticons carried in the text entered by the user, either by selecting parameters (same as in Embodiment 1) or by turning on the emoticon analysis switch.
[0257] S22. Correspondingly, the input of the emotion mapping module becomes the emoji icon identifier (ID). During the inference process, the emoji icon ID can be classified into several major categories of emotion tags. After classification, it can be mapped to a specified emotion implicit feature vector. The extraction method of the emotion implicit feature vector is the same as the inference process S16 in Example 1. The mapping corresponds to the emotion implicit feature vector, and after scale control, the implicit feature vector that can be recognized by the subsequent speech synthesis module is finally output.
[0258] It should be noted that in Implementation Example 1, the user directly selects a certain emotion tag; in Implementation Example 2, the emoji icons are classified into specific emotions, such as the angry emoji icon being classified into the "anger" emotion tag.
[0259] S23. Subsequent steps are the same as in Example 1.
[0260] Currently, no social media platform integrates with a TTS (Text-to-Speech) system to allow users to specify their own voice avatar to deliver audio comments with corresponding emotions. This application's embodiments also incorporate the characteristics of social media commenting by adding an emoticon-to-emotion mapping function, allowing audio comments to convey more intuitive emotions.
[0261] Embodiment 2 of this application integrates a virtual voice avatar customization system onto a social platform that primarily uses text-based communication, expanding user interaction from text to voice, thus enriching the communication experience. Furthermore, it adds emotional mappings from frequently used emoticons on social networks to audio comments, increasing the enjoyment of both communication and audio commentary.
[0262] Technical solution of embodiment three of this application
[0263] Example 3 is a virtual voice avatar generation method based on end-to-end voice conversion (VC). In Example 1, the method synthesizes a specified voice avatar speech by providing text to the user. In Example 3, the user inputs voice data and converts the voice data into the voice data of the specified voice avatar.
[0264] First, the training process of the acoustic model in the speech synthesis module will be introduced, such as... Figure 12 As shown:
[0265] Step S31: Prepare a corpus containing a variety of vocal images with different timbres, styles, and emotions;
[0266] Step S32: Unlike in Example 1, there is no text input. Instead, the voice data is input into the content feature extraction model to extract frame-level content features. The timbre encoder, style encoder, and emotion encoder are processed in the same way as in Example 1.
[0267] Step S33: Unlike in Example 1, the content features are at the frame level rather than the phoneme level, so there is no need for a duration prediction module and extended alignment. The implicit feature vectors such as timbre, style and emotion are input into the encoder along with the content features. The encoder outputs the energy, pitch and other features of the joint input and sends them to the decoder.
[0268] Step S34: The implicit feature vector after the alignment module is decoded to predict the corresponding speech features, and loss is calculated by comparing it with the speech features in the ground truth.
[0269] like Figure 13 As shown, the following describes the reasoning process based on the acoustic model, which mainly includes the following steps:
[0270] Step S35, the user, through, as follows Figure 13 The user parameter selection module shown allows adjustment of parameters in four areas: timbre, emotion, style, and others. It also supports exporting and importing adjusted parameters. There is no text input; only voice input captured by the voice acquisition system is supported.
[0271] Step S36 is the same as step S12 in the reasoning process of Example 1.
[0272] Step S37: Extract energy and pitch from the input speech, perform numerical scaling control, and input the encoder output and energy / pitch data into the decoder. The decoder outputs speech features, which are then fed into the vocoder, which outputs the speech data.
[0273] Current voice-changing technologies primarily use digital signal processing methods to modify and post-process user-input speech. This requires parameter adjustments to achieve the user's desired effect, and the same set of parameters cannot produce a consistent voice image for different users. Embodiment 3 of this application uses the same set of parameters, ensuring a consistent voice image for different users. Embodiment 3 allows adjustment of timbre, style, and emotion modules, resulting in a rich and three-dimensional target voice image. Furthermore, combined with Embodiment 1, it ensures that the user's virtual voice image remains consistent regardless of whether the input is text or voice.
[0274] As illustrated in Example 3, users can specify virtual voice avatar effects, transforming their own voice into the designated virtual voice avatar. In addition to changing the timbre, they can also change emotions, styles, etc. In voice systems such as game systems and live streaming systems, the consistency of the user's virtual voice avatar on the Internet can be maintained, while protecting user privacy.
[0275] Technical solution of embodiment four of this application
[0276] Embodiment 4 of this application is an extension of Embodiment 1. With the trend towards widespread entertainment, video self-media is emerging in large numbers. During video production, when voice-over and narration are needed, video producers typically choose to do the voice-over themselves, find suitable voice actors, extract from existing films and television shows, or synthesize from TTS services provided by various companies. Embodiment 4 of this application provides video producers with a convenient and quick voice-over tool based on Embodiment 1.
[0277] Figure 14 As shown, the following describes the reasoning process based on the acoustic model, which mainly includes the following steps:
[0278] Step S41: For dubbing scenes in video production, you can use, for example... Figure 14 The dubbing interface shown. The voice parameters for the style dimension can include a variety of styles such as: handsome idol, gentle goddess, professional customer service, famous storyteller, classical poetry, cute children's voice, charming, cartoon character 1, cartoon character 2, etc. Users can set the specific values of the voice parameters for the style dimension according to their needs. Additionally, it supports importing voice image templates. In this embodiment, adjustments can be made to style dimensions other than explicit voice parameters. For example, the style of the character to be displayed can be adjusted, thereby customizing the voice parameters for the style dimension according to the user's needs.
[0279] Step S42: The material window provides various parameter templates, which are preset virtual voice avatars that users can directly select and use. Customization is also available; for example, users can customize their avatars using... Figure 8 Users can customize their desired virtual voice avatar through the parameter selection interface, and save templates to import them into the material window as custom templates.
[0280] Step S43: Input text in the speech synthesis window and specify the materials in the material window to be used for each piece of text input. After real-time audio synthesis, the audio length can be counted. Figure 14 The example provided shows the text content entered by two users respectively. This is just an example, and there are no restrictions on the composition of the text content.
[0281] Step S44: Synthesize audio track window. For example, in the audio editing interface, you can set a single track or multiple tracks, and drag the audio of the specified sound image synthesized in the speech synthesis window into the audio track.
[0282] Step S45: The preview window allows you to preview the video dubbing effect.
[0283] Currently, finding voice actors is both time-consuming and costly. Voice-over tools based on virtual voice synthesis systems can save costs. Current speech synthesis methods lack the flexibility to freely adjust timbre, emotion, and style, offering limited options and resulting in mechanical synthesized effects. Embodiment 4 of this application can synthesize speech that closely resembles real human voices and fits the video's storyline.
[0284] Embodiment 4 of this application solves the difficulty for video producers in dubbing their videos by combining a TTS-based virtual voice avatar synthesis system, reducing the economic and time costs of dubbing. Furthermore, the template-based virtual voice avatar provides users with more choices and ways to engage.
[0285] Technical solution of embodiment five of this application
[0286] Embodiment 5 of this application is an extension of Embodiment 1. Currently, the demand for voice libraries is increasing in various directions and fields of speech synthesis technology, requiring voice libraries with different timbres, styles, and emotions. However, recording human voice libraries is time-consuming, costly, and cumbersome.
[0287] Example 5 proposes to construct a multi-scale sound library by leveraging the quantization functions of dialect scale, emotional scale, energy scale, and pitch scale, in conjunction with Example 1.
[0288] The training and inference processes are consistent with those in Example 1. Based on the requirements of the voice library, a voice library with a certain emotional scale of 0.3-1.2 is formulated for emotionally related speech synthesis scenarios.
[0289] Based on the expected parameters and scales specified by the user in the voice image, a large number of relevant specified audio files can be generated as voice library data for research on speech synthesis technology.
[0290] Recording live-action voice libraries is costly in terms of both time and money, and the process is cumbersome. Furthermore, voice actors cannot achieve precise numerical control over certain styles and emotions. Example 5 utilizes an end-to-end TTS-based virtual voice synthesis system to synthesize a rich and well-defined voice library at a low cost and in a short time.
[0291] As illustrated by the foregoing examples, the embodiments of this application decompose the voice avatar into multiple controllable parameters in different dimensions, allowing users to customize their own unique voice avatar like creating a character in a game. This enhances users' self-presentation and self-expression online while protecting their voice privacy. Furthermore, virtual voice avatars with multi-dimensional customization of emotion, style, and timbre can meet various user needs such as dubbing and voice library construction, and have wide applications.
[0292] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0293] To facilitate better implementation of the above-described solutions in the embodiments of this application, related apparatus for implementing the above-described solutions is also provided below.
[0294] Please see Figure 15 As shown in the embodiment of this application, a sound processing device, specifically a sound processing device 1500, may include: an acquisition module 1501, a feature mapping module 1502, and a sound processing module 1503, wherein...
[0295] The acquisition module is used to acquire sound parameter information, which includes: user input information, implicit sound parameters, and explicit sound parameters;
[0296] The feature mapping module is used to perform feature mapping on the implicit sound parameters to obtain implicit feature vectors;
[0297] The sound processing module is used to input the user input information, the implicit feature vector and the explicit sound parameters into the acoustic model, and output acoustic features through the acoustic model.
[0298] In some embodiments of this application, the implicit sound parameters include at least one of the following: sound parameters of timbre dimension, sound parameters of emotion dimension, and sound parameters of style dimension;
[0299] The implicit feature vector includes at least one of the following: timbre implicit feature vector, emotion implicit feature vector, and style implicit feature vector.
[0300] In some embodiments of this application, the explicit sound parameters include at least one of the following: sound parameters in the speech rate dimension, sound parameters in the energy dimension, and sound parameters in the pitch dimension.
[0301] In some embodiments of this application, the sound processing device further includes: a model training module, wherein,
[0302] The model training module is used to acquire a first training text and a first corpus before the acquisition module acquires the sound parameter information. The first corpus includes: first speech data corresponding to the first training text, wherein the first speech data is pre-configured with corresponding first explicit feature vectors and first acoustic features; inputting the first speech data into an initial acoustic model, and outputting a first implicit feature vector through the initial acoustic model; acquiring the text implicit feature vector corresponding to the first training text; combining the text implicit feature vector and the first implicit feature vector into a second implicit feature vector through the initial acoustic model; and applying the second implicit feature vector to the initial acoustic model. The initial acoustic model is used to predict the second explicit feature vector to obtain a second explicit feature vector; the second implicit feature vector is aligned with the second explicit feature vector to obtain a third implicit feature vector; the initial acoustic model is used to predict the second explicit feature vector and the third implicit feature vector to obtain a second acoustic feature; a loss is calculated between the second explicit feature vector and the first explicit feature vector to obtain a first loss calculation result; a loss is calculated between the second acoustic feature and the first acoustic feature to obtain a second loss calculation result; and the training of the initial acoustic model is terminated based on the first loss calculation result and the second loss calculation result.
[0303] In some embodiments of this application, the sound processing device further includes: a model training module, wherein,
[0304] The model training module is used to acquire a second corpus before the acquisition module acquires sound parameter information. The second corpus includes: second speech data, which is pre-configured with corresponding fourth explicit feature vectors and third acoustic features; inputting the second speech data into an initial acoustic model, and outputting a fourth implicit feature vector through the initial acoustic model; extracting content features from the second speech data to obtain a content implicit feature vector; combining the fourth implicit feature vector and the content implicit feature vector into a fifth implicit feature vector through the initial acoustic model; predicting the fourth explicit feature vector and the fifth implicit feature vector through the initial acoustic model to obtain a fourth acoustic feature; calculating a loss based on the third acoustic feature and the fourth acoustic feature to obtain a third loss calculation result; and determining whether to end the training of the initial acoustic model based on the third loss calculation result.
[0305] In some embodiments of this application, the acquisition module is specifically used to acquire the implicit sound parameters and the explicit sound parameters selected by the user; or, to acquire a sound parameter template, the sound parameter template including preset implicit sound parameters and preset explicit sound parameters; and to acquire the implicit sound parameters and the explicit sound parameters according to the sound parameter template.
[0306] In some embodiments of this application, the acquisition module is further configured to acquire the user input information through text data input by the user; or, acquire the user input information through voice data input by the user.
[0307] In some embodiments of this application, the implicit sound parameters include: sound parameters of a first dimension, wherein the first dimension includes one of the following: timbre dimension, emotion dimension, and style dimension;
[0308] The acquisition module is further configured to, when the sound parameters of the first dimension include multiple directions and preset amplitudes corresponding to the multiple directions respectively, acquire a first direction selected by the user from the multiple directions and a first scale selected by the user; adjust the preset amplitude corresponding to the first direction according to the first scale to obtain the adjusted amplitude corresponding to the first direction;
[0309] The sound processing module is further configured to perform amplitude transformation on the acoustic feature according to the adjusted amplitude corresponding to the first direction.
[0310] In some embodiments of this application, the implicit sound parameters include: sound parameters in the timbre dimension, and / or sound parameters in the style dimension;
[0311] The user input information includes: text information;
[0312] The feature mapping module is specifically used to perform feature mapping on the sound parameters of the timbre dimension to obtain a timbre implicit feature vector, and / or to perform feature mapping on the sound parameters of the style dimension to obtain a style implicit feature vector; and to obtain the emoticon icon corresponding to the text information and obtain the emotion implicit feature vector corresponding to the emoticon icon.
[0313] In some embodiments of this application, the sound processing module is further configured to restore the acoustic features to obtain restored speech data.
[0314] In some embodiments of this application, the sound processing module is further configured to add the restored speech data to one or more audio tracks; and play the restored speech data through the audio tracks.
[0315] In some embodiments of this application, the sound processing module is further configured to store the restored speech data in a sound library.
[0316] As illustrated by the foregoing embodiments, the process begins by acquiring sound parameter information, including user input information, implicit sound parameters, and explicit sound parameters. The implicit sound parameters are then feature-mapped to obtain implicit feature vectors. Finally, the user input information, implicit feature vectors, and explicit sound parameters are input into the acoustic model, which then outputs acoustic features. In this embodiment, user input information, explicit sound parameters, and implicit sound parameters are acquired. Since implicit sound parameters cannot be directly recognized by the acoustic model, they need to be mapped to implicit feature vectors. The acoustic model is input with user input information, implicit feature vectors, and explicit sound parameters. Therefore, the acoustic model can output acoustic features through sound parameters of multiple dimensions, resulting in richer sound processing effects.
[0317] It should be noted that the information interaction and execution process between the modules / units of the above-mentioned device are based on the same concept as the method embodiments of this application, and the resulting technical effects are the same as those of the method embodiments of this application. For details, please refer to the description in the method embodiments shown above in this application, and will not be repeated here.
[0318] This application also provides a computer storage medium storing a program that performs some or all of the steps described in the above method embodiments.
[0319] Next, we will introduce another sound processing device provided in the embodiments of this application. Please refer to [link to relevant documentation]. Figure 16 As shown, the sound processing device 1600 includes:
[0320] Receiver 1601, transmitter 1602, processor 1603, and memory 1604 (wherein the number of processors 1603 in the sound processing device 1600 can be one or more, Figure 16 (Taking a processor as an example). In some embodiments of this application, the receiver 1601, transmitter 1602, processor 1603, and memory 1604 can be connected via a bus or other means, wherein, Figure 16 Taking the example of a connection between China and Israel via a bus.
[0321] Memory 1604 may include read-only memory and random access memory, and provides instructions and data to processor 1603. A portion of memory 1604 may also include non-volatile random access memory (NVRAM). Memory 1604 stores operating system and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic business functions and handling hardware-based tasks.
[0322] Processor 1603 controls the operation of the sound processing device; processor 1603 can also be called a central processing unit (CPU). In specific applications, the various components of the sound processing device are coupled together through a bus system, which includes not only data buses but also power buses, control buses, and status signal buses. However, for clarity, all buses in the diagram are referred to as the bus system.
[0323] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 1603. The processor 1603 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 1603 or by instructions in the form of software. The processor 1603 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 1604. Processor 1603 reads the information in memory 1604 and, in conjunction with its hardware, completes the steps of the above method.
[0324] The receiver 1601 can be used to receive input digital or character information and generate signal input related to the settings and function control of the sound processing device. The transmitter 1602 may include a display device such as a display screen and can be used to output digital or character information through an external interface.
[0325] In this embodiment, processor 1603 is used to execute the aforementioned embodiments. Figure 1 The method shown is performed by the sound processing device.
[0326] In another possible design, when the sound processing device is a chip within the terminal, the chip includes a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, pins, or circuitry. The processing unit can execute computer-executable instructions stored in a storage unit to cause the chip within the terminal to perform the sound processing method described in any of the first aspects above. Optionally, the storage unit may be a storage unit within the chip, such as a register or cache. Alternatively, the storage unit may be a storage unit located outside the chip within the terminal, such as read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).
[0327] The processor mentioned above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of a program in the first aspect of the method.
[0328] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0329] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0330] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0331] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
Claims
1. A sound processing method, characterized in that, include: Acquire sound parameter information, which includes: user input information, implicit sound parameters, and explicit sound parameters; wherein, the implicit sound parameters include at least one of the following: sound parameters of timbre dimension, sound parameters of emotion dimension, and sound parameters of style dimension; the explicit sound parameters include at least one of the following: sound parameters of speech rate dimension, sound parameters of energy dimension, and sound parameters of pitch dimension. The implicit sound parameters are feature-mapped to obtain implicit feature vectors; wherein the implicit feature vectors include at least one of the following: timbre implicit feature vector, emotion implicit feature vector, and style implicit feature vector; The user input information, the implicit feature vector, and the explicit sound parameters are input into the acoustic model, and the acoustic features are output through the acoustic model.
2. The method according to claim 1, characterized in that, Before acquiring the sound parameter information, the method further includes: Obtain a first training text and a first corpus, the first corpus including: first speech data corresponding to the first training text, wherein the first speech data is pre-configured with a corresponding first explicit feature vector and a first acoustic feature; The first speech data is input into an initial acoustic model, and the first implicit feature vector is output through the initial acoustic model. Obtain the implicit feature vector of the text corresponding to the first training text; The text implicit feature vector and the first implicit feature vector are combined into a second implicit feature vector using the initial acoustic model. The second implicit feature vector is predicted using the initial acoustic model to obtain the second explicit feature vector. Align the second implicit feature vector with the second explicit feature vector to obtain the third implicit feature vector. The second acoustic feature is obtained by predicting the second explicit feature vector and the third implicit feature vector using the initial acoustic model. The loss is calculated based on the second explicit feature vector and the first explicit feature vector to obtain the first loss calculation result; Loss calculation is performed based on the second acoustic feature and the first acoustic feature to obtain the second loss calculation result; Based on the first loss calculation result and the second loss calculation result, determine whether to end the training of the initial acoustic model.
3. The method according to claim 1, characterized in that, Before acquiring the sound parameter information, the method further includes: Obtain a second corpus, which includes: second speech data, which is pre-configured with a corresponding fourth explicit feature vector and a third acoustic feature; The second speech data is input into the initial acoustic model, and the fourth implicit feature vector is output through the initial acoustic model. Content features are extracted from the second speech data to obtain implicit content feature vectors; The fourth implicit feature vector and the content implicit feature vector are combined into a fifth implicit feature vector using the initial acoustic model. The fourth acoustic feature is obtained by predicting the fourth explicit feature vector and the fifth implicit feature vector using the initial acoustic model. Loss calculation is performed based on the third acoustic feature and the fourth acoustic feature to obtain the third loss calculation result; The training of the initial acoustic model is terminated based on the result of the third loss calculation.
4. The method according to any one of claims 1 to 3, characterized in that, The acquisition of sound parameter information includes: Obtain the implicit and explicit sound parameters selected by the user; or, Obtain a sound parameter template, which includes preset implicit sound parameters and preset explicit sound parameters; obtain the implicit sound parameters and the explicit sound parameters according to the sound parameter template.
5. The method according to any one of claims 1 to 3, characterized in that, The acquisition of sound parameter information includes: The user input information is obtained by using the text data input by the user; or, The user input information is obtained by using the user's voice data.
6. The method according to any one of claims 1 to 3, characterized in that, The implicit sound parameters include: sound parameters of the first dimension, which includes one of the following: timbre dimension, emotion dimension, and style dimension; The step of obtaining sound parameter information includes: when the sound parameters of the first dimension include multiple directions and preset amplitudes corresponding to the multiple directions respectively, obtaining a first direction selected by the user from the multiple directions and a first scale selected by the user; adjusting the preset amplitude corresponding to the first direction according to the first scale to obtain the adjusted amplitude corresponding to the first direction; The method further includes: The acoustic feature is amplitude transformed according to the adjusted amplitude corresponding to the first direction.
7. The method according to any one of claims 1 to 3, characterized in that, The implicit sound parameters include: sound parameters in the timbre dimension, and / or sound parameters in the style dimension; The user input information includes: text information; The step of performing feature mapping on the implicit sound parameters to obtain implicit feature vectors includes: Feature mapping is performed on the sound parameters of the timbre dimension to obtain an implicit timbre feature vector, and / or feature mapping is performed on the sound parameters of the style dimension to obtain an implicit style feature vector; and, Obtain the emoji icon corresponding to the text information, and obtain the implicit emotional feature vector corresponding to the emoji icon.
8. The method according to any one of claims 1 to 3, characterized in that, The method further includes: The acoustic features are then restored to obtain the restored speech data.
9. The method according to claim 8, characterized in that, The method further includes: The restored voice data is added to one or more audio tracks; The restored voice data is played through the audio track.
10. The method according to claim 8, characterized in that, The method further includes: The restored voice data is stored in the audio library.
11. A sound processing device, characterized in that, The sound processing device includes: The acquisition module is used to acquire sound parameter information, which includes: user input information, implicit sound parameters, and explicit sound parameters; wherein, the implicit sound parameters include at least one of the following: sound parameters of timbre dimension, sound parameters of emotion dimension, and sound parameters of style dimension; the explicit sound parameters include at least one of the following: sound parameters of speech rate dimension, sound parameters of energy dimension, and sound parameters of pitch dimension. The feature mapping module is used to perform feature mapping on the implicit sound parameters to obtain implicit feature vectors; wherein, the implicit feature vectors include at least one of the following: timbre implicit feature vector, emotion implicit feature vector, and style implicit feature vector; The sound processing module is used to input the user input information, the implicit feature vector and the explicit sound parameters into the acoustic model, and output acoustic features through the acoustic model.
12. The sound processing apparatus according to claim 11, characterized in that, The sound processing device further includes: a model training module, wherein... The model training module is used to acquire a first training text and a first corpus before the acquisition module acquires the sound parameter information. The first corpus includes: first speech data corresponding to the first training text, wherein the first speech data is pre-configured with corresponding first explicit feature vectors and first acoustic features; inputting the first speech data into an initial acoustic model, and outputting a first implicit feature vector through the initial acoustic model; acquiring the text implicit feature vector corresponding to the first training text; combining the text implicit feature vector and the first implicit feature vector into a second implicit feature vector through the initial acoustic model; and applying the second implicit feature vector to the initial acoustic model. The initial acoustic model is used to predict the second explicit feature vector to obtain a second explicit feature vector; the second implicit feature vector is aligned with the second explicit feature vector to obtain a third implicit feature vector; the initial acoustic model is used to predict the second explicit feature vector and the third implicit feature vector to obtain a second acoustic feature; a loss is calculated between the second explicit feature vector and the first explicit feature vector to obtain a first loss calculation result; a loss is calculated between the second acoustic feature and the first acoustic feature to obtain a second loss calculation result; and the training of the initial acoustic model is terminated based on the first loss calculation result and the second loss calculation result.
13. The sound processing apparatus according to claim 11, characterized in that, The sound processing device further includes: a model training module, wherein... The model training module is used to acquire a second corpus before the acquisition module acquires sound parameter information. The second corpus includes: second speech data, which is pre-configured with corresponding fourth explicit feature vectors and third acoustic features; inputting the second speech data into an initial acoustic model, and outputting a fourth implicit feature vector through the initial acoustic model; extracting content features from the second speech data to obtain a content implicit feature vector; combining the fourth implicit feature vector and the content implicit feature vector into a fifth implicit feature vector through the initial acoustic model; predicting the fourth explicit feature vector and the fifth implicit feature vector through the initial acoustic model to obtain a fourth acoustic feature; calculating a loss based on the third acoustic feature and the fourth acoustic feature to obtain a third loss calculation result; and determining whether to end the training of the initial acoustic model based on the third loss calculation result.
14. The sound processing apparatus according to any one of claims 11 to 13, characterized in that, The acquisition module is specifically used to acquire the implicit sound parameters and the explicit sound parameters selected by the user; or, to acquire a sound parameter template, the sound parameter template including preset implicit sound parameters and preset explicit sound parameters; and to acquire the implicit sound parameters and the explicit sound parameters according to the sound parameter template.
15. The sound processing apparatus according to any one of claims 11 to 13, characterized in that, The acquisition module is further configured to acquire the user input information through text data input by the user; or, acquire the user input information through voice data input by the user.
16. The sound processing apparatus according to any one of claims 11 to 13, characterized in that, The implicit sound parameters include: sound parameters of the first dimension, which includes one of the following: timbre dimension, emotion dimension, and style dimension; The acquisition module is configured to, when the sound parameters of the first dimension include multiple directions and preset amplitudes corresponding to the multiple directions respectively, acquire a first direction selected by the user from the multiple directions and a first scale selected by the user; adjust the preset amplitude corresponding to the first direction according to the first scale to obtain the adjusted amplitude corresponding to the first direction; The sound processing module is further configured to perform amplitude transformation on the acoustic feature according to the adjusted amplitude corresponding to the first direction.
17. The sound processing apparatus according to any one of claims 11 to 13, characterized in that, The implicit sound parameters include: sound parameters in the timbre dimension, and / or sound parameters in the style dimension; The user input information includes: text information; The feature mapping module is specifically used to perform feature mapping on the sound parameters of the timbre dimension to obtain a timbre implicit feature vector, and / or to perform feature mapping on the sound parameters of the style dimension to obtain a style implicit feature vector; and to obtain the emoticon icon corresponding to the text information and obtain the emotion implicit feature vector corresponding to the emoticon icon.
18. The sound processing apparatus according to any one of claims 11 to 13, characterized in that, The sound processing module is also used to restore the acoustic features to obtain restored speech data.
19. The sound processing apparatus according to claim 18, characterized in that, The sound processing module is further configured to add the restored voice data to one or more audio tracks; and play the restored voice data through the audio tracks.
20. The sound processing apparatus according to claim 18, characterized in that, The sound processing module is also used to store the restored voice data in a sound library.
21. A terminal device, characterized in that, The terminal device includes: a processor and a memory; the processor and the memory communicate with each other. The memory is used to store instructions; The processor is configured to execute the instructions in the memory and perform the method as described in any one of claims 1 to 10.
22. A computer-readable storage medium comprising instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1-10.
23. A computer program product comprising instructions that, when run on a computer, cause the computer to perform the method as described in any one of claims 1-10.
Citation Information
Patent Citations
Expressive text-to-speech system and method
US20210225358A1