Method, device and equipment for voice emotion conversion and storage medium
By discretizing the speech signal and predicting prosodic features, combined with speaker identity tags, the problem of capturing non-text signals in existing technologies is solved, achieving richer and more realistic speech emotion conversion effects.
Patent Information
- Application Number
- CN202310152979.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-16
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2043-02-16
AI Technical Summary
Existing speech emotion conversion technologies struggle to effectively capture non-text-related signals, such as pauses, laughter, and lip-pursing sounds, resulting in unsatisfactory emotion conversion effects.
By encoding the input speech signal, a discrete articulation unit sequence is generated. The target emotion label and prosodic features are used for prediction, and the target emotion speech signal is synthesized by combining the speaker identity label. This breaks the dependence on text and captures the expressiveness of speech signals beyond the text.
It achieves speech emotion conversion without text while preserving vocabulary content and speaker timbre, thus enhancing the richness and authenticity of speech emotion conversion.
Smart Images

Figure CN116092478B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for voice emotion conversion. Background Technology
[0002] Emotion and prosody are a fusion of many factors in speech, such as paralinguistic information, intonation, stress, and style. Modeling emotional prosody in speech conversion aims to endow the model with the ability to select a speaking style appropriate for a given context. Prosodic style is difficult to define precisely, but it contains rich information, such as intention and emotion, and influences the speaker's choice of intonation and tone.
[0003] Existing emotion generation or emotion transfer techniques struggle to produce convincing results because they can only address a subset or one aspect of these problems. Signal-based emotion transfer methods primarily focus on manipulating parameters of the speech signal, only addressing variations at the speech and prosodic levels. Furthermore, current modeling methods struggle to model non-text-related signals in emotional speech, such as pauses, laughter, and lip-pursing sounds, making it difficult to capture expressive speech signals beyond the text and resulting in unsatisfactory speech emotion transfer effects. Summary of the Invention
[0004] To address the technical problem that existing voice emotion conversion solutions often suffer from unsatisfactory conversion results due to their limited scope or incomplete signal capture, this application provides a method, apparatus, device, and storage medium for voice emotion conversion, with the primary objective of improving the effectiveness of voice emotion conversion.
[0005] To achieve the above objectives, this application provides a method for speech emotion conversion, applied to a speech emotion conversion system, the method comprising:
[0006] The input raw speech signal is encoded to obtain a first discrete speech unit sequence with original emotion;
[0007] The first discrete phonological unit sequence is converted into a second discrete phonological unit sequence with the target emotion based on the target emotion tag, wherein the second discrete phonological unit sequence has the same lexical content as the first discrete phonological unit sequence;
[0008] Based on the target emotion tag, prosodic features are predicted for the second discrete articulation unit sequence to obtain predicted prosodic features, which include predicted duration and predicted fundamental frequency.
[0009] The target emotion speech signal is synthesized based on the target emotion tag, the second discrete articulation unit sequence, the predicted prosodic features, and the speaker identity tag of the original speech signal.
[0010] Furthermore, to achieve the above objectives, this application also provides a speech emotion conversion device, the device comprising:
[0011] The encoding module is used to encode the input raw speech signal to obtain a first discrete articulation unit sequence with the original emotion;
[0012] The emotion conversion module is used to convert a first discrete phonological unit sequence into a second discrete phonological unit sequence with the target emotion based on the target emotion tag, wherein the second discrete phonological unit sequence has the same lexical content as the first discrete phonological unit sequence;
[0013] The prediction module is used to predict prosodic features of the second discrete articulation unit sequence based on the target emotion tag, and obtain the predicted prosodic features, which include the predicted duration and the predicted fundamental frequency.
[0014] The synthesis module is used to synthesize the target emotion speech signal based on the target emotion tag, the second discrete articulation unit sequence, the predicted prosodic features, and the speaker identity tag of the original speech signal.
[0015] To achieve the above objectives, this application also provides a computer device, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, wherein the processor executes the computer-readable instructions to perform the steps of the speech emotion conversion method as described in any of the preceding claims.
[0016] To achieve the above objectives, this application also provides a computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, cause the processor to perform the steps of the speech emotion conversion method as described in any of the preceding claims.
[0017] The method, apparatus, device, and storage medium for speech emotion conversion proposed in this application break away from the traditional reliance on text. Instead of learning spoken content from written text, it uses discrete articulation units to learn the spoken content and emotional prosody of the original audio signal, capturing expressive speech signals beyond the text. It then converts the discrete articulation units of the original emotion into discrete articulation units with the target emotion, and finally concatenates the sequence of discrete articulation units with the target emotion based on predicted prosodic features to form a speech signal with the target emotion. This achieves textless discretization of spoken audio, capturing expressive speech signals beyond the text, resulting in richer and more realistic speech emotion conversion effects. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating a method for speech emotion conversion in one embodiment of this application;
[0019] Figure 2This is a structural block diagram of a speech emotion conversion device according to an embodiment of this application;
[0020] Figure 3 This is a block diagram of the internal structure of a computer device according to an embodiment of this application.
[0021] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0023] In daily communication, people often use nonverbal cues such as tone of voice, pauses, accents, rhythm, laughter, crying, shouting, and lip-pursing sounds to enhance the effectiveness of dialogue. For example, saying the same sentence when happy, angry, disappointed, or sleepy will sound very different, even though the content is the same. This demonstrates the strong relationship between the emotions conveyed in speech signals and the nonverbal cues within them.
[0024] Figure 1 This is a flowchart illustrating a speech emotion conversion method according to an embodiment of this application. The method is described using an example of its application in a speech emotion conversion system. The training method for the execution time prediction model includes the following steps S100-S400.
[0025] S100: Encode the input raw speech signal to obtain the first discrete pronunciation unit sequence with the original emotion.
[0026] Specifically, the input raw speech signal, i.e., audio, is encoded using textless NLP technology to obtain a first discrete articulation unit sequence. This is equivalent to decomposing and converting the audio, which is a continuous value sequence, into a discrete value sequence. The first discrete articulation unit sequence includes multiple first discrete articulation units. The first discrete articulation unit sequence possesses the original emotion of the original speech signal.
[0027] Using discrete unit sequences to represent speech is to capture non-linguistic pronunciations and also facilitates better modeling and sampling.
[0028] S200: Convert the first discrete phonological unit sequence into a second discrete phonological unit sequence with the target emotion based on the target emotion tag, wherein the second discrete phonological unit sequence has the same lexical content as the first discrete phonological unit sequence.
[0029] Specifically, the target emotion label is used to represent the target emotion. While maintaining the lexical content of the speech, the first discrete phonetic unit sequence with the original emotion is converted or translated into a second discrete phonetic unit sequence with the target emotion. The conversion process may involve deleting and / or replacing some of the first discrete phonetic units, as well as adding or inserting new discrete phonetic units, so that the final second discrete phonetic unit sequence possesses the target emotion. The second discrete phonetic unit sequence maintains the same lexical content as the first discrete phonetic unit sequence; this step changes the emotion from the perspective of altering the discrete phonetic unit sequence.
[0030] The second discrete articulation unit sequence contains second discrete articulation units that are not exactly the same as the first discrete articulation unit sequence contains first discrete articulation units; they may be the same or different in number.
[0031] This step aims to modify the first discrete articulation unit sequence by converting the original emotion into the target emotion.
[0032] S300: Based on the target emotion tag, the prosodic features of the second discrete articulation unit sequence are predicted to obtain the predicted prosodic features, which include the predicted duration and the predicted fundamental frequency.
[0033] Specifically, in speech, phoneme duration directly affects the length of articulation and overall prosody. Fundamental frequency, or pitch or F0, is related to perceived tone and intonation. The second discrete articulation unit sequence differs from the first discrete articulation unit sequence; therefore, prosodic feature prediction is required for each second discrete articulation unit in the newly generated second discrete articulation unit sequence. During prosodic feature prediction, the target sentiment label needs to be included as one of the prediction conditions to influence the prediction results.
[0034] The predicted prosodic features include the sub-predicted duration and sub-predicted fundamental frequency for each second discrete articulation unit.
[0035] Prosody is a feature that influences emotion. This step changes the emotion of the speech to be synthesized by re-predicting the prosodic features of the second discrete articulation unit.
[0036] S400: Synthesizes a target emotion speech signal based on the target emotion tag, the second discrete articulation unit sequence, the predicted prosodic features, and the speaker identity tag of the original speech signal.
[0037] Specifically, to ensure that the timbre of the speech remains unchanged—that is, it is still the speaker's timbre from the original speech signal—a speaker identity label needs to be input. Both the speaker identity label and the target emotion label are types of features. A vocoder concatenates the target emotion label, the second discrete articulation unit sequence, the predicted prosodic features, and the speaker identity label from the original speech signal to obtain a speech signal with the target emotion. This speech signal is a continuous signal. Thus, perceptible emotion modification of the speech corpus is achieved while preserving lexical content and speaker identity; this is known as speech emotion conversion.
[0038] This embodiment breaks away from the traditional reliance on text, not limiting itself to learning spoken content from written text. Instead, it uses discrete phonetic units (PPUs) to learn the spoken content and emotional prosody of the original audio signal, capturing expressive speech signals beyond the text. It then converts the original emotional discrete PPUs into discrete PPUs with the target emotion, and finally, based on predicted prosodic features, splices the sequence of discrete PPUs with the target emotion into a speech signal with the target emotion. This embodiment achieves textless discretization of spoken audio, capturing expressive speech signals beyond the text. While preserving the lexical content and speaker timbre of the speech signal, it performs textless speech emotion conversion based on the decomposed discrete speech representation, resulting in a richer, more realistic, and natural speech emotion conversion effect.
[0039] In one embodiment, the voice emotion conversion system includes a pre-trained voice model;
[0040] Step S100 specifically includes:
[0041] The input raw speech signal is sampled using a pre-trained speech model to obtain a spectral feature representation;
[0042] The spectral feature representation is numericalized, and the first discrete articulation unit sequence is obtained based on the multiple first discrete articulation units.
[0043] Specifically, the first discrete articulation unit sequence includes multiple first discrete articulation units.
[0044] The original speech signal is a speech waveform. The speech pre-trained model is obtained through self-supervised or unsupervised learning, specifically such as the wav2vec 2.0 model or the Hubert model, etc. This application is not limited to these.
[0045] Self-supervised learning is a machine learning method that directly extracts supervised information from large-scale unsupervised data for supervised learning and training (it can be seen as a special case of unsupervised learning). Self-supervised learning requires labels, but these labels come from the data itself, not from manual annotation. Self-supervised learning methods include context-based, time-series-based, and comparison-based methods.
[0046] Unsupervised learning refers to machine learning methods that learn predictive models from unlabeled data. Its essence is learning statistical patterns or latent structures within the data. Examples of unsupervised learning methods include clustering, K-means, and PCA.
[0047] The wav2vec 2.0 model encodes the original audio into a sequence of frame features, and then converts each frame feature sequence into corresponding discrete features. The frame feature sequence is represented by spectral features.
[0048] The HuBERT model is a self-supervised model for continuous audio signal masking prediction, similar to BERT. HuBERT obtains its training target by performing K-means clustering on MFCC features or HuBERT features, where MFCC features are spectral representations. The HuBERT model employs an iterative training approach. In the first iteration, the HuBERT model performs clustering on MFCC features; in the second iteration, it performs clustering on the intermediate layer features (HuBERT features) obtained from the first iteration, and so on. HuBERT borrows the loss function from the mask language model in BERT and applies Transformer to predict the discrete IDs of the masked positions to train the model. HuBERT uses an iterative approach to generate training metrics, i.e., the discrete IDs for each frame. K-means clustering of the MFCC features of the speech is used to generate discrete IDs for training the first-generation HuBERT model. Subsequently, the input of the already trained previous-generation model is clustered to generate new IDs for the next round of learning. The HuBERT model uses the K-means algorithm to convert continuous speech signals into discrete labels and uses these discrete labels as metrics for modeling. The K-means algorithm is a clustering algorithm used in unsupervised learning. The Hubert model can discretize the original speech signal into various types of discrete articulation units.
[0049] Speech pre-trained models are trained using large amounts of unsupervised data and can be applied to various downstream tasks. Using pre-trained speech pre-trained models to extract discrete representations from speech signals can greatly improve the efficiency and accuracy of discrete feature extraction, and reduce the tedious work of sample collection, annotation, and model training.
[0050] In one embodiment, the first discrete articulation unit sequence includes a language articulation unit and a first non-language articulation unit, and the speech emotion conversion system includes an emotion conversion model;
[0051] Step S200 specifically includes:
[0052] Based on the target emotion tag, the emotion conversion model is used to convert the first non-verbal pronunciation unit with the original emotion in the first discrete pronunciation unit sequence into the second non-verbal pronunciation unit with the target emotion, while retaining the verbal pronunciation units, thus obtaining the second discrete pronunciation unit sequence.
[0053] Specifically, language pronunciation units are discrete pronunciation units that correspond to text or vocabulary content, such as "hello" and "thank you." Non-language pronunciation units are discrete pronunciation units that do not correspond to text or vocabulary content, such as crying sounds and laughter sounds.
[0054] This method extracts information-rich discrete phonetic units from speech signals, capturing both linguistic and non-linguistic discrete phonetic units. Non-linguistic pronunciation, or non-textual pronunciation, carries rich emotional signals; therefore, this embodiment alters emotions by changing non-linguistic pronunciation.
[0055] The first discrete articulation unit sequence is judged from an overall perspective to possess target emotion, but some first discrete articulation units may have target emotion labels, while others may not.
[0056] The verbal pronunciation unit and the first nonverbal pronunciation unit are different first discrete pronunciation units. The sequence of first discrete pronunciation units contains at least one first nonverbal pronunciation unit. The first nonverbal pronunciation unit is modified by systematically performing at least one of the following operations—deletion, replacement, or insertion—on the first nonverbal pronunciation unit according to the target emotional label using an emotion conversion model, thereby obtaining a second nonverbal pronunciation unit. The second nonverbal pronunciation unit may contain some, none, or all of the first nonverbal pronunciation units, and may also contain newly inserted nonverbal pronunciation units, depending on the conversion process of the emotion conversion model. Specifically, the newly inserted nonverbal pronunciation unit is a pronunciation unit with the target emotional label, while the deleted first nonverbal pronunciation unit is a pronunciation unit with the original emotional label.
[0057] The model converts input audio into target emotion, which is essentially an end-to-end sequence translation problem. It is easier to convert emotions by inserting, deleting, or replacing some non-verbal audio signals, i.e., non-verbal vocal units.
[0058] This embodiment alters emotion by deleting, inserting, or replacing at least one of the non-verbal articulation units, achieving emotional transformation of non-textual, non-verbal signals in speech. This allows the model to not only detect changes in the signal's spectrum and parameters but also to model non-verbal vocalizations.
[0059] In one embodiment, the conversion process includes insertion, replacement, and deletion;
[0060] Using an emotion conversion model, the first non-verbal articulation unit in the first discrete articulation unit sequence, which carries the original emotion, is converted into a second non-verbal articulation unit carrying the target emotion through conversion processing, including:
[0061] An emotion transfer model was used to perform emotion recognition and labeling on the first discrete articulation unit, resulting in an emotion label for each first non-verbal articulation unit.
[0062] Obtain the target non-verbal pronunciation unit with the target emotion label, and perform at least one of the following operations: delete the first non-verbal pronunciation unit with the original emotion label, replace the original emotion label with the target non-verbal pronunciation unit, and insert the target non-verbal pronunciation unit to obtain the second non-verbal pronunciation unit with the target emotion.
[0063] Specifically, the sentiment transformation model possesses both sentiment labeling / recognition and sentiment transformation capabilities. The sentiment transformation model can be constructed using a sequence-to-sequence Transformer model structure. The generated sentiment labels can be used to control speech synthesis in novel ways, such as altering speed and emotional style, all independent of the lexical content of the speech.
[0064] In model training, the training data consists of discrete phonetic units of sample audio data with various emotional styles and tones. The sample audio data lacks specific emotional classification labels (such as joy, anger, sorrow, etc.), requiring the model to autonomously learn these labels, i.e., soft labels. Because speech contains various emotions and sounds such as laughter, crying, and pursing lips, some emotions cannot be manually classified. By using the model to autonomously learn high-dimensional features from the data and obtain multiple emotional labels, a model capable of representing various emotional styles and tones can be trained.
[0065] The sentiment transfer model's internal structure can generate interpretable soft labels from the data itself. These labels can be used to express various styles of control and delivery tasks, while improving the expression of long sentence synthesis. Sentiment style labels can be directly applied to noisy, unlabeled data, thereby enabling a highly scalable speech-to-speech system.
[0066] More specifically, the sentiment transfer model can classify discrete feature representations in speech, i.e., discrete articulation units, using learned label information, and perform potential segmentation and identification of sentiment style signals, i.e., first discrete articulation units. After the sentiment transfer model matches the target non-verbal articulation units corresponding to the target sentiment label, it uses the target non-verbal articulation units to replace or insert into the first discrete articulation units of the original sentiment label, or to delete the first non-verbal articulation units of the original sentiment label. This alters the first discrete articulation unit sequence through deletion, replacement, and insertion, resulting in a second discrete articulation unit sequence. Here, the target non-verbal articulation units are discrete articulation units that the sentiment transfer model has learned and classified during training.
[0067] This embodiment achieves emotional transformation of discrete articulation units by employing a learnable mechanism for inserting, deleting, and replacing non-speech phonations while preserving lexical content (e.g., deleting non-speech phonation units corresponding to laughter and inserting non-speech phonation units corresponding to crying while preserving speech content). Current speech synthesis methods struggle to handle emotional expression in long sentences and also perform poorly in modeling the emotional content of individual words within a sentence. This embodiment, through modeling discrete speech units and using soft-labeling of emotional embedding vectors, achieves better results in both fine-grained and long-range modeling.
[0068] In one embodiment, the speech emotion conversion system includes a duration prediction model, and the second discrete articulation unit sequence includes a plurality of second discrete articulation units;
[0069] Step S300 specifically includes:
[0070] Based on the target sentiment tag, the duration of each second discrete articulation unit is predicted using a duration prediction model. Based on the sub-predicted duration of the second discrete articulation unit, the predicted duration of the second discrete articulation unit sequence is obtained.
[0071] Specifically, the duration prediction model is a CNN model, which uses a convolutional neural network (CNN) to predict the duration of each second discrete vocal unit.
[0072] During training, the duration prediction model learns the mapping between discrete articulation units and duration. The discrete articulation units corresponding to the sample speech data are used as input, and the actual duration of the discrete articulation units is used as supervision labels. The duration prediction model is trained by minimizing the mean square error (MSE) between the predicted duration and the actual duration.
[0073] The actual duration of discrete phonation units can be obtained using external tools such as FFmpeg or librosa.
[0074] In one embodiment, the voice emotion conversion system also includes a fundamental frequency prediction model.
[0075] Step S300 specifically also includes:
[0076] The second discrete articulation unit sequence carrying the sub-prediction duration of all second discrete articulation units and the target sentiment tag are used as inputs to the fundamental frequency prediction model. The fundamental frequency of each second discrete articulation unit is predicted using the fundamental frequency prediction model. Based on the obtained sub-fundamental frequencies of the second discrete articulation units, the predicted fundamental frequency of the second discrete articulation unit sequence is obtained.
[0077] Specifically, the fundamental frequency estimation model can use a CNN model, where the fundamental frequency is f0. The fundamental frequency prediction model learns the mapping between discrete articulatory units and fundamental frequencies; it uses the discrete articulatory units corresponding to sample speech data as input, the true fundamental frequencies of the discrete articulatory units as supervision labels, and utilizes the mechanism of minimizing the binary cross-entropy (BCE) between the predicted fundamental frequency and the true fundamental frequency to train the fundamental frequency prediction model.
[0078] The YAAPT algorithm or the YIN algorithm can be used to extract the true fundamental frequency of the discrete vocal unit.
[0079] Because the mapping relationship between discrete articulation units and prosodic features is related to the target emotion, the target emotion label needs to be used as the input of the fundamental frequency prediction model. This allows the fundamental frequency prediction model to predict the fundamental frequency based on the target emotion label, resulting in higher accuracy.
[0080] In one embodiment, the voice emotion conversion system includes a vocoder;
[0081] Step S400 specifically includes:
[0082] The second discrete articulation unit sequence is dilated using the prediction duration to obtain the predicted discrete articulation unit sequence;
[0083] The predicted discrete phonetic unit sequence, the predicted fundamental frequency, the target emotion tag, and the speaker identity tag of the original speech signal are used as inputs to the vocoder;
[0084] The target emotional speech signal is obtained by concatenating the predicted discrete phonetic unit sequence, the predicted fundamental frequency, the target emotion tag, and the speaker identity tag along the time axis using a vocoder.
[0085] Specifically, the vocoder can use neural vocoders such as WaveNet or HiFi-GAN. HiFi-GAN, for example, consists of a generator G and a set of discriminators D. The generator G concatenates the predicted discrete phonetic unit sequence, predicted fundamental frequency, target emotion tag, and speaker identity tag along the time axis, and feeds this concatenation into a series of convolutional layers that output one-dimensional signals to obtain the target emotion speech signal.
[0086] The proposed emotional speech conversion scheme combines self-supervised representation learning and unsupervised learning. Based on a pre-trained speech model trained with a large amount of unlabeled data, it encodes continuous speech signals into discrete articulatory units (PAUs). This solves the problems of limited labeled data, difficulty in classifying emotional categories, and the challenge of modeling non-textual signals in emotional speech synthesis and conversion. Through deletion, replacement, and insertion mechanisms, it performs emotional conversion on non-textual, non-verbal PAUs. Then, it further re-emotionally transforms the discrete PAUs using prosodic feature prediction. This achieves efficient and rich emotional conversion and speech synthesis while preserving lexical content and speaker timbre.
[0087] Compared to existing text-based methods for emotion transfer, which often lose tone and nonverbal expression and thus some of the expressive power of spoken language, resulting in poor emotion transfer effectiveness, this application utilizes a self-supervised speech representation method that learns discrete units from the original audio. This eliminates reliance on text and conveys rich semantic and speech information beyond phonemes, leading to more natural and fluent speech emotion transfer.
[0088] Figure 2 This is a structural block diagram of a speech emotion conversion device according to an embodiment of this application. (Reference) Figure 2 The device includes:
[0089] The encoding module 100 is used to encode the input raw speech signal to obtain a first discrete pronunciation unit sequence with original emotion;
[0090] The emotion conversion module 200 is used to convert a first discrete phonological unit sequence into a second discrete phonological unit sequence with the target emotion according to the target emotion tag, wherein the second discrete phonological unit sequence has the same lexical content as the first discrete phonological unit sequence;
[0091] The prediction module 300 is used to predict prosodic features of the second discrete articulation unit sequence based on the target emotion tag, and obtain the predicted prosodic features, wherein the predicted prosodic features include the predicted duration and the predicted fundamental frequency.
[0092] The synthesis module 400 is used to synthesize a target emotional speech signal based on the target emotion tag, the second discrete articulation unit sequence, the predicted prosodic features, and the speaker identity tag of the original speech signal.
[0093] In one embodiment, the voice emotion conversion system includes a pre-trained voice model;
[0094] The encoding module 100 specifically includes:
[0095] The sampling module is used to sample the input raw speech signal using a pre-trained speech model to obtain a spectral feature representation;
[0096] The discrete module is used to quantify the spectral feature representation and obtain the first discrete articulation unit sequence based on the obtained multiple first discrete articulation units.
[0097] In one embodiment, the first discrete articulation unit sequence includes a language articulation unit and a first non-language articulation unit, and the speech emotion conversion system includes an emotion conversion model;
[0098] The emotion conversion module 200 is specifically used to convert the first non-verbal pronunciation unit with the original emotion in the first discrete pronunciation unit sequence into the second non-verbal pronunciation unit with the target emotion through conversion processing using the emotion conversion model, based on the target emotion label, while retaining the verbal pronunciation units, to obtain the second discrete pronunciation unit sequence.
[0099] In one embodiment, the conversion process includes insertion, replacement, and deletion;
[0100] The emotion conversion module 200 specifically includes:
[0101] The emotion recognition module is used to perform emotion recognition and labeling on the first discrete articulation units using an emotion transfer model, obtaining an emotion label for each first non-verbal articulation unit.
[0102] The conversion module is used to obtain the target non-verbal pronunciation unit with the target emotion label, and perform at least one operation to obtain the second non-verbal pronunciation unit with the target emotion label, such as deleting the first non-verbal pronunciation unit with the original emotion label, replacing the original emotion label with the target non-verbal pronunciation unit, and inserting the target non-verbal pronunciation unit.
[0103] In one embodiment, the speech emotion conversion system includes a duration prediction model, and the second discrete articulation unit sequence includes a plurality of second discrete articulation units;
[0104] Prediction module 300 specifically includes:
[0105] The duration prediction module is used to predict the duration of each second discrete articulation unit based on the target sentiment tag using a duration prediction model. Based on the sub-predicted duration of the second discrete articulation unit, the predicted duration of the second discrete articulation unit sequence is obtained.
[0106] In one embodiment, the voice emotion conversion system also includes a fundamental frequency prediction model.
[0107] The prediction module 300 also includes:
[0108] The fundamental frequency prediction module is used to take the second discrete articulation unit sequence carrying the sub-prediction duration of all second discrete articulation units and the target sentiment tag as input to the fundamental frequency prediction model, and use the fundamental frequency prediction model to predict the fundamental frequency of each second discrete articulation unit. Based on the obtained sub-fundamental frequencies of the second discrete articulation units, the predicted fundamental frequency of the second discrete articulation unit sequence is obtained.
[0109] In one embodiment, the voice emotion conversion system includes a vocoder;
[0110] The synthesis module 400 specifically includes:
[0111] The dilation module is used to dilate the second discrete articulation unit sequence using the prediction duration to obtain the predicted discrete articulation unit sequence.
[0112] The input module is used to take the predicted discrete phonetic unit sequence, the predicted fundamental frequency, the target emotion tag, and the speaker identity tag of the original speech signal as input to the vocoder;
[0113] The vocoder module is used to concatenate the predicted discrete speech unit sequence, the predicted fundamental frequency, the target emotion tag, and the speaker identity tag along the time axis to obtain the target emotion speech signal.
[0114] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0115] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0116] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0117] The terms "first" and "second" in the above-mentioned modules / units are only used to distinguish different modules / units and are not intended to specify which module / unit has a higher priority or any other limiting meaning. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those steps or modules explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or devices. The module divisions appearing in this application are merely logical divisions; in actual applications, different division methods may be used.
[0118] Specific limitations regarding the device for voice emotion conversion can be found in the limitations of the method for voice emotion conversion described above, and will not be repeated here. Each module in the aforementioned voice emotion conversion device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0119] Figure 3 This is a block diagram of the internal structure of a computer device according to an embodiment of this application. Specifically, this computer device may be a voice emotion conversion system. Figure 3 As shown, the computer device includes a processor, memory, network interface, input device, and display screen connected via a system bus. The processor provides computing and control capabilities. The memory includes storage media and internal memory. The storage media can be non-volatile or volatile. The storage media stores an operating system and may also store computer-readable instructions, which, when executed by the processor, enable the processor to implement methods for voice emotion conversion. The internal memory provides an environment for the operation of the operating system and computer-readable instructions in the storage media. The internal memory may also store computer-readable instructions, which, when executed by the processor, enable the processor to implement methods for voice emotion conversion. The network interface of the computer device is used for communication with an external server via a network connection. The display screen of the computer device can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, buttons, a trackball, or a touchpad located on the computer device casing, or an external keyboard, touchpad, or mouse, etc.
[0120] In one embodiment, a computer device is provided, including a memory, a processor, and computer-readable instructions (e.g., a computer program) stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, it implements the steps of the speech emotion conversion method described in the above embodiments, for example... Figure 1 The steps S100 to S400 shown, as well as other extensions and related steps of the method, are examples. Alternatively, the processor, when executing computer-readable instructions, implements the functions of each module / unit of the speech emotion conversion device in the above embodiments, for example... Figure 2 The functions of modules 100 to 400 are shown. To avoid repetition, they will not be described again here.
[0121] A processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of a computer device, connecting all parts of the computer device through various interfaces and lines.
[0122] Memory can be used to store computer-readable instructions and / or modules. The processor implements various functions of the computer device by running or executing the computer-readable instructions and / or modules stored in memory, and by accessing data stored in memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area can store data created based on the use of the mobile phone (such as audio data, video data, etc.).
[0123] The memory can be integrated into the processor or set up separately from the processor.
[0124] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0125] In one embodiment, a computer-readable storage medium is provided, on which computer-readable instructions are stored. When executed by a processor, the computer-readable instructions implement the steps of the speech emotion conversion method described in the above embodiments, for example... Figure 1 The steps S100 to S400 shown, as well as other extensions and related steps of the method, are examples. Alternatively, computer-readable instructions, when executed by a processor, implement the functions of each module / unit of the speech emotion conversion apparatus in the above embodiments, for example... Figure 2 The functions of modules 100 to 400 are shown. To avoid repetition, they will not be described again here.
[0126] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium, and when executed, they can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0127] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.
[0128] The sequence numbers of the embodiments in this application are merely for description and do not represent the superiority or inferiority of the embodiments. Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0129] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for speech emotion conversion, applied to a speech emotion conversion system, characterized in that, The method includes: The input raw speech signal is encoded to obtain a first discrete speech unit sequence with original emotion; The first discrete pronunciation unit sequence is converted into a second discrete pronunciation unit sequence with the target emotion based on the target emotion tag, wherein the second discrete pronunciation unit sequence has the same lexical content as the first discrete pronunciation unit sequence; the first discrete pronunciation unit sequence includes linguistic pronunciation units and first non-linguistic pronunciation units, and the speech emotion conversion system includes an emotion conversion model; the conversion of the first discrete pronunciation unit sequence into a second discrete pronunciation unit sequence with the target emotion based on the target emotion tag includes: based on the target emotion tag, using the emotion conversion model, converting the first non-linguistic pronunciation units with the original emotion in the first discrete pronunciation unit sequence into the second non-linguistic pronunciation units with the target emotion through conversion processing, while retaining the linguistic pronunciation units, to obtain the second discrete pronunciation unit sequence; the conversion processing includes insertion, replacement, and deletion; Based on the target emotion tag, prosodic features are predicted on the second discrete phonological unit sequence to obtain predicted prosodic features, wherein the predicted prosodic features include predicted duration and predicted fundamental frequency; The target emotion speech signal is synthesized based on the target emotion tag, the second discrete articulation unit sequence, the predicted prosodic features, and the speaker identity tag of the original speech signal.
2. The method according to claim 1, characterized in that, The speech emotion conversion system includes a pre-trained speech model; The process of encoding the input raw speech signal to obtain a first discrete articulation unit sequence with original emotion includes: The input raw speech signal is sampled using the aforementioned speech pre-training model to obtain a spectral feature representation; The spectral feature representation is numericalized, and a first discrete articulation unit sequence is obtained based on the multiple first discrete articulation units.
3. The method according to claim 1, characterized in that, The process of using the emotion conversion model to convert the first non-verbal articulation unit with the original emotion in the first discrete articulation unit sequence into a second non-verbal articulation unit with the target emotion includes: The emotion transfer model is used to perform emotion recognition and labeling on the first discrete articulation unit, resulting in an emotion label for each first non-verbal articulation unit. Obtain the target non-verbal pronunciation unit with the target emotion label, and perform at least one of the following operations: delete the first non-verbal pronunciation unit with the original emotion label, replace the original emotion label with the target non-verbal pronunciation unit, and insert the target non-verbal pronunciation unit to obtain the second non-verbal pronunciation unit with the target emotion.
4. The method according to claim 1, characterized in that, The speech emotion conversion system includes a duration prediction model, and the second discrete articulation unit sequence includes multiple second discrete articulation units; The step of predicting prosodic features on the second discrete articulation unit sequence based on the target emotion tag to obtain predicted prosodic features includes: Based on the target emotion tag, the duration of each second discrete articulation unit is predicted using the duration prediction model. Based on the sub-predicted duration of the second discrete articulation unit, the predicted duration of the second discrete articulation unit sequence is obtained.
5. The method according to claim 4, characterized in that, The speech emotion conversion system also includes a fundamental frequency prediction model. The step of predicting prosodic features of the second discrete phonological unit sequence based on the target emotion tag to obtain predicted prosodic features further includes: The second discrete articulation unit sequence carrying the sub-prediction durations of all second discrete articulation units and the target sentiment tag are used as inputs to the fundamental frequency prediction model. The fundamental frequency of each second discrete articulation unit is predicted using the fundamental frequency prediction model. Based on the obtained sub-fundamental frequencies of the second discrete articulation units, the predicted fundamental frequency of the second discrete articulation unit sequence is obtained.
6. The method according to claim 1, characterized in that, The voice emotion conversion system includes a vocoder; The process of synthesizing the target emotion speech signal based on the target emotion tag, the second discrete articulation unit sequence and predicted prosodic features, and the speaker identity tag of the original speech signal includes: The second discrete articulation unit sequence is dilated using the predicted duration to obtain a predicted discrete articulation unit sequence. The predicted discrete articulation unit sequence, the predicted fundamental frequency, the target emotion tag, and the speaker identity tag of the original speech signal are used as inputs to the vocoder; The vocoder is used to concatenate the predicted discrete pronunciation unit sequence, the predicted fundamental frequency, the target emotion tag, and the speaker identity tag along the time axis to obtain the target emotion speech signal.
7. A device for voice emotion conversion, characterized in that, The device includes: The encoding module is used to encode the input raw speech signal to obtain a first discrete articulation unit sequence with the original emotion; The emotion conversion module is used to convert the first discrete pronunciation unit sequence into a second discrete pronunciation unit sequence with the target emotion according to the target emotion tag, wherein the second discrete pronunciation unit sequence has the same lexical content as the first discrete pronunciation unit sequence; The prediction module is used to predict prosodic features of the second discrete phonological unit sequence based on the target emotion tag, so as to obtain predicted prosodic features, wherein the predicted prosodic features include predicted duration and predicted fundamental frequency; The synthesis module is used to synthesize a target emotion speech signal based on the target emotion tag, the second discrete articulation unit sequence, the predicted prosodic features, and the speaker identity tag of the original speech signal; The first discrete pronunciation unit sequence includes language pronunciation units and a first non-language pronunciation unit. The speech emotion conversion system includes an emotion conversion model. The step of converting the first discrete pronunciation unit sequence into a second discrete pronunciation unit sequence with a target emotion based on a target emotion label includes: based on the target emotion label, using the emotion conversion model to convert the first non-language pronunciation unit in the first discrete pronunciation unit sequence with the original emotion into the second non-language pronunciation unit with the target emotion through conversion processing, while retaining the language pronunciation unit, to obtain the second discrete pronunciation unit sequence. The conversion processing includes insertion, replacement, and deletion.
8. A computer device comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, characterized in that, When the processor executes the computer-readable instructions, it performs the steps of the voice emotion conversion method as described in any one of claims 1-6.
9. A computer-readable storage medium storing computer-readable instructions thereon, characterized in that, When the computer-readable instructions are executed by a processor, the processor performs the steps of the speech emotion conversion method as described in any one of claims 1-6.
Citation Information
Patent Citations
Voice synthesis method and device, computer equipment and computer readable storage medium
CN111108549A
Voice processing method and device, computer readable storage medium and electronic device
CN111883098A