Spectrum intensity adjustable speech synthesis methods, devices, computer equipment and media
By perturbing the formant frequencies and fusing features in the original speech spectrum, the problem of insufficient capture of emotional intensity differences in emotional speech synthesis is solved, achieving accuracy in emotional intensity and naturalness in speech, and can be applied to intelligent customer service scenarios in the financial sector.
Patent Information
- Application Number
- CN202411719880.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-11-26
AI Technical Summary
Current emotional speech synthesis models cannot capture the differences in emotional intensity when different speakers use the same emotion, resulting in low emotional accuracy of synthesized speech.
By acquiring the spectrum of the original speech, randomly perturbing the frequency of the formants, calculating the correlation between the perturbed spectrum features and the preset intensity matrix, determining the weight of the intensity matrix, and performing feature fusion to synthesize the target speech, the speaker information is destroyed while the emotional intensity information is preserved.
It improves the accuracy of emotional intensity in target speech, synthesizes natural and expressive speech, and enhances the service quality and efficiency of financial services.
Smart Images

Figure CN119559932B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence, and in particular to a speech synthesis method, apparatus, computer equipment, and medium with adjustable spectral intensity. Background Technology
[0002] The goal of speech synthesis technology is to transform text information into understandable, clear, natural, expressive, and near-human speech signals. Since human speech carries linguistic and emotional information, modeling emotional information is crucial for synthesizing natural and expressive speech.
[0003] In current research on emotional speech synthesis, most existing models learn the average emotional features of all speakers in the dataset. However, the emotional intensity of each speaker when using the same emotion to express the same semantics is often different. Existing models cannot capture this difference, resulting in low reliability of emotional intensity information and low emotional accuracy of synthesized speech.
[0004] Therefore, in the field of speech synthesis technology, how to improve the emotional accuracy of synthesized speech has become an urgent problem to be solved. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a speech synthesis method, apparatus, computer device and medium with adjustable spectral intensity to solve the problem that the reliability of emotional intensity information is low in current research on emotional speech synthesis, resulting in low emotional accuracy of synthesized speech.
[0006] In a first aspect, embodiments of the present invention provide a speech synthesis method with adjustable spectral intensity, the speech synthesis method comprising:
[0007] Obtain the original spectrum of the original speech, determine K formants in the original spectrum, randomly perturb the frequencies of the K formants to obtain the perturbation frequencies of the corresponding formants, and replace the frequencies of the corresponding formants in the original spectrum with the perturbation frequencies of each formant to obtain the perturbation spectrum, where K is an integer greater than 0.
[0008] The perturbation spectrum is feature-encoded to obtain perturbation spectrum features. The correlation between the perturbation spectrum features and N preset intensity matrices is calculated. Based on each correlation, the weight of the corresponding intensity matrix is determined, where N is an integer greater than 0.
[0009] The N intensity matrices are weighted and summed according to their respective weights to obtain intensity features. The intensity features and the disturbance spectrum features are then fused to obtain spectral intensity features.
[0010] The target text and target speech features to be synthesized are obtained. Based on the spectral intensity features, speech synthesis is performed on the target text. The synthesized speech is adjusted using the target speech features to obtain the corresponding target speech.
[0011] Secondly, embodiments of the present invention provide a speech synthesis device with adjustable spectral intensity, the speech synthesis device comprising:
[0012] The frequency perturbation module is used to acquire the original spectrum of the original speech, determine K formants in the original spectrum, randomly perturb the frequencies of the K formants to obtain the perturbation frequencies of the corresponding formants, and replace the frequencies of the corresponding formants in the original spectrum with the perturbation frequencies of each formant to obtain the perturbation spectrum, where K is an integer greater than 0.
[0013] The weight determination module is used to perform feature encoding on the disturbance spectrum to obtain disturbance spectrum features, calculate the correlation between the disturbance spectrum features and N preset intensity matrices, and determine the weight of the corresponding intensity matrix based on each correlation, where N is an integer greater than 0.
[0014] The feature fusion module is used to perform a weighted summation of the N intensity matrices according to the weights corresponding to the N intensity matrices to obtain intensity features, and to perform feature fusion of the intensity features and the disturbance spectrum features to obtain spectral intensity features;
[0015] The speech synthesis module is used to acquire the target text to be synthesized and the target speech features, perform speech synthesis on the target text according to the spectral intensity features, and adjust the synthesized speech using the target speech features to obtain the corresponding target speech.
[0016] Thirdly, embodiments of the present invention provide a computer device, the computer device including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the speech synthesis method as described in the first aspect.
[0017] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the speech synthesis method as described in the first aspect.
[0018] The beneficial effects of this invention compared to the prior art are as follows: This invention acquires the original spectrum of the original speech, determines K formants in the original spectrum, randomly perturbs the frequencies of the K formants to obtain the perturbation frequencies of the corresponding formants, replaces the frequencies of the corresponding formants in the original spectrum with the perturbation frequencies of each formant to obtain the perturbed spectrum, performs feature encoding on the perturbed spectrum to obtain the perturbed spectrum features, calculates the correlation between the perturbed spectrum features and N preset intensity matrices, determines the weights of the corresponding intensity matrices based on each correlation, performs a weighted summation of the N intensity matrices based on their weights to obtain the intensity features, performs feature fusion on the intensity features and the perturbed spectrum features to obtain the spectral intensity features, and then performs feature fusion on the intensity features and the perturbed spectrum features. Feature fusion is performed on dynamic spectral features to obtain spectral intensity features. The target text and target speech features to be synthesized are then obtained. Based on the spectral intensity features, speech synthesis is performed on the target text. The synthesized speech is adjusted using the target speech features to obtain the corresponding target speech. By perturbing the formants, the speaker information represented in the original spectrum is destroyed, and only the effective emotional information and emotional intensity information in the original spectrum are retained. The weight of the corresponding intensity matrix is determined by calculating the correlation between the perturbed spectral features and the preset N intensity matrices, which improves the accuracy of emotional intensity in the target speech. In the intelligent customer service scenario in the financial field, natural and expressive target speech is synthesized for robot customer service, improving the service quality and efficiency of financial business. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of an application environment for a speech synthesis method with adjustable spectral intensity provided in Embodiment 1 of the present invention;
[0021] Figure 2 This is a flowchart illustrating a speech synthesis method with adjustable spectral intensity provided in Embodiment 1 of the present invention;
[0022] Figure 3 This is a schematic diagram of the structure of a speech synthesis device with adjustable spectral intensity provided in Embodiment 2 of the present invention;
[0023] Figure 4 This is a schematic diagram of the structure of a computer device provided in Embodiment 3 of the present invention. Detailed Implementation
[0024] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0025] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0026] It should also be understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0027] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0028] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0029] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0030] The embodiments of this invention can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0031] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0032] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0033] To illustrate the technical solution of the present invention, specific embodiments are described below.
[0034] The first embodiment of this invention provides a speech synthesis method with adjustable spectral intensity, which can be applied to, for example... Figure 1 In this application environment, the client communicates with the server. Clients include, but are not limited to, handheld computers, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, and personal digital assistants (PDAs). The server can be implemented using a standalone server or a server cluster. This speech synthesis method can be applied to various fields such as the internet, finance, healthcare, education, and transportation. For example, due to the complexity and diversity of financial business, simple tasks such as numerous inquiries and after-sales services can severely consume the time and energy of sales personnel, reducing their work efficiency and quality. Adopting an intelligent conversational approach based on automatic speech synthesis can save significant labor costs. Furthermore, by effectively controlling the content of the voice conversation, it can provide customers with more natural, expressive, and richer voice customer service, thereby improving customer experience and ultimately enhancing the service quality of financial services.
[0035] See Figure 2 This is a flowchart illustrating a speech synthesis method with adjustable spectral intensity provided in Embodiment 1 of the present invention. The above speech synthesis method can be applied to... Figure 1In a client application, the speech synthesis method may include the following steps:
[0036] Step S201: Obtain the original spectrum of the original speech, determine the K formants in the original spectrum, randomly perturb the frequencies of the K formants to obtain the perturbation frequencies of the corresponding formants, and replace the frequencies of the corresponding formants in the original spectrum with the perturbation frequencies of each formant to obtain the perturbation spectrum.
[0037] Among them, the goal of speech synthesis technology is to transform text information into understandable, clear, natural, expressive, and near-human speech signals. Since human speech carries linguistic and emotional information, modeling emotional information is crucial for synthesizing natural and expressive speech.
[0038] The spectrum of speech reflects the distribution of energy and frequency of the speech signal in the frequency domain, revealing important speech features. Formants, located at the concentration of acoustic energy, are envelope peaks. In a two-dimensional spectrum, each envelope contains harmonics with significant amplitude, representing the peak value of the formants. As points of concentrated acoustic energy, formants are crucial features reflecting the resonance characteristics of the vocal tract and represent the most direct source of pronunciation information. Humans extensively utilize spectral and formant information in speech perception; therefore, the spectrum and formants, as vital feature parameters in speech signal processing, are widely used as key features for speech recognition and fundamental information for speech coding and transmission.
[0039] In current clinical applications, acoustic detection parameters for voice include fundamental frequency, intensity, harmonic-to-noise ratio, frequency perturbation, amplitude perturbation, formants, contact rate, contact rate perturbation, contact power, and contact power perturbation. Using these acoustic detection parameters, an objective acoustic evaluation of the speaker's voice can be performed. Therefore, when the frequency of the formants in the spectrum changes, the speaker information represented by the spectrum changes accordingly. Thus, in order to achieve the goal of speech synthesis with controllable emotional intensity, this embodiment randomly perturbs the frequencies of K formants in the original spectrum of the original speech to destroy the speaker information represented by the original spectrum, retaining only the effective emotional information and emotional intensity information in the original spectrum, thereby improving the reliability of the emotional and intensity information.
[0040] Specifically, in this embodiment, the frequencies of K formants in the original spectrum of the original speech are randomly perturbed to obtain the perturbed frequencies of the corresponding formants. The perturbed frequencies of each formant are then used to replace the frequencies of the corresponding formants in the original spectrum, thereby destroying the speaker information represented in the original spectrum and retaining only the effective emotional information and emotional intensity information. This perturbed spectrum is then obtained as a spectrum unrelated to the speaker but related to emotion, effectively representing the emotional information and emotional intensity information in the original spectrum. Here, K is an integer greater than 0. For example, in the intelligent customer service scenario in the financial field, the original speech can be the original speech of a human or robot customer service representative interacting with a customer. This serves as the basis for synthesizing speech with controllable emotional intensity, improving the naturalness and expressiveness of the synthesized speech from the robot customer service representative, thereby improving the service quality of financial services.
[0041] Optionally, determining the K resonance peaks in the original spectrum includes:
[0042] Obtain the spectral envelope of the original spectrum, identify the K local peaks in the spectral envelope, and determine the K local peaks as the K resonance peaks in the original spectrum.
[0043] The formant information is contained in the spectral envelope. It is generally believed that the maximum value in the spectral envelope is the formant. Therefore, in this embodiment, the spectral envelope of the original spectrum is obtained, and the K local peak values of the frequency in the spectral envelope are determined according to the frequency values at various points in the spectral envelope. The K local peak values are determined as the K formants in the original spectrum.
[0044] In this embodiment, K local peaks in the spectral envelope of the original spectrum are identified as K resonance peaks in the original spectrum, thereby improving the accuracy of resonance peak extraction.
[0045] Optionally, the frequencies of the K resonance peaks are randomly perturbed to obtain the perturbed frequencies of the corresponding resonance peaks, including:
[0046] For any formant, the original frequency of the formant is determined based on the spectrum of the original speech. Interference frequencies are randomly obtained within a preset frequency range. The sum of the original frequency and the interference frequency is determined as the perturbation frequency of the formant. The preset frequency range is [0, F], where F is the fundamental frequency of the spectrum.
[0047] By iterating through all the resonance peaks, the perturbation frequencies of all resonance peaks are obtained.
[0048] Wherein, the fundamental frequency is the frequency of the fundamental tone in the spectrum, which can be regarded as the basic pitch of the sound. Therefore, in this embodiment, when perturbing the formant, the preset frequency range of randomly selected interference frequency is set to [0, F], where F is the fundamental frequency of the spectrum.
[0049] For any resonance peak, the original frequency of the resonance peak is determined, and the interference frequency is randomly obtained within the preset frequency range. The sum of the original frequency and the interference frequency is determined as the perturbation frequency of the resonance peak, which serves as the frequency basis for perturbing the resonance peak.
[0050] In this embodiment, the preset frequency range for randomly selecting interference frequencies is set to [0, F], where F is the fundamental frequency of the spectrum, which improves the rationality of perturbing the resonance peak.
[0051] The steps described above—obtaining the original spectrum of the original speech, identifying K formants in the original spectrum, randomly perturbing the frequencies of the K formants to obtain the perturbed frequencies of the corresponding formants, and replacing the frequencies of the corresponding formants in the original spectrum with the perturbed frequencies of each formant to obtain the perturbed spectrum—destroy the speaker information represented in the original spectrum by perturbing the formants, retaining only the effective emotional information and emotional intensity information in the original spectrum, thus improving the reliability of emotional and intensity information.
[0052] Step S202: Perform feature encoding on the disturbance spectrum to obtain disturbance spectrum features, calculate the correlation between the disturbance spectrum features and the preset N intensity matrices, and determine the weight of the corresponding intensity matrix based on each correlation.
[0053] The preset N intensity matrices can be used to measure the intensity of emotion in the corresponding perturbation spectral features. N is an integer greater than 0. For example, N=3. The preset three intensity matrices are S, M and L. S represents the low intensity of emotion in the corresponding perturbation spectral features and is used to synthesize speech with low emotion intensity; M represents the medium intensity of emotion in the corresponding perturbation spectral features and is used to synthesize speech with medium emotion intensity; L represents the high intensity of emotion in the corresponding perturbation spectral features and is used to synthesize speech with high emotion intensity.
[0054] To further determine the emotional information contained in the perturbation spectrum and improve the accuracy of emotional intensity in the synthesized speech, this embodiment first performs feature encoding on the perturbation spectrum to obtain perturbation spectrum features, and calculates the correlation between the perturbation spectrum features and N preset intensity matrices. Based on each correlation, the weight of the corresponding intensity matrix is determined. Combining the N intensity matrices and their corresponding N weights improves the accuracy of emotional intensity in the synthesized speech. Correspondingly, the higher the correlation between the intensity matrix and the perturbation spectrum features, the larger the weight of the intensity matrix, thus dominating the emotional intensity in the synthesized speech.
[0055] Optionally, the correlation between the perturbation spectral characteristics and N preset intensity matrices is calculated, and the weights of the corresponding intensity matrices are determined based on each correlation, including:
[0056] Using a trained attention mechanism model, the correlation between the perturbation spectral features and the preset N intensity matrices is determined;
[0057] The correlation between the perturbation spectral characteristics and the corresponding intensity matrix is determined as the weight of the corresponding intensity matrix.
[0058] The attention mechanism model is an information filtering method. It introduces a task-relevant query vector as a benchmark for feature selection, uses a scoring function to calculate the correlation between input features and the query vector, obtains the probability distribution of selected input features, and finally filters out task-relevant features by weighted averaging of the input features according to the probability distribution. There are four common forms of scoring functions: additive model, dot product model, scaled dot product model, and bilinear model, which can be selected according to the specific circumstances.
[0059] Correspondingly, in this embodiment, the query vector is the perturbation spectrum feature, and the input feature is the preset N intensity matrices. The correlation between the perturbation spectrum feature and the preset N intensity matrices is calculated according to the scoring function, and the correlation between the perturbation spectrum feature and the corresponding intensity matrix is determined as the weight of the corresponding intensity matrix, which serves as the basis for filtering the intensity features in the perturbation spectrum feature.
[0060] This embodiment uses a trained attention mechanism model to determine the correlation between perturbation spectral features and N preset intensity matrices, and uses the correlation between perturbation spectral features and corresponding intensity matrices as the weights of the corresponding intensity matrices, thereby improving the accuracy of intensity information extraction.
[0061] Optionally, the training process of the attention mechanism model includes:
[0062] Obtain N preset intensity matrices, several original speech samples, M perturbation spectral feature samples corresponding to each original speech sample, and the target text sample and target speech feature sample to be synthesized, where M is an integer greater than 0;
[0063] For any perturbation spectral feature sample, an attention mechanism model is used to determine the correlation between the perturbation spectral feature sample and N intensity matrices. The correlation between the perturbation spectral feature sample and the corresponding intensity matrix is used to determine the weight sample of the corresponding intensity matrix.
[0064] The intensity feature samples are obtained by weighted summing of the N intensity matrices based on the weighted samples corresponding to the N intensity matrices.
[0065] The perturbation spectral feature sample and intensity feature sample are fused to obtain the spectral intensity feature sample. Based on the spectral intensity feature sample, the target text sample is used for speech synthesis. The synthesized speech is adjusted using the target speech features to obtain the corresponding target speech sample.
[0066] Calculate the emotional similarity between the original speech sample and the target speech sample corresponding to the perturbation spectrum feature sample, and determine the difference between the emotional similarity and the preset value as the corresponding model sub-loss;
[0067] By iterating through all the perturbation spectral feature samples, we obtain all the model sub-losses, calculate the sum of all the model sub-losses, and obtain the model loss.
[0068] The parameters of the attention mechanism model are adjusted based on the model loss until the model loss converges, resulting in a well-trained attention mechanism model.
[0069] To improve the accuracy of the weights corresponding to the intensity matrices, the attention mechanism model needs to be trained. In this embodiment, N preset intensity matrices, several original speech samples, M perturbation spectral feature samples corresponding to each original speech sample, and the target text sample and target speech feature sample to be synthesized are used as training samples, and the original speech samples are used as training labels to train the attention mechanism model.
[0070] Specifically, an attention mechanism model is used to determine the correlation between perturbation spectral feature samples and N intensity matrices. The correlation between the perturbation spectral feature samples and the corresponding intensity matrices is used to determine the weight samples of the corresponding intensity matrices. The N intensity matrices are then weighted and summed according to the weight samples corresponding to the N intensity matrices to obtain intensity feature samples, which represent the emotional intensity information in the original speech. Feature fusion is performed on the perturbation spectral feature samples and intensity feature samples to obtain spectral intensity feature samples, which represent the emotion and emotional intensity information in the original speech. Based on the spectral intensity feature samples, speech synthesis is performed on the target text sample. The synthesized speech is then adjusted using target speech features to obtain the corresponding target speech sample.
[0071] The higher the accuracy of the attention mechanism model, the higher the emotional similarity between the target synthesized speech and the corresponding original speech. Therefore, the emotional similarity between the original speech sample and the target speech sample corresponding to the perturbation spectrum feature sample is calculated, and the difference between the emotional similarity and the preset value is determined as the corresponding model sub-loss. The preset value can be set according to the actual situation. Correspondingly, the larger the sub-loss, the lower the accuracy of the attention mechanism model.
[0072] Then, iterate through all perturbation spectral feature samples, calculate the sum of all model sub-losses to obtain the model loss, and correct the parameters of the attention mechanism model according to the model loss until the model loss converges to obtain the trained attention mechanism model, which is used to determine the correlation between perturbation spectral features and the preset N intensity matrices to improve the accuracy of the weights of the corresponding intensity matrices.
[0073] In this embodiment, N preset intensity matrices, several original speech samples, M perturbation spectral feature samples corresponding to each original speech sample, and the target text sample and target speech feature sample to be synthesized are used as training samples. The original speech samples are used as training labels to train the attention mechanism model. The trained attention mechanism model is used to determine the correlation between the perturbation spectral features and the N preset intensity matrices, thereby improving the accuracy of the weights of the corresponding intensity matrices.
[0074] Optionally, obtaining the M perturbation spectral feature samples corresponding to each original speech sample includes:
[0075] For any original speech sample, obtain the original spectrum sample of the original speech sample, determine the K formants in the original spectrum sample, randomly perturb the frequencies of the K formants to obtain the perturbation frequencies of the corresponding formants, and replace the frequencies of the corresponding formants in the original spectrum sample with the perturbation frequencies of each formant to obtain the perturbed spectrum sample.
[0076] Repeat the above steps of randomly perturbing the frequencies of the K formants M times to obtain the perturbed frequencies of the corresponding formants. Replace the frequencies of the corresponding formants in the original spectrum samples with the perturbed frequencies of each formant to obtain perturbed spectrum samples. This will result in M perturbed spectrum samples corresponding to the original speech sample.
[0077] Encode each of the M perturbation spectrum samples to obtain the corresponding M perturbation spectrum feature samples;
[0078] By iterating through all the original speech samples, we obtain M perturbation spectral feature samples corresponding to each original speech sample.
[0079] To enhance the diversity of training samples, this embodiment randomly perturbs the frequencies of the K formants of any given original speech sample M times, resulting in M perturbated spectrum samples. These M perturbated spectrum samples are then encoded to obtain M corresponding perturbated spectrum feature samples. These M perturbated spectrum feature samples can be combined with N preset intensity matrices, the target text sample to be synthesized, and the target speech feature sample as training samples to train the attention mechanism model. This approach provides a large number of training samples from a small number of original speech samples, improving the training accuracy of the attention mechanism model.
[0080] In this embodiment, for any original speech sample, the frequencies of the K formants of the corresponding original spectrum sample are randomly perturbed M times, and the corresponding M perturbed spectrum feature samples are obtained. Based on a small number of original speech samples, a large number of training samples are obtained, which improves the training accuracy of the attention mechanism model.
[0081] Optionally, calculating the emotional similarity between the original speech sample and the target speech sample corresponding to the perturbation spectral feature sample includes:
[0082] The trained emotion encoder is used to encode the original speech sample and the target speech sample to obtain the first emotion feature sequence corresponding to the original speech sample and the second emotion feature sequence corresponding to the target speech sample.
[0083] Calculate the similarity between the first and second emotion feature sequences, and use it as the emotion similarity between the original speech sample and the target speech sample.
[0084] The accuracy of the attention mechanism model is measured by the emotional similarity between the target synthesized speech and the corresponding original speech. In order to accurately calculate the emotional similarity, this embodiment uses a trained emotion encoder to encode the original speech sample and the target speech sample to obtain the first emotional feature sequence of the corresponding original speech sample and the second emotional feature sequence of the corresponding target speech sample. The similarity between the first emotional feature sequence and the second emotional feature sequence can be used as the emotional similarity between the original speech sample and the target speech sample.
[0085] This embodiment uses a trained emotion encoder to encode the original speech sample and the target speech sample, obtaining a first emotion feature sequence corresponding to the original speech sample and a second emotion feature sequence corresponding to the target speech sample. The calculation of the emotion similarity between the original speech sample and the target speech sample is transformed into the calculation of the similarity between the corresponding first emotion feature sequence and the second emotion feature sequence, which improves the reliability and accuracy of the emotion similarity calculation.
[0086] The steps described above—encoding the perturbation spectrum to obtain perturbation spectrum features, calculating the correlation between the perturbation spectrum features and N preset intensity matrices, and determining the weight of the corresponding intensity matrix based on each correlation—in order to improve the accuracy of emotional intensity in the synthesized speech by combining the N intensity matrices and their corresponding N weights, thereby determining the weight of the corresponding intensity matrix through the calculation of the correlation between the perturbation spectrum features and the N preset intensity matrices.
[0087] Step S203: The N intensity matrices are weighted and summed according to their corresponding weights to obtain the intensity features. The intensity features and the disturbance spectrum features are then fused to obtain the spectrum intensity features.
[0088] The weights corresponding to the intensity matrices represent the degree to which the intensity matrix dominates the emotional intensity of the speech to be synthesized. Therefore, the N intensity matrices are weighted and summed according to their weights to obtain intensity features, which represent the emotional intensity information in the speech to be synthesized. Furthermore, the intensity features and perturbation spectral features are fused to obtain spectral intensity features, which simultaneously represent both the emotional information and the emotional intensity information in the speech to be synthesized.
[0089] The above steps, which involve weighted summation of the N intensity matrices based on their corresponding weights to obtain intensity features, and feature fusion of the intensity features and perturbation spectral features to obtain spectral intensity features, improve the accuracy of emotional information and emotional intensity information in the synthesized speech.
[0090] Step S204: Obtain the target text and target speech features to be synthesized; synthesize speech from the target text based on the spectral intensity features; adjust the synthesized speech using the target speech features to obtain the corresponding target speech.
[0091] In this process, the target text to be synthesized is used to represent the text information in the speech to be synthesized, the target speech features to be synthesized can be used to represent the timbre information in the speech to be synthesized, and the spectral intensity features can be used to represent the emotional information and emotional intensity information in the speech to be synthesized. Then, the target text can be synthesized based on the spectral intensity features to synthesize speech containing text information, emotional information, and emotional intensity information. Then, the synthesized speech is adjusted using the target speech features to give the synthesized speech target timbre information, thus obtaining the corresponding target speech.
[0092] For example, in the intelligent customer service scenario in the financial field, the target text can be the script text, and the target speech feature can be the extracted voice information of human customer service. Then, by combining the spectral intensity features, the target text, and the target speech features, a natural and expressive target speech can be synthesized for the robot customer service, thereby assisting human customer service in communicating with customers and improving the service quality and efficiency of financial services.
[0093] The steps described above—obtaining the target text and target speech features to be synthesized, performing speech synthesis on the target text based on spectral intensity features, adjusting the synthesized speech using the target speech features to obtain the corresponding target speech—increase the emotional accuracy of the synthesized target speech.
[0094] This invention acquires the original spectrum of the original speech, identifies K formants in the original spectrum, randomly perturbs the frequencies of the K formants to obtain the perturbation frequencies of the corresponding formants, replaces the frequencies of the corresponding formants in the original spectrum with the perturbation frequencies of each formant to obtain the perturbed spectrum, performs feature encoding on the perturbed spectrum to obtain the perturbed spectrum features, calculates the correlation between the perturbed spectrum features and N preset intensity matrices, determines the weights of the corresponding intensity matrices based on each correlation, performs a weighted summation of the N intensity matrices according to their corresponding weights to obtain the intensity features, performs feature fusion on the intensity features and the perturbed spectrum features to obtain the spectral intensity features, and performs feature fusion on the intensity features and the perturbed spectrum features. The process involves obtaining spectral intensity features, acquiring target text and target speech features, synthesizing speech based on the spectral intensity features, adjusting the synthesized speech using the target speech features, and obtaining the corresponding target speech. By perturbing the formants, the speaker information represented in the original spectrum is destroyed, retaining only the effective emotional information and emotional intensity information in the original spectrum. The weights of the corresponding intensity matrices are determined by calculating the correlation between the perturbed spectral features and N preset intensity matrices, thus improving the accuracy of emotional intensity in the target speech. In the intelligent customer service scenario in the financial field, this process synthesizes natural and expressive target speech for robot customer service, improving the service quality and efficiency of financial services.
[0095] Corresponding to the speech synthesis method in the above embodiments, Figure 3 A structural block diagram of the speech synthesis device with adjustable spectral intensity provided in Embodiment 2 of the present invention is given. For ease of explanation, only the parts related to the embodiments of the present invention are shown.
[0096] See Figure 3 The speech synthesis device includes:
[0097] The frequency perturbation module 31 is used to acquire the original spectrum of the original speech, determine the K formants in the original spectrum, randomly perturb the frequencies of the K formants to obtain the perturbation frequencies of the corresponding formants, and replace the frequencies of the corresponding formants in the original spectrum with the perturbation frequencies of each formant to obtain the perturbation spectrum, where K is an integer greater than 0.
[0098] The weight determination module 32 is used to perform feature encoding on the disturbance spectrum to obtain disturbance spectrum features, calculate the correlation between the disturbance spectrum features and N preset intensity matrices, and determine the weight of the corresponding intensity matrix based on each correlation, where N is an integer greater than 0.
[0099] The feature fusion module 33 is used to perform a weighted summation of the N intensity matrices according to the weights corresponding to the N intensity matrices to obtain intensity features, and to perform feature fusion of the intensity features and the perturbation spectrum features to obtain spectral intensity features.
[0100] The speech synthesis module 34 is used to acquire the target text to be synthesized and the target speech features, perform speech synthesis on the target text based on the spectral intensity features, and adjust the synthesized speech using the target speech features to obtain the corresponding target speech.
[0101] Optionally, the frequency disturbance module 31 mentioned above includes:
[0102] The first frequency perturbation submodule is used to determine the original frequency of any formant based on the spectrum of the original speech, randomly acquire the interference frequency within a preset frequency range, and determine the perturbation frequency of the formant by the sum of the original frequency and the interference frequency. The preset frequency range is [0, F], where F is the fundamental frequency of the spectrum.
[0103] The second frequency perturbation submodule is used to traverse all resonance peaks and obtain the perturbation frequencies of all resonance peaks.
[0104] Optionally, the aforementioned weight determination module 32 includes:
[0105] The correlation determination submodule is used to determine the correlation between the perturbation spectral features and the preset N intensity matrices using a trained attention mechanism model.
[0106] The weight determination submodule is used to determine the correlation between the perturbation spectral features and the corresponding intensity matrix as the weight of the corresponding intensity matrix.
[0107] Optionally, the aforementioned correlation determination submodule includes:
[0108] The data acquisition unit is used to acquire N preset intensity matrices, several original speech samples, M perturbation spectrum feature samples corresponding to each original speech sample, and target text samples and target speech feature samples to be synthesized, where M is an integer greater than 0;
[0109] The weight determination unit is used to determine the correlation between the perturbation spectrum feature sample and N intensity matrices for any perturbation spectrum feature sample using an attention mechanism model, and to determine the weight sample of the corresponding intensity matrix based on the correlation between the perturbation spectrum feature sample and the corresponding intensity matrix.
[0110] The weighted summation unit is used to sum the N intensity matrices based on the weight samples corresponding to the N intensity matrices to obtain intensity feature samples;
[0111] The speech synthesis unit is used to fuse perturbation spectral feature samples and intensity feature samples to obtain spectral intensity feature samples, perform speech synthesis on target text samples based on spectral intensity feature samples, and adjust the synthesized speech using target speech features to obtain the corresponding target speech sample.
[0112] The sub-loss calculation unit is used to calculate the emotional similarity between the original speech sample and the target speech sample corresponding to the perturbation spectrum feature sample, and to determine the difference between the emotional similarity and the preset value as the corresponding model sub-loss.
[0113] The model loss calculation unit is used to traverse all perturbation spectrum feature samples, obtain all model sub-losses, calculate the sum of all model sub-losses, and obtain the model loss.
[0114] The parameter correction unit is used to correct the parameters of the attention mechanism model based on the model loss until the model loss converges, thus obtaining the trained attention mechanism model.
[0115] Optionally, the above data acquisition unit includes:
[0116] The first frequency perturbation subunit is used to obtain the original spectrum sample of any original speech sample, determine the K formants in the original spectrum sample, randomly perturb the frequencies of the K formants to obtain the perturbation frequency of the corresponding formant, and replace the frequency of the corresponding formant in the original spectrum sample with the perturbation frequency of each formant to obtain the perturbation spectrum sample.
[0117] The second frequency perturbation subunit is used to repeatedly perform the above steps of randomly perturbing the frequencies of K formants M times to obtain the perturbation frequency of the corresponding formant, and replacing the frequency of the corresponding formant in the original spectrum sample with the perturbation frequency of each formant to obtain the perturbation spectrum sample, thereby obtaining M perturbation spectrum samples corresponding to the original speech sample.
[0118] The feature encoding subunit is used to encode the M perturbation spectrum samples respectively to obtain the corresponding M perturbation spectrum feature samples;
[0119] The data acquisition subunit is used to traverse all the original speech samples and obtain M perturbation spectral feature samples corresponding to each original speech sample.
[0120] Optionally, the above-mentioned sub-loss calculation unit includes:
[0121] The feature encoding subunit is used to encode the original speech sample and the target speech sample using the trained emotion encoder to obtain the first emotion feature sequence corresponding to the original speech sample and the second emotion feature sequence corresponding to the target speech sample.
[0122] The similarity calculation subunit is used to calculate the similarity between the first emotional feature sequence and the second emotional feature sequence, which is used as the emotional similarity between the original speech sample and the target speech sample.
[0123] Optionally, the frequency disturbance module 31 mentioned above includes:
[0124] The formant extraction submodule is used to obtain the spectral envelope of the original spectrum, determine the K local peaks in the spectral envelope, and identify the K local peaks as the K formants in the original spectrum.
[0125] It should be noted that the information interaction and execution process between the above modules are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.
[0126] Figure 4 This is a schematic diagram of the structure of a computer device provided in Embodiment 3 of the present invention. Figure 4 As shown, the computer device of this embodiment includes: at least one processor ( Figure 4 Only one is shown in the diagram), a memory, and a computer program stored in the memory and executable on at least one processor, which, when executed by the processor, implements the steps in any of the above-described speech synthesis method embodiments.
[0127] This computer device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 4 The examples of computer devices are merely examples and do not constitute a limitation on computer devices. Computer devices may include more or fewer components than shown in the illustration, or combinations of certain components, or different components, such as network interfaces, displays, and input devices.
[0128] The processor referred to can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0129] Memory includes readable storage media, internal memory, etc., wherein internal memory can be the RAM of a computer device, providing an environment for the operation of the operating system and computer-readable instructions stored in the readable storage media. The readable storage media can be the hard drive of a computer device, or in other embodiments, it can be an external storage device of the computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, memory can include both internal storage units and external storage devices of the computer device. Memory is used to store the operating system, applications, bootloader, data, and other programs, such as program code for computer programs. Memory can also be used to temporarily store data that has been output or will be output.
[0130] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the functions described above can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the processes in the methods of the above embodiments by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code, a recording medium, a computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0131] The present invention can implement all or part of the processes in the methods of the above embodiments, or it can be accomplished by a computer program product. When the computer program product is run on a computer device, the computer device executes the steps in the above method embodiments.
[0132] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0133] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0134] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0135] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0136] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A speech synthesis method with adjustable spectral intensity, characterized in that, The speech synthesis method includes: The original spectrum of the original speech is obtained, K formants in the original spectrum are determined, the frequencies of the K formants are randomly perturbed to obtain the perturbed frequencies of the corresponding formants, and the frequencies of the corresponding formants in the original spectrum are replaced by the perturbed frequencies of each formant to obtain the perturbed spectrum. K is an integer greater than 0. The perturbed spectrum is a spectrum that is independent of the speaker but related to emotion, and is used to characterize the emotional information and the intensity information of emotion in the original spectrum. The perturbation spectrum is feature-encoded to obtain perturbation spectrum features. The correlation between the perturbation spectrum features and N preset intensity matrices is calculated. Based on each correlation, the weight of the corresponding intensity matrix is determined. N is an integer greater than 0. The intensity matrix is used to measure the intensity information of emotion in the corresponding perturbation spectrum features. The N intensity matrices are weighted and summed according to their respective weights to obtain intensity features. The intensity features and the disturbance spectrum features are then fused to obtain spectral intensity features. The target text and target speech features to be synthesized are obtained. Based on the spectral intensity features, speech synthesis is performed on the target text. The synthesized speech is adjusted using the target speech features to obtain the corresponding target speech.
2. The speech synthesis method according to claim 1, characterized in that, The step of randomly perturbing the frequencies of the K resonance peaks to obtain the perturbed frequencies of the corresponding resonance peaks includes: For any formant, the original frequency of the formant is determined based on the spectrum of the original speech. An interference frequency is randomly obtained within a preset frequency range. The sum of the original frequency and the interference frequency is determined as the perturbation frequency of the formant. The preset frequency range is [0, F], where F is the fundamental frequency of the spectrum. By iterating through all the resonance peaks, the perturbation frequencies of all resonance peaks are obtained.
3. The speech synthesis method according to claim 1, characterized in that, The calculation of the correlation between the disturbance spectrum features and the preset N intensity matrices, and the determination of the weight of the corresponding intensity matrix based on each correlation, includes: Using a trained attention mechanism model, the correlation between the perturbation spectral features and the preset N intensity matrices is determined; The correlation between the perturbation spectral features and the corresponding intensity matrix is determined as the weight of the corresponding intensity matrix.
4. The speech synthesis method according to claim 3, characterized in that, The training process of the attention mechanism model includes: Obtain N preset intensity matrices, several original speech samples, M perturbation spectral feature samples corresponding to each original speech sample, and the target text sample and target speech feature sample to be synthesized, where M is an integer greater than 0; For any perturbation spectrum feature sample, an attention mechanism model is used to determine the correlation between the perturbation spectrum feature sample and the N intensity matrices, and the weight sample of the corresponding intensity matrix is determined by the correlation between the perturbation spectrum feature sample and the corresponding intensity matrix. The N intensity matrices are weighted and summed based on the weighted samples corresponding to the N intensity matrices to obtain intensity feature samples; The perturbation spectral feature sample and the intensity feature sample are fused to obtain a spectral intensity feature sample. Based on the spectral intensity feature sample, the target text sample is used for speech synthesis. The synthesized speech is then adjusted using the target speech features to obtain the corresponding target speech sample. Calculate the emotional similarity between the original speech sample and the target speech sample corresponding to the perturbation spectral feature sample, and determine the difference between the emotional similarity and the preset value as the corresponding model sub-loss; By iterating through all the perturbation spectral feature samples, we obtain all the model sub-losses, calculate the sum of all the model sub-losses, and obtain the model loss. The parameters of the attention mechanism model are adjusted based on the model loss until the model loss converges, resulting in a trained attention mechanism model.
5. The speech synthesis method according to claim 4, characterized in that, The M perturbation spectral feature samples corresponding to each original speech sample are obtained as follows: For any original speech sample, obtain the original spectrum sample of the original speech sample, determine K formants in the original spectrum sample, randomly perturb the frequencies of the K formants to obtain the perturbation frequencies of the corresponding formants, and replace the frequencies of the corresponding formants in the original spectrum sample with the perturbation frequencies of each formant to obtain the perturbed spectrum sample. Repeat the above steps of randomly perturbing the frequencies of the K formants M times to obtain the perturbed frequencies of the corresponding formants, and replace the frequencies of the corresponding formants in the original spectrum sample with the perturbed frequencies of each formant to obtain perturbed spectrum samples, thereby obtaining M perturbed spectrum samples corresponding to the original speech sample. The M perturbation spectrum samples are encoded respectively to obtain the corresponding M perturbation spectrum feature samples; By iterating through all the original speech samples, we obtain M perturbation spectral feature samples corresponding to each original speech sample.
6. The speech synthesis method according to claim 4, characterized in that, The calculation of the emotional similarity between the original speech sample and the target speech sample corresponding to the perturbation spectral feature sample includes: The original speech sample and the target speech sample are encoded using a trained emotion encoder to obtain a first emotion feature sequence corresponding to the original speech sample and a second emotion feature sequence corresponding to the target speech sample. The similarity between the first emotional feature sequence and the second emotional feature sequence is calculated and used as the emotional similarity between the original speech sample and the target speech sample.
7. The speech synthesis method according to claim 1, characterized in that, Determining the K resonance peaks in the original spectrum includes: Obtain the spectral envelope of the original spectrum, determine K local peaks in the spectral envelope, and identify the K local peaks as K resonance peaks in the original spectrum.
8. A speech synthesis device with adjustable spectral intensity, characterized in that, The speech synthesis device includes: The frequency perturbation module is used to acquire the original spectrum of the original speech, determine K formants in the original spectrum, randomly perturb the frequencies of the K formants to obtain the perturbation frequencies of the corresponding formants, and replace the frequencies of the corresponding formants in the original spectrum with the perturbation frequencies of each formant to obtain the perturbation spectrum. K is an integer greater than 0. The perturbation spectrum is a spectrum that is independent of the speaker but related to emotion, and is used to characterize the emotional information and the intensity information of the emotion in the original spectrum. The weight determination module is used to perform feature encoding on the disturbance spectrum to obtain disturbance spectrum features, calculate the correlation between the disturbance spectrum features and N preset intensity matrices, and determine the weight of the corresponding intensity matrix based on each correlation, where N is an integer greater than 0. The intensity matrix is used to measure the intensity information of emotion in the corresponding disturbance spectrum features. The feature fusion module is used to perform a weighted summation of the N intensity matrices according to the weights corresponding to the N intensity matrices to obtain intensity features, and to perform feature fusion of the intensity features and the disturbance spectrum features to obtain spectral intensity features; The speech synthesis module is used to acquire the target text to be synthesized and the target speech features, perform speech synthesis on the target text according to the spectral intensity features, and adjust the synthesized speech using the target speech features to obtain the corresponding target speech.
9. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the speech synthesis method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech synthesis method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Voice synthesis method and device, computer equipment and computer readable storage medium
CN111108549A
Multi-emotion speech synthesis method and device, electronic equipment and storage medium
CN116129864A