A speech synthesis method based on electroencephalogram emotion

By constructing an EEG emotion measurement model and an emotional speech synthesis model, and by utilizing EEG signal acquisition and neural network technology to optimize speech generation, the problem of insufficient emotional expression in existing technologies has been solved, and a speech synthesis effect with rich emotions has been achieved.

CN119724146BActive Publication Date: 2025-10-24GUANGZHOU UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411905530.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-15
Publication Date
2025-10-24
Estimated Expiration
2044-05-15

AI Technical Summary

Technical Problem

Existing speech synthesis methods based on deep neural networks lack the ability to mine and model fine-grained factors of timbre and emotion, and cannot differentiate expressions for specific audiences, resulting in poor emotional effects in the speech synthesis of film and television dramas.

Method used

By constructing an EEG emotion measurement model and an emotional speech synthesis model, the test subjects' emotional data are acquired using an EEG signal acquisition device, and emotion annotation and preprocessing are performed. A convolutional neural network is trained, and combined with an attention mechanism and a fully connected neural network, the speech generation is optimized to meet the audience's empathy needs.

Benefits of technology

It achieves optimized speech generation based on emotion measurement results, synthesizes emotionally rich speech that meets the audience's empathy needs, and improves the emotional expression effect of film and television dubbing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119724146B_ABST
    Figure CN119724146B_ABST
Patent Text Reader

Abstract

The application discloses a speech synthesis method based on electroencephalogram emotion, and relates to the field of intelligent speech synthesis. The method comprises the following steps: obtaining the electroencephalogram emotion data of a tester after the tester hears a speech segment by using an electroencephalogram signal collector, and performing emotion labeling to obtain the labeled emotion extreme value group data and then performing pretreatment; performing electroencephalogram emotion measurement model training on a convolutional neural network through the pretreated emotion extreme value group data to obtain a trained electroencephalogram emotion measurement model; outputting the recognition result of the trained electroencephalogram emotion measurement model as the input of a vits model to perform emotion speech synthesis model training and obtain a trained emotion speech synthesis model; and performing speech synthesis on a to-be-voiced film and television work through the trained emotion speech synthesis model and outputting a final speech synthesis result. The electroencephalogram emotion measurement model and the emotion speech synthesis model are proposed, speech generation can be optimized under the emotion measurement result, and emotion-rich speech meeting the audience's empathy demand can be synthesized.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The application is a divisional application, the original application number is 202410603270.5, the application date is May 15, 2024, and the invention name is "Cantonese speech intelligent synthesis method and system based on electroencephalogram emotion measurement". TECHNICAL FIELD

[0002] The present application relates to the technical field of intelligent speech synthesis, in particular to a speech synthesis method based on electroencephalogram emotion. BACKGROUND

[0003] Speech synthesis refers to automatically generating speech from text using a computer. In the film and television industry, automatic dubbing programs can automatically generate speech from text, emotional expression, timbre and other elements in the script, and then match the speech to the picture, thereby greatly reducing the cost of dubbing. However, since the film and television production needs to achieve empathy, the generated speech has very high requirements for emotion, so the emotional factor becomes the key research content of film and television speech synthesis.

[0004] Due to the development of deep neural networks, intelligent speech synthesis has made great progress. Speech synthesis based on deep neural networks usually builds a parameterized neural network, uses sample data in the form of text-speech pairs to model the mapping relationship from text to speech, and thus realizes the conversion from text to speech. This simple data-driven speech synthesis method lacks the mining and modeling of finer-grained factors such as timbre and emotion.

[0005] The usual way of emotional speech synthesis is to make the model learn the emotional style of the speech through explicit labels, that is, manually label the speech with text labels representing emotion, and the model learns the prosody of the speech in the process of speech synthesis, thereby achieving the effect of generating emotional speech. However, text explicit labels can only make the model learn the average style of the sample data and the basic prosody of the speech, still lacking fine-grained analysis of emotional prosody. And this form of labeling speech through text labels is highly dependent on the way the model designer models the emotion measure and the subjective will of the person, and the effect of learning and expressing emotional style needs to be improved.

[0006] Existing emotion measure modeling has two ways: discrete and continuous. The mainstream discrete way contains several basic emotional states and related extensions, such as "anger", "anticipation", "fear", "sadness", "trust", "surprise", "joy", etc. In addition, there is a method based on the palette theory, which can further generate other emotions from the basic emotional states as primary colors. And emotion wheel representation method, emotion quantification method based on attributes or hierarchical.

[0007] Compared with discrete mode, continuous mode can represent more detailed emotional state with higher accuracy. Continuous mode usually uses several basic coordinate axes to represent emotion, and the commonly used method is the valence-arousal bipolar emotional quadrant system, which describes emotion from two dimensions of valence and arousal.

[0008] Studies have shown that changes in the potential information of the cerebral cortex can represent a lot of human cognitive-related information. When people listen to the emotional characteristics in the audio, they will imagine, which will cause changes in the potential of the cerebral cortex. This potential change can be represented by electroencephalogram (EEG), and more detailed information can be extracted from the electroencephalogram by signal processing method, such as the information of the emotional changes in the audio.

[0009] In general, due to the different emotional empathy points of different audiences for film and television content, the speech synthesis method based on the pre-training of explicit labels cannot be differentiated for specific audiences. Based on this problem, a more detailed emotional modeling method is needed to guide the optimization of speech synthesis effect.

[0010] Therefore, it is urgent to propose a speech synthesis method based on electroencephalogram emotion to solve the limitations of the prior art. SUMMARY

[0011] Therefore, the present application provides a speech synthesis method based on electroencephalogram emotion, which constructs an electroencephalogram emotion measurement model and an emotional speech synthesis model to optimize speech generation under the emotion measurement result and synthesize emotional speech that meets the empathy needs of the audience.

[0012] In order to achieve the above purpose, the technical scheme adopted by the present application is as follows:

[0013] A speech synthesis method based on electroencephalogram emotion, comprising the following steps:

[0014] S1. Data acquisition: using an electroencephalogram signal acquisition instrument to acquire the electroencephalogram emotion data of the tester after listening to the speech segment;

[0015] S2. Data labeling: labeling the collected electroencephalogram emotion data to obtain labeled emotion extreme value group data;

[0016] S3. Data preprocessing: preprocessing the emotion extreme value group data to obtain preprocessed emotion extreme value group data;

[0017] S4. Electroencephalogram emotion measurement model training: training the convolutional neural network through the preprocessed emotion extreme value group data to obtain the trained electroencephalogram emotion measurement model;

[0018] S5. Emotional speech synthesis model training: output the recognition result of the trained electroencephalogram emotion measurement model as the input of the vits model, so as to perform emotional speech synthesis model training, and obtain a trained emotional speech synthesis model;

[0019] S6. Speech synthesis: performing speech synthesis on the to-be-dubbed film and television works by the trained emotional speech synthesis model, and outputting the final speech synthesis result.

[0020] The above method, optionally, the specific content of acquiring the electroencephalogram emotional data of the tester after hearing the voice segment in S1 is:

[0021] The tester wears the electroencephalogram acquisition helmet, calibrates the electrode position, turns on the recorder and data acquisition software in turn, the tester completes the calibration process of opening eyes, closing eyes and mouse clicking according to the instruction of the data acquisition software, the computer calculates the signal calibration time deviation in the calibration process, the tester will hear the voice segment, the electroencephalogram acquisition helmet collects the electroencephalogram data of the tester and marks the start and end time of the voice segment in the electroencephalogram data.

[0022] The above method, optionally, the specific content of marking the collected electroencephalogram emotional data in S2 is:

[0023] First, the electroencephalogram emotional data obtained in S1 is processed, and the electroencephalogram emotional data representing the emotional extreme value is screened out, and the electroencephalogram emotional data is marked with emotional polarity.

[0024] The above method, optionally, the emotional extreme value group data is classified into a extreme value group, b extreme value group and c extreme value group according to the emotional two-pole opposite characteristics of six emotions.

[0025] The above method, optionally, the emotional extreme value group data is preprocessed to remove noise in S3 to obtain emotional extreme value group data after removing noise.

[0026] The above method, optionally, the specific content of training the electroencephalogram emotion measurement model of the convolutional neural network by the preprocessed emotional extreme value group data in S4 is:

[0027] S41. Input signal: the input signal has two segments, including anchor sample data anchor and real-time input electroencephalogram data input, the anchor sample data is the emotional extreme value group data after removing noise obtained in S3;

[0028] S42. Feature extraction: the electroencephalogram emotion measurement model mainly captures the emotional fluctuation period, therefore, the window sliding method is adopted to frame the emotional extreme value group data after removing noise, and the framed data is represented as:

[0029]

[0030]

[0031] wherein, C is the number of channels, T is the length of time, is the real field, M is the window size, is the splicing operation, N is the number of frames;

[0032] Each frame is represented by a convolutional neural network for feature extraction:

[0033]

[0034] wherein, K is a two-dimensional convolution kernel, the convolution is operated longitudinally along the channel and transversely along the time axis, S n is each frame signal, is a feature map representing the output result after convolution of each frame signal, D is a D-dimensional vector, i is the longitudinal position index of the feature image pixel, j is the transverse position index of the feature image pixel, c is the longitudinal position index of the convolution kernel, and m is the transverse position index of the convolution kernel;

[0035] S43. Anchor removal coherent noise module based on attention mechanism: applying attention mechanism to eliminate features similar to anchor data, taking the features of anchor sample data as K, and the features of real-time input electroencephalogram data as Q and V, represented as:

[0036]

[0037]

[0038] The distance is represented as:

[0039]

[0040] The distance attention mechanism is represented as:

[0041]

[0042] wherein, d ij is the distance between the ith feature and the jth feature;

[0043] S44. The classifier classifies: A input After compressing the features through multiple layers of convolutional neural network, the full connection neural network is inputted to predict the emotion measure:

[0044]

[0045]

[0046] Among them, Z is the feature vector flattened after the convolution operation, the vector dimension is O, ANN is the fully connected neural network function, It is a scalar value between 0 and 1 representing the emotion measure. The closer the predicted value is to 0, the more similar the emotion of the input EEG signal input_trail is to the EEG signal anchor_trail used for anchoring. Conversely, the more dissimilar it is, the more similar the emotion measure of the input EEG signal is determined.

[0047] S45. Train each sample labeled with the emotion scale separately to obtain three trained EEG emotion measurement models

[0048] In the above method, optionally, the recognition result of the trained EEG emotion measurement model is output in S5 as the input of the VIT model, thereby training the emotional speech synthesis model. The specific content of the trained emotional speech synthesis model is:

[0049] The emotional speech synthesis model is divided into model learning speech reconstruction training and model emotion enhancement learning training;

[0050] S51. Model learning speech reconstruction: Use the pre-processed emotional extreme value group data obtained in S3 to train the model and learn speech reconstruction. Emotional information is embedded in the text encoder in the form of a vector; the loss function is the generated speech mel-spectrogram Mel spectrogram with real samples X mel The gap between them, that is, whether the speech can be reconstructed, the loss value for each sample is expressed as:

[0051]

[0052] S52. Model-enhanced emotion learning: Use the trained EEG emotion measurement model obtained in S4 to perform emotion recognition on the AI-generated Cantonese speech. Input the voice actor's emotional feedback signal about the AI ​​dubbing through the device to calibrate the encoding loss function of the AI ​​dubbing emotion encoder. The loss value is then calculated to complete the model-enhanced learning training for the emotional speech synthesis model. The total loss value calculation formula is expressed as:

[0053] loss stage2 =αloss recon +βloss emotion

[0054]

[0055] Among them, α and β are the weight coefficients of reconstruction loss and sentiment loss in the total loss of model reinforcement learning.

[0056] The application discloses a Cantonese speech intelligent synthesis system based on an electroencephalogram emotion measure, and applies a speech synthesis method based on the electroencephalogram emotion, which comprises the following modules: a data acquisition module, a data labeling module, a data preprocessing module, an electroencephalogram emotion measure model training module, an emotional speech synthesis model training module and a speech synthesis module.

[0057] The data acquisition module is connected with the input end of the data labeling module and is used for acquiring the electroencephalogram emotion data of a tester after hearing a speech segment by using an electroencephalogram signal acquisition instrument.

[0058] The data labeling module is connected with the input end of the data preprocessing module and is used for labeling the acquired electroencephalogram emotion data to obtain labeled emotion extreme value group data.

[0059] The data preprocessing module is connected with the input end of the electroencephalogram emotion measure model training module and is used for preprocessing the emotion extreme value group data to obtain preprocessed emotion extreme value group data.

[0060] The electroencephalogram emotion measure model training module is connected with the input end of the emotional speech synthesis model training module and is used for training a convolutional neural network by using the preprocessed emotion extreme value group data to obtain a trained electroencephalogram emotion measure model.

[0061] The emotional speech synthesis model training module is connected with the input end of the speech synthesis module and is used for outputting the recognition result of the trained electroencephalogram emotion measure model as the input of a vits model, so as to train an emotional speech synthesis model and obtain a trained emotional speech synthesis model.

[0062] The speech synthesis module is used for performing speech synthesis on a to-be-voiced movie or TV play by using the trained emotional speech synthesis model and outputting a final speech synthesis result.

[0063] Compared with the prior art, the application provides a speech synthesis method based on an electroencephalogram emotion, and has the following beneficial effects: the application proposes an electroencephalogram emotion measure model and an emotional speech synthesis model, the emotional speech synthesis model can convert the text in a script into speech, a listener can listen to the synthesized speech while wearing a non-invasive electroencephalogram device, generate an electroencephalogram, and generate an emotion measure by using the electroencephalogram emotion measure model, which is beneficial to optimizing speech generation according to the emotion measure result and synthesizing emotional speech that meets the empathy demand of the listener. BRIEF DESCRIPTION OF DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings described below only constitute a part of the embodiments of the present application, and all other drawings obtained by those of ordinary skill in the art without creative effort based on the provided drawings also belong to the protection scope of the present application.

[0065] Figure 1 A flow chart of a Cantonese speech intelligent synthesis method based on electroencephalogram emotion measurement is provided for the present application.

[0066] Figure 2 An electroencephalogram emotion data labeling schematic diagram is provided for the present application.

[0067] Figure 3 An electroencephalogram six-emotion measurement schematic diagram is provided for the present application.

[0068] Figure 4 A flow chart of noise removal preprocessing on emotion extreme value group data is provided for the present application.

[0069] Figure 5 An electroencephalogram emotion measurement model structure schematic diagram is provided for the present application.

[0070] Figure 6 A feature distribution schematic diagram is provided for the present application.

[0071] Figure 7 A sentiment speech synthesis model framework diagram is provided for the present application.

[0072] Figure 8 A dubbing flow chart based on the electroencephalogram emotion measurement model and the sentiment speech synthesis model is provided for the embodiments of the present application. DETAILED DESCRIPTION

[0073] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort belong to the protection scope of the present application.

[0074] Referring to Figure 1 The present application discloses a speech synthesis method based on electroencephalogram emotion, comprising the following steps:

[0075] S1. Data acquisition: acquiring electroencephalogram emotion data of a tester after hearing a speech segment by using an electroencephalogram signal acquisition instrument;

[0076] S2. Data labeling: The collected electroencephalogram emotion data is labeled with emotions to obtain labeled emotion extreme value group data;

[0077] S3. Data preprocessing: The emotion extreme value group data is preprocessed to obtain preprocessed emotion extreme value group data;

[0078] S4. Electroencephalogram emotion metric model training: The convolutional neural network is trained for electroencephalogram emotion metric model training through the preprocessed emotion extreme value group data, and a trained electroencephalogram emotion metric model is obtained;

[0079] S5. Emotional speech synthesis model training: The recognition result of the trained electroencephalogram emotion metric model is output as the input of the vits model, so as to perform emotional speech synthesis model training, and a trained emotional speech synthesis model is obtained;

[0080] S6. Speech synthesis: The trained emotional speech synthesis model is used to synthesize speech for the to-be-dubbed film and television drama, and the final speech synthesis result is output.

[0081] Further, the specific content of acquiring the electroencephalogram emotion data of the tester after hearing the voice segment in S1 is as follows:

[0082] The tester wears an electroencephalogram acquisition helmet, calibrates the electrode position, and turns on the recorder and data acquisition software in sequence. The tester completes the calibration process of opening eyes, closing eyes and mouse clicking according to the instructions of the data acquisition software. The computer calculates the signal calibration time deviation during the calibration process. The tester will hear a voice segment. The electroencephalogram acquisition helmet collects the electroencephalogram data of the tester and marks the time of starting and ending playing the voice segment in the electroencephalogram data.

[0083] Specifically, physiological saline is added to the electrodes of the multi-point electroencephalogram device, and the electrodes are installed. The tester wears an electroencephalogram acquisition helmet and calibrates the electrode position. The recorder and data acquisition software are turned on in sequence, and the software calibration stage is entered. The tester completes the calibration process of opening eyes, closing eyes and mouse clicking according to the instructions of the software. The computer calculates the signal calibration time deviation T offset during the calibration process. Then the tester listens to a voice segment with emotions. This process collects electroencephalogram data and marks the timestamps T begin , T end of starting and ending playing the voice segment in the data. Finally, the calculated signal calibration time deviation T offset is marked on the electroencephalogram data. The time index of the electroencephalogram data segment is: T begin + T offset : T end + T offset . The human electroencephalogram data timestamp marking diagram is shown in Figure 2 . The electroencephalogram and the corresponding voice labeled sample are (Si ,X i ,c),S i is the ith electroencephalogram sample, X i is the ith audio sample, wherein the emotion category is c.

[0084] Further, the collected electroencephalogram emotion data in S2 is labeled with emotions to obtain the specific content of the labeled emotion extreme group data:

[0085] First, the electroencephalogram emotion data obtained in S1 is processed for features, and the electroencephalogram emotion data representing the emotion extreme is screened out, and the electroencephalogram emotion data is labeled with emotional polarity.

[0086] Further, the emotion extreme group data is classified into a extreme group, b extreme group and c extreme group according to the emotion two-pole opposite feature.

[0087] Specifically, according to the emotion two-pole opposite feature, six basic emotions are classified into three groups of extreme values, including [sadness, joy], [anger, expectation], [fear, surprise]. Each group is called an emotion scale, and the center intersection point represents a neutral emotion, and the continuous value between the two extreme values represents the intensity of the emotion, as shown in Figure 3 . The sample is represented as (S i ,X i ,e y ), wherein S represents electroencephalogram data, X represents speech data, e represents emotional scale value, y∈{a,b,c}, a represents [sadness, joy] extreme group, b represents [anger, expectation] extreme group, c represents [fear, surprise] extreme group, e∈[0,1], and e usually takes 1 or 0 when labeling.

[0088] Further, the emotion extreme group data in S3 is preprocessed to remove noise to obtain the emotion extreme group data after removing noise.

[0089] Specifically, the brain electrical data collected by the safer non-invasive electroencephalogram acquisition device will inevitably introduce more noise, so the process of analyzing the electroencephalogram focuses on removing noise, which includes incoherent noise and coherent noise. Incoherent noise refers to noise whose frequency characteristics differ greatly from those of useful signals. This type of noise is additive in the frequency domain and is relatively easy to remove. Coherent noise refers to noise whose frequency characteristics are similar to those of useful signals. This type of noise is easily mixed in and difficult to remove. The noise removal process is shown in Figure 4 .

[0090] Further, as shown in Figure 5 , the preprocessed emotion extreme group data in S4 is used to train a convolutional neural network for electroencephalogram emotion scale model to obtain the specific content of the trained electroencephalogram emotion scale model:

[0091] S41. Input signal: The input signal has two segments, including anchor sample data anchor and real-time input EEG data input, the anchor sample data is the emotional extreme value group data obtained after removing noise in S3;

[0092] S42. Feature extraction: The EEG emotion measurement model mainly captures the emotional fluctuation period, therefore, the window sliding method is adopted to frame the emotional extreme value group data after removing noise, and the framed data is expressed as:

[0093]

[0094]

[0095] wherein, C is the number of channels, T is the time length, is a real field, M is the window size, is a splicing operation, N is the number of frames;

[0096] As shown in Figure 6 , each frame is subjected to feature extraction by a convolutional neural network and is expressed as:

[0097]

[0098] wherein, K is a two-dimensional convolution kernel, the convolution is operated longitudinally along the channel and transversely along the time axis, S n is each frame signal, is a feature map representing the output result after convolution of each frame signal, D is a D-dimensional vector, i is the longitudinal position index of the feature image pixel, j is the transverse position index of the feature image pixel, c is the longitudinal position index of the convolution kernel, and m is the transverse position index of the convolution kernel;

[0099] S43. Anchor removal coherent noise module based on attention mechanism: the attention mechanism is applied to eliminate features similar to the anchor data, the features of the anchor sample data are taken as K, the features of the real-time input EEG data are taken as Q and V, and are expressed as:

[0100]

[0101]

[0102] The distance is expressed as:

[0103]

[0104] The distance attention mechanism is expressed as:

[0105]

[0106] wherein d ij is the distance between the ith feature and the jth feature;

[0107] S44. The classifier classifies: A input After the feature compression by the multi-layer convolutional neural network, the full connection neural network is inputted to perform the emotion measure prediction:

[0108]

[0109]

[0110] wherein Z is the flattened feature vector after the convolution operation, the vector dimension is O, ANN is a full connection neural network function, is a 0-1 scalar value representing the emotion measure, the closer the predicted value to 0 indicates that the input EEG signal input_trail is more similar to the EEG signal anchor_trail used for anchoring, and vice versa, so as to determine the emotion measure value of the input EEG signal;

[0111] S45. Each emotion measure scale labeled sample is trained respectively to obtain three trained EEG emotion measure models

[0112] Specifically, the emotion extreme group data after noise removal in S42 feature extraction contains a variety of frequency components. According to research, the EEG in the frequency band of 0Hz to 64Hz has the greatest impact on emotion, so most of the incoherent background noise such as eye movement and muscle movement generated electrical signals can be removed after filtering the emotion extreme group data after noise removal.

[0113] Further, the recognition result of the trained EEG emotion measure model output in S5 is inputted as the vits model, so as to perform the emotion speech synthesis model training, and the specific content of the trained emotion speech synthesis model is:

[0114] The emotion speech synthesis model is divided into model learning speech reconstruction training and model emotion enhancement learning training.

[0115] S51. Model learning speech reconstruction: the preprocessed emotion extreme group data obtained in S3 is used to train the model to learn speech reconstruction, and the emotional information is embedded in the form of a vector in the text encoder; the loss function is the difference between the generated speech mel spectrum X and the mel spectrum X mel of the real sample, i.e. whether the speech can be reconstructed, and the loss value of each sample is represented as:

[0116]

[0117] S52. Model emotion reinforcement learning: using the trained electroencephalogram emotion measurement model obtained in S4 to perform emotion recognition on the Cantonese voice generated by AI, importing the feedback signal of the voice actor to the emotion of the AI dubbing through the device, for calibrating the coding loss function of the AI dubbing emotion encoder, then calculating the loss value, completing the model reinforcement learning training of the emotional speech synthesis model, and the total loss value calculation formula is:

[0118] loss stage2 =αloss recon +βloss emotion

[0119]

[0120] Wherein, α, β are the weight coefficients of the reconstruction loss and the emotion loss in the total loss of the model reinforcement learning.

[0121] Specifically, the specific steps of S52 model emotion reinforcement learning are: the subject (professional voice actor) wears the electroencephalogram acquisition device, evaluates the model generated voice and performs emotion selection operation, selects from {sadness, joy, anger, expectation, fear, surprise}, selects the emotion category and measure l. The subject performs the selection operation at the same time, the system will automatically record the time stamp T begin Then calculate the correction value T offset Then use the corrected label time stamp T label , intercept the electroencephalogram signal as the input of the electroencephalogram emotion measurement model , and finally get the output of the model Then calculate the loss value and perform secondary training on the Cantonese voice generation model.

[0122] The loss function includes the error between the model generated voice and the real voice, and the error of the electroencephalogram device monitoring the professional voice actor to the AI dubbing emotion measure and the real emotion label. Here, the cross-entropy form is adopted, and the loss value for each sample can be expressed as:

[0123] loss stage2 =αloss recon +βloss emotion

[0124]

[0125] Wherein α and β are hyperparameters, representing the weight coefficients of the reconstruction loss and the emotion loss in the total loss of the model reinforcement learning.

[0126] Specifically, the basic framework of the Cantonese speech emotion speech synthesis model formed by the present invention is VITSS, which is an end-to-end speech synthesis model using adversarial learning conditional variational autoencoder. Speech emotion information is embedded in the text encoder of the VITSS model. The speech emotion information can come from manual annotation or EEG emotion recognition results. The emotion speech synthesis model is divided into model learning speech reconstruction training and model emotion enhancement learning training. The training process is as follows: Figure 7 shown.

[0127] In a specific embodiment, in order to solve the problem of insufficient emotional expression in traditional Cantonese speech synthesis of film and television dramas trained by explicit text emotion labels, a method for intelligent synthesis of Cantonese speech of film and television dramas based on EEG emotion measurement is proposed. Figure 8 As shown, the specific method is: prepare Cantonese dialogue text according to the content of the film and television drama, and then the dubbing actor wears the EEG device; the dubbing program is started and time calibration is performed; the text content is input into the emotional speech synthesis model; the dubbing actor evaluates the generation effect of the model; if satisfied with the generation effect, the speech generation result is saved; if not satisfied, the dubbing actor reselects the emotion, the system will automatically obtain the EEG data and input the EEG emotion measurement model to calculate the measurement to optimize the emotional speech synthesis model, and then regenerate the speech. The dubbing actor continues to evaluate until he is satisfied with the generation effect, then ends the speech generation and saves the speech generation result.

[0128] and Figure 1 Corresponding to the method described above, the embodiment of the present invention also provides a Cantonese speech intelligent synthesis system based on EEG emotion measurement for Figure 1 The specific implementation of the method includes: data acquisition module, data annotation module, data preprocessing module, EEG emotion measurement model training module, emotion speech synthesis model training module, and speech synthesis module;

[0129] The data acquisition module is connected to the input end of the data annotation module and is used to use an EEG signal collector to obtain the EEG emotional data of the tester after hearing the voice clip;

[0130] The data labeling module is connected to the input end of the data preprocessing module and is used to label the collected EEG emotion data to obtain the labeled emotion extreme value group data;

[0131] A data preprocessing module is connected to the input end of the EEG emotion measurement model training module and is used to preprocess the emotion extreme value group data to obtain preprocessed emotion extreme value group data;

[0132] The EEG emotion measurement model training module is connected to the input end of the emotional speech synthesis model training module and is used to train the EEG emotion measurement model on the convolutional neural network using the preprocessed emotion extreme value group data to obtain a trained EEG emotion measurement model;

[0133] The emotional speech synthesis model training module is connected with the input end of the speech synthesis module, is used for outputting the recognition result of the trained electroencephalogram emotion measurement model as the input of the vits model, thereby performing emotional speech synthesis model training, and obtaining the trained emotional speech synthesis model;

[0134] The speech synthesis module is used for performing speech synthesis on the to-be-voiced film and television drama through the trained emotional speech synthesis model, and outputting the final speech synthesis result.

[0135] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same or similar parts of various embodiments can be referred to each other. For the device disclosed by the embodiments, since it corresponds to the method disclosed by the embodiments, the description is relatively simple, and the related parts can be referred to the method part.

[0136] The above description of the disclosed embodiments enables a person skilled in the art to implement or use the present application. Various modifications to the embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for emotion-based speech synthesis based on electroencephalography, characterized by, The method comprises the following steps: S1. Data acquisition: use an electroencephalogram signal acquisition instrument to acquire the electroencephalogram emotional data of the tester after hearing the voice segment; S2. Data labeling: label the collected electroencephalogram emotional data, and obtain the labeled emotional extreme value group data, specifically: first, perform feature processing on the electroencephalogram emotional data obtained in S1, and screen out the electroencephalogram emotional data that can represent the emotional extreme value, and perform emotional polarity labeling on the electroencephalogram emotional data; S3. Data preprocessing: preprocessing the emotional extreme value group data to obtain the preprocessed emotional extreme value group data; S4. Electroencephalogram emotional measure model training: training the convolutional neural network through the preprocessed emotional extreme value group data to obtain the trained electroencephalogram emotional measure model; S5. Emotional speech synthesis model training: output the recognition result of the trained electroencephalogram emotional measure model as the input of the vits model, and then perform emotional speech synthesis model training to obtain the trained emotional speech synthesis model; S6. Speech synthesis: performing speech synthesis on the video to be dubbed through the trained emotional speech synthesis model, and outputting the final speech synthesis result; The specific content of S5 is: The emotional speech synthesis model is divided into model learning speech reconstruction training and model emotion enhancement learning training; S51. Model learning speech reconstruction: Use the pre-processed emotional extreme value group data obtained in S3 to train the model and learn speech reconstruction. Emotional information is embedded in the text encoder in the form of a vector; the loss function is the generated speech mel-spectrogram Mel spectrogram with real samples The gap between them, that is, whether the speech can be reconstructed, the loss value for each sample is expressed as: S52. Model emotion enhancement learning: using the trained electroencephalogram emotional measure model obtained in S4 to perform emotion recognition on the Cantonese voice generated by AI, importing the feedback signal of the voice actor to the AI dubbing emotion through the device, which is used to calibrate the coding loss function of the AI dubbing emotion encoder, then calculating the loss value, and completing the model enhancement learning training of the emotional speech synthesis model; Wherein, the loss function contains the error between the model generated voice and the real voice, the error between the professional voice actor's AI dubbing emotion measure monitored by the electroencephalogram device and the real emotional label, adopts the form of cross entropy, and the loss value of each sample can be expressed as: wherein, , is a weight coefficient of the reconstruction loss and the emotion loss in the total loss of model reinforcement learning.

2. The emotion-based speech synthesis method using electroencephalogram according to claim 1, wherein, The specific content of S1 is: Electrodes of multi-point electroencephalogram device are filled with physiological saline, and the electrodes are installed; the tester wears the electroencephalogram acquisition helmet, the electrode positions are calibrated, the recorder and the data acquisition software are sequentially turned on, and the software calibration stage is entered; the tester completes the calibration process of opening eyes, closing eyes and mouse clicking according to the instruction of the data acquisition software, and the computer calculates the signal calibration time deviation in the calibration process , the tester will hear a voice segment with emotion, the electroencephalogram acquisition helmet collects electroencephalogram data of the tester, and time stamps of starting playing and ending playing of the voice segment are marked in the electroencephalogram data 、 ; finally, the electroencephalogram data are marked in combination with the calculated signal calibration time deviation , and the time index of the electroencephalogram data segment is .

3. The emotion-based speech synthesis method using electroencephalogram according to claim 1, wherein, The data of the emotional pole groups are classified according to the two-pole opposite characteristics of the six emotions pole groups, pole groups and pole groups; each pole group is called an emotional scale, the central intersection point represents neutral emotion, and the continuous values between the two poles represent the intensity of the group emotion; The sample is represented as ,in, represents EEG data, Indicates the i EEG samples, Represents voice data, Indicates the i audio samples, represents the sentiment measurement value, , Indicates the extreme value group of [sadness, joy], Expressing [anger, expectation] extreme value group, represents the extreme value group of [fear, surprise], , when marking Takes 1 or 0.

4. The emotion-based speech synthesis method using electroencephalogram according to claim 1, wherein, In S3, the emotional extreme value group data is preprocessed to remove noise to obtain the emotional extreme value group data after removing noise; wherein, the noise includes incoherent noise and coherent noise.

5. The emotion-based speech synthesis method using electroencephalogram according to claim 1, wherein, The specific content of S4 is: S41. Input signal: The input signal has two parts, including anchor sample data , and real-time input EEG data The anchor sample data is the emotional extreme group data obtained after removing noise in S3; S42. Feature extraction: the electroencephalogram emotional measure model mainly captures the emotional fluctuation period, therefore, the window sliding method is adopted to frame the emotional extreme value group data after removing noise, and the framed data is expressed as: ; wherein, , , is the number of channels, is the time length, is the real number field, is the window size, is the stitching operation, is the number of frames; Each frame is subjected to feature extraction by a convolutional neural network and is expressed as: wherein, is a two-dimensional convolution kernel, the convolution is operated along the channel in the longitudinal direction and along the time axis in the transverse direction, is a two-dimensional convolution kernel, the convolution is operated along the channel in the longitudinal direction and along the time axis in the transverse direction, is a two-dimensional convolution kernel, the convolution is operated along the channel in the longitudinal direction and along the time axis in the transverse direction, is a two-dimensional convolution kernel, the convolution is operated along the channel in the longitudinal direction and along the time axis in the transverse direction, is a two-dimensional convolution kernel, the convolution is operated along the channel in the longitudinal direction and along the time axis in the transverse direction, is a two-dimensional convolution kernel, the convolution is operated along the channel in the longitudinal direction and along the time axis in the transverse direction, is a two-dimensional convolution kernel, the convolution is operated along the channel in the longitudinal direction and along the time axis in the transverse direction, is a two-dimensional convolution kernel, the convolution is operated along the channel in the longitudinal direction and along the time axis in the transverse direction, is a two-dimensional convolution kernel, the convolution is operated along the channel in the longitudinal direction and along the time axis in the transverse direction, S43. Anchor removal coherent noise module based on attention mechanism: the attention mechanism is applied to eliminate and anchor the features similar to the anchor data, the features of the anchor sample data are taken as K, the features of the real-time input electroencephalogram data are taken as Q and V, and are expressed as: The distance is expressed as: The distance attention mechanism is expressed as: wherein is the distance of the first feature and the second feature is the distance of the first feature and the second feature is the distance of the first feature and the second feature S44. The classifier classifies: the After the feature is compressed by the multi-layer convolutional neural network, the full connection neural network is input to predict the emotion measure: wherein, is the flattened feature vector after convolution operation, vector dimension is , is the fully connected neural network function, is a 0-1 scalar value representing the emotion measure, the closer the predicted value to 0 indicates the input EEG signal Emotion and EEG signal for anchoring is more similar, and vice versa, to determine the emotion measure of the input EEG signal; S45. Train each sample marked on the emotion scale respectively to obtain three trained electroencephalogram emotion scale models .

Citation Information

Patent Citations

  • Electroencephalogram signal recognition system and method

    CN111973178A

  • Speech synthesis method based on cross-subject multi-mode and related equipment

    CN113421546A