A speech generation method, apparatus, device, and storage medium
By generating emotional guidance information for recordings to guide the speaker in reading the recorded text, the problem of lack of emotional color in speech synthesis technology is solved, emotionally rich speech synthesis in the speech library is achieved, and the naturalness and authenticity of human-computer interaction are improved.
Patent Information
- Application Number
- CN202210867412.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-22
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2042-07-22
AI Technical Summary
The speech generated by existing speech synthesis technology lacks emotional color, resulting in a dull human-computer interaction effect and unable to achieve the natural and realistic effect of empathetic interaction with users.
By generating emotional guidance information for recordings, guiding the speaker to read the recording text, collecting voice data with emotional colors, and building a voice library to adapt to speech synthesis in different scenarios, styles and emotions.
The generated speech has emotional color, which improves the naturalness and authenticity of human-computer interaction and enhances the applicability and expressiveness of the speech synthesis system.
Smart Images

Figure CN115472185B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech synthesis technology, and in particular to a speech generation method, apparatus, device and storage medium. Background Art
[0002] Speech synthesis is typically based on speech materials from a speech library, synthesizing speech that suits the interaction scenario. With the development and application of speech-based human-computer interaction technology, speech synthesis is increasingly being used in smart devices equipped with human-computer interaction functions, such as smart homes and smart robots.
[0003] Currently, a common method for constructing a voice database is to have a human speaker read a text aloud using standard pronunciation, then record the speaker's speech. This recorded speech is then stored as voice material in the database. The speech generated using this method is monotonous and straightforward. The synthesized speech based on this material lacks any emotional overtones, resulting in a very rigid and inflexible synthesized speech. This fails to achieve empathetic interaction with the user, and certainly fails to make human-computer interaction as natural and authentic as human-to-human interaction. Summary of the Invention
[0004] Based on the above technical status, this application proposes a speech generation method, device, equipment and storage medium, which can generate speech that meets the target emotional effect. By using these speech to perform speech synthesis, synthesized speech with emotional color can be obtained, which is conducive to improving the human-computer interaction effect.
[0005] The first aspect of the present application provides a speech generation method, comprising: generating recording emotion guidance information based on a recording text and a target speech emotional effect; outputting the recording emotion guidance information so that a target speaker reads the recording text under the guidance of the recording emotion guidance information; and collecting the target speaker's reading voice of the recording text to obtain speech data corresponding to the recording text.
[0006] The second aspect of the present application provides a speech generation device, including: an information generation unit, used to generate recording emotion guidance information based on the recording text and the target speech emotional effect; a data output unit, used to output the recording emotion guidance information so that the target speaker reads the recording text under the guidance of the recording emotion guidance information; a data collection unit, used to collect the target speaker's reading voice of the recording text to obtain speech data corresponding to the recording text.
[0007] The third aspect of the present application provides a speech generation device, comprising: a processor, and a memory, a microphone and an output device respectively connected to the processor; wherein the memory is used to store data and computer programs; the processor is used to generate recording emotion guidance information according to the recording text and the target speech emotion effect by running the computer program in the memory, and send the generated recording emotion guidance information to the output device; the output device is used to output the recording emotion guidance information sent by the processor, so that the target speaker reads the recording text under the guidance of the recording emotion guidance information; the microphone is connected to the memory, and is used to collect the target speaker's reading voice of the recording text, obtain voice data corresponding to the recording text, and store the voice data in the memory.
[0008] A fourth aspect of the present application provides a storage medium having a computer program stored thereon. When the computer program is executed by a processor, the above-mentioned speech generation method is implemented.
[0009] The speech generation method proposed in this application can generate and output recording emotion guidance information based on the recorded text and the target speech emotion effect when generating speech, so that the speaker can read the recorded text under the guidance of the recording emotion guidance information. At this time, the speaker's reading of the recorded text is collected to obtain speech data corresponding to the recorded text.
[0010] In the above-mentioned speech generation process, recording emotion guidance information can be generated in real time according to the recording text and the target speech emotional effect, thereby providing the speaker with a recording emotional reference, so that the speaker can more intuitively and accurately know with what emotion the recording text should be read aloud, and then can collect speech data with various emotional colors. On the one hand, this method can provide convenience for the speaker to pronounce, that is, it can automatically generate recording emotion guidance information, which is convenient for the speaker to know how to adjust the pronunciation emotion; on the other hand, any speech emotion is used as the target speech emotional effect, and then by executing the technical solution of this application, speech data with various emotions can be generated. These speech data can be used as speech materials to synthesize speech with emotional colors, which can improve the speech synthesis effect, and thus help to improve the human-computer interaction effect based on speech. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0012] Figure 1 A flowchart of a speech generation method provided in an embodiment of the present application;
[0013] Figure 2 A schematic diagram of the processing process of the spoken text generation model provided in an embodiment of the present application;
[0014] Figure 3 A schematic diagram of the text error correction process of the text error correction model provided in an embodiment of the present application;
[0015] Figure 4 A schematic diagram of iterative training of a speech recognition system provided in an embodiment of the present application;
[0016] Figure 5 A schematic diagram of the structure of a speech generation device provided in an embodiment of the present application;
[0017] Figure 6 A schematic diagram of the structure of the speech generation device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0018] The technical solutions proposed in the embodiments of this application are applicable to speech generation applications, specifically to generating speech to construct a speech library. By using the technical solutions in the embodiments of this application, the generated speech can have an emotional effect, so that the speech library constructed using these speech can be used to synthesize speech with emotional color and applicable to a wider range of scenarios and language styles.
[0019] Humans communicate in a variety of ways in daily life, but voice is the most direct, understandable, and natural mode of communication. The rapid development of computers and internet technology has profoundly changed our lifestyles, making the relationship between humans and computers inseparable. Speech synthesis is now widely used in interactive fields such as smart homes and intelligent robots. In recent years, numerous development technologies related to speech synthesis have been continuously innovated, gradually realizing the dream of human-machine interaction through voice. Successful examples of voice development technologies and application products are also emerging, such as mobile voice assistants and voice input methods.
[0020] However, when using intelligent voice devices, compared to the monotonous, straightforward speech synthesis technology used by machines, personalized speech synthesis tailored to different scenarios and styles is becoming increasingly necessary. This personalized speech can make human-computer interaction systems more like human-to-human communication. For example, in mobile voice assistants, the machine can select the appropriate emotion to communicate with the owner based on their mood, achieving empathy. In car voice assistants, when the vehicle's energy is low, the machine can switch to a weak tone to communicate with the owner. In audio novels, the narrative style, tone, and emotion can be selected based on the plot, greatly enhancing the expressiveness. Therefore, in these situations, speech synthesis systems with greater expressiveness in different scenarios, styles, and emotions are extremely urgent. To train such speech synthesis systems, building a raw speech library is particularly important.
[0021] The current conventional method for constructing a voice library involves having a human speaker read a text aloud using standard pronunciation. This recording is then recorded and stored as voice material in the voice library. The speech generated using this method is monotonous, straightforward, and devoid of emotion. The synthesized speech based on this material also lacks any emotional overtones. This results in a very rigid, ineffective synthesis of speech, unable to achieve empathetic interaction with the user, and even less in making human-computer interaction as natural and authentic as human-to-human interaction.
[0022] In response to the above-mentioned technical status, the embodiments of the present application propose a new speech generation method, which can make the generated speech have emotional color. The speech generated by this method is used for speech library construction, which can enable the speech library to be used to generate speech suitable for different scenarios, different styles, and different emotions, so that speech synthesis can support more natural and realistic human-computer interaction.
[0023] The following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of this application, and not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0024] Exemplary method
[0025] See also Figure 1 The present application embodiment first proposes a speech generation method, which includes:
[0026] S101. Generate recording emotion guidance information based on the recording text and the target voice emotion effect.
[0027] The above-mentioned recording text refers to a text used for recording. The recording text can be a text content of any language, any content, and any length.
[0028] The above-mentioned voice emotion effect refers to an emotion effect possessed by a voice. As a preferred implementation, the above-mentioned target voice emotion effect refers to an emotion effect achieved by carrying any one of all known emotions in a voice, such as any one of various known emotions including joy, anger, sadness, fear, tension, and excitement.
[0029] The embodiments of the present application can be used to generate a voice to construct a voice library, which should theoretically be able to support the construction of a voice with any emotion effect. Therefore, the above-mentioned target voice emotion effect can be to traverse each possible voice emotion effect. That is, each voice emotion effect is taken as the above-mentioned target voice emotion effect, and the technical solutions of the embodiments of the present application are executed to generate a voice corresponding to each voice emotion effect. The voice thus generated and the voice library constructed using the generated voice can be used to generate a voice with any emotion effect.
[0030] The above-mentioned voice emotion guiding information refers to information used to guide a sound producer to pronounce according to the above-mentioned target voice emotion effect. The voice emotion guiding information can be information in any form, such as text form, audio form, or video form. The specific information content can be emotion prompt text, emotion prompt voice, audio or video with the same emotion tone as the target voice emotion effect, and the like.
[0031] In the implementation of the technical solutions of the embodiments of the present application, the recording emotion guiding information generated according to the recording text and the target voice emotion effect can be one or more of a recording emotion guiding video, a recording emotion guiding audio, and a recording emotion guiding text.
[0032] The above-mentioned recording emotion guiding video can be a video with the same emotion as the target voice emotion effect.
[0033] The above-mentioned recording emotion guiding audio can be a recording emotion guiding voice and / or a recording emotion guiding music, such as an emotion guiding prompt voice that prompts the sound producer to adjust the emotion of pronunciation, or music that matches the emotion of the text content of the recording text or the emotion of the target voice emotion effect.
[0034] As an optional implementation, when the recording text and the target voice emotion effect are determined, the recording text and the target voice emotion effect are analyzed to determine the context of the recording text and the emotion tone of the target voice emotion effect.
[0035] Then, recording emotion guidance information is generated that matches the context of the recording text and / or has the same emotional tone as the target speech emotion effect.
[0036] For example, when the recorded text is a novel text, the emotion that matches the novel content is determined based on the context of the novel content, and then guidance information is generated to guide the speaker to read the novel with this emotion, such as generating contextual prompt text or contextual prompt sound, or searching for music, videos, etc. with corresponding contexts as recording emotion guidance information.
[0037] For example, when the target speech emotion is sadness, you can search for movies, TV shows, or music with strong sadness to generate recording emotion guidance information. This recording emotion guidance information will help guide the speaker to read the recording text with sadness.
[0038] As an optional situation, when the context of the recorded text is different from the target voice emotional effect, the recorded emotional guidance information can be generated based on the context of the recorded text or based on the target voice emotional effect according to the setting.
[0039] It can be understood that embodiments of the present application can automatically generate recording emotion guidance information that matches the recording text and the target speech emotion effect based on the recording text and the target speech emotion effect. This method of generating recording emotion guidance information eliminates the need to deliberately require the speaker to follow standard pronunciation or pronounce according to a certain emotion. Instead, corresponding recording emotion guidance information can be generated in real time based on the content of the recording text and the required speech emotion effect. This allows for more real-time and dynamic guidance of the speaker's pronunciation emotion.
[0040] S102: Outputting the recorded emotion guidance information so that the target speaker can read the recorded text under the guidance of the recorded emotion guidance information.
[0041] Specifically, after generating recording emotion guidance information that matches the recording text and / or the target voice emotion effect, the generated recording emotion guidance information is output to the target speaker. The target speaker reads the recording text emotionally under the guidance of the recording emotion guidance information.
[0042] For example, assuming that the above-mentioned recorded emotion guidance information is a recorded emotion prompt text, when the speaker reads the recorded text, the recorded emotion prompt text is output in real time. After the speaker sees the recorded emotion prompt text, he reads according to the recorded emotion prompted by the text, and the speaker can obtain the emotional reading voice issued by the speaker.
[0043] For another example, assuming that the above-mentioned recorded emotional guidance information is an audio or video with the same emotional tone as the target voice emotional effect, when the speaker reads the recorded text, the audio or video is played to immerse the speaker in the emotional atmosphere of the target voice emotional effect and read aloud, thereby guiding the speaker's reading to also carry the same emotional tone.
[0044] The embodiment of the present application outputs the recording emotion guidance information generated in real time according to the recording text and the target voice emotional effect to the speaker, so that the speaker can be guided by the recording emotion information in real time during the process of reading the recording text, thereby enabling the speaker to subtly adjust the reading emotion according to the recording emotion guidance information in real time during the entire process of reading the recording text, and obtain an emotionally rich reading voice.
[0045] S103: Collect the target speaker's reading voice of the recorded text to obtain voice data corresponding to the recorded text.
[0046] Specifically, when the target speaker reads the recorded text aloud, the target speaker's reading voice is collected and recorded, that is, voice data corresponding to the recorded text is obtained.
[0047] As can be seen from the above description, when generating speech, the embodiment of the present application can generate and output recording emotion guidance information based on the recorded text and the target speech emotion effect, so that the speaker can read the recorded text under the guidance of the recorded emotion guidance information. At this time, the speaker's reading of the recorded text is collected to obtain speech data corresponding to the recorded text.
[0048] In the above-mentioned speech generation process, recording emotion guidance information can be generated in real time according to the recording text and the target speech emotional effect, thereby providing the speaker with a recording emotion reference, so that the speaker can more intuitively and accurately know with what emotion the recording text should be read aloud, and then can collect speech data with various emotional colors. On the one hand, this method can provide convenience for the speaker to pronounce, that is, it can automatically generate recording emotion guidance information, which is convenient for the speaker to know how to adjust the pronunciation emotion; on the other hand, the technical solution of the embodiment of the present application can be used to generate speech data with various emotions. These speech data can be used as speech materials for synthesizing speech with emotional colors, which can improve the speech synthesis effect, and thus help improve the human-computer interaction effect based on speech.
[0049] Generally, when a speaker expresses strong emotion through speech, the speaker will change the timbre, so that the timbre of the speech with strong emotion deviates greatly from the normal timbre of the speaker, and if the timbre of the speech of the same speaker is greatly different, it will affect the effect of training the speech synthesis system using the speech of the speaker as material, and will lead to instability of the speech synthesis system.
[0050] Therefore, it should be ensured that the timbres of the speeches of various emotions of the same speaker are consistent. In order to achieve the above-mentioned purpose, after outputting the recording emotion guide information to guide the target speaker to read the recording text with emotion to obtain the speech data corresponding to the recording text, as a preferred embodiment, the embodiment of the application also performs timbre consistency verification on the speech data with emotional color to determine whether it conforms to the normal timbre of the target speaker.
[0051] For example, the embodiment of the application takes the timbre of the speech of the target speaker under the set emotion as the normal timbre of the target speaker.
[0052] On this basis, after collecting the reading speech of the target speaker on the recording text, the timbre of the reading speech of the target speaker on the recording text is detected to determine whether it is consistent with the timbre of the reading speech of the target speaker under the set emotion.
[0053] If they are consistent, the reading speech of the target speaker on the recording text can be stored.
[0054] If they are not consistent, it means that the timbre of the reading speech of the target speaker on the recording text has deviated from the normal timbre of the target speaker, and if the reading speech of the target speaker on the recording text is used to train the speech synthesis model, it may lead to instability of the speech synthesis model. Therefore, the embodiment of the application directly discards the reading speech of the target speaker on the recording text.
[0055] As a preferred embodiment, the reading speech of the target speaker under the set emotion is preferably the reading speech of the target speaker under the neutral emotion, because the speech of the speaker under the neutral emotion state is the most representative of the timbre of the speaker.
[0056] Therefore, the embodiment of the application determines whether the phonemes of the reading speech of the target speaker on the recording text are consistent with the normal timbre of the speaker by detecting whether the timbre of the reading speech of the target speaker on the recording text is consistent with the timbre of the reading speech of the target speaker under the neutral emotion.
[0057] As an optional detection method, first, speaker representation information is extracted from the reading speech of the target speaker on the recording text and the reading speech of the target speaker under the set emotion (preferably the reading speech under the neutral emotion), respectively.
[0058] For example, feature extraction is performed on the target speaker's reading speech of the recorded text and the target speaker's reading speech with set emotions, and then the extracted speech features are input into the speaker model respectively to identify the corresponding speaker representation features.
[0059] Then, the similarity between the speaker representation information extracted from the target speaker's reading of the recorded text and the speaker representation information extracted from the target speaker's reading of the set emotion is calculated. In other words, the similarity between the speaker representation features extracted from each of the two aforementioned speech is calculated.
[0060] If the calculated similarity is greater than the preset similarity threshold, it can be determined that the timbre of the target speaker's reading voice of the recorded text is consistent with the timbre of the target speaker's reading voice with the set emotion; otherwise, it can be determined that the timbre of the target speaker's reading voice of the recorded text is inconsistent with the timbre of the target speaker's reading voice with the set emotion.
[0061] When generating speech, the embodiments of the present application primarily rely on a speaker reading aloud the recorded text to produce speech corresponding to the recorded text. However, different speakers have different timbre and vocal expressiveness. Therefore, to obtain speech with a specific emotional tone, a speaker whose timbre matches that emotional tone and who has good vocal expressiveness for that emotional tone should be selected.
[0062] In order to facilitate the screening of speakers, the embodiment of the present application pre-collects detailed portrait information of each candidate speaker, and then when selecting the target speaker, the appropriate target speaker is selected from the candidate speakers based on the recording text and the target voice emotional effect, combined with the portrait information of each candidate.
[0063] Among them, the above-mentioned portrait information of the candidate speaker includes the basic information of the speaker, such as gender, age, nationality, accent, etc., and also includes the speaker's personality information, such as personality, occupation, audience, etc. In addition, it also includes the speaker's pronunciation characteristics information, such as timbre, pronunciation style, desired pronunciation role and / or pronunciation style, undesirable pronunciation role and / or pronunciation style, etc.
[0064] Table 1 shows an example of speaker portrait information.
[0065] Table 1
[0066]
[0067] The embodiment of the present application profiles the candidate translators from a more comprehensive perspective, so that the characteristics, abilities and appropriate voice emotions of the candidate speakers can be comprehensively and intuitively determined through the portrait information of the candidate speakers, which is conducive to selecting a suitable speaker from multiple candidate speakers.
[0068] On the basis of clarifying the portrait information of each candidate speaker, when determining the recording text and the target voice emotional effect, the candidate speaker who can read the recording text and express the target voice emotional effect can be screened from the candidate speakers according to the portrait information of each candidate speaker as the first candidate speaker.
[0069] Then, the pronunciation effects of the first candidate speakers are evaluated based on their audition voices, and a target speaker is selected from the first candidate speakers based on the evaluation results of the pronunciation effects of the first candidate speakers.
[0070] As an optional embodiment, this embodiment evaluates the pronunciation effect of each first candidate speaker based on at least one of timbre, articulation, breath, rhythm, voice expressiveness, and speech synthesis effect. Preferably, the pronunciation effect of the first candidate speaker can be evaluated based on each of the above aspects to improve the comprehensiveness of the evaluation result.
[0071] The above-mentioned voice expressiveness includes emotional expressiveness and / or stylistic expressiveness.
[0072] As an exemplary implementation, when evaluating the pronunciation of a first candidate speaker, a pre-trained pronunciation evaluation model can be used to automatically evaluate the first candidate speaker's pronunciation. The evaluation results for each candidate speaker in various aspects are then synthesized, and the speaker with the best overall evaluation result is selected from the first candidate speakers as the target speaker.
[0073] The above-mentioned speech synthesis effect includes the effect of the speech synthesized based on the trial speech of the first candidate speaker.
[0074] In this embodiment of the present application, when evaluating the speech synthesis performance of the first candidate speaker, a simple speech synthesis system corresponding to the first candidate speaker is first trained using the first candidate speaker's test speech. This speech synthesis system is generally obtained by training an acoustic model, which includes but is not limited to an HMM model, a neural network model, and the like.
[0075] Then, the trial pronunciation voice of the first candidate pronunciation person is input into a corresponding speech synthesis system, speech is synthesized by using the trial pronunciation voice of the pronunciation person, and then speech synthesis effect evaluation is performed on the synthesized speech to obtain a speech synthesis effect evaluation result.
[0076] Exemplarily, the speech synthesis effect evaluation on the synthesized speech can also be realized by means of a pre-trained speech synthesis effect evaluation model, that is, the synthesized speech is input into the pre-trained speech synthesis effect evaluation model to obtain a speech synthesis effect evaluation result output by the model.
[0077] According to the speech synthesis effect evaluation result, a pronunciation effect evaluation result of the first candidate pronunciation person can be determined. For example, when the pronunciation effect evaluation of the first candidate pronunciation person is performed, if the evaluation is performed only from the aspect of speech synthesis effect, the speech synthesis effect evaluation result can be directly taken as the pronunciation effect evaluation result of the first candidate pronunciation person; if the pronunciation effect evaluation of the first candidate pronunciation person is performed from multiple aspects, the speech synthesis effect evaluation result determined through the above processing is comprehensively taken with evaluation results obtained through other aspects as the pronunciation effect evaluation result of the first candidate pronunciation person.
[0078] The pronunciation person is evaluated by means of speech synthesis effect in the embodiment of the application, the effect of the pronunciation of the pronunciation person when applied in speech synthesis can be more truly evaluated, thereby helping to select a pronunciation person with better speech synthesis effect.
[0079] On the other hand, when speech is generated according to a recording text and a speech library is constructed by using the generated speech, the quality of the recording text also directly affects the quality of the speech library.
[0080] In the production of the recording text, in addition to selecting texts of various scenes and various styles and moods, the phoneme coverage of the text also needs to be considered, for example, the Ngram coverage of the phoneme can be taken as an index to ensure that the produced text can effectively cover most phoneme combinations and ensure the effect of subsequent speech synthesis system training.
[0081] In addition, in some scenes or some moods, the speech synthesis system needs to have more colloquial expressions and can realize more personified communication, so a large number of colloquial texts are needed to produce the speech library or express certain speech moods through colloquial texts.
[0082] In view of the above situation, the colloquial recording text corresponding to the recording text is generated according to the recording text before the recording mood guiding information is generated according to the recording text and the target speech mood effect.
[0083] As an exemplary implementation, the embodiment of the present application pre-trains a spoken text generation model, which can generate spoken text by adding semantic words to the text input to the model.
[0084] Based on the above-mentioned spoken text generation model, the acquired recorded text is input into the spoken text generation model to obtain the spoken recorded text output by the model corresponding to the recorded text.
[0085] See also Figure 2 The spoken text generation model processing process shown in the figure first performs word segmentation on the input text, and then extracts features for each word W (for example, feature extraction can be performed through a recurrent neural network) to obtain a feature representation E for each word. Then, based on the feature representation of each word, the hidden layer h and the fully connected layer are processed to determine whether an interjection is inserted after each word. Among them, the colloquial interjections are all pre-set interjections.
[0086] The spoken text generation model can automatically generate spoken text from written text by conducting a certain amount of training on parallel text pairs of written and spoken text, thereby improving the efficiency of spoken text generation.
[0087] Furthermore, generating a spoken audio text corresponding to the audio text is more conducive to the speaker reading the audio text to obtain a speech with the target speech emotional effect.
[0088] Furthermore, after the spoken audio recording text is generated through the above-mentioned processing, the embodiment of the present application also performs text error correction on the generated spoken audio recording text to ensure the correctness of the spoken audio recording text.
[0089] As an optional implementation, the embodiment of the present application pre-trains a text correction model for performing text correction on spoken audio recordings.
[0090] The text error correction model can detect various types of text errors such as missing characters, typos, and misspellings from the text input into the model, and correct the detected text errors.
[0091] The above-mentioned text error correction model can use but is not limited to a text error correction solution based on a deep neural network for text error correction.
[0092] This text correction uses four editing tag types: C (copy the current input text to the output), D (delete the current input text), XC (copy the current input text to the output and add text X in front of it), and XH (replace the current input text with X).
[0093] Figure 3The text correction process of the text correction model is shown as an example. In this text correction process, in addition to retaining most of the original text content, the character "大" is added before the character "大", and the character "文" is replaced by "闻". Figure 3 The encoder and decoder of the text error correction model shown can generally adopt various deep neural network models. Through the above models, missing words, typos, misspellings, etc. in the recorded text can be detected and automatically corrected, greatly improving the efficiency of text error correction.
[0094] The voice data corresponding to the recorded text generated by the embodiment of the present application can be directly used to construct a voice library.
[0095] When building a speech database, it is usually necessary to annotate the speech data, such as word-phonetic alignment annotation and rhythm annotation.
[0096] In order to improve the efficiency of annotating speech data, the embodiment of the present application pre-trains a speech recognition model for performing phoneme recognition on the speech data, so that the speech data can be annotated with phonemes based on the phoneme recognition results output by the model.
[0097] See also Figure 4 As shown, the above-mentioned speech recognition model can be obtained by iteratively training a speaker-independent speech recognition system through the speech of the speaker.
[0098] The iterative process of the speaker-independent speech recognition system is as follows:
[0099] (1) The speaker's speech is first segmented by a speaker-independent speech recognition system to obtain the result of phoneme alignment, that is, the various phonemes recognized from the speaker's speech and the start and end times of the phonemes are obtained.
[0100] The speaker-independent speech recognition system may be, but is not limited to, a hidden Markov model or a deep neural network model.
[0101] (2) The phoneme segmentation results are then used to fine-tune the acoustic model of the speaker-independent speech recognition system. For example, the hidden Markov model can be trained using the maximum a posteriori probability criterion, while the deep neural network model can use the cross-entropy criterion to update some parameters of the neural network.
[0102] (3) Using the model obtained in step (2), repeat the operation in step 1 and iterate until the phoneme segmentation result no longer changes or changes very little.
[0103] By inputting the speech data obtained by the processing in the above step S103 into the above speech recognition model, a phoneme recognition result output by the model can be obtained. The phoneme recognition result includes the start and end position information of the phoneme.
[0104] Then, the speech data may be annotated with phonemes according to the phoneme recognition result. For example, according to the start and end position information of each phoneme in the phoneme recognition result, the speech data may be annotated with phonemes at the corresponding start and end positions.
[0105] On the other hand, when performing prosody annotation on speech data, the prosody of the speech data can be predicted with the help of a prosody prediction model, and then the speech data can be prosody-annotated according to the prosody prediction result.
[0106] The specific process of training the prosody prediction model, using the prosody prediction model to predict prosody, and performing prosody tagging based on the prosody prediction results can all be referred to the process of phoneme tagging, which will not be described in detail in this embodiment.
[0107] This embodiment, with the help of a speech recognition model and a prosody prediction model, can automatically obtain the phoneme recognition results and prosody prediction results of speech data, so that the speech data can be annotated with phonemes and prosody based on the phoneme recognition results and prosody prediction results. Compared with manual phoneme and prosody annotation, the annotation efficiency of this embodiment is higher.
[0108] Exemplary apparatus
[0109] Corresponding to the above-mentioned speech generation method, the present application embodiment also provides a speech generation device, see Figure 5 As shown, the device includes:
[0110] The information generating unit 100 is used to generate recording emotion guidance information based on the recording text and the target speech emotion effect;
[0111] The data output unit 110 is used to output the recorded emotion guidance information so that the target speaker can read the recorded text under the guidance of the recorded emotion guidance information;
[0112] The data collection unit 120 is configured to collect the target speaker's reading of the recorded text to obtain voice data corresponding to the recorded text.
[0113] As an optional implementation, generating recording emotion guidance information based on the recording text and the target voice emotion effect includes:
[0114] According to the recording text and the target voice emotional effect, recording emotion guidance information is generated that matches the context of the recording text and / or has the same emotional tone as the target voice emotional effect.
[0115] As an optional implementation manner, the recorded emotion guidance information includes at least one of a recorded emotion guidance video, a recorded emotion guidance audio, and a recorded emotion guidance text;
[0116] The recorded emotion-guiding audio includes at least one of recorded emotion-guiding voice and recorded emotion-guiding music.
[0117] As an optional implementation, the device further includes:
[0118] A data processing unit, configured to detect whether the timbre of the target speaker's reading of the recorded text is consistent with the timbre of the target speaker's reading of the set emotion;
[0119] If the timbre of the target speaker's reading voice of the recorded text is inconsistent with the timbre of the target speaker's reading voice with the set emotion, the target speaker's reading voice of the recorded text is discarded.
[0120] As an optional implementation, detecting whether the timbre of the target speaker's reading of the recorded text is consistent with the timbre of the target speaker's reading of the set emotion includes:
[0121] Extracting speaker representation information from the target speaker's reading of the recorded text and the target speaker's reading with set emotions;
[0122] Calculating the similarity between the speaker representation information extracted from the target speaker's reading of the recorded text and the speaker representation information extracted from the target speaker's reading of the recorded text with a set emotion;
[0123] If the calculated similarity is greater than a preset similarity threshold, it is determined that the timbre of the target speaker's reading voice of the recorded text is consistent with the timbre of the target speaker's reading voice with the set emotion.
[0124] As an optional implementation, the device further includes:
[0125] The speaker selection unit is used to select the target speaker from the candidate speakers based on the recording text, the target speech emotional effect, and the profile information of the candidate speakers;
[0126] The candidate speaker's profile information includes basic information about the speaker, speaker's personality information, and speaker's pronunciation characteristics;
[0127] The basic information of the voice actor includes at least one of gender, age, nationality and accent; the personal information of the voice actor includes at least one of personality, occupation and audience; and the voice characteristic information of the voice actor includes at least one of tone, voice style, desired voice role and / or voice style, undesired voice role and / or voice style.
[0128] As an optional implementation, the target voice actor is selected from the candidate voice actors according to the recorded text, the target voice emotion effect and the portrait information of the candidate voice actors, and the method comprises the following steps of:
[0129] selecting a first candidate voice actor from the candidate voice actors according to the recorded text, the target voice emotion effect and the portrait information of the candidate voice actors, the first candidate voice actor being capable of reading the recorded text and expressing the target voice emotion effect;
[0130] performing voice effect evaluation on each first candidate voice actor according to the trial voice of each first candidate voice actor;
[0131] selecting the target voice actor from the first candidate voice actors according to the voice effect evaluation results of the first candidate voice actors.
[0132] As an optional implementation, the voice effect evaluation on each first candidate voice actor according to the trial voice of each first candidate voice actor comprises the following steps of:
[0133] performing voice effect evaluation on each first candidate voice actor from at least one of tone, pronunciation, breath, rhythm, voice expressiveness and voice synthesis effect according to the trial voice of each first candidate voice actor;
[0134] The voice expressiveness includes at least one of emotion expressiveness and style expressiveness; and the voice synthesis effect includes the effect of the synthesized voice based on the trial voice.
[0135] As an optional implementation, the voice effect evaluation on the first candidate voice actor from the voice synthesis effect according to the trial voice of the first candidate voice actor comprises the following steps of:
[0136] training a voice synthesis system corresponding to the first candidate voice actor by using the trial voice of the first candidate voice actor;
[0137] performing voice synthesis effect evaluation on the synthesized voice output by the voice synthesis system corresponding to the first candidate voice actor to obtain a voice synthesis effect evaluation result;
[0138] determining the voice effect evaluation result of the first candidate voice actor according to the voice synthesis effect evaluation result.
[0139] As an optional implementation, the apparatus further comprises:
[0140] a data preprocessing unit configured to generate a spoken text corresponding to the recorded text according to the recorded text.
[0141] As an optional implementation, generating a spoken text corresponding to the recorded text according to the recorded text comprises:
[0142] inputting the recorded text into a pre-trained spoken text generation model to obtain the spoken text corresponding to the recorded text;
[0143] wherein the spoken text generation model has a function of generating a spoken text by adding a mood word to an input text.
[0144] As an optional implementation, the data preprocessing unit is further configured to:
[0145] inputting the spoken text into a pre-trained text correction model to perform text correction processing on the generated spoken text;
[0146] wherein the text correction model has at least a function of detecting at least one text error of a missing word, a wrong word, and an incorrect word from an input text, and correcting the detected text error.
[0147] As an optional implementation, the apparatus further comprises:
[0148] a first speech annotation unit configured to input speech data corresponding to the recorded text into a pre-trained speech recognition model to obtain a phoneme recognition result of the speech data;
[0149] performing phoneme annotation on the speech data according to the phoneme recognition result.
[0150] As an optional implementation, the apparatus further comprises:
[0151] a second speech annotation unit configured to input speech data corresponding to the recorded text into a pre-trained prosody prediction model to obtain a prosody prediction result of the speech data;
[0152] performing prosody annotation on the speech data according to the prosody prediction result.
[0153] The speech generation device provided in this embodiment is based on the same concept as the speech generation method provided in the above embodiments of this application. It can execute the speech generation method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects. For technical details not fully described in this embodiment, please refer to the specific processing content of the speech generation method provided in the above embodiments of this application, and will not be repeated here.
[0154] Exemplary electronic device
[0155] Another embodiment of the present application also provides a speech generating device, see Figure 6 As shown, the device includes:
[0156] Memory 200 and processor 210;
[0157] The memory 200 is connected to the processor 210 and is used to store computer programs and data;
[0158] The processor 210 is configured to generate recording emotion guidance information according to the recording text and the target speech emotion effect by running the computer program stored in the memory 200 .
[0159] Specifically, the above-mentioned speech generating device further includes: a bus, a communication interface 220 , an input device 230 and an output device 240 .
[0160] The processor 210, the memory 200, the communication interface 220, the input device 230 and the output device 240 are interconnected via a bus.
[0161] A bus may include a pathway that transfers information between components of a computer system.
[0162] Processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, or the like, or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of programs in accordance with the present invention. Alternatively, it can be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware components.
[0163] The processor 210 may include a main processor, and may also include a baseband chip, a modem, and the like.
[0164] The memory 200 stores a program for executing the technical solution of the present invention, and may also store an operating system and other key businesses. Specifically, the program may include program code, and the program code includes computer operating instructions. More specifically, the memory 200 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, disk storage, flash, and the like. The memory 200 has a data storage function and can be used to store the recorded emotional guidance information generated by the processor, as well as to store various types of data collected or received by the input device 230.
[0165] The output device 240 may include a device that allows information to be output to a user, such as a display screen, a printer, a speaker, etc. The processor 210 may send the generated recorded emotion guidance information to the output device 240, and the output device 240 may output the recorded emotion guidance information to the user, for example, to the target speaker, so that the target speaker reads the recorded text under the guidance of the recorded emotion guidance information.
[0166] The input device 230 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor. The user can input the recorded text and select the target voice emotion effect through the input device 230. The input device 230 sends the received recorded text and the target voice emotion effect information to the processor 210, so that the processor 210 generates recorded emotion guidance information based on the recorded text and the target voice emotion effect information.
[0167] In the embodiment of the present application, the input device 230 includes a microphone, which is used to collect the target speaker's voice reading of the recorded text to obtain voice data corresponding to the recorded text. The voice data collected by the input device 230 is ultimately stored in the memory 200.
[0168] The communication interface 220 may include any device such as a transceiver to communicate with other devices or communication networks, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.
[0169] In addition, the processor 210 executes the program stored in the memory 200 and calls other devices to implement the various steps of any speech generation method provided in the above embodiments of the present application.
[0170] Specifically, the processor 210 generates recording emotion guidance information based on the recording text and the target voice emotion effect, including:
[0171] According to the recording text and the target voice emotional effect, recording emotion guidance information is generated that matches the context of the recording text and / or has the same emotional tone as the target voice emotional effect.
[0172] As an optional implementation manner, the recorded emotion guidance information includes at least one of a recorded emotion guidance video, a recorded emotion guidance audio, and a recorded emotion guidance text;
[0173] The recorded emotion-guiding audio includes at least one of recorded emotion-guiding voice and recorded emotion-guiding music.
[0174] As an optional implementation manner, the processor 210 is further configured to:
[0175] Detecting whether the timbre of the target speaker's reading of the recorded text is consistent with the timbre of the target speaker's reading of the set emotion;
[0176] If the timbre of the target speaker's reading voice of the recorded text is inconsistent with the timbre of the target speaker's reading voice with the set emotion, the target speaker's reading voice of the recorded text is discarded.
[0177] As an optional implementation, detecting whether the timbre of the target speaker's reading of the recorded text is consistent with the timbre of the target speaker's reading of the set emotion includes:
[0178] Extracting speaker representation information from the target speaker's reading of the recorded text and the target speaker's reading with set emotions;
[0179] Calculating the similarity between the speaker representation information extracted from the target speaker's reading of the recorded text and the speaker representation information extracted from the target speaker's reading of the recorded text with a set emotion;
[0180] If the calculated similarity is greater than a preset similarity threshold, it is determined that the timbre of the target speaker's reading voice of the recorded text is consistent with the timbre of the target speaker's reading voice with the set emotion.
[0181] As an optional implementation manner, before outputting the recorded emotion guidance information, the processor 210 is further configured to:
[0182] The target speaker is selected from the candidate speakers based on the recording text, the target speech emotional effect, and the profile information of the candidate speakers;
[0183] The candidate speaker's profile information includes basic information about the speaker, speaker's personality information, and speaker's pronunciation characteristics;
[0184] The basic information of the speaker includes at least one of gender, age, nationality and accent; the personal information of the speaker includes at least one of personality, occupation and audience; the pronunciation characteristic information of the speaker includes at least one of timbre, pronunciation style, desired pronunciation role and / or pronunciation style, and undesired pronunciation role and / or pronunciation style.
[0185] As an optional implementation, based on the recording text, the target speech emotional effect, and the profile information of the candidate speakers, the target speaker is selected from the candidate speakers, including:
[0186] Based on the recorded text, the target speech emotional effect, and the portrait information of the candidate speakers, selecting a first candidate speaker who can read the recorded text and express the target speech emotional effect from the candidate speakers;
[0187] Evaluating the pronunciation effect of each first candidate speaker based on the audition voice of each first candidate speaker;
[0188] A target speaker is selected from the first candidate speakers based on the pronunciation effect evaluation results of the first candidate speakers.
[0189] As an optional implementation, based on the audition voice of each first candidate speaker, the pronunciation effect of each first candidate speaker is evaluated, including:
[0190] Based on the audition speech of each first candidate speaker, evaluating the pronunciation effect of each first candidate speaker in terms of at least one of timbre, articulation, breath, rhythm, vocal expressiveness, and speech synthesis effect;
[0191] The voice expressiveness includes at least one of emotional expressiveness and style expressiveness; and the speech synthesis effect includes the effect of the speech synthesized based on the test speech.
[0192] As an optional implementation, based on the trial voice of the first candidate speaker, the pronunciation effect of the first candidate speaker is evaluated from the perspective of speech synthesis effect, including:
[0193] Using the trial voice of the first candidate speaker, training a speech synthesis system corresponding to the first candidate speaker;
[0194] performing a speech synthesis effect evaluation on the synthesized speech output by the speech synthesis system corresponding to the first candidate speaker to obtain a speech synthesis effect evaluation result;
[0195] Determine a pronunciation effect evaluation result for the first candidate speaker based on the speech synthesis effect evaluation result.
[0196] As an optional implementation, before generating the recording emotion guidance information based on the recording text and the target voice emotion effect, the processor 210 is further configured to:
[0197] Based on the audio recording text, a spoken audio recording text corresponding to the audio recording text is generated.
[0198] As an optional implementation, generating a spoken audio text corresponding to the audio text according to the audio text includes:
[0199] Input the recorded text into a pre-trained spoken text generation model to obtain a spoken recorded text corresponding to the recorded text;
[0200] The spoken text generation model has the function of generating spoken text by adding modal particles to the text of the input model.
[0201] As an optional implementation manner, the processor 210 is further configured to:
[0202] Inputting the spoken audio recording text into a pre-trained text error correction model to perform text error correction processing on the generated spoken audio recording text;
[0203] The text error correction model has at least the function of detecting at least one text error among missing characters, wrong characters and misspellings from the text of the input model, and correcting the detected text errors.
[0204] As an optional implementation manner, the processor 210 is further configured to:
[0205] Inputting the voice data corresponding to the recorded text into a pre-trained voice recognition model to obtain a phoneme recognition result for the voice data;
[0206] The speech data is phoneme-tagged according to the phoneme recognition result.
[0207] As an optional implementation manner, the processor 210 is further configured to:
[0208] Inputting the speech data corresponding to the recorded text into a pre-trained prosody prediction model to obtain a prosody prediction result for the speech data;
[0209] Prosody annotation is performed on the speech data according to the prosody prediction result.
[0210] The speech generation device provided in this embodiment is based on the same concept as the speech generation method provided in the above embodiments of this application. It can execute the speech generation method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects. For technical details not fully described in this embodiment, please refer to the specific processing content of the speech generation method provided in the above embodiments of this application, and will not be repeated here.
[0211] Exemplary computer program product and storage medium
[0212] In addition to the above-mentioned methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions. When the computer program instructions are executed by a processor, the processor executes the steps of the speech generation method described in the above-mentioned "Exemplary Method" section of this specification and can achieve corresponding technical effects.
[0213] The computer program product may be written in any combination of one or more programming languages to implement the program code for performing the operations of the embodiments of the present application, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0214] In addition, an embodiment of the present application may also be a storage medium on which a computer program is stored. The computer program is used by a processor to execute the steps of the speech generation method described in the above "Exemplary Method" section of this specification and achieve corresponding technical effects.
[0215] For the sake of simplicity, the aforementioned method embodiments are described as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0216] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similarities between the various embodiments can be referred to in conjunction with each other. For device embodiments, since they are generally similar to method embodiments, their description is relatively simple, and for relevant details, reference can be made to the description of the method embodiments.
[0217] The steps in the methods of each embodiment of the present application can be adjusted in sequence, merged, and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.
[0218] The modules and sub-modules in the devices and terminals of the various embodiments of the present application can be merged, divided, and deleted according to actual needs.
[0219] In the several embodiments provided in this application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or submodules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or module, which can be electrical, mechanical or other forms.
[0220] Modules or submodules described as separate components may or may not be physically separate, and components described as modules or submodules may or may not be physical modules or submodules. That is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules may be selected to achieve the objectives of this embodiment based on actual needs.
[0221] In addition, each functional module or submodule in each embodiment of the present application may be integrated into a processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into a single module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or software functional modules or submodules.
[0222] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0223] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, software executed by a processor, or a combination of the two. The software may be stored in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0224] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0225] The above description of the disclosed embodiments will enable those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.
Claims
1. A speech generation method, characterized in that: include: Analyze the recorded text and the target voice emotional effect, determine the context of the recorded text and the emotional tone of the target voice emotional effect, and generate recorded emotional guidance information in real time that matches the context of the recorded text and has the same emotional tone as the target voice emotional effect; Outputting the recorded emotion guidance information so that the target speaker reads the recorded text under the guidance of the recorded emotion guidance information; The target speaker's voice reading of the recorded text is collected to obtain voice data corresponding to the recorded text.
2. The method according to claim 1, characterized in that The recorded emotion guidance information includes at least one of a recorded emotion guidance video, a recorded emotion guidance audio, and a recorded emotion guidance text; The recorded emotion-guiding audio includes at least one of recorded emotion-guiding voice and recorded emotion-guiding music.
3. The method according to claim 1, characterized in that The method further comprises: Detecting whether the timbre of the target speaker's reading of the recorded text is consistent with the timbre of the target speaker's reading of the set emotion; If the timbre of the target speaker's reading voice of the recorded text is inconsistent with the timbre of the target speaker's reading voice with the set emotion, the target speaker's reading voice of the recorded text is discarded.
4. The method according to claim 1, wherein Before outputting the recorded emotion guidance information, the method further includes: The target speaker is selected from the candidate speakers based on the recording text, the target speech emotional effect, and the profile information of the candidate speakers; The candidate speaker's profile information includes basic information about the speaker, speaker's personality information, and speaker's pronunciation characteristics; The basic information of the speaker includes at least one of gender, age, nationality and accent; the personal information of the speaker includes at least one of personality, occupation and audience; the pronunciation characteristic information of the speaker includes at least one of timbre, pronunciation style, desired pronunciation role and / or pronunciation style, and undesired pronunciation role and / or pronunciation style.
5. The method according to claim 1, characterized in that Before generating recording emotion guidance information based on the recording text and the target speech emotion effect, the method further includes: Based on the audio recording text, a spoken audio recording text corresponding to the audio recording text is generated.
6. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Inputting the voice data corresponding to the recorded text into a pre-trained voice recognition model to obtain a phoneme recognition result for the voice data; The speech data is phoneme-tagged according to the phoneme recognition result.
7. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Inputting the speech data corresponding to the recorded text into a pre-trained prosody prediction model to obtain a prosody prediction result for the speech data; Prosody annotation is performed on the speech data according to the prosody prediction result.
8. A speech generating device, characterized in that: include: An information generation unit is used to analyze the recorded text and the target voice emotional effect, determine the context of the recorded text and the emotional tone of the target voice emotional effect, and generate recorded emotional guidance information in real time that matches the context of the recorded text and has the same emotional tone as the target voice emotional effect; A data output unit, configured to output the recorded emotion guidance information so that the target speaker can read the recorded text under the guidance of the recorded emotion guidance information; The data collection unit is used to collect the target speaker's reading voice of the recorded text to obtain voice data corresponding to the recorded text.
9. A speech generating device, characterized in that: include: a processor, and a memory, a microphone, and an output device connected to the processor respectively; Wherein, the memory is used to store data and computer programs; The processor is configured to parse the recorded text and the target voice emotion effect by running the computer program in the memory, determine the context of the recorded text and the emotional tone of the target voice emotion effect, generate in real time recorded emotion guidance information that matches the context of the recorded text and has the same emotional tone as the target voice emotion effect, and send the generated recorded emotion guidance information to the output device; The output device is used to output the recorded emotion guidance information sent by the processor, so that the target speaker reads the recorded text under the guidance of the recorded emotion guidance information; The microphone is connected to the memory and is used for collecting the target speaker's reading voice of the recorded text, obtaining voice data corresponding to the recorded text, and storing the voice data in the memory.
10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the speech generation method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Speech synthesis method, device and system, and storage medium
CN109616094A
Sample data set generation method, device and equipment
CN114203160A