Data processing method and device in speech generation and electronic equipment

By annotating the locations of non-linguistic events in phoneme sequences and performing frame-level prosodic prediction and feature data processing, the problem of insufficient diversity and realism of non-linguistic events in existing technologies is solved, generating more natural and context-appropriate speech.

CN119785756BActive Publication Date: 2025-11-28NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411825028.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-11-28
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

Existing text-to-speech technologies lack diversity and realism when generating non-verbal events, resulting in unnatural and unrealistic generated speech. In particular, non-verbal events such as laughter and crying cannot exhibit diversity in different contexts.

Method used

By determining the location of non-linguistic events in the phoneme sequence, frame-level prosodic prediction and non-linguistic feature data prediction or extraction are performed to generate speech signals containing diverse, realistic and natural non-linguistic events.

Benefits of technology

It achieves precise control over the diversity and authenticity of non-verbal events during speech generation, resulting in more natural and context-appropriate speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785756B_ABST
    Figure CN119785756B_ABST
Patent Text Reader

Abstract

The application provides a speech synthesis method and device, electronic equipment and a computer readable storage medium, wherein the method comprises: determining a phoneme sequence corresponding to text of to-be-generated speech, wherein a position of a non-language event to be added is marked in the phoneme sequence; performing frame number prediction and prosody prediction on each phoneme and the non-language event according to the phoneme sequence to obtain first phoneme feature data of a frame level with added prosody information; determining frame-level non-language feature data to be added in the to-be-generated speech; and processing the first phoneme feature data into speech signals with the added non-language event according to the non-language feature data. Therefore, the data processing method provided in the speech generation embodiment of the application can generate speech containing diversified and natural non-language events.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a data processing method and device in voice generation, electronic equipment and computer readable storage medium. BACKGROUND

[0002] Text-to-Speech (TTS) technology is a technology that converts written text into spoken language expression, aiming to enable machines to imitate human speech in a natural and fluent manner, thereby realizing high-quality voice communication in human-computer interaction. Although existing TTS technology has made remarkable achievements in naturalness of sound quality and emotional expression, it performs poorly in handling non-verbal non-verbal events (such as laughter, crying, coughing, etc.). These sound elements, although not a direct component of language information, are related to whether the generated voice has rich levels of communication and emotional expression. Taking laughter as an example, it is not only a natural expression of human emotion, but also reflects the speaker's psychological state and interactive manner.

[0003] Currently, the industry usually represents non-verbal events as a specific language unit, for example, represents laughter as a phoneme, marks a specific language unit in the text to indicate the position of laughter, and then inserts the corresponding laughter when generating voice. However, non-verbal events are diverse and complex, such as laughter can be divided into smile, sneer, wry smile, and wild laughter, etc. In this way, non-verbal events are represented as an independent phoneme, which leads to the same non-verbal event in different context texts showing the same pronunciation, lacking the inherent diversity of non-verbal events, making the generated voice lack naturalness and authenticity.

[0004] Therefore, how to accurately control and generate diverse non-verbal events while generating natural voice has become a technical problem to be solved. SUMMARY

[0005] The present application provides a voice synthesis method, device, electronic equipment and computer readable storage medium, which can generate voice containing diverse and authentic non-verbal events. The specific scheme is as follows:

[0006] In a first aspect, the present application provides a data processing method in voice generation, the method comprising:

[0007] determining a phoneme sequence corresponding to the text of the voice to be generated, wherein the phoneme sequence is marked with the position of the non-verbal event to be added;

[0008] According to the phoneme sequence, frame number prediction and prosody prediction are performed on each phoneme and the non-language event respectively to obtain first phoneme feature data with frame level and added prosody information;

[0009] The second determining unit is configured to determine frame level non-language feature data to be added in the speech to be generated.

[0010] The processing unit is configured to process the first phoneme feature data into a speech signal with the non-language event added according to the non-language feature data.

[0011] In a second aspect, an embodiment of the present application provides a data processing apparatus in speech generation, and the apparatus comprises:

[0012] The first determining unit is configured to determine a phoneme sequence corresponding to a text of speech to be generated, wherein the phoneme sequence is marked with a position of a non-language event to be added;

[0013] The prediction unit is configured to perform frame number prediction and prosody prediction on each phoneme and the non-language event respectively according to the phoneme sequence to obtain first phoneme feature data with frame level and added prosody information.

[0014] The second determining unit is configured to determine frame level non-language feature data to be added in the speech to be generated.

[0015] The processing unit is configured to process the first phoneme feature data into a speech signal with the non-language event added according to the non-language feature data.

[0016] In a third aspect, the present application further provides an electronic device, comprising:

[0017] a processor; and

[0018] a memory configured to store a data processing program, and the electronic device is powered on and runs the program through the processor to execute the method in the first aspect.

[0019] In a fourth aspect, an embodiment of the present application further provides a computer readable storage medium storing a data processing program, and the program is run by a processor to execute the method in the first aspect.

[0020] Compared with the prior art, the present application has the following advantages:

[0021] The data processing method in speech generation provided by the embodiment of the present application comprises the following steps: determining a phoneme sequence corresponding to a text of speech to be generated, wherein the phoneme sequence is marked with a position of a non-language event to be added; performing prediction of frame numbers occupied by each phoneme and prosody prediction on each phoneme and the non-language event according to the phoneme sequence to obtain first phoneme feature data of frame level with added prosody information; determining frame-level non-language feature data to be added in the speech to be generated; and processing the first phoneme feature data into speech signals with the added non-language event according to the non-language feature data.

[0022] It can be seen that the data processing method in speech generation provided by the embodiment of the present application can predict the frame numbers occupied by each phoneme and the frame numbers occupied by the non-language event based on the phoneme sequence, and can predict the prosody information of each frame, thereby obtaining the first phoneme feature data of frame level with added prosody information, because the phoneme sequence comprises each phoneme and is marked with the position of the non-language event to be added. Then, the frame-level non-language feature data to be added in the speech to be generated is determined, because the non-language feature data is fine-grained data representing fine distribution information of the non-language event. Therefore, the first phoneme feature data can be processed into speech signals containing real and natural non-language events according to the non-language feature data. In addition, the non-language feature data can be predicted from the text or extracted from the preset audio, and different non-language events can be expressed by different texts or different preset audios, thereby realizing the diversity of non-language events. Therefore, the data processing method in speech generation provided by the embodiment of the present application can generate speech containing diversified and real natural non-language events. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 is an architecture diagram of a speech generation system provided by the embodiment of the present application;

[0024] Figure 2 is a flowchart of a data processing method in speech generation provided by the embodiment of the present application;

[0025] Figure 3 is a structure diagram of a prosody prediction model applied by the data processing method in speech generation provided by the embodiment of the present application;

[0026] Figure 4 is a structure block diagram of an example of a data processing apparatus in speech generation provided by the embodiment of the present application;

[0027] Figure 5 is a structure block diagram of an example of an electronic device for data processing provided by the embodiment of the present application. DETAILED DESCRIPTION

[0028] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be practiced without the specific details. In other instances, well-known methods have not been described in detail in order not to unnecessarily obscure aspects of the present application.

[0029] It should be noted that the terms "first", "second", "third", and the like in the description and in the claims of the present application are used for distinguishing between similar objects and not necessarily for describing a sequential or chronological order. The use of such terms in the description is therefore not to be construed as implying a specific order or chronology. Further, the terms "comprises", "comprising", "has", "having", "includes", "including", and the like, are to be construed open- ended, i.e., to mean including but not limited to, to indicate the presence of what follows, illustrative and non-exclusive listing of elements, and allow for other elements to be present or to be added, and are not to be construed as indicating a limitation of what follows. In addition, it is to be understood that the application can be practiced without the specific details which are set forth in the description and in the claims.

[0030] It should be understood that, in the embodiments of the present application, "at least one" means one or more, and "multiple" means two or more than two. "And / or" is only a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. The character " / " generally represents an "or" relationship between the front and rear associated objects. "Including A, B and / or C" means including any one or any two or three of A, B and C.

[0031] It should be understood that, in the embodiments of the present application, "B corresponding to A", "B corresponding to A", "A corresponding to B" or "B corresponding to A" means that B is associated with A, and B can be determined according to A. Determining B according to A does not mean that B is determined only according to A, but also can be determined according to A and / or other information.

[0032] Before the embodiments of the present application are described in detail, the prior art is further described first.

[0033] Text-to-Speech (TTS) technology aims to synthesize text into natural and understandable speech. With the development of deep learning, modern neural network-based TTS models can generate speech signals with highly similar timbre and certain emotions. However, human speech not only contains text-related information, but also contains non-verbal events such as laughter, crying and coughing. These non-verbal events can convey various emotions or intentions in communication.

[0034] In the related art, the following two methods are commonly used to generate non-verbal events during speech synthesis. Method one: these non-verbal events are simplified into specific language units (such as phonemes or special markers). Method two: a pre-training-fine-tuning scheme based on a language model is used to pre-train a conditional language model on large-scale data and fine-tune it on small-scale high-quality labeled data to achieve better modeling effect of non-verbal events.

[0035] However, in human speech, non-verbal events are diverse and complex, such as laughter can be subdivided into smiling, mocking, wry smile, and wild laughter. The above method one ignores the diversity and complexity of non-verbal events themselves, and the same non-verbal event (such as laughter) is expressed as the same pronunciation in different context texts, resulting in a lack of diversity, authenticity and expressiveness of the generated non-verbal events. Although the above method two can improve the performance of the model to some extent, obtaining accurate labeled high-quality fine-tuning data requires a lot of manpower and resources, and the cost is high. Moreover, the complexity of the model and the demand for computing resources increase the actual difficulty of technology landing.

[0036] Therefore, how to generate speech containing diverse and authentic non-verbal events at a lower cost becomes extremely important.

[0037] Based on the above reasons, in order to be able to generate speech containing diverse and authentic non-verbal events at a lower cost, the first embodiment of the present application provides a data processing method in speech generation, which is applied to an electronic device. The electronic device can be a desktop computer, a notebook computer, a mobile phone, a tablet computer, an electronic watch, etc., or other electronic devices capable of data processing in speech synthesis. The present application does not make specific limitations.

[0038] The speech generation method provided by the present application can be used for intelligent customer service, intelligent robots, telephone customer service, speech generation and intelligent question answering in education and training, speech reply to players and game audio generation in game interaction, and answers to user questions by life assistants, etc. It can also be used for speech generation in other scenarios, and the present application does not make specific limitations.

[0039] For example, the scheme provided by the present application can be applied to intelligent dialogue of artificial intelligence voice assistants, and the data processing method in speech generation provided by the present application can be applied to add context-compliant and diverse laughter, crying or coughing sounds to the speech to be answered by artificial intelligence voice assistants.

[0040] Before introducing the data processing method in speech generation provided by the first embodiment of the present application, the speech generation system architecture to which the data processing method in speech generation provided by the present embodiment is applied is first introduced. As shown in Figure 1The diagram shown is an architecture diagram of the speech generation system provided in this application embodiment. The speech generation system 10 includes an acoustic feature determination module 11 and a non-linguistic feature determination module 12. The acoustic feature determination module 11 includes a text encoder 111, a prosody prediction model 112, a decoder 113, a non-linguistic event modeling model 116, and a vocoder 117. The non-linguistic feature determination module 12 includes a non-linguistic feature prediction model 121 and a non-linguistic feature extraction model 122. Specifically, the phoneme embedding sequence contains markers indicating the location of the non-linguistic event to be added. The text encoder 111 converts the phoneme embedding sequence into a phoneme hidden sequence. The prosody prediction model 112 predicts the duration and prosody of each phoneme and the duration and prosody of the non-linguistic event based on the phoneme hidden sequence, thereby converting the phoneme hidden sequence into frame-level first phoneme feature data with added prosody information. The decoder 113 decodes the first phoneme feature data to obtain a first acoustic spectrum. The non-linguistic feature prediction model 121 predicts frame-level non-linguistic feature data from the first phoneme feature data. The non-linguistic feature extraction model 122 extracts frame-level non-linguistic feature data from the recording (which is a preset audio during the inference stage). The non-linguistic event modeling model 116 processes the first acoustic spectrum based on the non-linguistic feature data predicted by the non-linguistic feature prediction model 121 or the non-linguistic feature data extracted by the non-linguistic feature extraction model 122 to obtain a second acoustic spectrum with added non-linguistic events. The vocoder 117 converts the second acoustic spectrum into a speech signal in the time domain. It is important to note that Figure 1 The various models described herein will be detailed in the description of the first embodiment of this application.

[0041] It should be noted that, in the embodiments of this application, the executing entity of the data processing method in speech generation can be a terminal device or a server, wherein the terminal device can be a local terminal device. The embodiments of this application do not limit the type of executing entity.

[0042] The technical solution of this application will be described in detail below through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0043] The following, combined with Figure 2 and Figure 3 This application introduces a data processing method for speech generation provided in its embodiments.

[0044] like Figure 2 As shown, the data processing method for speech generation provided in this application includes the following steps S101 to S104.

[0045] In step S101, a phoneme sequence corresponding to the text of the speech to be generated is determined, wherein the phoneme sequence is marked with a position of a non-verbal event to be added.

[0046] The text of the speech to be generated can be a piece of news, a piece of broadcast script, a piece of lecture notes, a piece of advertising script, a piece of announcement, a piece of commentary, or a reply text to user input content, and can be in Chinese, English, Japanese, Russian, French, or other languages.

[0047] The text has a corresponding phoneme sequence, which is marked with a position of a non-verbal event to be added. The non-verbal event refers to an event in the text to be generated that does not have corresponding text. The non-verbal event can include one of a laughter event, a crying event, a coughing event, and the like. A corresponding event marker can be added at the corresponding position of the phoneme sequence. For example, when the non-verbal event is a laughter event, the event marker can be <laugh>.

[0048] It can be understood that a phoneme is the smallest unit of speech divided according to the natural properties of speech, and is divided according to the pronunciation action in a syllable. One pronunciation action constitutes one phoneme. For example, the Chinese syllable "ah (a1)" has only one phoneme, "dai" has two phonemes d and ai4, and "a1" represents the tone of "a" as 1 tone, and "ai4" represents the tone of "ai" as 4 tone. Therefore, each text can be split into a corresponding phoneme sequence.

[0049] Optionally, step S101 can be implemented according to the following steps:

[0050] obtaining a text corresponding to the speech to be generated;

[0051] annotating the text at a position of the non-verbal event to obtain an annotated text;

[0052] converting the annotated text into the phoneme sequence.

[0053] In this embodiment, the text corresponding to the speech to be generated can be input to the electronic device by a user or a background, so that the electronic device obtains the text of the speech to be generated.

[0054] When the method provided in the present application is applied to an artificial intelligence voice assistant, a question can be input by a user. The question can be in the form of voice or text. After receiving the question input by the user, a reply text corresponding to the question can be determined, and the reply text is the text of the speech to be generated.

[0055] After obtaining the text corresponding to the speech to be generated, the text can be annotated at a position of the non-verbal event to obtain an annotated text.

[0056] In a specific implementation, the text corresponding to the speech to be generated can be input into a pre-trained BERT model to add annotations of positions of non-verbal events in the text corresponding to the speech to be generated through the pre-trained BERT model, so as to obtain an annotated text.

[0057] In another specific implementation, annotations of positions of non-verbal events in the text corresponding to the speech to be generated can be made through manual annotation, so as to obtain an annotated text.

[0058] It can be understood that the non-verbal event can be located at the end, the beginning, or any position in the middle of the text. For example, the non-verbal event is the event of clearing the throat before a piece of text is read, or the event of laughing loudly after the text is read.

[0059] After that, the labeled text can be converted into a phoneme sequence.

[0060] For example, the text to be generated into speech is "Today is really happy.", and the corresponding phonemes are shown in Table 1:

[0061] Table 1.

[0062]

[0063] In the above Table 1, the non-verbal event is located at the end of the text to be generated into speech.

[0064] Step S102: According to the phoneme sequence, the frame number occupied by each phoneme and the non-verbal event is predicted and prosody prediction is performed, respectively, to obtain frame-level first phoneme feature data with added prosody information.

[0065] The above phoneme sequence is a text representation at the phoneme level. This step is used to convert the text representation at the phoneme level into frame-level first phoneme feature data with added prosody information.

[0066] Human language usually has prosody information, which represents the intonation in speech. If the prosody information is not considered when generating speech, the generated speech will not have intonation, and will be monotonous, mechanical and not real and natural.

[0067] In this step, the frame number occupied by each phoneme in the phoneme sequence and the non-verbal event can be predicted, and the frame number occupied by the corresponding non-verbal event marked by the event marker can be predicted, that is, the speech frame number corresponding to each phoneme and the speech frame number corresponding to the non-verbal event are predicted, and the prosody information of each frame is predicted.

[0068] Since the phoneme sequence contains phonemes and event markers corresponding to non-verbal events, based on the predicted speech frame number and the prosody information corresponding to each frame, frame-level first phoneme feature data with added prosody information can be obtained. The first phoneme feature data is frame-level data that integrates the phoneme information and prosody information of each frame.

[0069] It should be noted that speech is a quasi-stationary signal, that is, short-time stationary, and the short-time length is generally 10-30 ms, so when processing speech signals, the overall non-stationary and time-varying effects of the speech signal are reduced, and the signal is processed in frames. In the embodiments of the present application, by predicting the speech frame number corresponding to each phoneme and non-verbal event, and predicting the prosody information of each frame, a basis is provided for generating realistic short-time stationary speech signals.

[0070] Step S103: Determine the frame-level non-verbal feature data to be added to the speech to be generated.

[0071] The non-verbal feature data can represent fine distribution information of the non-verbal event. The fine distribution information can include, but is not limited to, frequency of the non-verbal event, sub-type of the non-verbal event, duration of the non-verbal event (number of speech frames), prosody information of the non-verbal event.

[0072] The frequency of the non-verbal event can reflect the state of the speaker. For example, high-frequency laughter indicates that the speaker is in a relaxed and happy state, and low-frequency laughter indicates that the speaker is in a serious state. The sub-type represents the emotional type corresponding to the non-verbal event. For example, the sub-type of laughter can include smiling, mocking, wry smile, hysterical laughter, contemptuous smile, and awkward smile. The duration of the non-verbal event is used to convey different emotional intensity. For example, long continuous laughter usually expresses more intense happiness or excitement than short laughter. The prosody information of the non-verbal event represents the emotional state of the speaker. For example, laughter with a higher pitch indicates that the speaker is in a happy and excited state, and laughter with a lower pitch indicates that the speaker is in a nervous and awkward state.

[0073] In the embodiments of the present application, the non-verbal feature data can be determined in different ways based on different speech generation tasks.

[0074] In an optional specific embodiment, step S103 can be implemented by the following steps:

[0075] If the speech to be generated is speech to be generated based on text to predict non-verbal events, frame-level non-verbal feature data is predicted according to the first phoneme feature data.

[0076] In the embodiments, the speech generation task is a task of predicting non-verbal events based on text. For example, the text is "Today, I am really happy.", and the non-verbal feature data representing the fine distribution information of the non-verbal event can be predicted based on the context information of the text.

[0077] Specifically, the first phoneme feature data can be input into a pre-trained non-verbal feature prediction model to output frame-level non-verbal feature data through the non-verbal feature prediction model.

[0078] The pre-trained non-verbal feature prediction model is used to predict matching non-verbal feature data based on text data. In the case of a laughter event as the non-verbal event, the pre-trained non-verbal feature prediction model is a laughter prediction model. After inputting the frame-level first phoneme feature data into the pre-trained laughter prediction model, the laughter prediction model can output frame-level laughter feature data, which represents the fine distribution information of the laughter.

[0079] In another optional embodiment, step S103 can be implemented by the following steps:

[0080] If the to-be-generated speech is speech to be generated for migrating non-verbal events in the preset audio, frame-level non-verbal feature data is extracted from the preset audio.

[0081] In this embodiment, the speech generation task is a task of migrating non-verbal events in the preset audio. For example, the preset audio is "Today, I am in a good mood, ha ha ha!" and the corresponding audio can extract non-verbal feature data representing fine distribution information of non-verbal events from the audio.

[0082] Specifically, the preset audio can be input into a pre-trained non-verbal feature extraction model to output frame-level non-verbal feature data through the non-verbal feature extraction model.

[0083] The pre-trained non-verbal feature extraction model is used to identify and separate features corresponding to laughter from an audio signal. In the case of a non-verbal event being a laughter event, the pre-trained non-verbal feature extraction model is a laughter extraction model. After inputting the preset audio containing the laughter event into the pre-trained laughter extraction model, the laughter extraction model can extract frame-level laughter feature data representing fine distribution information of laughter in the preset audio.

[0084] Optionally, the step of "extracting frame-level non-verbal feature data from the preset audio" can include the following steps:

[0085] According to a first frame rate, initial non-verbal feature data is extracted from the preset audio;

[0086] According to the second frame rate corresponding to the first phoneme feature data, a linear interpolation is performed on the non-verbal vectors corresponding to each sampling frame in the initial non-verbal feature data to obtain frame-level non-verbal feature data matching the second frame rate.

[0087] When extracting non-verbal feature data from the preset audio, the preset audio is usually down-sampled to a first frame rate through the convolution layer, the pooling layer, etc. of the pre-trained non-verbal feature extraction model, and the data of the 32-dimensional hidden layer with the first frame rate is extracted through the last layer of the non-verbal feature extraction model as non-verbal feature data. In actual applications, the first frame rate is smaller than the second frame rate corresponding to the first phoneme feature data predicted from the text by the prosody prediction model. In this embodiment, a linear interpolation is performed on the non-verbal vectors corresponding to each sampling frame to obtain non-verbal feature data aligned with the first phoneme feature data.

[0088] Specifically, the interpolation points to be added between every two adjacent sampling frames can be determined according to the first frame rate and the second frame rate, and then the non-language vector corresponding to the interpolation point is determined based on the non-language vectors corresponding to the two adjacent sampling frames, so that the non-language vectors between the two adjacent sampling frames are smoothly transitioned. Then, the non-language vector corresponding to the interpolation point is added to the initial non-language feature data, so as to obtain the non-language feature data aligned with the first phoneme feature data.

[0089] For example, the laughter extraction model extracts initial laughter feature data from the reference audio at a frame rate of 10 frames per second (i.e., 10 Hz). The prosody prediction model predicts the number of frames at a frame rate of 50 frames per second (i.e., 50 Hz). Assuming that the initial laughter feature data is [f1, f2, f3, f4, f5], in order to achieve a frame rate of 50 frames per second, 4 interpolation points need to be added between every two adjacent sampling frames, and the non-language vectors of the interpolation points are calculated by interpolation, to obtain the laughter feature data [f1, f 12 , f 14 , f 16 , f 18 , f2, f 22 , f 24 , f 26 , f 28 , f3, f 32 , f 34 , f 36 , f 38 , f4, f 42 , f 44 , f 46 , f 48 , f5].

[0090] Step S104: processing the first phoneme feature data into a speech signal with the non-language event added according to the non-language feature data.

[0091] It should be noted that the speech frame number and prosody information corresponding to the non-language event in the first phoneme feature data predicted in step S102 are initial and relatively rough information, that is, the first phoneme feature data includes rough information of the non-language event, and cannot well reflect the fine distribution information of the non-language event.

[0092] It should be noted that, since the non-verbal feature data represents the fine distribution information of the non-verbal event, taking the non-verbal feature data as an additional non-verbal condition can finely adjust the first phoneme feature data including the coarse information of the non-verbal event, so that the number of speech frames, prosody information, frequency and subdivision type corresponding to the non-verbal event are more consistent with the context. By finely reconstructing the non-verbal event through the non-verbal feature data, rich spectral details related to the non-verbal event can be modeled, and through conversion processing, speech containing real and natural non-verbal events can be obtained.

[0093] In the case of a laughter event as the non-verbal event, the laughter feature data predicted by the pre-trained laughter prediction model or the laughter feature data extracted by the pre-trained laughter extraction model can be taken as an additional laughter condition, so as to perform laughter reconstruction on the first phoneme feature data including the coarse information of the laughter event, and generate speech containing real and natural laughter.

[0094] Optionally, the above step S104 can be implemented by the following steps:

[0095] decoding the first phoneme feature data to obtain a first acoustic spectrum corresponding to the target text, wherein the first acoustic spectrum reflects coarse information of the non-verbal event;

[0096] processing the first acoustic spectrum according to the non-verbal feature data to obtain a second acoustic spectrum with the non-verbal event added, wherein the second acoustic spectrum reflects fine distribution information of the non-verbal event;

[0097] converting the second acoustic spectrum into a speech signal in time domain.

[0098] The first acoustic spectrum and the second acoustic spectrum are both mel spectrum, which can reflect the acoustic characteristics of the speech signal to be generated.

[0099] It can be understood that the mel spectrum (Mel spectrum diagram) represents the distribution of the speech signal at different frequencies, and the mel scale is a scale based on the perceptual judgment of the pitch by listeners who are equidistant from each other. Since the human ear is more sensitive to low-frequency signals and less sensitive to high-frequency signals, for equal distances on the normal frequency, the human ear is more likely to identify the frequency on the low-frequency band than the frequency on the high-frequency band. Based on this, the mel scale is proposed, so that the frequencies on the low-frequency band and the frequencies on the high-frequency band with equal distances on the new scale are the same to the human ear. Therefore, the mel spectrum is a spectral image consistent with the human auditory characteristics.

[0100] In this embodiment, decoding the first phoneme feature data to obtain the first acoustic spectrum corresponding to the target text means: predicting the acoustic feature of each frame according to the frame-level first phoneme feature data to obtain the first acoustic spectrum. Since the first phoneme feature data is frame-level data integrating the phoneme information and prosody information of each frame, and the first phoneme feature data includes the rough information of the non-language event, the first acoustic spectrum obtained by decoding is the Mel spectrum reflecting the phoneme information and prosody information of each frame and the rough information of the non-language event.

[0101] In a specific implementation, the first phoneme feature data can be input into the decoder 113 shown in the figure, and the first phoneme feature data is decoded by the decoder to convert the first phoneme feature data into the first acoustic spectrum. Figure 1

[0102] Since the non-language feature data represents the fine distribution information of the non-language event, the second acoustic spectrum obtained by processing the first acoustic spectrum through the non-language feature data is the Mel spectrum that can reflect the phoneme information and prosody information of each frame and the fine distribution information of the non-language event.

[0103] In a specific implementation, the above step "processing the first acoustic spectrum according to the non-language feature data to obtain the second acoustic spectrum" can include the following steps:

[0104] inputting the first acoustic spectrum and the non-language feature data into the pre-trained non-language event modeling model to output the second acoustic spectrum with the non-language event added by the non-language event modeling model.

[0105] In this implementation, the first acoustic spectrum and the non-language feature data are input into the non-language event modeling model 116 shown in the figure, so that the non-language event modeling model reconstructs the non-language event in the first acoustic spectrum based on the non-language feature data to output the second acoustic spectrum reflecting the phoneme information and prosody information of each frame and the fine distribution information of the non-language event. Figure 1

[0106] Then, the second acoustic spectrum can be converted to convert the data in the frequency domain into the speech signal (i.e., speech waveform) in the time domain. The speech signal in the time domain focuses on the change of sound in time, refers to the graphical representation of the audio signal changing with time, can present the amplitude change of sound on the time axis, and can intuitively show the intensity, periodicity and other time-related characteristics of sound.

[0107] ​​In a specific implementation, the mel spectrum can be converted into a speech signal in the time domain by a HIFI-GAN model (a generative adversarial network model for efficient and high-fidelity speech synthesis). The input of the HIFI-GAN is the mel spectrum, which is up-sampled through multiple convolutional layers until the output speech waveform, which is the speech signal in the time domain.

[0108] The data processing method in speech generation provided by the embodiments of the present application includes the following steps: determining a phoneme sequence corresponding to text of speech to be generated, wherein the phoneme sequence is marked with a position of a non-verbal event to be added; performing prediction of frame numbers occupied by each phoneme and the non-verbal event and prosody prediction on the phoneme and the non-verbal event respectively according to the phoneme sequence to obtain first phoneme feature data at a frame level with added prosody information; determining non-verbal feature data at a frame level to be added in the speech to be generated; and processing the first phoneme feature data into speech signals with the non-verbal event added according to the non-verbal feature data.

[0109] It can be seen that, since the phoneme sequence includes each phoneme and is marked with a position of a non-verbal event to be added, the data processing method in speech generation provided by the embodiments of the present application can predict the number of speech frames occupied by each phoneme and the number of speech frames occupied by the non-verbal event based on the phoneme sequence, and can predict prosody information of each frame, thereby obtaining first phoneme feature data at a frame level with added prosody information. Then, non-verbal feature data at a frame level to be added in the speech to be generated is determined. Since the non-verbal feature data is fine-grained data representing fine distribution information of the non-verbal event, the first phoneme feature data can be processed into speech signals containing real and natural non-verbal events according to the non-verbal feature data. In addition, the non-verbal feature data can be predicted from text or extracted from a preset audio, and different texts or different preset audios can represent different pronunciations of non-verbal events, thereby realizing the diversity of non-verbal events. Therefore, the data processing method in speech generation provided by the embodiments of the present application can generate speech containing diversified and real natural non-verbal events.

[0110] In an optional implementation, the above step S102 can include the following steps S1021-S1023:

[0111] Step S1021: converting the phoneme sequence into a phoneme embedding sequence, wherein the phoneme embedding sequence contains a first feature vector of a first dimension corresponding to each phoneme and the non-verbal event;

[0112] Step S1022: Encode the phoneme embedding sequence to obtain a phoneme hidden sequence containing context information, wherein the phoneme hidden sequence contains a second feature vector of the second dimension corresponding to each phoneme and the non-linguistic event;

[0113] Step S1023: Based on the phoneme hiding sequence, predict the number of frames occupied by each phoneme and the non-linguistic event and predict the prosody, respectively, to obtain frame-level first phoneme feature data with added prosody information.

[0114] In this embodiment, converting the phoneme sequence into a phoneme embedding sequence is a process of converting the phoneme sequence into a low-dimensional vector sequence, which is the phoneme embedding sequence. In specific implementation, the phoneme sequence can be input into a pre-trained embedding layer to output a phoneme embedding sequence through the pre-trained embedding layer. Each phoneme corresponds to a first feature vector (low-dimensional vector) of the first dimension, and non-linguistic events also correspond to a first feature vector of the first dimension.

[0115] After that, it can be attached Figure 1 The text encoder 111 shown encodes the phoneme embedding sequence, converting the low-dimensional vector sequence into a high-dimensional vector sequence, which is the phoneme hidden sequence. In the phoneme hidden sequence, each phoneme corresponds to a second feature vector (high-dimensional vector) in the second dimension. These vectors capture the contextual information and features of the preceding and following phonemes. Therefore, the phoneme hidden sequence contains rich contextual information and features.

[0116] After obtaining the phoneme hidden sequence, the number of frames occupied by each phoneme and non-language event distribution can be predicted and prosody can be predicted based on the context information and features contained in the phoneme hidden sequence. This allows for the prediction of the number of speech frames occupied by each phoneme and the number of speech frames occupied by non-language events. Prosody information is then predicted frame by frame, thereby transforming the phoneme-level phoneme hidden sequence, which does not contain prosody information, into frame-level first phoneme feature data that incorporates prosody information.

[0117] This approach converts the phoneme sequence into a phoneme embedding sequence, thereby representing each phoneme and non-linguistic event in the phoneme sequence as a first feature vector of the first dimension. Then, the phoneme embedding sequence is encoded to obtain a phoneme hidden sequence. Since the phoneme hidden sequence contains rich contextual information and features, it is easier to capture the relationships between phonemes and between phonemes and non-linguistic events when predicting the number of frames occupied and prosody of each phoneme and non-linguistic event. This allows for the extraction of richer features from the phoneme hidden sequence, providing a good foundation for generating realistic and natural speech that includes contextually appropriate non-linguistic events.

[0118] Optionally, step S1023 above can be achieved through the following steps:

[0119] The phoneme hidden sequence is input into a pre-trained prosodic prediction model, which predicts the number of frames occupied by each phoneme and the non-linguistic event in the phoneme hidden sequence and the prosodicity, respectively, to obtain frame-level first phoneme feature data with added prosodic information.

[0120] In a specific implementation, the phoneme-level hidden phoneme sequence can be input into the appendix. Figure 1 In the pre-trained prosody prediction model 112 shown, the number of frames occupied and the prosody of the phoneme hidden sequence are predicted by the pre-trained prosody prediction model, respectively, to obtain frame-level first phoneme feature data with added prosody information.

[0121] Optionally, step S1023 may include steps S20 to S22:

[0122] Step S20: Based on the phoneme hiding sequence, predict the number of frames occupied by each phoneme and the non-linguistic event to obtain frame-level second phoneme feature data;

[0123] Step S21: Based on the second phoneme feature data, perform prosody prediction on each frame to obtain the prosody information corresponding to each frame;

[0124] Step S22: Process the second phoneme feature data according to the prosody information to obtain frame-level first phoneme feature data with added prosody information.

[0125] Step S20 described above is used to convert the phoneme-level phoneme representation (phoneme hidden sequence) into a frame-level phoneme representation (second phoneme feature data). Specifically, by predicting the number of frames occupied by each phoneme and non-language event, the number of speech frames occupied by each phoneme and non-language event is obtained, thereby converting the phoneme hidden sequence into frame-level second phoneme feature data. This second phoneme feature data is frame-level data that integrates phoneme information and contextual information.

[0126] In one optional implementation, step S20 may include steps S20a and S20b:

[0127] Step S20a: Based on the phoneme hiding sequence, predict the number of frames occupied by each phoneme and the non-language event to obtain the number of speech frames occupied by each phoneme and the non-language event respectively;

[0128] Step S20b: Copy each second feature vector according to the corresponding number of speech frames to obtain frame-level second phoneme feature data.

[0129] In a specific implementation, the step S20a-S20b refers to inputting the phoneme hidden sequence into the pre-trained duration predictor to output the frame-level second phoneme feature data through the duration predictor.

[0130] As shown in Figure 3 FIG. 1 is a structural diagram of a prosody prediction model applied by a data processing method in speech generation provided by an embodiment of the present application, and the prosody prediction model 112 includes a duration predictor 1121.

[0131] The duration predictor 1121 is configured to predict the number of speech frames occupied by each phoneme and non-linguistic event according to the phoneme hidden sequence, and copy the second feature vector of the second dimension corresponding to each phoneme and non-linguistic event according to the corresponding number of speech frames, so as to convert the phoneme-level phoneme hidden sequence into frame-level second phoneme feature data.

[0132] In an example of an optional specific implementation, the step S21 can include the following steps S21a and S21b:

[0133] Step S21a: predicting the pitch information and energy information corresponding to each frame according to the second phoneme feature data;

[0134] Step S21b: integrating the pitch information and energy information belonging to the same frame to obtain the prosody information corresponding to each frame.

[0135] It should be noted that the prosody information usually includes pitch information and energy information, and the pitch information represents the high and low of the sound, and the energy information represents the strength of the sound.

[0136] In a specific implementation, the step S21a refers to inputting the second phoneme feature data into a pre-trained pitch predictor to output the pitch information corresponding to each frame through the pitch predictor, and inputting the second phoneme feature data into a pre-trained energy predictor to output the energy information corresponding to each frame through the energy predictor.

[0137] It should be noted that the prediction of the pitch information and the prediction of the energy information do not have a prior order in execution, and in a specific implementation, the pitch information of each frame can be predicted first, and then the energy information of each frame is predicted, or the energy information of each frame can be predicted first, and then the pitch information of each frame is predicted, or the pitch information and the energy information of each frame can be predicted simultaneously.

[0138] As shown in Figure 3 The prosody prediction model 112 includes a pitch predictor 1122 and an energy predictor 1123.

[0139] The pitch predictor 1122 is configured to predict a pitch value of each frame according to the second phoneme feature data, and embed the pitch value into a pitch embedding vector, which is the pitch information mentioned above. The energy predictor 1123 is configured to predict an energy value of each frame according to the second phoneme feature data, and embed the energy value into an energy embedding vector, which is the energy information mentioned above.

[0140] After that, the pitch information and the energy information belonging to the same frame can be integrated to obtain prosody information corresponding to each frame, which can reflect the height of each frame of sound and the strength of each frame of sound.

[0141] Since the second phoneme feature data is frame-level data integrating phoneme information and context information, based on the prosody information capable of accurately reflecting the height of each frame of sound and the strength of each frame of sound, the second phoneme feature data can be processed into frame-level first phoneme feature data capable of reflecting the height of each frame of sound and the strength of each frame of sound.

[0142] The training of the prosody prediction model, the non-language feature prediction model, and the non-language event modeling model in the data processing method in the speech generation provided by the embodiments of the present application is described in detail as follows:

[0143] Optionally, the data processing method in the speech generation provided by the embodiments of the present application can include the following steps:

[0144] Obtaining at least one sample audio, wherein the sample audio includes at least one sample non-language event;

[0145] According to the sample audio data, determining a sample phoneme sequence marked with the position of the sample non-language event;

[0146] According to the sample audio data and the sample phoneme sequence, training the prosody prediction model to be trained to obtain a pre-trained prosody prediction model; or

[0147] According to the sample audio data and the sample phoneme sequence, training the non-language feature prediction model to be trained to obtain a pre-trained non-language feature prediction model; or

[0148] According to the sample audio data and the sample phoneme sequence, training the non-language event modeling model to be trained to obtain a pre-trained non-language event modeling model.

[0149] In this embodiment, the sample audio is a real recording containing a sample non-linguistic event. For example, the sample audio can be the corresponding audio of "Today is really happy! Ha ha ha…". After obtaining the sample audio, each phoneme in the sample audio can be determined, thereby obtaining a corresponding sample phoneme sequence, and the sample non-linguistic event position can be marked in the sample phoneme sequence. In addition, based on the sample audio, the real speech frame number of each sample phoneme in the sample audio and the real speech frame number of the sample non-linguistic event in the sample audio can also be determined.

[0150] As shown in Table 2, it is an example table of an example of a sample phoneme corresponding to a sample phoneme audio in a data processing method in speech generation provided by the embodiments of the present application.

[0151] Table 2.

[0152]

[0153] In the above Table 2, the sample phonemes contained in the sample audio are j, in1, t, ian1, zh, en1, k, ai1, x, in1, and the sample non-linguistic event is at the end position, and the sample event label is laugh. Among them, the speech frame number of the sample phoneme j is 4 frames, the speech frame number of the sample phoneme in1 is 7 frames, the speech frame number of the sample phoneme t is 4 frames, the speech frame number of the sample phoneme ian1 is 13 frames, the speech frame number of the sample phoneme zh is 6 frames, the speech frame number of the sample phoneme en1 is 24 frames, the speech frame number of the sample phoneme k is 8 frames, the speech frame number of the sample phoneme ai1 is 9 frames, the speech frame number of the sample phoneme x is 9 frames, the speech frame number of the sample phoneme in1 is 18 frames, and the speech frame number of the sample non-linguistic event is 18 frames.

[0154] Through the sample audio data and the sample phoneme sequence, the prosody prediction model to be trained, or the non-linguistic feature prediction model to be trained, or the non-linguistic event modeling model to be trained can be trained.

[0155] In an example optional specific embodiment, the above-mentioned step "training the prosody prediction model to be trained according to the sample audio data and the sample phoneme sequence to obtain a pre-trained prosody prediction model" can include the following steps:

[0156] determining real variance information corresponding to the sample audio, wherein the real variance information includes real speech frame numbers corresponding to each sample phoneme and the sample non-linguistic event, and real prosody information corresponding to each sample phoneme and the sample non-linguistic event;

[0157] inputting the sample phoneme sequence into the prosody prediction model to be trained to output predicted variance information corresponding to the sample phoneme sequence by the prosody prediction model, wherein the predicted variance information comprises predicted speech frame numbers corresponding to each of the sample phonemes and the sample non-linguistic events respectively, and predicted prosody information corresponding to each of the sample phonemes and the sample non-linguistic events respectively;

[0158] adjusting model parameters of the prosody prediction model so that a difference between the predicted variance information and the real variance information is less than a first threshold value, to obtain a pre-trained prosody prediction model.

[0159] In an optional implementation, the duration predictor, the pitch predictor and the energy predictor included in the prosody prediction model can be trained respectively.

[0160] Specifically, in the foregoing introduction, the real speech frame numbers corresponding to each of the sample phonemes and the sample non-linguistic events in the sample audio have been determined, and in this step, real pitch information and real energy information corresponding to each of the sample phonemes and the sample non-linguistic events in the sample audio can be determined.

[0161] Then, the sample phoneme sequence can be embedded to convert the sample phoneme sequence into a sample phoneme embedding sequence, and the sample phoneme embedding sequence can be encoded to obtain a sample phoneme hidden sequence. Then, the sample phoneme hidden sequence is input into the prosody prediction model to be trained, so that the predicted speech frame numbers corresponding to each of the sample phonemes and the sample non-linguistic events in the sample phoneme sequence are predicted by the duration predictor in the prosody prediction model, and the predicted pitch information corresponding to each of the sample phonemes and the sample non-linguistic events in the sample phoneme sequence is predicted by the pitch predictor in the prosody prediction model, and the predicted energy information corresponding to each of the sample phonemes and the sample non-linguistic events in the sample phoneme sequence is predicted.

[0162] When training the duration predictor to be trained, the model parameters of the duration predictor to be trained can be adjusted based on a training strategy that the difference between the predicted speech frame number and the real speech frame number is within a first preset range, so as to obtain a pre-trained duration predictor. In this way, the difference between the predicted speech frame number predicted by the pre-trained duration predictor and the real speech frame number is small, and the pre-trained duration predictor can accurately predict the speech frame number based on the text.

[0163] When training the pitch predictor to be trained, the model parameters of the pitch predictor to be trained can be adjusted based on a training strategy that the difference between the predicted pitch information and the real pitch information is within a second preset range, so as to obtain a pre-trained pitch predictor. In this way, the difference between the predicted pitch information predicted by the pre-trained pitch predictor and the real pitch information is small, and the pre-trained pitch predictor can accurately predict the pitch information based on the text.

[0164] In training the to-be-trained energy predictor, a model parameter of the to-be-trained energy predictor can be adjusted based on a training strategy that a difference between the predicted energy information and the real energy information is within a third preset range, so as to obtain a pre-trained energy predictor. In this way, the difference between the predicted energy information predicted by the pre-trained energy predictor and the real energy information is small, and the energy information can be accurately predicted based on the text.

[0165] In an optional embodiment, the prosody prediction model can be trained as a whole.

[0166] Specifically, the real pitch information and the real energy information can be integrated into real prosody information, and the real prosody information and the real number of speech frames can be determined as real variance information. The predicted pitch information and the predicted energy information can be integrated into predicted prosody information, and the predicted prosody information and the predicted number of speech frames can be determined as predicted variance information.

[0167] In training the to-be-trained prosody prediction model, a model parameter of the to-be-trained prosody prediction model can be adjusted based on a training strategy that a difference between the predicted variance information and the real variance information is less than a first threshold, so as to obtain a pre-trained prosody prediction model. In this way, the difference between the predicted variance information predicted by the pre-trained prosody prediction model and the real variance information is small, and the prosody information can be accurately predicted based on the text.

[0168] In an optional embodiment, the step of "training a to-be-trained non-language feature prediction model based on the sample audio data and the sample phoneme sequence, to obtain a pre-trained non-language feature prediction model" can include the following steps:

[0169] extracting real non-language feature data at a frame level corresponding to the sample audio;

[0170] performing, according to the sample phoneme sequence, prediction of a number of frames occupied by each of the sample phonemes and prosody prediction on the sample non-language events, to obtain predicted phoneme feature data at a frame level with prosody information added;

[0171] inputting the predicted phoneme feature data into the to-be-trained non-language feature prediction model, to output predicted non-language feature data at a frame level through the non-language feature prediction model;

[0172] adjusting a model parameter of the non-language feature prediction model, so that a difference between the predicted non-language feature data and the real non-language feature data is less than a second threshold, to obtain a trained non-language feature prediction model.

[0173] In the embodiment, the prosody prediction model can be trained by attaching the prosody prediction model to the non-language feature prediction model. Figure 1 The pre-trained non-verbal feature extraction model 122 in the pre-training stage extracts the frame-level real non-verbal feature data corresponding to the sample audio. The sample phoneme hidden sequence is input into the pre-trained prosody prediction model, and the pre-trained prosody prediction model is used to predict the frame number and prosody of each sample phoneme and sample non-verbal event, and output frame-level predicted phoneme feature data with prosody information.

[0174] The model parameters of the non-verbal feature prediction model to be trained can be adjusted based on a training strategy that the difference between the predicted non-verbal feature data and the real non-verbal feature data is less than a second threshold, so as to obtain a pre-trained non-verbal feature prediction model. In this way, the difference between the predicted non-verbal feature data predicted by the pre-trained non-verbal feature prediction model and the real non-verbal feature data is small, and the non-verbal feature data can be accurately predicted based on the text.

[0175] In the specific implementation, the training of the non-verbal feature prediction model can be optimized by using the MSE function. In the embodiment of the present application, the optimization function used by the non-verbal feature prediction model is shown in the following formula (1):

[0176]

[0177] In the above formula (1), L laughte is the predicted non-verbal feature data, L ref is the real non-verbal feature data, and L laughter represents the difference between the predicted non-verbal feature data and the real non-verbal feature data in the MSE loss function.

[0178] In an example of an optional specific implementation, the above step "training the non-verbal event modeling model to be trained based on the sample audio data and the sample phoneme sequence to obtain a pre-trained non-verbal event modeling model" can include the following steps:

[0179] extracting frame-level real non-verbal feature data from the sample audio;

[0180] determining the real acoustic spectrum corresponding to the sample audio;

[0181] determining the first predicted acoustic spectrum corresponding to the sample phoneme sequence, wherein the first predicted acoustic spectrum reflects the rough information of the sample non-verbal event;

[0182] The real non-verbal feature data and the predicted acoustic spectrum are input into the non-verbal event modeling model to be trained, so as to output a second predicted acoustic spectrum through the non-verbal event modeling model, wherein the first predicted acoustic spectrum reflects the fine distribution information of the sample non-verbal events;

[0183] The model parameters of the non-verbal event modeling model are adjusted so that the difference between the second predicted acoustic spectrum and the real acoustic spectrum is less than a third threshold, thereby obtaining the trained non-verbal event modeling model.

[0184] In the foregoing introduction, through appendix Figure 1 The pre-trained non-verbal feature extraction model 122 extracts the frame-level real non-verbal feature data corresponding to the sample audio. Then, the real acoustic spectrum can be extracted from the sample audio. For example, at a sampling rate of 16kHz, with a frame length of 50ms and a frame shift of 12.5ms, a Mel spectrum can be extracted from the sample audio using Fast Fourier Transform; this Mel spectrum is the real acoustic spectrum corresponding to the sample audio.

[0185] Then, the first predicted acoustic spectrum corresponding to the sample phoneme sequence can be determined. The step "determining the first predicted acoustic spectrum corresponding to the sample phoneme sequence" specifically refers to: predicting the number of frames occupied by each sample phoneme and the sample non-linguistic event according to the sample phoneme sequence, and predicting the prosody, to obtain frame-level predicted phoneme feature data with added prosody information; decoding the predicted phoneme feature data to obtain the first predicted acoustic spectrum.

[0186] In other words, through the attachment Figure 1 The pre-trained prosodic prediction model 112 converts the hidden sequence of sample phonemes corresponding to the sample phoneme sequence into predicted phoneme feature data, and then... Figure 1 The decoder 113 in the middle decodes the predicted phoneme feature data to obtain the first predicted acoustic spectrum. Then, the first predicted acoustic spectrum can be refined and reconstructed by using the non-language event model to be trained, with the real non-language feature data extracted from the sample audio as a reference, thereby outputting the second predicted acoustic spectrum.

[0187] A training strategy can be adopted to adjust the model parameters of the non-verbal event reconstruction model to be trained, based on the premise that the difference between the second predicted acoustic spectrum and the real acoustic spectrum is less than a third threshold, thereby obtaining a pre-trained non-verbal event reconstruction model. In this way, the difference between the second predicted acoustic spectrum reconstructed by the pre-trained non-verbal event reconstruction model and the real acoustic spectrum is small, enabling accurate reconstruction of the acoustic spectrum containing fine distribution information of non-verbal events based on non-verbal feature data.

[0188] In a specific implementation, an optimization objective of the non-verbal event reconstruction model can be to maximize a likelihood of the training data. An optimization function used by the non-verbal event reconstruction model is shown in the following equation (2):

[0189]

[0190] In the above equation (2), L glow is a value of the optimization function of the non-verbal event reconstruction model, and an optimization objective is to minimize the value to obtain better model performance. x (i) is a real acoustic spectrum corresponding to the i-th sample audio, is a first predicted acoustic spectrum corresponding to the i-th sample audio, is real non-verbal feature data corresponding to the i-th sample audio, and N is a total number of sample audios, is a log-likelihood function.

[0191] Corresponding to the data processing method in speech generation provided by the first embodiment of the present application, the second embodiment of the present application further provides a data processing apparatus in speech generation, as shown in Figure 4 The data processing apparatus 400 in speech generation includes:

[0192] A first determination unit 401 is configured to determine a phoneme sequence corresponding to text of speech to be generated, wherein a position of a non-verbal event to be added is marked in the phoneme sequence.

[0193] A prediction unit 402 is configured to perform frame number prediction and prosody prediction on each phoneme and the non-verbal event respectively according to the phoneme sequence, to obtain first phoneme feature data of frame level with added prosody information.

[0194] A second determination unit 403 is configured to determine frame-level non-verbal feature data to be added in the speech to be generated.

[0195] A processing unit 404 is configured to process the first phoneme feature data into speech signals with the added non-verbal event according to the non-verbal feature data.

[0196] Optionally, the first determination unit 401 is specifically configured to:

[0197] Obtain text corresponding to the speech to be generated;

[0198] Mark a position of the non-verbal event in the text to obtain marked text;

[0199] Convert the marked text into the phoneme sequence.

[0200] Optionally, the prediction unit 402 is specifically configured to:

[0201] convert the phoneme sequence into a phoneme embedding sequence, wherein each of the phonemes and the non-language event correspond to a first feature vector of a first dimension in the phoneme embedding sequence;

[0202] encode the phoneme embedding sequence to obtain a phoneme hidden sequence containing context information, wherein each of the phonemes and the non-language event correspond to a second feature vector of a second dimension in the phoneme hidden sequence;

[0203] predict, according to the phoneme hidden sequence, the number of frames occupied by each of the phonemes and the non-language event, and prosody prediction, to obtain first phoneme feature data of a frame level with added prosody information.

[0204] Optionally, the prediction unit 402 is specifically configured to:

[0205] predict, according to the phoneme hidden sequence, the number of frames occupied by each of the phonemes and the non-language event, to obtain second phoneme feature data of a frame level;

[0206] perform prosody prediction on each frame according to the second phoneme feature data to obtain prosody information corresponding to each frame;

[0207] perform processing on the second phoneme feature data according to the prosody information to obtain first phoneme feature data of a frame level with added prosody information.

[0208] Optionally, the prediction unit 402 is specifically configured to:

[0209] predict, according to the phoneme hidden sequence, the number of frames occupied by each of the phonemes and the non-language event, to obtain the number of speech frames occupied by each of the phonemes and the non-language event;

[0210] copy each second feature vector according to the corresponding number of speech frames to obtain second phoneme feature data of a frame level.

[0211] Optionally, the prediction unit 402 is specifically configured to:

[0212] predict, according to the second phoneme feature data, pitch information and energy information corresponding to each frame;

[0213] integrate the pitch information and the energy information belonging to the same frame to obtain prosody information corresponding to each frame.

[0214] Optionally, the second determination unit 403 is specifically configured to:

[0215] If the speech to be generated is speech to be generated based on text prediction of non-verbal events, according to the first phoneme feature data, frame-level non-verbal feature data is predicted.

[0216] If the speech to be generated is speech to be generated for migration of non-verbal events in preset audio, frame-level non-verbal feature data is extracted from the preset audio.

[0217] Optionally, the processing unit 404 is specifically configured to:

[0218] The first phoneme feature data is decoded to obtain a first acoustic spectrum corresponding to the target text, wherein the first acoustic spectrum reflects rough information of the non-verbal event;

[0219] The first acoustic spectrum is processed according to the non-verbal feature data to obtain a second acoustic spectrum, wherein the second acoustic spectrum reflects fine distribution information of the non-verbal event;

[0220] The second acoustic spectrum is converted into a speech signal in the time domain.

[0221] Optionally, the second determination unit 403 is specifically configured to:

[0222] According to a first frame rate, initial non-verbal feature data is extracted from the preset audio;

[0223] According to a second frame rate corresponding to the first phoneme feature data, a non-verbal vector corresponding to each sampling frame in the initial non-verbal feature data is linearly interpolated to obtain frame-level non-verbal feature data matched with the second frame rate.

[0224] Optionally, the prediction unit 402 is specifically configured to:

[0225] The phoneme hidden sequence is input into a pre-trained prosody prediction model, so that the prosody prediction model is used to predict the number of frames occupied by each phoneme and the prosody in the phoneme hidden sequence, respectively, to obtain first phoneme feature data with frame level and added prosody information.

[0226] Optionally, the second determination unit 403 is specifically configured to:

[0227] The first phoneme feature data is input into a pre-trained non-verbal feature prediction model, so that the non-verbal feature prediction model outputs frame-level non-verbal feature data.

[0228] Optionally, the processing unit 404 is specifically configured to:

[0229] inputting the first acoustic spectrum and the non-verbal feature data into a pre-trained non-verbal event modeling model to output a second acoustic spectrum through the non-verbal event modeling model.

[0230] Optionally, the second determining unit 403 is specifically configured to:

[0231] inputting the preset audio into a pre-trained non-verbal feature extraction model to output frame-level non-verbal feature data through the non-verbal feature extraction model.

[0232] Optionally, the voice-generated data processing apparatus further comprises a training unit, and the training unit is configured to:

[0233] obtain at least one sample audio, wherein the sample audio comprises at least one sample non-verbal event;

[0234] determine a sample phoneme sequence marked with a position of the sample non-verbal event according to the sample audio data;

[0235] train a prosody prediction model to be trained according to the sample audio data and the sample phoneme sequence to obtain a pre-trained prosody prediction model; or

[0236] train a non-verbal feature prediction model to be trained according to the sample audio data and the sample phoneme sequence to obtain a pre-trained non-verbal feature prediction model; or

[0237] train a non-verbal event modeling model to be trained according to the sample audio data and the sample phoneme sequence to obtain a pre-trained non-verbal event modeling model.

[0238] Optionally, the training unit is specifically configured to:

[0239] determine real variance information corresponding to the sample audio, wherein the real variance information comprises real speech frame numbers corresponding to each sample phoneme and the sample non-verbal event respectively, and real prosody information corresponding to each sample phoneme and the sample non-verbal event respectively;

[0240] input the sample phoneme sequence into a prosody prediction model to be trained to output predicted variance information corresponding to the sample phoneme sequence through the prosody prediction model, wherein the predicted variance information comprises predicted speech frame numbers corresponding to each sample phoneme and the sample non-verbal event respectively, and predicted prosody information corresponding to each sample phoneme and the sample non-verbal event respectively;

[0241] adjust model parameters of the prosody prediction model to make a difference between the predicted variance information and the real variance information less than a first threshold value, to obtain a pre-trained prosody prediction model.

[0242] Optionally, the training unit is specifically configured to:

[0243] extract frame-level real non-language feature data corresponding to the sample audio;

[0244] perform frame number prediction and prosody prediction on each of the sample phonemes and the sample non-language events according to the sample phoneme sequence, to obtain frame-level predicted phoneme feature data with prosody information added;

[0245] input the predicted phoneme feature data into a non-language feature prediction model to be trained, so as to output frame-level predicted non-language feature data through the non-language feature prediction model;

[0246] adjust model parameters of the non-language feature prediction model, so that a difference between the predicted non-language feature data and the real non-language feature data is less than a second threshold value, to obtain a trained non-language feature prediction model.

[0247] Optionally, the training unit is specifically configured to:

[0248] extract frame-level real non-language feature data from the sample audio;

[0249] determine real acoustic spectrum corresponding to the sample audio;

[0250] determine a first predicted acoustic spectrum corresponding to the sample phoneme sequence, wherein the first predicted acoustic spectrum reflects rough information of the sample non-language event;

[0251] input the real non-language feature data and the predicted acoustic spectrum into a non-language event modeling model to be trained, so as to output a second predicted acoustic spectrum through the non-language event modeling model, wherein the first predicted acoustic spectrum reflects fine distribution information of the sample non-language event;

[0252] adjust model parameters of the non-language event modeling model, so that a difference between the second predicted acoustic spectrum and the real acoustic spectrum is less than a third threshold value, to obtain a trained non-language event modeling model.

[0253] Corresponding to the data processing method in speech generation provided by the first embodiment of the present application, the third embodiment of the present application further provides an electronic device for data processing in speech generation.

[0254] As shown in Figure 5 , it is a structural block diagram of an example of the electronic device for data processing provided by the embodiment of the present application.

[0255] In this embodiment, an optional hardware structure of the electronic device 500 can be as shown in the figure, including at least one processor 501, at least one memory 502 and at least one communication bus 505; the memory 502 contains a program 503 and data 504. Figure 5

[0256] The bus 505 can be a communication device for transmitting data between components inside the electronic device 500, such as an internal bus (for example, a CPU-memory bus, where the processor is a central processing unit, CPU), an external bus (for example, a universal serial bus port, a peripheral component interconnect express port), etc.

[0257] In addition, the electronic device also includes at least one network interface 506 and at least one peripheral interface 507. The network interface 506 provides wired or wireless communication related to an external network 508 (for example, the Internet, an intranet, a local area network, a mobile communication network, etc.); in some embodiments, the network interface 506 can include any combination of any number of network interface controllers (English: network interface controller, NIC), radio frequency (English: Radio Frequency, RF) modules, transponders, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (English: Near Field Communication, NFC) adapters, cellular network chips, etc.

[0258] The peripheral interface 507 is used to connect with peripherals, which can be peripherals 1 (509) in the figure, peripherals 2 (510) and peripherals 3 (511). The peripheral is a peripheral device, which can include but is not limited to a cursor control device (for example, a mouse, a touchpad or a touch screen), a keyboard, a display (for example, a cathode ray tube display, a liquid crystal display), a video input device (for example, a video camera or an input interface communicatively coupled to a video archive), etc. Figure 5 Figure 5 Figure 5

[0259] The processor 501 can be a CPU, or an application specific integrated circuit ASIC, or one or more integrated circuits configured to implement one or more embodiments of the present application.

[0260] ​​​​The memory 502 can include a high-speed RAM (Random Access Memory) memory, and can also include a non-volatile memory such as at least one disk memory.

[0261] The processor 501 calls the program and data stored in the memory 502 to perform the following steps:

[0262] Determine the phoneme sequence corresponding to the text of the speech to be generated, wherein the phoneme sequence is marked with the position of the non-verbal event to be added;

[0263] According to the phoneme sequence, the frame number occupied by each phoneme and the non-verbal event is predicted and prosody prediction is performed respectively to obtain first phoneme feature data with frame-level and added prosody information;

[0264] Determine the frame-level non-verbal feature data to be added in the speech to be generated;

[0265] According to the non-verbal feature data, the first phoneme feature data is processed into a speech signal with the added non-verbal event.

[0266] Corresponding to the data processing method in speech generation provided by the first embodiment of the present application, the fourth embodiment of the present application provides a computer readable storage medium, which stores a program of a data processing method in speech generation. The program is run by a processor to perform the following steps:

[0267] Determine the phoneme sequence corresponding to the text of the speech to be generated, wherein the phoneme sequence is marked with the position of the non-verbal event to be added;

[0268] According to the phoneme sequence, the frame number occupied by each phoneme and the non-verbal event is predicted and prosody prediction is performed respectively to obtain first phoneme feature data with frame-level and added prosody information;

[0269] Determine the frame-level non-verbal feature data to be added in the speech to be generated;

[0270] According to the non-verbal feature data, the first phoneme feature data is processed into a speech signal with the added non-verbal event.

[0271] It should be noted that the detailed description of the apparatus, electronic device and computer readable storage medium provided by the second embodiment, the third embodiment and the fourth embodiment of the present application can refer to the related description of the first embodiment of the present application, which will not be repeated here.

[0272] Although the present application is disclosed with reference to the preferred embodiments above, it is not intended to limit the present application, and any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application, and the scope of protection of the present application should be defined by the scope of claims.

[0273] In one typical configuration, a node device in the blockchain includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0274] The memory can include non-persistent memory in computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0275] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other properties of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage media or any other non-transmission medium that can be used to store information that can be accessed by a computing device. According to the definition herein, computer-readable media does not include non-transitory computer-readable media such as modulated data signals and carriers.

[0276] 2. Those skilled in the art should understand that the embodiments of the present application can be provided as a method, system or computer program product. Therefore, the present application can be in the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can be in the form of a computer program product implemented on one or more computer-readable storage media containing computer-usable program code (including but not limited to disk storage, CD-ROM, optical storage, etc.).

[0277] Although the present application is disclosed with reference to the preferred embodiments above, it is not intended to limit the present application, and any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application, and the scope of protection of the present application should be defined by the scope of claims.< / laugh>

Claims

1. A data processing method for speech generation, characterized in that, The method includes: Determine the phoneme sequence corresponding to the text to be generated as speech, wherein the phoneme sequence is marked with the location of the non-linguistic event to be added; Based on the phoneme sequence, the number of frames occupied by each phoneme and the non-linguistic event are predicted and the prosody is predicted respectively, to obtain frame-level first phoneme feature data with added prosody information; Identify the frame-level non-linguistic feature data to be added to the speech to be generated; Based on the non-linguistic feature data, the first phoneme feature data is processed into a speech signal incorporating the non-linguistic event.

2. The method according to claim 1, characterized in that, The step of determining the phoneme sequence corresponding to the text to be generated as speech includes: Obtain the text corresponding to the speech to be generated; The text is annotated with the location of the non-verbal event to obtain the annotated text; The annotated text is converted into the phoneme sequence.

3. The method according to claim 1, characterized in that, The step involves predicting the number of frames occupied by each phoneme and the non-linguistic event based on the phoneme sequence, and predicting the prosody, to obtain frame-level first phoneme feature data with added prosodic information, including: The phoneme sequence is converted into a phoneme embedding sequence, wherein the phoneme embedding sequence contains a first feature vector of a first dimension corresponding to each phoneme and the non-linguistic event; The phoneme embedding sequence is encoded to obtain a phoneme hidden sequence containing contextual information, wherein the phoneme hidden sequence contains a second feature vector of the second dimension corresponding to each phoneme and the non-linguistic event; Based on the phoneme hiding sequence, the number of frames occupied by each phoneme and the non-linguistic event are predicted and the prosody is predicted respectively, to obtain frame-level first phoneme feature data with added prosody information.

4. The method according to claim 3, characterized in that, The step involves predicting the number of frames occupied by each phoneme and the non-linguistic event based on the phoneme hidden sequence, and predicting prosody to obtain frame-level first phoneme feature data with added prosodic information, including: Based on the phoneme hiding sequence, the number of frames occupied by each phoneme and the non-linguistic event is predicted to obtain the second phoneme feature data at the frame level. Based on the second phoneme feature data, prosody prediction is performed on each frame to obtain the prosody information corresponding to each frame. The second phoneme feature data is processed according to the prosodic information to obtain frame-level first phoneme feature data with added prosodic information.

5. The method according to claim 4, characterized in that, The step of predicting the number of frames occupied by each phoneme and the non-linguistic event based on the phoneme hidden sequence to obtain frame-level second phoneme feature data includes: Based on the phoneme hiding sequence, the number of frames occupied by each phoneme and the non-verbal event is predicted to obtain the number of speech frames occupied by each phoneme and the non-verbal event respectively. Each second feature vector is copied according to the corresponding number of speech frames to obtain frame-level second phoneme feature data.

6. The method according to claim 4, characterized in that, The step of predicting the prosody of each frame based on the second phoneme feature data to obtain the prosody information corresponding to each frame includes: Based on the second phoneme feature data, predict the pitch and energy information corresponding to each frame; The pitch information and energy information belonging to the same frame are integrated to obtain the prosodic information corresponding to each frame.

7. The method according to claim 1, characterized in that, The process of determining the frame-level non-linguistic feature data to be added to the speech to be generated includes: If the speech to be generated is the speech to be generated based on the prediction of non-linguistic events from text, then the non-linguistic feature data at the frame level is predicted based on the first phoneme feature data. If the speech to be generated is the speech to be generated by transferring non-linguistic events in a preset audio, then frame-level non-linguistic feature data is extracted from the preset audio.

8. The method according to claim 1, characterized in that, The step of processing the first phoneme feature data into a speech signal incorporating the non-verbal event based on the non-verbal feature data includes: The first phoneme feature data is decoded to obtain the first acoustic spectrum corresponding to the text, wherein the first acoustic spectrum reflects coarse information about the non-verbal event; The first acoustic spectrum is processed based on the non-verbal feature data to obtain a second acoustic spectrum, wherein the second acoustic spectrum reflects the fine distribution information of the non-verbal events; The second acoustic spectrum is converted into a speech signal in the time domain.

9. The method according to claim 7, characterized in that, The step of extracting frame-level non-linguistic feature data from the preset audio includes: According to the first frame rate, extract initial non-linguistic feature data from the preset audio; Based on the second frame rate corresponding to the first phoneme feature data, linear interpolation is performed on the non-language vectors corresponding to each sampled frame in the initial non-language feature data to obtain frame-level non-language feature data that matches the second frame rate.

10. The method according to claim 3, characterized in that, The step involves predicting the number of frames occupied by each phoneme and the non-linguistic event based on the phoneme hidden sequence, and predicting prosody to obtain frame-level first phoneme feature data with added prosodic information, including: The phoneme hidden sequence is input into a pre-trained prosodic prediction model, which predicts the number of frames occupied by each phoneme and the non-linguistic event in the phoneme hidden sequence and the prosodicity, respectively, to obtain frame-level first phoneme feature data with added prosodic information.

11. The method according to claim 7, characterized in that, The step of predicting frame-level non-linguistic feature data based on the first phoneme feature data includes: The first phoneme feature data is input into a pre-trained non-linguistic feature prediction model to output frame-level non-linguistic feature data.

12. The method according to claim 8, characterized in that, The step of processing the first acoustic spectrum based on the non-linguistic feature data to obtain the second acoustic spectrum includes: The first acoustic spectrum and the non-linguistic feature data are input into a pre-trained non-linguistic event modeling model to output a second acoustic spectrum through the non-linguistic event modeling model.

13. The method according to claim 7, characterized in that, The step of extracting frame-level non-linguistic feature data from the preset audio includes: The preset audio input is used to pre-train a non-linguistic feature extraction model, which outputs frame-level non-linguistic feature data.

14. The method according to any one of claims 10 to 12, characterized in that, The method further includes: Acquire at least one sample audio, wherein the sample audio includes at least one sample non-verbal event; Based on the sample audio data, determine the sample phoneme sequence that marks the location of the sample non-linguistic event; Based on the sample audio data and the sample phoneme sequence, the prosodic prediction model to be trained is trained to obtain a pre-trained prosodic prediction model; or Based on the sample audio data and the sample phoneme sequence, the non-verbal feature prediction model to be trained is trained to obtain a pre-trained non-verbal feature prediction model; or Based on the sample audio data and the sample phoneme sequence, the non-verbal event modeling model to be trained is trained to obtain a pre-trained non-verbal event modeling model.

15. The method according to claim 14, characterized in that, The step of training the prosodic prediction model to be trained based on the sample audio data and the sample phoneme sequence to obtain a pre-trained prosodic prediction model includes: Determine the true variance information corresponding to the sample audio, wherein the true variance information includes the true speech frame number corresponding to each sample phoneme and the sample non-verbal event, and the true prosodic information corresponding to each sample phoneme and the sample non-verbal event; The sample phoneme sequence is input into the prosody prediction model to be trained, so that the prosody prediction model outputs the prediction variance information corresponding to the sample phoneme sequence. The prediction variance information includes the predicted speech frame number corresponding to each sample phoneme and the sample non-verbal event, and the predicted prosody information corresponding to each sample phoneme and the sample non-verbal event. The model parameters of the prosody prediction model are adjusted so that the difference between the predicted variance information and the true variance information is less than a first threshold, thereby obtaining a pre-trained prosody prediction model.

16. The method according to claim 14, characterized in that, The step of training the non-linguistic feature prediction model to be trained based on the sample audio data and the sample phoneme sequence to obtain a pre-trained non-linguistic feature prediction model includes: Extract the frame-level real non-linguistic feature data corresponding to the sample audio; Based on the sample phoneme sequence, the number of frames occupied by each sample phoneme and the sample non-linguistic event are predicted and the prosody is predicted, respectively, to obtain frame-level predicted phoneme feature data with added prosody information. The predicted phoneme feature data is input into the non-linguistic feature prediction model to be trained, so that the non-linguistic feature prediction model outputs frame-level predicted non-linguistic feature data. The model parameters of the non-linguistic feature prediction model are adjusted so that the difference between the predicted non-linguistic feature data and the real non-linguistic feature data is less than a second threshold, thereby obtaining the trained non-linguistic feature prediction model.

17. The method according to claim 14, characterized in that, The step of training the non-verbal event modeling model to be trained based on the sample audio data and the sample phoneme sequence to obtain a pre-trained non-verbal event modeling model includes: Extract frame-level real non-verbal feature data from the sample audio; Determine the true acoustic spectrum corresponding to the sample audio; Determine the first predicted acoustic spectrum corresponding to the sample phoneme sequence, wherein the first predicted acoustic spectrum reflects coarse information about the sample non-linguistic events; The real non-verbal feature data and the predicted acoustic spectrum are input into the non-verbal event modeling model to be trained, so as to output a second predicted acoustic spectrum through the non-verbal event modeling model, wherein the first predicted acoustic spectrum reflects the fine distribution information of the sample non-verbal events; The model parameters of the non-verbal event modeling model are adjusted so that the difference between the second predicted acoustic spectrum and the real acoustic spectrum is less than a third threshold, thereby obtaining the trained non-verbal event modeling model.

18. A data processing apparatus for speech generation, characterized in that, The device includes: The first determining unit is used to determine the phoneme sequence corresponding to the text to be generated speech, wherein the phoneme sequence is marked with the location of the non-linguistic event to be added; The prediction unit is used to predict the number of frames occupied by each phoneme and the non-linguistic event according to the phoneme sequence, and to predict the prosody, so as to obtain frame-level first phoneme feature data with added prosody information. The second determining unit is used to determine the frame-level non-linguistic feature data to be added to the speech to be generated; The processing unit is configured to process the first phoneme feature data into a speech signal incorporating the non-language event based on the non-language feature data.

19. An electronic device, characterized in that, include: processor; as well as A memory for storing a data processing program, which, when the electronic device is powered on and runs through the processor, performs the method as described in any one of claims 1-17.

20. A computer-readable storage medium, characterized in that, The system contains a data processing program that is executed by a processor to perform the method as described in any one of claims 1-17.

Citation Information

Patent Citations

  • Speech synthesis method and device, readable medium and electronic equipment

    CN116189652A

  • Speech synthesis method and device, electronic equipment and storage medium

    CN118711560A