Audio generation methods, apparatus, electronic devices and storage media

By automatically generating laughter synthesized audio using a spectrum generation model, the problems of fluctuating quality and low efficiency in existing laughter synthesized audio technologies are solved, achieving efficient and automated laughter audio generation.

CN119380692BActive Publication Date: 2026-04-03VIVO MOBILE COMM CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-24
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, the generation of laughter-synthesized audio relies on manual editing, resulting in large quality fluctuations and low efficiency.

Method used

By acquiring the laughter feature data of the reference object, a spectrum generation model is used to automatically generate a pre-defined style of synthesized laughter audio, including the text content of the reference text and laughter related to the emotions represented by the text content.

Benefits of technology

It simplifies the process of synthesizing laughter audio, improves generation efficiency and quality, and reduces the need for manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380692B_ABST
    Figure CN119380692B_ABST
Patent Text Reader

Abstract

This application discloses an audio generation method, apparatus, electronic device, and storage medium, belonging to the field of electronic device technology. The method includes acquiring a reference object and laughter feature data related to the reference object. The reference object includes reference text and reference audio. The reference text is text used for laughter synthesis, and the reference audio is audio used to instruct the generation of a preset style. The method further includes determining laughter speech spectrum data based on the reference object and the laughter feature data using a spectrum generation model. The method also includes converting the laughter speech spectrum data into laughter synthesized audio, wherein the laughter synthesized audio is audio of a preset style and includes the text content of the reference text and laughter related to the emotion represented by the text content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to an audio generation method, apparatus, electronic device, and storage medium. Background Technology

[0002] Laughter, as a powerful and positive sound expression, has broad application prospects in film and television production, game development, virtual reality experiences, and online social interaction. Controllable laughter audio generation technology has great potential.

[0003] In related technologies, laughter audio clips can be manually edited from a laughter library to obtain synthesized laughter audio. However, manual editing relies too heavily on human editing skills and experience. This not only leads to significant fluctuations in the quality of the synthesized laughter audio due to differences in the skill levels of different operators, but also increases the number of steps required to generate the synthesized laughter audio, reducing its efficiency. Summary of the Invention

[0004] The purpose of this application is to provide an audio generation method, apparatus, electronic device, and storage medium that can reduce the operation of editing laughter-synthesized audio and improve the efficiency and quality of generating laughter-synthesized audio.

[0005] In a first aspect, embodiments of this application provide an audio generation method, including:

[0006] Acquire reference objects and laughter feature data related to the reference objects. The reference objects include reference text and reference audio. The reference text is the text used for laughter synthesis, and the reference audio is the audio used to indicate the generation of a preset style.

[0007] Laughter speech spectrum data are determined using a spectrum generation model based on a reference object and laughter feature data.

[0008] Laughter speech spectrum data is converted into laughter synthesized audio, wherein the laughter synthesized audio is audio of a preset style, and the laughter synthesized audio includes the text content of the reference text and laughter related to the emotion represented by the text content.

[0009] Secondly, embodiments of this application provide an audio generation apparatus, including:

[0010] The acquisition module is used to acquire reference objects and laughter feature data related to the reference objects. The reference objects include reference text and reference audio. The reference text is the text used for laughter synthesis, and the reference audio is the reference audio used to indicate the generation of audio of a preset style or type.

[0011] The determination module is used to determine the spectral data of laughter speech based on the reference object and laughter feature data through the spectrum generation model;

[0012] The conversion module is used to convert laughter speech spectrum data into laughter synthesized audio, wherein the laughter synthesized audio is of a preset style or type and includes text content of a reference text and audio of laughter related to the emotion represented by the text content.

[0013] Thirdly, embodiments of this application provide an electronic device, which includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the steps of the audio generation method as described in the first aspect.

[0014] Fourthly, embodiments of this application provide a readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the audio generation method as described in the first aspect.

[0015] Fifthly, embodiments of this application provide a chip, which includes a processor and a display interface, the display interface and the processor being coupled together, the processor being used to run programs or instructions to implement the steps of the audio generation method as described in the first aspect.

[0016] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the steps of the audio generation method as described in the first aspect.

[0017] In this embodiment, a spectrum generation model can be used to determine laughter speech spectrum data based on a reference object and laughter feature data associated with the reference object. The reference object includes reference text and reference audio; the reference text is the text used for laughter synthesis, and the reference audio is the audio used to instruct the generation of a preset style. Then, the laughter speech spectrum data is converted into a synthesized laughter audio of a preset style, including the text content of the reference text and laughter emotion-related to the text content. This allows for the automatic generation of synthesized laughter audio based on the reference text, reference audio, and laughter feature data, eliminating the need for users to spend significant time manually editing, simplifying the process, reducing the difficulty, and improving the efficiency and quality of synthesized laughter audio generation. Attached Figure Description

[0018] Figure 1 Flowcharts of audio generation methods provided for some embodiments of this application;

[0019] Figure 2 A schematic diagram of the interface for an audio generation method provided in some embodiments of this application;

[0020] Figure 3 Schematic diagrams of the audio generation apparatus provided for some embodiments of this application;

[0021] Figure 4 Schematic diagrams of the structure of electronic devices provided for some embodiments of this application;

[0022] Figure 5 A schematic diagram of the hardware structure of an electronic device provided for some embodiments of this application. Detailed Implementation

[0023] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0024] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0025] This application provides an audio generation method, apparatus, and electronic device. The following description, in conjunction with the accompanying drawings, further details their implementation. Figures 1 to 5 The audio generation method provided in this application will be described in detail through specific embodiments and application scenarios.

[0026] First, combined Figure 1 This application provides a detailed description of an audio generation method based on an embodiment.

[0027] Figure 1 A flowchart of an audio generation method provided for some embodiments of this application.

[0028] like Figure 1 As shown, the audio generation method provided in this application embodiment can be applied to electronic devices. Based on this, the audio generation method may include steps 110 and 130, as detailed below.

[0029] Step 110: Obtain reference objects and laughter feature data related to the reference objects. The reference objects include reference text and reference audio. The reference text is the text used for laughter synthesis, and the reference audio is the audio used to indicate the generation of a preset style.

[0030] Step 120: Using a spectrum generation model, determine the spectrum data of laughter speech based on the reference object and laughter feature data;

[0031] Step 130: Convert the laughter speech spectrum data into laughter synthesized audio, wherein the laughter synthesized audio is audio of a preset style, and the laughter synthesized audio includes the text content of the reference text and laughter related to the emotion represented by the text content.

[0032] Therefore, based on reference text, reference audio, and laughter feature data, laughter synthesized audio can be automatically generated without requiring users to spend a lot of time manually editing. This simplifies the process of generating laughter synthesized audio, reduces the difficulty of generating laughter synthesized audio, and improves the efficiency and quality of generating laughter synthesized audio.

[0033] The steps described above are explained in detail below.

[0034] First, regarding step 110, in some embodiments of this application, the reference object can be obtained in the following manner, and the audio generation method may also include steps 1401 to 1405.

[0035] Step 1401: Display the laughter synthesis interface, which includes a first area and a second area.

[0036] Step 1402: Receive the user's first input for the first area.

[0037] Step 1403: In response to the first input, determine the text entered by the user in the first area as the reference text.

[0038] Step 1404: Receive the user's second input for the second area.

[0039] Step 1405: In response to the second input, determine the preset audio selected by the user in the second area as the reference audio; or, determine the user audio added by the user in the second area as the reference audio.

[0040] For example, the order of the first and second inputs is not limited in this embodiment. Here, the user can touch the second area first, and then touch the first area. For example, Figure 2As shown, the electronic device first receives user input to the second area and obtains the reference audio uploaded by the user. Then, the electronic device receives user input to the first area and obtains the reference text entered by the user: "Do you remember when we went to the zoo last time? That monkey actually imitated your expression, hahaha."

[0041] It should be noted that the preset style of the reference audio in this application embodiment can be the style of the audio content, that is, the emotional style expressed through the melody and harmony of the audio, such as light and cheerful, or melancholic autumn. Its preset style can also be the style of the audio type, such as rock, folk, electronic, or light music.

[0042] Furthermore, the reference audio in this embodiment can be associated with a reference audio prompt. The prompt is used to guide the speech spectrum generation model to generate audio in a preset style corresponding to the reference audio. The prompt is a key component of human-computer interaction and can effectively control and optimize the model's response quality and relevance.

[0043] In some other embodiments of this application, laughter feature data can be obtained in the following two ways.

[0044] In some embodiments, the laughter synthesis interface further includes a third area for displaying at least one of the following preset options: preset laughter audio options and preset laughter tag options. Based on this, the audio generation method may further include steps 1406 to 1407 before step 110.

[0045] Step 1406: Receive the user's third input for the third area.

[0046] Step 1407: In response to the third input, the laughter feature data corresponding to the target preset option selected by the user in the third area is determined as laughter feature data related to the reference object.

[0047] For example, you can still refer to Figure 2 Users can input into the third area to select a target preset option from the three preset options displayed in the third area: joyful laughter, arrogant laughter, and mocking laughter. The laughter feature data corresponding to the target preset option is then determined as the laughter feature data related to the reference object.

[0048] In some other embodiments, the audio generation method may also include step 1408 prior to step 110.

[0049] Based on the first textual semantics of the reference text and the second textual semantics of the text corresponding to the reference audio, laughter feature data related to the reference object is generated.

[0050] For example, a semantic recognition model can be used to perform semantic recognition on the reference text to obtain the first text semantics. Similarly, a semantic recognition model can be used to perform semantic recognition on the text corresponding to the reference audio to obtain the second text semantics. Then, the first and second text semantics are input into a large model to obtain the relevant laughter feature data.

[0051] Next, regarding step 120, in some embodiments of this application, the spectrum generation model includes a data processing model and a speech spectrum generation model. Based on this, step 120 may specifically include steps 1201 and 1202.

[0052] Step 1201: The reference object and laughter feature data are processed through a data processing model to obtain a spectrum generation object. The spectrum generation object includes phoneme augmentation data corresponding to the reference text, speech spectrum data of the reference audio, and emotional data of the laughter feature data. The phoneme augmentation data includes the data of the assembled phoneme data of the assembled text after augmentation phoneme data frames. The assembled text is the text concatenated from the reference text and the audio text of the reference audio.

[0053] For example, the input to the data processing model may include three inputs: reference text, reference audio, and laughter feature data. The reference audio can be associated with a prompt; for instance, if the reference audio is "Good Luck," the prompt can be an instruction to guide the speech spectrum generation model to generate "joyful" audio corresponding to the reference audio. The laughter feature data is associated with a laughter prompt; for example, the laughter feature data may include joyful laughter corresponding to a preset option for joyful laughter. In this case, the laughter prompt can be used to guide the speech spectrum generation model to generate audio including "joyful laughter."

[0054] Based on this, the data processing model extracts the Mel spectrum features of the reference audio and feeds them into the Automatic Speech Recognition (ASR) model to extract the audio text. Then, the audio text and the reference text are concatenated along the time dimension to obtain the compiled text, which is then converted into phoneme representation (phoneme augmentation data) through a text front-end. Additionally, the data processing model extracts the Mel spectrum features corresponding to the reference audio and feeds them into a Mel encoder for feature extraction and dimensionality reduction to obtain speech spectrum data. Finally, the laughter audio and laughter prompt are fed into a laughter emotion detection model to extract emotional data, such as laughter embedding.

[0055] In some embodiments, the data processing model includes a first data processing model for processing a reference object. The first data processing model includes a text encoding model and a phoneme duration prediction model. The spectrum generation object includes phoneme augmentation data. Based on this, before step 1201 above, the audio generation method may also include steps 2101 and 2102.

[0056] Step 2101: Perform speech recognition processing on the audio spectrum data of the reference audio to obtain the audio text.

[0057] For example, ASR can be used to perform speech recognition processing on the audio spectrum data of the reference audio to obtain audio text, such as the lyrics of the reference audio, or the speech text of the lyrics and harmony parts of the reference audio.

[0058] Step 2102: Concatenate the reference text and audio text to obtain the assembled text.

[0059] For example, audio text can be appended to the reference text to improve recognition efficiency.

[0060] Based on this, step 1201 may specifically include steps 12011 to 12013.

[0061] Step 12011: Using a text encoding model, the assembled text is processed to perform phoneme conversion to obtain assembled phoneme data. The assembled phoneme data includes phoneme sequences corresponding to text sentences in the assembled text. Here, in this embodiment, a phoneme refers to the smallest sound unit in human language that can distinguish meaning.

[0062] For example, the assembled phoneme data can be the smallest unit of the actual pronunciation of each text in the assembled text, such as " / h / " or " / ǎo / ".

[0063] Step 12012: Using the phoneme duration prediction model, calculate the second phoneme duration of each phoneme based on the first phoneme duration of each phoneme in the phoneme sequence. The second phoneme duration is used to represent the expected duration of the phoneme in the reference audio.

[0064] For example, a trained phoneme duration predictor can predict the second phoneme duration of each phoneme based on the first phoneme duration of each phoneme. If the first phoneme duration of / ǎo / is 0.2s, the phoneme duration prediction model can calculate that the expected duration of / ǎo / in the reference audio is 0.5s.

[0065] Step 12013: Adjust the first phoneme duration of each phoneme in the assembled phoneme data according to the second phoneme duration to obtain phoneme augmentation data. The number of data frames in the phoneme augmentation data is greater than the number of data frames in the assembled phoneme data.

[0066] For example, since the second phoneme duration of / ǎo / is longer than the first phoneme duration, the duration of the first phoneme of / ǎo / in the assembled phoneme data can be adjusted according to the second phoneme duration of / ǎo / . That is, the duration of / ǎo / in the assembled phoneme data is extended from 0.2s to 0.5s. In this way, phoneme augmentation data will be obtained. At this time, if the second phoneme duration of at least one phoneme is longer than the first phoneme duration, the number of data frames of the phoneme augmentation data obtained by adjusting the phoneme duration is usually greater than the number of data frames of the assembled phoneme data.

[0067] In other embodiments, the data processing model includes a second data processing model for processing a reference object, the second data processing model being determined by a spectrum feature encoder, and the spectrum generation object including speech spectrum data of the reference audio. Based on this, step 1201 may specifically include step 12014, in which the second data processing model is used to extract features from the audio spectrum data of the reference audio to obtain speech spectrum data, wherein the data dimension of the speech spectrum data is lower than that of the audio spectrum data.

[0068] For example, the audio spectrum data of the reference audio may include pure music spectrum, pure human voice spectrum, and mixed spectrum of music and human voice. Therefore, the second data processing model can be used to extract the spectrum of human voice, i.e., pure human voice spectrum, and mixed spectrum of music and human voice as speech spectrum data from the audio spectrum data of the reference audio.

[0069] In some other embodiments, the data processing model includes a third data processing model for processing laughter feature data, and the spectrum generation object includes emotional data of the laughter feature data. The third data processing model is trained by the laughter emotion recognition model. Based on this, step 1201 may specifically include step 12015, which uses the third data processing model to perform emotion recognition on the laughter feature data to obtain emotional data.

[0070] For example, the third data processing model can extract laughter feature data, such as pitch, rhythm, and intensity, from laughter. These laughter feature data can reflect a person's emotional state. Then, through machine learning or deep learning algorithms, the extracted laughter feature data is analyzed. By analyzing the laughter feature data, the emotional data behind the laughter, such as happiness, embarrassment, and irony, can be identified. This identification helps to understand a person's emotional state more deeply.

[0071] Thus, when identifying emotional data through the third data processing model, techniques such as laughter feature data and analysis of laughter feature data are used to improve the accuracy and efficiency of emotional data identification.

[0072] Step 1202: Using the speech spectrum generation model, determine the speech spectrum data of laughter based on the spectrum generation object.

[0073] Specifically, step 1202 may include steps 12021 to 12025.

[0074] Step 12021: According to the second phoneme duration of each phoneme in the phoneme augmentation data, align each phoneme in the phoneme augmentation data with the spectral features in the speech spectrum data of the reference audio to obtain aligned audio spectrum data. The aligned audio spectrum data includes speech spectrum data and mask spectrum data of the mask region.

[0075] For example, based on the second phoneme duration of each phoneme, the spectral features, i.e. Mel features, in the speech spectrum data of the reference audio are aligned with the corresponding text phoneme features.

[0076] This ensures consistency between audio features and text features in the duration dimension.

[0077] Step 12022: Reconstruct the audio spectrum data of the masked region using emotional data and phoneme augmentation data to obtain the reconstructed audio spectrum data of the masked region.

[0078] For example, by using emotional data and phoneme augmentation data, the audio spectrum data of the masked region, such as Mel features, is reconstructed so that the phoneme features and Mel features are precisely aligned in the time dimension.

[0079] This ensures that phoneme features and Mel features remain consistent over time, thereby improving the accuracy of feature encoding and the generation effect of the model.

[0080] Step 12023: Replace the mask spectrum data in the aligned audio spectrum data with the reconstructed audio spectrum data to obtain the extended mask spectrum data.

[0081] Step 12024: Based on the emotion data, determine the position of the laughter feature data in the extended mask spectral data.

[0082] For example, different emotional data will affect the position of laughter feature data in the extended mask spectrum data. For instance, if the emotional data represents happiness, the laughter feature data will be distributed at the beginning and end of the extended mask spectrum data, while if the emotional data represents mockery, the laughter feature data will be distributed in a more central position in the extended mask spectrum data.

[0083] Step 12025: According to the location, perform data fusion on the laughter feature data and the extended mask spectrum data to obtain laughter speech spectrum data.

[0084] For example, the location can be determined according to step 12024, and the laughter feature data and the extended mask spectrum data can be fused to obtain laughter speech spectrum data.

[0085] The spectrum generation model mentioned above can be trained using the following steps.

[0086] Therefore, prior to step 120, the audio generation method may include steps 3101 to 3106, as detailed below.

[0087] Step 3101: Obtain the sample reference object and sample laughter feature data. The sample reference object includes the sample reference audio and the sample reference text.

[0088] For example, due to the difficulty in obtaining proprietary sample laughter feature data and sample reference audio, a two-stage training approach is adopted to address the shortage of sample laughter feature data and sample reference audio: a pre-training process and a fine-tuning process. For the pre-training process, the Libri-Light dataset is selected. Libri-Light is a benchmark platform for unsupervised or partially supervised automatic speech recognition. This dataset contains 60,000 hours of data. The sample reference audio in Libri-Light covers various types, including novels, essays, poems, and documents, fully demonstrating diverse text and language expression features. This makes the dataset highly natural and diverse for speech synthesis tasks. In the fine-tuning training stage, 300 hours of sample laughter feature data were collected online. Unlike the pre-training dataset, which mainly consists of natural data from everyday conversations and reading aloud, the sample laughter feature data used for fine-tuning includes laughter from various scenarios related to the aforementioned types of sample reference audio.

[0089] Since the Libri-Light dataset does not provide corresponding text information for the audio, a series of data preprocessing steps are required to generate sample reference text that can be used for model training. The specific process is as follows: First, each audio segment is transcribed using a pre-trained speech recognition model (Whisper) to generate the corresponding sample reference text. To ensure the accuracy of model inference, all audio segments must first be standardized, for example, adjusted to a sampling rate of 16kHz.

[0090] Step 3102: Based on the sample reference object and sample laughter feature data, train the first preprocessing model until the first preset training conditions are met, and obtain the second preprocessing model.

[0091] In some embodiments of this application, the first preprocessing model includes a first sample preprocessing model for processing sample reference objects. The first sample preprocessing model includes a text encoder and a phoneme duration predictor. The second preprocessing model includes a first data processing model for processing reference objects in the data processing model. The sample reference text is the text after the sample reference audio has been processed by speech recognition. The sample reference audio and the sample reference text are time-aligned at the phoneme level. Based on this, step 3102 may specifically include steps 31021 to 31023.

[0092] Step 31021: The sample reference text is processed by a text encoder to obtain sample phoneme data, which includes the sample phoneme sequence corresponding to the text sentence in the sample reference text.

[0093] For example, a text encoder can be used to compress high-dimensional sample reference text into a compact vector representation, enabling the model to efficiently process large-scale data without losing key information. It also serves to align information from different modalities, such as sample reference text and sample reference audio, ensuring that information from different modalities can be effectively fused and work together.

[0094] Specifically, a text encoder using a multi-layered deep learning model (Transformer) architecture for processing sequence data employs a multi-head self-attention mechanism, enabling it to simultaneously focus on different locations within the text and capture both global and local semantic information. This can be expressed as shown in the following formula (1):

[0095] Word embedding representation: Each text word in the text information Convert to its corresponding word embedding vector By embedding matrix To achieve:

[0096] (1)

[0097] in, It is a pre-trained or randomly initialized embedding matrix. It is an index of text words.

[0098] Encoding process: For text information with a sequence length of... The text sentence, the text encoder generates the hidden state To capture contextual relationship data in the text. For the Transformer model, it relies on a self-attention mechanism, as shown in Equation (2):

[0099] (2)

[0100] Output representation: The output generated by the text encoder is sentence-level sample phoneme data, which is represented as... ,Right now .

[0101] Step 31022: Using a phoneme duration predictor, calculate the second sample phoneme duration for each sample phoneme based on the first sample phoneme duration of each sample phoneme in the sample phoneme sequence. The second sample phoneme duration is used to represent the expected duration of the sample phoneme in the sample reference audio.

[0102] For example, if the duration of the first sample phoneme / ǎo / is 3s, the phoneme duration predictor can calculate the expected duration of the sample phoneme / ǎo / in the sample reference audio, i.e., the duration of the second sample phoneme, which is, for example, 4s, based on the first sample phoneme duration of 3s. Step 31023: If the difference between the duration of the second sample phoneme and the duration of the third sample phoneme is less than or equal to a second threshold, the text encoder is determined as the text encoding model in the first data processing model, and the phoneme duration predictor is determined as the phoneme duration prediction model in the first data processing model. The duration of the third sample phoneme is the actual duration of the sample phoneme in the sample reference audio.

[0103] For example, if the second sample phoneme duration of the sample phoneme / ǎo / is, for example, 4 seconds, and the actual duration of the sample phoneme / ǎo / in the sample reference audio is, for example, the third sample phoneme duration, is, 5 seconds, then the difference between the two is less than or equal to the second threshold 3. In this case, the text encoder can be determined as the text encoding model in the first data processing model, and the phoneme duration predictor can be determined as the phoneme duration prediction model in the first data processing model.

[0104] In some other embodiments of this application, the first preprocessing model includes a second sample preprocessing model for processing the sample reference object. The second sample preprocessing model includes a spectral feature encoder. The second preprocessing model includes a second data processing model for processing the reference object in the data processing model. Based on this, step 3102 may specifically include steps 31024 to 31025.

[0105] Step 31024: Using a spectral feature encoder, feature extraction is performed on the sample audio spectrum data of the sample reference audio to obtain sample speech spectrum data. The data dimension of the sample speech spectrum data is lower than that of the sample audio spectrum data.

[0106] Step 31025: If the proportion of the sample speech spectrum data in the sample audio spectrum data is greater than or equal to the third threshold, the spectrum feature encoder is determined as the second data processing model.

[0107] For example, a spectral feature encoder, also known as a Mel spectrogram encoder, extracts the spectral features of an audio signal, preserving the time-frequency information of speech, thus enabling the model to perceive the frequency distribution and temporal variations of the audio. To improve processing efficiency, the encoder typically uses a one-dimensional average pooling layer to downsample the input Mel features multiple times, effectively reducing the dimensionality of the features and saving computational resources for subsequent processing. Structurally, a Mel spectrogram encoder typically includes a one-dimensional convolutional layer, layer normalization, activation functions, and average pooling layers. It can extract higher-level Mel spectrogram features layer by layer, ensuring sufficient capture of temporal dependencies and spectral information.

[0108] In some other embodiments of this application, the first preprocessing model includes a laughter emotion recognition model for processing sample laughter feature data, and the second preprocessing model includes a third data processing model for processing laughter feature data in the data processing model. Based on this, step 3102 may specifically include steps 31026 to 31028.

[0109] Step 31026: Extract sample laughter feature data from the sample audio.

[0110] Step 31027: Using the laughter emotion recognition model, the sample laughter feature data is identified to obtain the first sample emotion data of the sample laughter feature data.

[0111] Step 31028: If the difference between the first sample emotion data and the second sample emotion data corresponding to the sample laughter feature data is less than or equal to the fourth threshold, the laughter emotion recognition model is determined as the third data processing model.

[0112] For example, the extraction of emotional data, specifically the latent variables of laughter, is a crucial step in the controllable laughter generation model, aiming to extract the emotional control variables from the audio. Specifically, this process employs an open-source speech emotion recognition model similar to emotion2vec to extract emotional vectors from the audio. After extraction, the length of the emotional vectors is adjusted by connecting linear layers to align with the model's other input features. During training, to ensure the stability of the extracted emotional vectors, the weights of the laughter latent variable extraction model are typically frozen, and only subsequent linear layers are trained. This design guarantees the effectiveness of the emotional vectors in capturing the emotional features of laughter and allows for adaptive adjustments during generation, thus becoming a core element in achieving controllable laughter generation.

[0113] Therefore, the embodiments of this application can utilize a laughter detection model to effectively extract emotional latent variables in laughter, and through fine-tuning training on a proprietary laughter dataset, can achieve accurate and effective generation of laughter.

[0114] Step 3103: The reference object and sample laughter feature data are processed through the second preprocessing model to obtain a sample spectrum generation object. The sample spectrum generation object includes sample phoneme augmentation data corresponding to the sample reference text, sample speech spectrum data of the sample reference audio, and sample emotion data of the sample laughter feature data. The sample phoneme augmentation data includes the sample assembly phoneme data of the sample assembly text after augmentation phoneme data frames. The sample assembly text is the text concatenated from the sample reference text and the sample audio text of the sample reference audio.

[0115] Step 3104: Construct the first sample laughter speech spectrum data by generating an object based on the sample spectrum using the sample spectrum decoder.

[0116] For example, in this embodiment of the application, the sample spectrum decoder and the sample spectrum discriminator can form a generative adversarial network (GAN). The sample spectrum decoder attempts to generate realistic data to fool the discriminator, while the sample spectrum discriminator is responsible for distinguishing between real data and generated data. The two continuously compete against each other, adjusting parameters through backpropagation. Ultimately, the generator produces highly realistic data, achieving tasks such as data synthesis and simulation.

[0117] Specifically, step 3104 may include steps 31041 to 31045.

[0118] Step 31041: Extract the sample audio spectrum data of the sample reference audio using the sample spectrum decoder.

[0119] Step 31042: Perform masking processing on the sample audio spectrum data to obtain sample mask spectrum data, which includes sample mask spectrum data of at least one sample mask region.

[0120] For example, during the training phase, random masking is applied to the Mel spectrogram to enhance the model's robustness and generalization ability. Specifically, a starting position is randomly selected on the time axis of the Mel spectrogram, and then a time frame following that starting position is masked. The mask length is randomly determined, within a defined range, and a maximum mask length is set to prevent obscuring excessive information. Finally, based on the generated mask position and size, the values ​​of the corresponding regions in the Mel spectrogram are set to zero.

[0121] The specific steps are as follows: First, given the input Mel sequence... ,in This represents the frame number of Mel, and each element... Represents the Mel spectrum. Generate mask length. By setting the mask length within a range Between, the length of the random mask It can be represented as: Generate mask starting position and mask area , , The masking operation masks certain frames of the input sequence, resulting in a masked sequence. It is represented by the following formula (3):

[0122] (3)

[0123] Step 31043: According to the second sample phoneme duration of each sample phoneme in the sample phoneme augmentation data, align each sample phoneme in the sample phoneme augmentation data with the sample spectral features in the sample speech spectrum data of the sample reference audio to obtain sample aligned audio spectrum data. The sample aligned audio spectrum data includes sample speech spectrum data and sample mask spectrum data of the sample mask region.

[0124] For example, the Montreal Forced Aligner (MFA) tool is used to perform phoneme-level temporal alignment of the sample spectral features and sample reference phonemes in the sample reference audio. The MFA tool uses an acoustic model to annotate each phoneme in the audio signal of the sample spectral features, generating the start and end times of each phoneme, thereby obtaining accurate phoneme durations and improving the accuracy of the alignment between sample phonemes and sample spectral features.

[0125] Step 31044: Using sample emotion data and sample phoneme augmentation data, the audio spectrum data of the sample mask region is reconstructed to obtain the sample reconstructed audio spectrum data of the sample mask region.

[0126] For example, a phoneme duration predictor and LR phoneme embedding frame augmentation are used. The function of the phoneme duration predictor is to predict the duration of each phoneme. In the data preprocessing stage of training, the actual duration of each phoneme has been obtained by the MFA alignment tool. In the frame augmentation stage of phoneme embedding, the duration predictor is trained by using the real phoneme duration, and the phoneme embedding is augmented by using this real duration information. In terms of model structure, the duration predictor adopts a multi-layer Transformer stacked structure. This design fully extracts contextual information, thereby improving the accuracy of phoneme duration prediction. This method not only enhances the model's ability to predict phoneme duration, but also ensures the naturalness and coherence of the generated speech. In order to prevent the training of the phoneme duration predictor from affecting the training of the entire module, the embodiment of this application adopts the gradient truncation method to train the module separately. Specifically, the mean squared error loss can be used to promote model convergence. The specific formula (4) is as follows:

[0127] (4)

[0128] in, This represents the predicted duration of the i-th phoneme. This represents the actual duration corresponding to the i-th phoneme. Assuming there are three phoneme duration samples, the actual value... Predicted value Then its loss can be expressed as: .

[0129] Step 31045: Replace the sample mask spectrum data in the sample aligned audio spectrum data with the sample reconstructed audio spectrum data to obtain the sample expanded mask spectrum data.

[0130] Step 31046: Based on the sample emotion data, determine the sample position of the sample laughter feature data in the sample augmented mask spectrum data.

[0131] Step 31047: According to the sample location, perform data fusion on the sample laughter feature data and the sample augmented mask spectrum data to obtain the first sample laughter speech spectrum data.

[0132] For example, after receiving sample text information, sample audio information, and sample laughter feature data, the Mel spectrogram decoder in this embodiment fuses these three types of information to obtain the first sample laughter speech spectrum data. Based on the fused information, the Mel generator starts the feature extraction and transformation process. It uses a series of neural networks and mathematical transformations to convert the input encoded information into a spectrum representation on the Mel frequency scale. In terms of model structure, unlike other models that use the Transformer architecture, it adopts the more efficient ConvNeXt V2 architecture. The ConvNeXtV2 architecture significantly improves the performance of pure convolutional networks on various recognition benchmarks by using the Fully Convolutional Masked Autoencoder (FCMAE) framework and a new Global Response Normalization (GRN) layer. For the Mel decoder, this embodiment can use the Mean Squared Error (MSE) loss to guide model optimization, and the specific calculation formula (5) is as follows:

[0133] (5)

[0134] in This represents the Mel spectrum, which represents the first sample of laughter speech spectrum data generated. This represents the second sample of laughter speech spectrum data, which is the actual Mel spectrum. This represents the difference between the generated Mel spectrum (which is randomly masked) and the true Mel spectrum.

[0135] Therefore, this application embodiment can use a random masking scheme to randomly mask the input Mel data, thereby creating different training samples and enhancing data diversity. Simultaneously, by introducing random masks, the model is forced to learn invariant and robust features in the data, which helps improve the model's generalization ability when facing new data and reduces the risk of overfitting. Furthermore, by pre-training on large-scale big data and fine-tuning on relatively limited small-scale laughter data, the difficulty of obtaining laughter data can be effectively overcome without changing model performance.

[0136] Step 3105: Using a sample spectrum discriminator, determine the difference between the first sample laughter speech spectrum data and the second sample laughter speech spectrum data. The second sample laughter speech spectrum data is the reference laughter speech spectrum data corresponding to the sample reference object and the sample laughter feature data.

[0137] Step 3106: If the difference value is less than or equal to the first threshold, the second preprocessing model is determined as the data processing model, and the sample spectrum decoder is determined as the speech spectrum generation model.

[0138] For example, the sample spectrum discriminator can be used to determine the difference between the Mel spectra generated by the generative model and the Mel spectra obtained from the real audio conversion. During training, an adversarial training approach is adopted, in which the generative model and the discriminator compete with each other. The generative model attempts to generate Mel spectra that can fool the discriminator, while the discriminator tries to distinguish between real and generated Mel spectra. This adversarial process prompts both sides to continuously improve their abilities, ultimately enabling the generative model to generate high-quality Mel spectra. If the real Mel spectrum classification is set to 1, the loss function of the discriminator is as shown in the following formula (6):

[0139] (6)

[0140] The loss function of the generator is shown in the following formula (7):

[0141] (7)

[0142] in This represents the generated Mel spectrum, which is the spectral data of the first sample laughter speech. This represents the true Mel-spectral representation of the second sample laughter speech spectrum data. This is the output of the discriminator to the generated Mel spectrum data (or audio), i.e., the probability that the discriminator considers it to be real data. After the model training is completed, the Mel discriminator is discarded, and the trained vocoder is used as a spectrum generator to convert the Mel spectrogram into synthesized laughter audio.

[0143] Then, in step 130, in some embodiments of this application, the laughter speech spectrum data generated by the Mel encoder can be sent to a vocoder to synthesize the corresponding laughter synthesized audio. Here, a common vocoder in speech synthesis, such as HiFiGAN, BigVGAN, etc., can be selected.

[0144] Thus, the audio generation method in this application embodiment can be applied to scenarios in virtual games where interesting scenes or tasks are set up to generate synthesized laughter audio, so that laughter is triggered when the user completes the task, or in virtual reality experiences, the occurrence of laughter can be controlled according to the atmosphere and plot changes of the scene.

[0145] It should be noted that the audio generation method provided in this application can be executed by electronic devices such as mobile phones, tablets, laptops, PDAs, and wearable devices. Some embodiments of this application use electronic devices as the executing entity to illustrate the audio generation method provided in this application.

[0146] The audio generation method provided in this application can be executed by an audio generation device. This application uses an audio generation device executing the audio generation method as an example to illustrate the device for the audio generation method provided in this application.

[0147] This application also provides an audio generation apparatus. (Specifically combined with...) Figure 3 Please provide a detailed explanation.

[0148] Figure 3 This is a schematic diagram of the structure of an audio generation apparatus provided for some embodiments of this application.

[0149] like Figure 3 As shown, the audio generation device 30 can be applied to electronic devices, and the audio generation device 30 may specifically include:

[0150] The acquisition module 301 is used to acquire a reference object and laughter feature data related to the reference object. The reference object includes reference text and reference audio. The reference text is the text used for laughter synthesis, and the reference audio is the reference audio used to indicate the generation of audio with a preset style.

[0151] The determination module 302 is used to determine the laughter speech spectrum data based on the reference object and laughter feature data through the spectrum generation model;

[0152] The conversion module 303 is used to convert laughter speech spectrum data into laughter synthesized audio, wherein the laughter synthesized audio is in a preset style and includes the text content of the reference text and the audio of laughter related to the emotion represented by the text content.

[0153] The audio generation device 30 in the embodiments of this application will be described in detail below.

[0154] In some embodiments of this application, the determining module 302 can be specifically used to, when the spectrum generation model includes a data processing model and a speech spectrum generation model, process the reference object and laughter feature data through the data processing model to obtain a spectrum generation object; wherein, the spectrum generation object includes phoneme augmentation data corresponding to the reference text, speech spectrum data of the reference audio, and emotional data of laughter feature data, the phoneme augmentation data includes the data of the assembled phoneme data of the assembled text after augmentation phoneme data frames, and the assembled text is the text concatenated from the reference text and the audio text of the reference audio;

[0155] Laughter speech spectrum data is determined based on the spectrum generation object using a speech spectrum generation model.

[0156] In some embodiments of this application, the audio generation apparatus 30 may further include a processing module for performing speech recognition processing on the audio spectrum data of the reference audio to obtain audio text, provided that the data processing model includes a first data processing model for processing a reference object, the first data processing model includes a text encoding model and a phoneme duration prediction model, and the spectrum generation object includes phoneme augmentation data.

[0157] The audio generation device 30 may also include a splicing module for splicing the reference text and the audio text to obtain the assembled text;

[0158] The audio generation device 30 may further include a processing module for performing phoneme conversion processing on the assembled text through a text encoding model to obtain assembled phoneme data, wherein the assembled phoneme data includes a phoneme sequence corresponding to the assembled text;

[0159] The audio generation device 30 may also include a calculation module for calculating the second phoneme duration of each phoneme in the phoneme sequence based on the first phoneme duration of each phoneme in the phoneme sequence using a phoneme duration prediction model. The second phoneme duration is used to represent the expected duration of the phoneme in the reference audio.

[0160] The audio generation device 30 may further include an adjustment module for adjusting the first phoneme duration of each phoneme in the assembled phoneme data according to the second phoneme duration, to obtain phoneme augmentation data, wherein the number of data frames of the phoneme augmentation data is greater than the number of data frames of the assembled phoneme data.

[0161] In some embodiments of this application, the audio generation apparatus 30 may further include a processing module for extracting features from the audio spectrum data of the reference audio through the second data processing model, wherein the data processing model includes a second data processing model for processing a reference object, the second data processing model is determined by a spectrum feature encoder, and the spectrum generation object includes speech spectrum data of the reference audio, to obtain speech spectrum data, wherein the data dimension of the speech spectrum data is lower than that of the audio spectrum data.

[0162] In some embodiments of this application, the determining module 302 may be specifically used to align each phoneme in the phoneme augmentation data with the spectral features in the speech spectrum data of the reference audio according to the second phoneme duration of each phoneme in the phoneme augmentation data, so as to obtain aligned audio spectrum data, which includes speech spectrum data and mask spectrum data of the mask region.

[0163] By using emotional data and phoneme augmentation data, the audio spectrum data of the masked region is reconstructed to obtain the reconstructed audio spectrum data of the masked region.

[0164] Replace the mask spectrum data in the aligned audio spectrum data with the reconstructed audio spectrum data to obtain the extended mask spectrum data;

[0165] Based on emotion data, determine the location of laughter feature data in the augmented masked spectral data;

[0166] Based on location, the laughter feature data and the extended mask spectrum data are fused to obtain the laughter speech spectrum data.

[0167] In some embodiments of this application, the acquisition module 301 can also be used to acquire sample reference objects and sample laughter feature data, wherein the sample reference objects include sample reference audio and sample reference text.

[0168] The audio generation device 30 may also include a training module for training the first preprocessing model based on the sample reference object and sample laughter feature data until the first preset training conditions are met, thereby obtaining the second preprocessing model.

[0169] The audio generation device 30 may further include a processing module for processing the reference object and sample laughter feature data through a second preprocessing model to obtain a sample spectrum generation object. The sample spectrum generation object includes sample phoneme augmentation data corresponding to the sample reference text, sample speech spectrum data of the sample reference audio, and sample emotion data of the sample laughter feature data. The sample phoneme augmentation data includes the data of the sample assembly phoneme data of the sample assembly text after augmentation phoneme data frames. The sample assembly text is the text concatenated from the sample reference text and the sample audio text of the sample reference audio.

[0170] The audio generation device 30 may also include a building module for constructing first sample laughter speech spectrum data from the sample spectrum generation object using a sample spectrum decoder.

[0171] The determining module 302 can also be used to determine the difference between the first sample laughter speech spectrum data and the second sample laughter speech spectrum data through the sample spectrum discriminator, wherein the second sample laughter speech spectrum data is the reference laughter speech spectrum data corresponding to the sample reference object and the sample laughter feature data.

[0172] The determination module 302 can also be used to determine the second preprocessing model as the data processing model when the difference value is less than or equal to the first threshold, and to determine the sample spectrum decoder as the speech spectrum generation model.

[0173] In some embodiments of this application, the audio generation apparatus 30 may further include a training module for processing sample reference objects in a first preprocessing model, which includes a text encoder and a phoneme duration predictor; and a second preprocessing model, which includes a first data processing model in a data processing model for processing reference objects. The sample reference text is the text of the sample reference audio after speech recognition processing. With the sample reference audio and sample reference text aligned temporally at the phoneme level, the text encoder performs phoneme conversion processing on the sample reference text to obtain sample phoneme data. The sample phoneme data includes sample phoneme sequences corresponding to text sentences in the sample reference text.

[0174] The phoneme duration predictor calculates the second sample phoneme duration for each sample phoneme based on the first sample phoneme duration in the sample phoneme sequence. The second sample phoneme duration is used to represent the expected duration of the sample phoneme in the sample reference audio.

[0175] If the difference between the duration of the second sample phoneme and the duration of the third sample phoneme is less than or equal to the second threshold, the text encoder is determined as the text encoding model in the first data processing model, and the phoneme duration predictor is determined as the phoneme duration prediction model in the first data processing model. The duration of the third sample phoneme is the actual duration of the sample phoneme in the sample reference audio.

[0176] In some embodiments of this application, the audio generation apparatus 30 may further include a training module for extracting sample laughter feature data from sample audio, provided that the first preprocessing model includes a laughter emotion recognition model for processing sample laughter feature data and the second preprocessing model includes a third data processing model for processing laughter feature data in the data processing model.

[0177] The laughter emotion recognition model is used to identify the feature data of sample laughter and obtain the first sample emotion data of the sample laughter feature data.

[0178] If the difference between the first sample emotion data and the second sample emotion data corresponding to the sample laughter feature data is less than or equal to the fourth threshold, the laughter emotion recognition model is determined as the third data processing model.

[0179] In some embodiments of this application, the audio generation apparatus 30 may further include a building module for extracting sample audio spectrum data of the sample reference audio through a sample spectrum decoder.

[0180] The sample audio spectrum data is masked to obtain sample mask spectrum data, which includes sample mask spectrum data of at least one sample mask region.

[0181] According to the second sample phoneme duration of each sample phoneme in the sample phoneme augmentation data, each sample phoneme in the sample phoneme augmentation data is aligned with the sample spectral features in the sample speech spectral data of the sample reference audio to obtain sample aligned audio spectral data. The sample aligned audio spectral data includes sample speech spectral data and sample mask spectral data of the sample mask region.

[0182] By using sample emotion data and sample phoneme augmentation data, the audio spectrum data of the sample mask region is reconstructed to obtain the sample reconstructed audio spectrum data of the sample mask region.

[0183] Replace the sample mask spectrum data in the sample aligned audio spectrum data with the sample reconstructed audio spectrum data to obtain the sample augmented mask spectrum data.

[0184] Based on the sample emotion data, determine the sample position of the sample laughter feature data in the sample augmented mask spectrum data;

[0185] Based on the sample location, the sample laughter feature data and the sample augmented mask spectrum data are fused to obtain the first sample laughter speech spectrum data.

[0186] In some embodiments of this application, the audio generation device 30 may further include a display module for displaying a laughter synthesis interface, the laughter synthesis interface including a first area and a second area.

[0187] The audio generating device 30 may further include a receiving module for receiving a first input from a user to the first region;

[0188] The determination module 302 can also be used to determine the text entered by the user in the first area as the reference text in response to the first input;

[0189] The audio generating device 30 may further include a receiving module for receiving a second input from the user to the second region;

[0190] The determining module 302 can also be used to, in response to the second input, determine the preset audio selected by the user in the second area as the reference audio; or, determine the user audio added by the user in the second area as the reference audio.

[0191] In some embodiments of this application, the audio generation device 30 may further include a receiving module for receiving a third input from the user to the third area when the laughter synthesis interface further includes a third area, the third area being used to display at least one of the following preset options: preset laughter audio option, preset laughter tag option;

[0192] The determination module 302 can also be used, in response to the third input, to determine the laughter feature data corresponding to the target preset option selected by the user in the third area as laughter feature data related to the reference object.

[0193] In some embodiments of this application, the audio generation apparatus 30 may further include a generation module for generating laughter feature data related to the reference object based on a first textual semantics of the reference text and a second textual semantics of the text corresponding to the reference audio.

[0194] The audio generation device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0195] The audio generation device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.

[0196] The audio generation device provided in this application embodiment can achieve... Figures 1 to 2 The various processes implemented in the audio generation method embodiments shown achieve the same technical effect, and will not be described again here to avoid repetition.

[0197] Based on this, the audio generation apparatus provided in this application can determine laughter speech spectrum data through a spectrum generation model, based on a reference object and laughter feature data related to the reference object. The reference object includes reference text and reference audio; the reference text is the text used for laughter synthesis, and the reference audio is the reference audio used to instruct the generation of audio in a preset style. Then, the laughter speech spectrum data is converted into a synthesized laughter audio in a preset style, including the text content of the reference text and laughter related to the emotion represented by the text content. In this way, synthesized laughter audio can be automatically generated based on the reference text, reference audio, and laughter feature data, eliminating the need for users to spend a significant amount of time manually editing, simplifying the operation of generating synthesized laughter audio, reducing the difficulty of generating synthesized laughter audio, and improving the efficiency and quality of generating synthesized laughter audio.

[0198] Optional, such as Figure 4 As shown, this application embodiment also provides an electronic device 40, including a processor 401 and a memory 402. The memory 402 stores a program or instructions that can run on the processor 401. When the program or instructions are executed by the processor 401, they implement the various steps of the above-described audio generation method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0199] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0200] Figure 5 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application.

[0201] The electronic device 500 includes, but is not limited to, components such as: radio frequency unit 501, network module 502, audio output unit 503, input unit 504, sensor 505, display unit 506, user input unit 507, interface unit 508, memory 509, and processor 510.

[0202] Those skilled in the art will understand that the electronic device 500 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 510 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system.

[0203] Figure 5 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0204] In this embodiment, the processor 510 is configured to acquire a reference object and laughter feature data related to the reference object. The reference object includes reference text and reference audio. The reference text is text used for laughter synthesis, and the reference audio is audio used to instruct the generation of a preset style. The processor 510 is further configured to determine laughter speech spectrum data based on the reference object and the laughter feature data using a spectrum generation model. The processor 510 is further configured to convert the laughter speech spectrum data into laughter synthesized audio, wherein the laughter synthesized audio is audio of a preset style, and the laughter synthesized audio includes the text content of the reference text and laughter related to the emotion represented by the text content.

[0205] The electronic device 500 will be described in detail below.

[0206] In some embodiments of this application, processor 510 is used to process reference object and laughter feature data through data processing model when the spectrum generation model includes data processing model and spectrum generation model to obtain spectrum generation object; wherein, spectrum generation object includes phoneme augmentation data corresponding to reference text, speech spectrum data of reference audio and emotional data of laughter feature data, phoneme augmentation data includes data of assembled phoneme data of assembled text after augmentation phoneme data frames, and assembled text is text concatenated from reference text and audio text of reference audio;

[0207] Laughter speech spectrum data are determined based on the spectrum generation object using a spectrum generation model.

[0208] In some embodiments of this application, processor 510 is configured to perform speech recognition processing on audio spectrum data of reference audio to obtain audio text, provided that the data processing model includes a first data processing model for processing a reference object, the first data processing model includes a text encoding model and a phoneme duration prediction model, and the spectrum generation object includes phoneme augmentation data.

[0209] The reference text and audio text are concatenated to obtain the assembled text;

[0210] The assembly text is processed by phoneme conversion using a text encoding model to obtain assembly phoneme data, which includes phoneme sequences corresponding to text sentences in the assembly text.

[0211] The phoneme duration prediction model calculates the second phoneme duration of each phoneme based on the first phoneme duration of each phoneme in the phoneme sequence. The second phoneme duration is used to represent the expected duration of the phoneme in the reference audio.

[0212] Based on the second phoneme duration, the first phoneme duration of each phoneme in the assembled phoneme data is adjusted to obtain phoneme augmentation data. The number of data frames in the phoneme augmentation data is greater than the number of data frames in the assembled phoneme data.

[0213] In some embodiments of this application, processor 510 is configured to perform feature extraction on audio spectrum data of reference audio through the second data processing model when the data processing model includes a second data processing model for processing a reference object, the second data processing model being determined by a spectrum feature encoder, and the spectrum generation object including speech spectrum data of reference audio, to obtain speech spectrum data, wherein the data dimension of the speech spectrum data is lower than that of the audio spectrum data.

[0214] In some embodiments of this application, the processor 510 is configured to align each phoneme in the phoneme extension data with the spectral features in the speech spectrum data of the reference audio according to the second phoneme duration of each phoneme in the phoneme extension data, thereby obtaining aligned audio spectrum data, which includes speech spectrum data and mask spectrum data of the mask region.

[0215] By using emotional data and phoneme augmentation data, the audio spectrum data of the masked region is reconstructed to obtain the reconstructed audio spectrum data of the masked region.

[0216] Replace the mask spectrum data in the aligned audio spectrum data with the reconstructed audio spectrum data to obtain the extended mask spectrum data;

[0217] Based on emotion data, determine the location of laughter feature data in the augmented masked spectral data;

[0218] Based on location, the laughter feature data and the extended mask spectrum data are fused to obtain the laughter speech spectrum data.

[0219] In some embodiments of this application, processor 510 is used to acquire sample reference objects and sample laughter feature data, wherein the sample reference objects include sample reference audio and sample reference text;

[0220] Based on the sample reference object and sample laughter feature data, the first preprocessing model is trained until the first preset training condition is met, and the second preprocessing model is obtained.

[0221] The second preprocessing model processes the reference object and sample laughter feature data to obtain a sample spectrum generation object. The sample spectrum generation object includes sample phoneme augmentation data corresponding to the sample reference text, sample speech spectrum data of the sample reference audio, and sample emotion data of the sample laughter feature data. The sample phoneme augmentation data includes the sample assembly phoneme data of the sample assembly text after augmentation phoneme data frames. The sample assembly text is the text concatenated from the sample reference text and the sample audio text of the sample reference audio.

[0222] The first sample laughter speech spectrum data is constructed by generating an object based on the sample spectrum using a sample spectrum decoder.

[0223] The difference between the first sample laughter speech spectrum data and the second sample laughter speech spectrum data is determined by the sample spectrum discriminator. The second sample laughter speech spectrum data is the reference laughter speech spectrum data corresponding to the sample reference object and the sample laughter feature data.

[0224] If the difference value is less than or equal to the first threshold, the second preprocessing model is determined as the data processing model, and the sample spectrum decoder is determined as the spectrum generation model.

[0225] In some embodiments of this application, the processor 510 is configured to: first preprocessing model including a first sample preprocessing model for processing sample reference objects, the first sample preprocessing model including a text encoder and a phoneme duration predictor; second preprocessing model including a first data processing model for processing reference objects in the data processing model; sample reference text being the text of sample reference audio after speech recognition processing; sample reference audio and sample reference text being temporally aligned at the phoneme level; and sample reference text being processed by a text encoder to perform phoneme conversion processing to obtain sample phoneme data, the sample phoneme data including sample phoneme sequences corresponding to text sentences in the sample reference text;

[0226] The phoneme duration predictor calculates the second sample phoneme duration for each sample phoneme based on the first sample phoneme duration in the sample phoneme sequence. The second sample phoneme duration is used to represent the expected duration of the sample phoneme in the sample reference audio.

[0227] If the difference between the duration of the second sample phoneme and the duration of the third sample phoneme is less than or equal to the second threshold, the text encoder is determined as the text encoding model in the first data processing model, and the phoneme duration predictor is determined as the phoneme duration prediction model in the first data processing model. The duration of the third sample phoneme is the actual duration of the sample phoneme in the sample reference audio.

[0228] In some embodiments of this application, processor 510 is configured to extract sample laughter feature data from sample audio when the first preprocessing model includes a laughter emotion recognition model for processing sample laughter feature data and the second preprocessing model includes a third data processing model for processing laughter feature data in the data processing model.

[0229] The laughter emotion recognition model is used to identify the feature data of sample laughter and obtain the first sample emotion data of the sample laughter feature data.

[0230] If the difference between the first sample emotion data and the second sample emotion data corresponding to the sample laughter feature data is less than or equal to the fourth threshold, the laughter emotion recognition model is determined as the third data processing model.

[0231] In some embodiments of this application, processor 510 is configured to extract sample audio spectrum data of sample reference audio using a sample spectrum decoder.

[0232] The sample audio spectrum data is masked to obtain sample mask spectrum data, which includes sample mask spectrum data of at least one sample mask region.

[0233] According to the second sample phoneme duration of each sample phoneme in the sample phoneme augmentation data, each sample phoneme in the sample phoneme augmentation data is aligned with the sample spectral features in the sample speech spectral data of the sample reference audio to obtain sample aligned audio spectral data. The sample aligned audio spectral data includes sample speech spectral data and sample mask spectral data of the sample mask region.

[0234] By using sample emotion data and sample phoneme augmentation data, the audio spectrum data of the sample mask region is reconstructed to obtain the sample reconstructed audio spectrum data of the sample mask region.

[0235] Replace the sample mask spectrum data in the sample aligned audio spectrum data with the sample reconstructed audio spectrum data to obtain the sample augmented mask spectrum data.

[0236] Based on the sample emotion data, determine the sample position of the sample laughter feature data in the sample augmented mask spectrum data;

[0237] Based on the sample location, the sample laughter feature data and the sample augmented mask spectrum data are fused to obtain the first sample laughter speech spectrum data.

[0238] In some embodiments of this application, the display unit 506 is used to display a laughter synthesis interface, which includes a first area and a second area.

[0239] User input unit 507 is used to receive the user's first input to the first area;

[0240] The processor 510 is configured to, in response to the first input, determine the text entered by the user in the first area as the reference text;

[0241] User input unit 507 is used to receive a second input from the user to the second area;

[0242] The processor 510 is configured to, in response to the second input, determine a preset audio selected by the user in the second area as a reference audio; or, determine a user audio added by the user in the second area as a reference audio.

[0243] In some embodiments of this application, the user input unit 507 is used to receive third input from the user to the third area when the laughter synthesis interface further includes a third area, the third area being used to display at least one of the following preset options: preset laughter audio option, preset laughter tag option;

[0244] The processor 510 can also be used, in response to a third input, to determine the laughter feature data corresponding to the target preset option selected by the user in the third area as laughter feature data related to the reference object.

[0245] In some embodiments of this application, processor 510 is configured to generate laughter feature data related to a reference object based on a first textual semantics of the reference text and a second textual semantics of the text corresponding to the reference audio.

[0246] It should be understood that the input unit 504 may include a graphics processing unit (GPU) 5041 and a microphone 5042. The GPU 5041 processes image data of still images or videos acquired by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 506 may include a display panel, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 507 includes at least one of a touch panel 5071 and other input devices 5072. The touch panel 5071 is also called a touch screen. The touch panel 5071 may include two parts: a touch detection device and a touch display. Other input devices 5072 may include, but are not limited to, a physical keyboard, function keys (such as volume display buttons, power buttons, etc.), a trackball, a mouse, and a joystick, which will not be described in detail here.

[0247] The memory 509 can be used to store software programs and various data. The memory 509 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 509 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 509 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0248] Processor 510 may include one or more processing units; optionally, processor 510 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless display signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 510.

[0249] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described audio generation method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0250] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0251] In addition, this application embodiment provides another chip, which includes a processor and a display interface. The display interface and the processor are coupled. The processor is used to run programs or instructions to implement the various processes of the above-described audio generation method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0252] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0253] This application provides a computer program product that is stored in a storage medium and executed by at least one processor to implement the various processes of the audio generation method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0254] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0255] Furthermore, it should be noted that the scope of the methods and apparatus in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. In addition, features described with reference to certain examples may be combined in other examples.

[0256] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0257] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An audio generation method, characterized in that, include: Obtain a reference object and laughter feature data related to the reference object. The reference object includes reference text and reference audio. The reference text is text used for laughter synthesis, and the reference audio is audio used to indicate the generation of a preset style. The reference object and the laughter feature data are processed by the data processing model in the spectrum generation model to obtain a spectrum generation object; wherein, the spectrum generation object includes phoneme augmentation data corresponding to the reference text, speech spectrum data of the reference audio, and emotional data of the laughter feature data, the phoneme augmentation data includes the data of the assembled phoneme data of the assembled text after augmentation phoneme data frames, and the assembled text is the text concatenated from the reference text and the audio text of the reference audio; Using the speech spectrum generation model in the spectrum generation model, each phoneme in the phoneme augmentation data is aligned with the spectral features in the speech spectrum data of the reference audio according to the second phoneme duration of each phoneme in the phoneme augmentation data, to obtain aligned audio spectrum data. The aligned audio spectrum data includes the speech spectrum data and the mask spectrum data of the mask region. The second phoneme duration is used to represent the expected duration of the phoneme in the reference audio. The audio spectrum data of the mask region is reconstructed using the emotion data and the phoneme augmentation data to obtain reconstructed audio spectrum data of the mask region. The mask spectrum data in the aligned audio spectrum data is replaced with the reconstructed audio spectrum data to obtain augmented mask spectrum data. Based on the emotion data, the position of the laughter feature data in the augmented mask spectrum data is determined. According to the position, the laughter feature data and the augmented mask spectrum data are fused to obtain laughter speech spectrum data. The laughter speech spectrum data is converted into laughter synthesized audio, wherein the laughter synthesized audio is audio of the preset style, and the laughter synthesized audio includes the text content of the reference text and laughter related to the emotion represented by the text content.

2. The method according to claim 1, characterized in that, The data processing model includes a first data processing model for processing the reference object, the first data processing model includes a text encoding model and a phoneme duration prediction model, and the spectrum generation object includes the phoneme augmentation data; Before processing the reference object and the laughter feature data using the data processing model to obtain the spectrum generation object, the method further includes: The audio spectrum data of the reference audio is subjected to speech recognition processing to obtain the audio text; The reference text and the audio text are concatenated to obtain the assembled text; The process of processing the reference object and the laughter feature data using the data processing model to obtain a spectrum generation object includes: The assembly text is processed by phoneme conversion using the text encoding model to obtain the assembly phoneme data, which includes the phoneme sequence corresponding to the assembly text. Using the phoneme duration prediction model, the second phoneme duration of each phoneme in the phoneme sequence is calculated based on the first phoneme duration. According to the second phoneme duration, the first phoneme duration of each phoneme in the assembled phoneme data is adjusted to obtain the phoneme augmentation data, wherein the number of data frames of the phoneme augmentation data is greater than the number of data frames of the assembled phoneme data.

3. The method according to claim 1, characterized in that, The data processing model includes a second data processing model for processing the reference object, the second data processing model being determined by a spectral feature encoder, and the spectrum generation object including the speech spectrum data of the reference audio. The process of processing the reference object and the laughter feature data using the data processing model to obtain a spectrum generation object includes: The second data processing model is used to extract features from the audio spectrum data of the reference audio to obtain the speech spectrum data, wherein the data dimension of the speech spectrum data is lower than that of the audio spectrum data.

4. The method according to claim 1, characterized in that, The method further includes: Acquire sample reference objects and sample laughter feature data, wherein the sample reference objects include sample reference audio and sample reference text; Based on the sample reference object and the sample laughter feature data, the first preprocessing model is trained until the first preset training condition is met, and the second preprocessing model is obtained. The second preprocessing model processes the reference object and the sample laughter feature data to obtain a sample spectrum generation object. The sample spectrum generation object includes sample phoneme augmentation data corresponding to the sample reference text, sample speech spectrum data of the sample reference audio, and sample emotion data of the sample laughter feature data. The sample phoneme augmentation data includes the sample assembly phoneme data of the sample assembly text after being augmented with phoneme data frames. The sample assembly text is the text concatenated from the sample reference text and the sample audio text of the sample reference audio. Using a sample spectrum decoder, the first sample laughter speech spectrum data is constructed based on the sample spectrum generation object. The difference between the first sample laughter speech spectrum data and the second sample laughter speech spectrum data is determined by the sample spectrum discriminator. The second sample laughter speech spectrum data is the reference laughter speech spectrum data corresponding to the sample reference object and the sample laughter feature data. If the difference value is less than or equal to the first threshold, the second preprocessing model is determined as the data processing model, and the sample spectrum decoder is determined as the speech spectrum generation model.

5. The method according to claim 4, characterized in that, The first preprocessing model includes a first sample preprocessing model for processing the sample reference object. The first sample preprocessing model includes a text encoder and a phoneme duration predictor. The second preprocessing model includes the first data processing model in the data processing model for processing the reference object. The sample reference text is the text of the sample reference audio after speech recognition processing. The sample reference audio and the sample reference text are time-aligned at the phoneme level. The step of training the first preprocessing model based on the sample reference object and the sample laughter feature data until the first preset training condition is met, to obtain the second preprocessing model, includes: The text encoder performs phoneme conversion processing on the sample reference text to obtain sample phoneme data, which includes sample phoneme sequences corresponding to text sentences in the sample reference text. The phoneme duration predictor calculates the second sample phoneme duration for each sample phoneme based on the first sample phoneme duration in the sample phoneme sequence. The second sample phoneme duration is used to represent the expected duration of the sample phoneme in the sample reference audio. If the difference between the duration of the second sample phoneme and the duration of the third sample phoneme is less than or equal to the second threshold, the text encoder is determined as the text encoding model in the first data processing model, and the phoneme duration predictor is determined as the phoneme duration prediction model in the first data processing model, wherein the duration of the third sample phoneme is the actual duration of the sample phoneme in the sample reference audio.

6. The method according to claim 4, characterized in that, The first preprocessing model includes a laughter emotion recognition model for processing the laughter feature data of the samples, and the second preprocessing model includes a third data processing model in the data processing model for processing the laughter feature data; The step of training the first preprocessing model based on the sample reference object and the sample laughter feature data until the first preset training condition is met, to obtain the second preprocessing model, includes: Extract sample laughter feature data from the sample audio; The laughter emotion recognition model is used to identify the sample laughter feature data to obtain the first sample emotion data of the sample laughter feature data. If the difference between the first sample emotion data and the second sample emotion data corresponding to the sample laughter feature data is less than or equal to the fourth threshold, the laughter emotion recognition model is determined as the third data processing model.

7. The method according to claim 4, characterized in that, The step of constructing the first sample laughter speech spectrum data by using a sample spectrum decoder and generating an object based on the sample spectrum includes: The sample audio spectrum data of the sample reference audio is extracted by the sample spectrum decoder. The sample audio spectrum data is masked to obtain sample mask spectrum data, which includes sample mask spectrum data of at least one sample mask region. According to the second sample phoneme duration of each sample phoneme in the sample phoneme augmentation data, each sample phoneme in the sample phoneme augmentation data is aligned with the sample spectral features in the sample speech spectral data of the sample reference audio to obtain sample aligned audio spectral data. The sample aligned audio spectral data includes the sample speech spectral data and the sample mask spectral data of the sample mask region. The audio spectrum data of the sample mask region is reconstructed using the sample emotion data and the sample phoneme augmentation data to obtain the sample reconstructed audio spectrum data of the sample mask region. Replace the sample mask spectrum data in the sample aligned audio spectrum data with the sample reconstructed audio spectrum data to obtain the sample augmented mask spectrum data; Based on the sample emotion data, determine the sample position of the sample laughter feature data in the sample augmented mask spectrum data; According to the sample location, the sample laughter feature data and the sample expanded mask spectrum data are fused to obtain the first sample laughter speech spectrum data.

8. The method according to claim 1, characterized in that, Before acquiring the reference object and the laughter feature data associated with the reference object, the method further includes: The laughter synthesis interface is displayed, and the laughter synthesis interface includes a first area and a second area; Receive the user's first input to the first area; In response to the first input, the text entered by the user in the first area is determined as the reference text; Receive the user's second input to the second area; In response to the second input, the preset audio selected by the user in the second area is determined as the reference audio; or, the user audio added by the user in the second area is determined as the reference audio.

9. An audio generation device, characterized in that, include: The acquisition module is used to acquire a reference object and laughter feature data related to the reference object. The reference object includes reference text and reference audio. The reference text is text used for laughter synthesis, and the reference audio is audio used to indicate the generation of a preset style. The determination module is used to process the reference object and the laughter feature data through the data processing model in the spectrum generation model to obtain a spectrum generation object; wherein, the spectrum generation object includes phoneme augmentation data corresponding to the reference text, speech spectrum data of the reference audio, and emotional data of the laughter feature data, the phoneme augmentation data includes the data of the assembled phoneme data of the assembled text after augmentation phoneme data frames, and the assembled text is the text concatenated from the reference text and the audio text of the reference audio; Using the speech spectrum generation model in the spectrum generation model, each phoneme in the phoneme augmentation data is aligned with the spectral features in the speech spectrum data of the reference audio according to the second phoneme duration of each phoneme in the phoneme augmentation data, to obtain aligned audio spectrum data. The aligned audio spectrum data includes the speech spectrum data and the mask spectrum data of the mask region. The second phoneme duration is used to represent the expected duration of the phoneme in the reference audio. The audio spectrum data of the mask region is reconstructed using the emotion data and the phoneme augmentation data to obtain reconstructed audio spectrum data of the mask region. The mask spectrum data in the aligned audio spectrum data is replaced with the reconstructed audio spectrum data to obtain augmented mask spectrum data. Based on the emotion data, the position of the laughter feature data in the augmented mask spectrum data is determined. According to the position, the laughter feature data and the augmented mask spectrum data are fused to obtain laughter speech spectrum data. A conversion module is used to convert the laughter speech spectrum data into laughter synthesized audio, wherein the laughter synthesized audio is audio of the preset style, and the laughter synthesized audio includes the text content of the reference text and laughter related to the emotion represented by the text content.

10. The apparatus according to claim 9, characterized in that, The determining module is specifically used to, when the spectrum generation model includes a data processing model and a speech spectrum generation model, process the reference object and the laughter feature data through the data processing model to obtain a spectrum generation object; wherein, the spectrum generation object includes phoneme augmentation data corresponding to the reference text, speech spectrum data of the reference audio, and emotional data of the laughter feature data, the phoneme augmentation data includes the data after the assembled phoneme data of the assembled text has been augmented with phoneme data frames, and the assembled text is the text concatenated from the reference text and the audio text of the reference audio; The laughter speech spectrum data is determined using the speech spectrum generation model and the spectrum generation object.

11. An electronic device, characterized in that, include: A processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the audio generation method as described in any one of claims 1-8.

12. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the audio generation method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Speech synthesis method and device, medium and electronic equipment

    CN114255738A

  • Voice processing method, device and equipment and computer readable storage medium

    CN114582315A