Voice generation method and device, storage medium and electronic equipment
By identifying the voice output and sound effect features of the target scene in the metaverse and adjusting the voice output method, the problem of consistent digital human voice in different scenes is solved, achieving better scene adaptation and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MIGU CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-17
AI Technical Summary
When synthesizing speech in the metaverse using existing technologies, the digital human's voice remains consistent across different scenarios, resulting in poor integration with the person and the scene, which negatively impacts the user experience.
By acquiring image information of the target scene, recognizing voice output features and sound effect features, adjusting the voice output method to adapt to different scenes, and generating suitable voice output.
It improves the adaptability of digital human voices to different scenarios, thus enhancing the user experience.
Smart Images

Figure CN121884769A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing technology, and in particular to a speech generation method, apparatus, storage medium and electronic device. Background Technology
[0002] Text-to-Speech (TTS) is an artificial intelligence technology that automatically converts text into natural and fluent human speech. It is a key component of human-computer interaction (HCI) and is widely used in scenarios such as intelligent assistants, audio reading, navigation broadcasts, and accessibility services.
[0003] Currently, speech synthesis mainly adopts an encoder-decoder-vocoder structure, training the speaker's voice from a speech library. The data in the speech library is usually recorded in a high-quality recording studio, and the consistency and stability of the sound are guaranteed as much as possible, so as to synthesize speech.
[0004] However, since digital humans can exist in various different scenarios in the metaverse, when the synthesized speech obtained using this speech synthesis method is applied to the digital human in the metaverse, the digital human's voice will be the same in different scenarios. This results in poor integration between the digital human's voice and the person and the scene, affecting the user experience. Summary of the Invention
[0005] In view of this, this application provides a speech generation method, apparatus, storage medium and electronic device, the main purpose of which is to improve the technical problem that the synthesized speech of the existing technology is applied to the metaverse digital human, in which the digital human voice is the same in different scenarios, resulting in poor integration between the digital human's voice and the person and the scene, which affects the user experience.
[0006] Firstly, this application provides a speech generation method, including: Obtain the text information to be output corresponding to the target character element in the target scene; Determine the speech output features corresponding to the target scene based on the scene image corresponding to the target scene; Based on the voice output features, determine the voice output method corresponding to the target character element; The speech to be output is generated according to the speech output method, and the speech to be output is used by the target character element to output speech in the target scene.
[0007] Secondly, this application provides a speech generation apparatus, comprising: The acquisition module is configured to acquire the text information to be output corresponding to the target character element in the target scene. The determination module is configured to determine the speech output features corresponding to the target scene based on the scene image corresponding to the target scene; The determination module is also configured to determine the voice output mode corresponding to the target person element based on the voice output features; The generation module is configured to generate the speech to be output corresponding to the text information to be output according to the speech output method. The speech to be output is used by the target character element to output speech in the target scene.
[0008] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the speech generation method described in the first aspect.
[0009] Fourthly, this application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement the speech generation method described in the first aspect.
[0010] Fifthly, this application provides a computer program product, which includes a computer program that, when executed by a processor, implements the speech generation method described in the first aspect.
[0011] By employing the above technical solutions, this application provides a speech generation method, apparatus, storage medium, and electronic device, comprising: acquiring text information to be output corresponding to a target person element in a target scene; determining speech output features corresponding to the target scene based on a scene image corresponding to the target scene; determining the speech output mode corresponding to the target person element based on the speech output features; and generating speech to be output corresponding to the text information to be output according to the speech output mode, wherein the speech to be output is used for speech output by the target person element in the target scene. Compared with the prior art, this application can determine the speech output features of the target person element in the target scene through the scene image corresponding to the target person element, and thus determine the speech output mode of the target person element in the target scene, thereby generating the speech to be output for the target person element in the target scene. This allows the application to output different speech to be output when the target person element is in different scenes, thereby adapting the voice of the target person element to the target scene, improving the adaptability of the voice of the target person element to the target scene, and enhancing the user experience. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0013] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 A flowchart illustrating a speech generation method provided in an embodiment of this application is shown; Figure 2 A flowchart illustrating a speech generation method provided in an embodiment of this application is shown; Figure 3 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 4 This illustration shows a structural diagram of an example provided in an embodiment of this application; Figure 5 This illustration shows a structural diagram of an example provided in an embodiment of this application; Figure 6 This illustration shows a structural diagram of an example provided in an embodiment of this application; Figure 7 This illustration shows a structural diagram of an example provided in an embodiment of this application; Figure 8 A flowchart illustrating an example provided in an embodiment of this application is shown; Figure 9 This paper shows a schematic diagram of the structure of a speech generation device provided in an embodiment of this application; Figure 10 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0015] The embodiments of this application will now be described in more detail with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0016] To address the technical issue of existing synthesized speech technology exhibiting identical voices across different scenarios when applied to metaverse digital humans, resulting in poor integration of the digital human's voice with the character and the surrounding environment and negatively impacting user experience, this embodiment provides a speech generation method, such as... Figure 1 As shown, the method includes: Step 101: Obtain the text information to be output corresponding to the target character element in the target scene.
[0017] In the embodiments of this application, the target scene can be the virtual space or environment in which the target character element is located in the metaverse or virtual environment; specifically, the metaverse is not a single product or technology, but a persistent, immersive, user-co-created 3D digital world that integrates virtual reality, artificial intelligence, blockchain, digital identity and economic systems. In the metaverse, the character element can work, socialize, entertain, create and even own assets, etc., to provide users with a three-dimensional Internet experience.
[0018] In some examples, the target character element can be a character element in the metaverse. For example, the target character element can be, but is not limited to, digital humans, cartoon characters, animals, etc. Correspondingly, the text to be output can be the text content that the target character element needs to speak, that is, the text that needs to be synthesized into speech in this embodiment of the application.
[0019] As an alternative approach, if digital human 1 is in a library, the target scene in this application embodiment can be a library scene, and digital human 1 can be the target human element in this application embodiment. Therefore, the text content 1 that digital human 1 needs to say in the library scene can be determined, that is, the text content 1 is the text information to be output in this application embodiment.
[0020] As an alternative approach, if the digital human 2 is in a classroom, the target scenario in this application embodiment can be a classroom teaching scenario, and the digital human 2 can be the target human element in this application embodiment. Therefore, the text content 2 that the digital human 2 needs to speak in the classroom teaching scenario can be determined, that is, the text content 2 is the text information to be output in this application embodiment.
[0021] Step 102: Determine the speech output features corresponding to the target scene based on the scene image corresponding to the target scene.
[0022] In this embodiment, the scene image of the target scene can specifically be an environmental image of the current scene captured by a virtual camera or sensor; the scene image of the target scene in this embodiment can be used to characterize the acoustic environmental features of the target scene. In some examples, the speech output features can be the speech output features of the target character element in the target scene. The speech output features can be feature vectors extracted from the scene image of the target scene to control speech synthesis. For example, they can include vocalization features and sound effect features.
[0023] As an alternative approach, if digital human 1 is in a library, the target scene in this application embodiment can be a library scene, and digital human 1 can be the target human element in this application embodiment. Then, the text content 1 that digital human 1 needs to speak in the library scene can be determined, that is, the text content 1 is the text information to be output in this application embodiment. Then, the voice output feature 1 corresponding to the library scene can be determined based on the scene image corresponding to the library scene.
[0024] As an alternative approach, if the digital human 2 is in a classroom, the target scene in this application embodiment can be a classroom teaching scene, and the digital human 2 can be the target human element in this application embodiment. In turn, the text content 2 that the digital human 2 needs to speak in the classroom teaching scene can be determined, that is, the text content 2 is the text information to be output in this application embodiment. Then, the voice output feature 2 corresponding to the classroom teaching scene can be determined based on the scene image corresponding to the classroom teaching scene.
[0025] It should be noted that different target scenarios in this application embodiment may correspond to different speech output features or the same speech features, and no specific limitation is made in this application embodiment.
[0026] Step 103: Determine the voice output mode corresponding to the target character element based on the voice output features.
[0027] In the embodiments of this application, the voice output method can be the voice synthesis parameters and sound effect processing method determined according to the voice output characteristics. For example, in a noisy environment, a higher volume, clearer pronunciation, and slight reverberation can be used.
[0028] As an alternative approach, if digital human 1 is in a library, the target scene in this application embodiment can be a library scene, and digital human 1 can be the target human element in this application embodiment. Then, the text content 1 that digital human 1 needs to speak in the library scene can be determined, that is, the text content 1 is the text information to be output in this application embodiment. Then, the voice output feature 1 corresponding to the library scene can be determined based on the scene image corresponding to the library scene, and the voice output mode 1 of digital human 1 in the library scene can be determined based on the voice output feature 1.
[0029] As an alternative approach, if the digital human 2 is in a classroom, the target scenario in this application embodiment can be a classroom teaching scenario, and the digital human 2 can be the target human element in this application embodiment. Therefore, the text content 2 that the digital human 2 needs to speak in the classroom teaching scenario can be determined, that is, the text content 2 is the text information to be output in this application embodiment. Then, the voice output feature 2 corresponding to the classroom teaching scenario can be determined based on the scene image corresponding to the classroom teaching scenario, and the voice output mode 2 of the digital human 2 in the classroom teaching scenario can be determined based on the voice output feature 2.
[0030] Step 104: Generate the speech to be output corresponding to the text information to be output according to the speech output method.
[0031] The voice to be output is used for the target character element to output voice in the target scene.
[0032] In this embodiment of the application, the speech to be output can be the final synthesized and output speech signal, which is used to play in the target scene through the target task element.
[0033] As an alternative approach, if digital human 1 is in a library, the target scene in this embodiment can be determined to be a library scene, and digital human 1 can be the target human element in this embodiment. Then, the text content 1 that digital human 1 needs to speak in the library scene can be determined, that is, the text content 1 is the text information to be output in this embodiment. Then, the voice output feature 1 corresponding to the library scene can be determined based on the scene image corresponding to the library scene, and the voice output mode 1 of digital human 1 in the library scene can be determined based on the voice output feature 1. Then, the voice to be output 1 of digital human 1 in the library scene can be synthesized through the voice output mode 1, so that digital human 1 can output voice according to the voice to be output 1 in the library scene.
[0034] As an alternative approach, if the digital human 2 is in a classroom, the target scenario in this application embodiment can be a classroom teaching scenario, and the digital human 2 can be the target human element in this application embodiment. Therefore, the text content 2 that the digital human 2 needs to speak in the classroom teaching scenario can be determined, i.e., the text content 2 is the text information to be output in this application embodiment. Then, based on the scene image corresponding to the classroom teaching scenario, the speech output feature 2 corresponding to the classroom teaching scenario can be determined, and based on the speech output feature 2, the speech output method 2 of the digital human 2 in the classroom teaching scenario can be determined. Then, the speech to be output 2 of the digital human 2 in the classroom teaching scenario can be synthesized through the speech output method 2, so that the digital human 2 can output speech according to the speech to be output 2 in the classroom teaching scenario.
[0035] Compared with existing technologies, this embodiment can determine the voice output features of the target character element in the target scene by using the scene image corresponding to the target character element. This allows the embodiment to determine the voice output mode of the target character element in the target scene, thereby generating the voice to be output for the target character element in the target scene. This enables the embodiment to output different voices when the target character element is in different scenes, thus adapting the voice of the target character element to the target scene, improving the adaptability of the voice of the target character element to the target scene, and enhancing the user experience.
[0036] As a refinement and extension of the above embodiments, this embodiment provides a speech generation method, such as... Figure 2 As shown, the method includes: Step 201: Obtain the text information to be output corresponding to the target character element in the target scene.
[0037] For example, such as Figure 3 As shown, the system acquires in real-time images of the environment (i.e., the target scene in this embodiment) where the digital human (i.e., the target human element in this embodiment) is located and the text to be spoken (i.e., the text information to be output in this embodiment), and then can synthesize the text Z to be synthesized by the digital human. text Select the current digital human speaker spk, and acquire an image Y of the environment in which the current digital human is located. image The text is a series of seemingly unrelated phrases and sentences, making it impossible to translate coherently. It appears to be a collection of fragments from various sources, possibly related to speech synthesis, text, and technology.
[0038] Step 202: Identify the scene sound information corresponding to the target scene based on the scene image corresponding to the target scene.
[0039] For example, based on the example in step 201, after the digital human enters the target scene, the sound in the target scene can also be identified based on the image. For example, when the digital human in the metaverse (i.e. the target human element in this application embodiment) enters the virtual market, the system will capture the scene image of the current environment through the sensors or cameras in the virtual scene. The image will be input into a pre-trained convolutional neural network model. The model will analyze the environmental image features of the current environment and generate acoustic features (i.e., scene sound information in this application embodiment) based on the environmental image features, and identify that this is a noisy environment.
[0040] Step 203: Generate target pronunciation features and target sound effect features corresponding to the target scene based on the scene sound information, and determine the target pronunciation features and target sound effect features as speech output features.
[0041] As an alternative, based on the example in step 202, embodiments of this application can also process noisy market scene images using convolutional neural networks (CNNs) to extract vocal features (i.e., target vocal features in embodiments of this application) and sound effect features (i.e., target sound effect features in embodiments of this application).
[0042] For example, in the case of a noisy environment, the present application embodiment can generate a vocal feature suitable for use in a noisy environment (i.e., the target vocal feature in the present application embodiment). This feature tells the TTS model that it needs to synthesize a higher volume and clearer sound, and may need to adjust the vocalization method to make the sound more penetrating.
[0043] For example, in the case of a large amount of background noise in the environment, such as echoes and reverberation, an audio effect feature suitable for processing sound effects in noisy environments (i.e., the target audio effect feature in the embodiments of this application) can be generated so that the TTS model can add some environmental noise suppression or audio effect enhancement processing during synthesis.
[0044] Step 204: Determine the voice output mode corresponding to the target character element based on the voice output features.
[0045] Optionally, when performing the "determining the voice output mode corresponding to the target character element based on voice output features", the following methods may be used, but are not limited to: determining the target pronunciation mode corresponding to the target character element based on target pronunciation features; determining the target sound effect corresponding to the target character element based on target sound effect features; and determining the voice output mode based on the target pronunciation mode and the target sound effect.
[0046] As an alternative approach, based on the example in step 203, this embodiment of the application can also input these two environmental features (voice vector and sound effect vector) into the TTS model of the digital human. The encoder of the TTS model will adjust the human's voice production method (i.e., the target pronunciation method in this embodiment) according to the voice vector (i.e., the target pronunciation method in this embodiment). For example, due to a noisy environment, the TTS model will synthesize a higher volume, clearer, and slightly slower voice to ensure that the voice can be clearly heard in a noisy environment. At the same time, the frequency range of the voice will be adjusted so that the voice can penetrate the background noise.
[0047] As an optional approach, the decoder of the TTS model will adjust the sound effects in the speech (i.e., the target sound effects in this embodiment) according to the sound effect vector (i.e., the target sound effect in this embodiment). For example, due to the noisy environment, the TTS model in this embodiment can also perform some noise suppression processing on the sound to eliminate the interference of background noise and ensure clear speech output. In addition, the model may also add a small amount of environmental reverberation effect to the sound to make the sound blend into the scene more naturally.
[0048] Step 205: Generate the speech to be output corresponding to the text information to be output according to the speech output method.
[0049] The voice to be output is used for the target character element to output voice in the target scene.
[0050] Optionally, before performing "generating the speech to be output corresponding to the text information to be output according to the speech output method", the following methods may be used, but not limited to: performing semantic recognition on the text information to be output to obtain the text semantic information corresponding to the text information to be output; determining the basic sound information corresponding to the target person element; and generating the first speech spectrum information corresponding to the text information to be output based on the basic sound information.
[0051] As an alternative, based on the example in step 204, embodiments of this application can also use a vocoder to convert the processed speech spectrum into actual sound output. In this noisy environment, the digital human's voice becomes louder and clearer, and after system optimization, background noise will not affect your understanding of its speech.
[0052] For example, if a digital human from the metaverse enters a virtual marketplace, a place bustling with people and filled with various noises such as conversations, car horns, and background music, the digital human would need to synthesize a louder, clearer voice to match this environment. It might also need to adjust its pronunciation to make the voice more penetrating. Simultaneously, a subtle amount of environmental reverberation would be added to the sound to make it blend more naturally into the scene.
[0053] Optionally, when performing the "generating the speech to be output corresponding to the text information to be output according to the speech output method", the following methods can be used, but are not limited to: adjusting the first speech spectrum information according to the speech output method to obtain the target speech spectrum information corresponding to the text information to be output; and converting the target speech spectrum information into speech to obtain the speech to be output.
[0054] As an optional approach, embodiments of this application can use an acoustic environment feature extraction model to extract key scenes from acoustic environment images. The model can use a classic CNN model to extract features, such as... Figure 4 The diagram shows the structure of an image feature extraction model. Specifically, convolutional layers (conv) are used to extract information from the image. The kernel size is 3*3, and the convolutional channels are selected as 64, 128, 256, and 512 dimensions, respectively. A pooling layer is used to remove redundant information from the image and improve the robustness of the model. A 1024 fully connected layer is used to upscale the learned features to a higher dimension for feature integration. Finally, a softmax layer is used to output the category of the current image. Based on the above example, images can be classified into scenes that affect the digital human's vocalization and scenes that generate sound effects. The image feature extraction model described above is trained separately for these two scenes, and feature vectors are extracted. These two vectors are then named the vocalization vector V. s i (i.e., the target pronunciation features in the embodiments of this application) and sound effect vector V e i (i.e., the target sound effect features in the embodiments of this application).
[0055] As an alternative approach, during the training of the image feature extraction model, a large number of images of the metaverse scene, containing various scenes, can be prepared. The categories of sound impact from these scenes can be labeled to form a dataset represented by Formula 1. The scenes affecting the digital human's voice can include, but are not limited to, normal, quiet, noisy, warm, formal, and pleasant sounds. The scenes producing sound effects for the digital human's voice can include, but are not limited to, normal, echo, reverberation, and significant reverberation. Formula 1 is detailed below: (Formula 1) For example, embodiments of this application may further include X in the training set. i image The model is extracted using a CNN, and then the vocal vector V is obtained after passing it through a fully connected layer. s i (i.e., the target pronunciation features in the embodiments of this application) (1*1024 dimensions) or sound effect vector V e i (i.e., the target sound effect features in the embodiments of this application) ((1*1024 dimensions)), sound vector V s i (i.e., the target pronunciation features in the embodiments of this application) or sound effect vector V e i (i.e., the target sound effect features in this embodiment) are then processed through softmax and then LOSS is calculated until the model converges. Specifically, LOSS can be calculated using cross-entropy formula two, as shown below: (Formula 2) For example, in the embodiments of this application, the Adam optimizer can be used during training, with the learning rate set to 0.0001, the learning rate decay set to 0.00099, the model batch size set to 32, and the model training stopped after 400,000 steps, thereby obtaining two feature extraction models, namely the vocalization vector V. s i Feature extraction model and sound effect vector V e i Feature extraction model.
[0056] As an optional approach, such as Figure 5 As shown, this is a schematic diagram of a TTS model employing an encoding / decoding structure; specifically, the vocal vector V s i and sound effect vector V e i It can be added to the TTS model in a zero-shot manner for training. The main role of the encoder in TTS is to encode the speaker's timbre, prosodic information, and semantic information. Therefore, we add the vocal vectors from the environment to V. s i This is then used in the Encoder to control the speech of the TTS model.
[0057] Optionally, when performing the action of "adjusting the first speech spectrum information according to the speech output method to obtain the target speech spectrum information corresponding to the text information to be output", the following methods may be used, but are not limited to: adjusting the first speech spectrum information according to the target pronunciation method in the speech output method to obtain the second speech spectrum information corresponding to the text information to be output; adjusting the second speech spectrum information according to the target sound effect in the speech output method to obtain the target speech spectrum information corresponding to the text information to be output.
[0058] It should be noted that the main function of the decoder is to synthesize the Mel spectrum of speech. Therefore, the sound effect vector (Vei) can be added to the decoder, allowing the decoder to learn and control the Mel spectrum that can produce sound effects. The encoder mainly maps multimodal information into the speech modality space, such as... Figure 6 As shown, ENCODER is a sequence-to-sequence model. Specifically, the encoder can fuse the semantics of the text sequence, the timbre information of the selected speaker, and the information affecting the vocalization from the acoustic environment, and map them into the speech modal space. The Transformer Encoder then maps the fused information to speech state information. The txtEmd module encodes text information and maps it to a 256-dimensional vector text. infoThe spkEmd module encodes the speaker's timbre information and maps it to a 256-dimensional vector spk. info The linear layer MLP maps the 1024-dimensional information of the sound vector to the 256-dimensional information of the sound. info Therefore, Formula 3 can be used to calculate the fused information enc. input Formula 3 is shown below: (Formula 3) For example, the decoder decodes information containing speech modalities into a speech spectrum, such as... Figure 7 As shown, the DECODER is a sequence-to-sequence model, consisting of a 6-layer Transformer-Decoder model and a Postnet model. In the decoder, the Transformer Decoder takes a clean Mel spectrum and an audio vector as input. The Mel spectrum output by the decoder is processed and added to the PostNet model to obtain a Mel spectrum containing sound effects. Finally, the two spectra are added together to obtain the final output Mel spectrum containing sound effects. The vocoder is a model that synthesizes the Mel spectrum into speech. In this embodiment, the vocoder can be Hifigan.
[0059] As an optional approach, the TTS model training process can be divided into two parts: encoding / decoding and vocoder training. Specifically, acoustic environment images, approximately 2-10 seconds of speech data, and the text and corresponding speaker of the speech data can be prepared to form the training set shown in Formula 4, which is detailed below: (Formula 4) For example, embodiments of this application may further incorporate Y i image The origin sound vector V is extracted using an acoustic environment feature extraction model. s i and sound effect vector V e i Enter Z i text Spk i V s i V e i The speech spectrum X is obtained by feeding it into the encoder and decoder modules of the TTS model. i mel And thus from real audio X i audio Solving for the true speech spectrum X i melThe loss is calculated on both the true Mel spectrum and the predicted Mel spectrum. Training is then performed with the goal of minimizing the loss using Equation 5, which is shown below: (Formula 5) As an alternative approach, embodiments of this application can further train the vocoder through a self-supervised generalization process. The input to the vocoder is a Mel spectrum, which is synthesized into a predicted sound X. i audio ’ X i audio ’ Will be with Real Voice X i audio The model is trained using adversarial training (GA), with the model optimizer being Adam, the learning rate set to 0.0002, a learning rate decay method with a decay ratio of 0.9995, a batch size of 64, and a training step count of 800,000.
[0060] As an optional approach, such as Figure 8 As shown in the embodiments of this application, the following examples can also be provided: In the case of a digital human entering a virtual library in the metaverse, where the library is a quiet environment with minimal echo and reverberation, and the overall atmosphere needs to be kept quiet, the digital human needs to produce a softer voice according to this environment. In the case of a digital human entering a virtual market in the metaverse, where people are coming and going and the surroundings are filled with various noisy sounds, such as conversations, vehicle horns, and background music, the digital human needs to produce a higher volume and clearer voice according to this environment, and may need to adjust its pronunciation to make the voice more penetrating. Simultaneously, a slight environmental reverberation effect is added to the sound to make it blend more naturally into the scene.
[0061] It's important to note that text-to-speech (TTS) technology is widely used in various aspects of life, capable of synthesizing a near-human voice from given user text. In real life, the environment influences how a person speaks. First, people choose different vocalization methods based on the environment; for example, on a quiet night, people usually don't need to raise their volume to ensure their voice is heard, resulting in a softer, more moderate tone; on a busy street, people often need to raise their volume to cope with external noise, resulting in a louder, higher-pitched voice. Furthermore, the environment also alters the voice a person has already produced. For instance, in a spacious cathedral, a person's voice will reverberate due to the environment; in a tunnel, a person's voice will echo.
[0062] In related technologies, TTS models typically employ an encoder-decoder-vocoder structure, training the speaker's voice from a speech database. The data in the speech database is usually recorded in a high-quality recording studio, ensuring the consistency and stability of the voice actors' voices as much as possible. However, the application of this speech synthesis technology in the metaverse scenario has the following drawbacks: 1. In the metaverse digital human domain, digital humans exist in a virtual world constructed within the metaverse, encompassing various scenarios. Existing solutions do not consider the impact of the acoustic environment on the digital human's vocalization, resulting in a somewhat rigid and lifeless digital human image; 2. Existing solutions do not consider the potential for sound effects such as echo and reverberation caused by the acoustic environment, failing to provide players with an immersive experience and sacrificing some realism; 3. Existing solutions train the speaker's voice from speech data, ensuring the TTS-synthesized voice maintains consistency with the acoustic environment during recording. To create speakers for different scenarios, data needs to be re-recorded and retrained, a cumbersome process with high costs.
[0063] It should be noted that this application embodiment acquires acoustic environment images of the digital human in the metaverse scene and embeds the extracted features into the TTS model, achieving the effect that the TTS synthesized voice can be adjusted to adapt to different acoustic environments. Furthermore, this application embodiment can extract features of the scene image of the digital human's environment in real time. These features are divided into two parts: those affecting vocalization and those affecting sound effects. Incorporating these features into the TTS process allows the TTS model to synthesize voices that adapt to different acoustic environments, resulting in a more realistic interactive experience. Compared to existing TTS voice synthesis methods, this application embodiment eliminates the need for re-acquiring data and retraining the TTS model when creating voice actors for different scenarios, improving efficiency and saving costs. In summary, this application embodiment can adaptively control the digital human's vocalization method according to the acoustic environment, and the sound will generate sound effects based on the environment. It is applicable to mainstream TTS systems, enhances the realism of the metaverse digital human's voice interaction, and can bring greater user stickiness.
[0064] Compared with existing technologies, this embodiment can determine the voice output features of the target character element in the target scene by using the scene image corresponding to the target character element. This allows the embodiment to determine the voice output mode of the target character element in the target scene, thereby generating the voice to be output for the target character element in the target scene. This enables the embodiment to output different voices when the target character element is in different scenes, thus adapting the voice of the target character element to the target scene, improving the adaptability of the voice of the target character element to the target scene, and enhancing the user experience.
[0065] Furthermore, as Figure 1 and Figure 2 To illustrate the specific implementation of the method shown, this embodiment provides a speech generation device, such as... Figure 9 As shown, the device includes: an acquisition module 31, a determination module 32, and a generation module 33.
[0066] The acquisition module 31 is configured to acquire the text information to be output corresponding to the target character element in the target scene; The determining module 32 is configured to determine the speech output features corresponding to the target scene based on the scene image corresponding to the target scene; The determining module 32 is also configured to determine the voice output mode corresponding to the target person element based on the voice output features; The generation module 33 is configured to generate the speech to be output corresponding to the text information to be output according to the speech output method. The speech to be output is used by the target character element to output speech in the target scene.
[0067] In some examples of this embodiment, the determining module 32 is specifically configured to identify scene sound information corresponding to the target scene based on the scene image corresponding to the target scene; generate target pronunciation features and target sound effect features corresponding to the target scene based on the scene sound information; and determine the target pronunciation features and the target sound effect features as the speech output features.
[0068] In some examples of this embodiment, the determining module 32 is further configured to determine the target pronunciation mode corresponding to the target character element based on the target pronunciation features; determine the target sound effect corresponding to the target character element based on the target sound effect features; and determine the voice output mode based on the target pronunciation mode and the target sound effect.
[0069] In some examples of this embodiment, the generation module 33 is also configured to perform semantic recognition on the text information to be output to obtain text semantic information corresponding to the text information to be output; determine the basic sound information corresponding to the target person element; and generate first speech spectrum information corresponding to the text information to be output based on the basic sound information.
[0070] In some examples of this embodiment, the generation module 33 is specifically configured to adjust the first speech spectrum information according to the speech output method to obtain the target speech spectrum information corresponding to the text information to be output; and to convert the target speech spectrum information into speech to obtain the speech to be output.
[0071] In some examples of this embodiment, the generation module 33 is further configured to adjust the first speech spectrum information according to the target pronunciation mode in the speech output mode to obtain the second speech spectrum information corresponding to the text information to be output; and to adjust the second speech spectrum information according to the target sound effect in the speech output mode to obtain the target speech spectrum information corresponding to the text information to be output.
[0072] It should be noted that other corresponding descriptions of the functional units involved in the speech generation device provided in this embodiment can be found in [reference needed]. Figure 1 and Figure 2 The corresponding descriptions in [the document] will not be repeated here.
[0073] Based on the above, Figure 1 and Figure 2 Accordingly, this embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. Figure 1 and Figure 2 The method shown.
[0074] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.
[0075] like Figure 10 The diagram shown is a hardware structure schematic of an electronic device according to the present invention, comprising: At least one processor 401; and, A memory 402 is communicatively connected to at least one of the processors 401; wherein, The memory 402 stores instructions that can be executed by at least one of the processors to enable at least one of the processors to perform the speech generation method as described above.
[0076] Figure 10 Take a processor 401 as an example.
[0077] The electronic device may also include an input device 403 and a display device 404.
[0078] The processor 401, memory 402, input device 403, and display device 404 can be connected via a bus or other means. Figure 10 Taking the example of a connection between China and Israel via a bus.
[0079] Memory 402, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the speech generation method in the embodiments of this application, for example, Figure 1 and Figure 2 The method flow is shown. The processor 401 executes various functional applications and data processing by running non-volatile software programs, instructions, and modules stored in the memory 402, thereby implementing the speech generation method in the above embodiments.
[0080] Memory 402 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the speech generation method, etc. Furthermore, memory 402 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 402 may optionally include memory remotely located relative to processor 401, and these remote memories may be connected to the apparatus performing the speech generation method via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0081] Input device 403 can receive user clicks and generate signal inputs related to user settings and function control of the voice generation method. Display device 404 may include display devices such as a display screen.
[0082] When one or more modules are stored in the memory 402, and are run by one or more processors 401, the speech generation method in any of the above method embodiments is executed.
[0083] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0084] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0085] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0086] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms, or it can be implemented by hardware. By applying the solution of this embodiment, compared with the prior art, this embodiment can determine the voice output characteristics of the target character element in the target scene through the scene image corresponding to the target character element, and then determine the voice output mode of the target character element in the target scene, thereby generating the voice to be output of the target character element in the target scene. This allows this embodiment to output different voices to be output when the target character element is in different scenes, thereby adapting the voice of the target character element to the target scene, improving the adaptability of the voice of the target character element to the target scene, and enhancing the user experience.
[0087] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0088] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A voice generation method characterized by, include: Obtain the text information to be output corresponding to the target character element in the target scene; Determine the speech output features corresponding to the target scene based on the scene image corresponding to the target scene; Based on the voice output features, determine the voice output method corresponding to the target character element; The speech to be output is generated according to the speech output method, and the speech to be output is used by the target character element to output speech in the target scene.
2. The method according to claim 1, characterized in that, Determining the speech output features corresponding to the target scene based on the scene image corresponding to the target scene includes: Based on the scene image corresponding to the target scene, the scene sound information corresponding to the target scene is identified; Based on the scene sound information, target pronunciation features and target sound effect features corresponding to the target scene are generated, and the target pronunciation features and target sound effect features are determined as the speech output features.
3. The method according to claim 2, characterized in that, The step of determining the voice output method corresponding to the target person element based on the voice output features includes: Based on the target pronunciation features, determine the target pronunciation mode corresponding to the target character element; Based on the target sound effect features, determine the target sound effect corresponding to the target character element; The speech output method is determined based on the target pronunciation method and the target sound effect.
4. The method according to claim 1, characterized in that, Before generating the speech corresponding to the text information to be output according to the speech output method, the method further includes: Perform semantic recognition on the text information to be output to obtain the text semantic information corresponding to the text information to be output. Determine the basic sound information corresponding to the target character element; The first speech spectrum information corresponding to the text information to be output is generated based on the basic sound information.
5. The method according to claim 4, characterized in that, The step of generating the speech corresponding to the text information to be output according to the speech output method includes: The first speech spectrum information is adjusted according to the speech output method to obtain the target speech spectrum information corresponding to the text information to be output. The target speech spectrum information is converted into speech to obtain the speech to be output.
6. The method according to claim 5, characterized in that, The step of adjusting the first speech spectrum information according to the speech output method to obtain the target speech spectrum information corresponding to the text information to be output includes: The first speech spectrum information is adjusted according to the target pronunciation method in the speech output method to obtain the second speech spectrum information corresponding to the text information to be output; The second speech spectrum information is adjusted according to the target sound effect in the speech output method to obtain the target speech spectrum information corresponding to the text information to be output.
7. A speech generation device, characterized in that, include: The acquisition module is configured to acquire the text information to be output corresponding to the target character element in the target scene. The determination module is configured to determine the speech output features corresponding to the target scene based on the scene image corresponding to the target scene; The determination module is also configured to determine the voice output mode corresponding to the target person element based on the voice output features; The generation module is configured to generate the speech to be output corresponding to the text information to be output according to the speech output method. The speech to be output is used by the target character element to output speech in the target scene.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.
9. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.
10. A computer program product, the computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.