Voice generation method and device based on image indication, equipment and medium

Through the speech generation method of image indication, image encoding and echo classifiers are used to generate echo speech matching the scene, solving the problem that the speech generation model in the prior art cannot be integrated into the acoustic environment, and real immersive and real voice experience is achieved.

CN120279882APending Publication Date: 2025-07-08PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510441773.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing speech generation model cannot effectively integrate into the acoustic environment of the specified scenarios, resulting in the inability to provide immersive and authentic voice experiences, especially in applications in the medical and financial fields.

Method used

By obtaining the prompt image and performing image encoding processing, acoustic embedding features are extracted, and inputting them into the pre-trained speech generation model as environmental echo conditions. Combined with the echo classifier, an echo speech classifier is used to identify the category of the target echo speech, and echo speech matching the scene is generated.

Benefits of technology

It realizes the embodiment of the scene acoustic environment in the speech generation process, so that the generated speech not only matches the text, but also matches the visual scenes in the image, providing an immersive and real voice experience, adapting to the reverberation effect of different scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279882A_ABST
    Figure CN120279882A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business system platforms of financial science and technology, medical treatment and health and the like, and discloses a voice generation method, device and equipment based on image indication and a medium, and the method comprises the steps: obtaining a prompt image and a target text of a to-be-generated voice; carrying out image coding processing on the prompt image to obtain an acoustic embedding feature matched with the environment in the prompt image; inputting the target text and the acoustic embedded features into a pre-trained voice generation model, and performing voice generation processing of environment fusion on the target text by taking the acoustic embedded features as environment reverberation conditions to generate corresponding target reverberation voice; and performing echo recognition on the target echo voice through a pre-trained echo classifier, and determining the echo category of the target echo voice. The scene echo is embedded into the speech synthesis process through image prompt, so that the generated speech is matched with the text and the scene in the image, the reverberation effect is adaptively adjusted, and the speech immersion and reality are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a method, apparatus, device and medium for generating speech based on image indication. Background Art

[0002] Text-To-Speech (TTS) technology refers to the process of generating speech from text. Currently, the work in the field of speech generation mainly focuses on generating natural speech, such as natural intonation, natural speaking rhythm, etc., and uses some specific cues to control the emotion and prosody of the synthesized speech. However, with the continuous improvement of the demand for speech quality, in various fields such as the medical field and the financial field, an immersive experience needs to be provided, and higher and higher requirements are put forward for the spatial perception characteristics of the synthesized speech. For example, in the field of medical health, when a medical robot conveys a treatment plan to a patient through TTS technology, adding the echo in the ward can simulate the spatial environment of the ward, making the patient feel more immersive. Or in the field of fintech business, when a financial company's wealth management consulting room uses a TTS system to introduce financial products to customers, integrating the echo effect of the consulting room can create a more professional and focused communication atmosphere.

[0003] However, the current speech generation models are all designed for neutral or echo-free environments, which does not match the demand for providing an immersive experience. Therefore, how to reflect the acoustic environment of a given scene in the process of speech synthesis is still an urgent problem to be solved. Summary of the Invention

[0004] In view of the above deficiencies of the prior art, the purpose of the present invention is to provide a method, apparatus, device and medium for generating speech based on image indication that can be applied to the medical field, fintech or other related fields. Its main purpose is to integrate the acoustic characteristics of a specified scene into the generated speech, so that the generated speech matches the corresponding scene and provides a more immersive and realistic speech generation effect.

[0005] The technical solution of the present invention is as follows:

[0006] The first aspect of the present invention provides a method for generating speech based on image indication, including:

[0007] Obtain a prompt image and the target text of the speech to be generated;

[0008] Perform image encoding processing on the prompt image to obtain an acoustic embedding feature matching the environment in the prompt image;

[0009] Input the target text and the acoustic embedding features into a pre-trained speech generation model, and perform environment-fused speech generation processing on the target text with the acoustic embedding features as the environmental reverberation conditions to generate corresponding target reverberant speech;

[0010] Perform reverberation recognition on the target reverberant speech through a pre-trained reverberation classifier to confirm the reverberation category of the target reverberant speech.

[0011] The second aspect of the present invention provides a speech generation device based on image indication, including:

[0012] An acquisition module for acquiring a prompt image and a target text of the speech to be generated;

[0013] An image processing module for performing image encoding processing on the prompt image to obtain acoustic embedding features matching the environment in the prompt image;

[0014] A speech generation module for inputting the target text and the acoustic embedding features into a pre-trained speech generation model, and performing environment-fused speech generation processing on the target text with the acoustic embedding features as the environmental reverberation conditions to generate corresponding target reverberant speech;

[0015] A reverberation classification module for performing reverberation recognition on the target reverberant speech through a pre-trained reverberation classifier to confirm the reverberation category of the target reverberant speech.

[0016] The third aspect of the present invention provides a computer device, including at least one processor; and,

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned speech generation method based on image indication.

[0019] The fourth aspect of the present invention provides a non-volatile computer-readable storage medium, and the non-volatile computer-readable storage medium stores computer-executable instructions, which when executed by one or more processors, can enable the one or more processors to execute the above-mentioned speech generation method based on image indication.

[0020] Beneficial effects: The present invention discloses a method, apparatus, device, and medium for image-guided speech generation. Compared with the prior art, in the embodiments of the present invention, a prompt image and a target text for which speech is to be generated are obtained; the prompt image is subjected to image encoding processing to obtain an acoustic embedding feature that matches the environment in the prompt image; the target text and the acoustic embedding feature are input into a pre-trained speech generation model, and the target text is subjected to environment fusion speech generation processing with the acoustic embedding feature as the environmental reverberation condition to generate a corresponding target reverberant speech; the target reverberant speech is subjected to reverberation recognition through a pre-trained reverberation classifier to confirm the reverberation category of the target reverberant speech. By using the prompt of the image modality to guide speech generation and embedding the scene reverberation perception into the speech synthesis process, the acoustic environment of the corresponding scene can be reflected in the speech synthesis process, so that the generated speech can not only match the text, but also match the visual scene reflected in the image, and the reverberation effect can be adaptively adjusted according to different scenes, providing a more immersive and realistic speech generation effect. Description of the Drawings

[0021] In order to more clearly illustrate the solutions in the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a schematic diagram of an application environment for the method for image-guided speech generation provided by the embodiments of the present invention;

[0023] Figure 2 It is a flowchart of the method for image-guided speech generation provided by the embodiments of the present invention;

[0024] Figure 3 It is a flowchart of step S202 in the method for image-guided speech generation provided by the embodiments of the present invention;

[0025] Figure 4 It is a flowchart of step S203 in the method for image-guided speech generation provided by the embodiments of the present invention;

[0026] Figure 5 It is another flowchart of the method for image-guided speech generation provided by the embodiments of the present invention;

[0027] Figure 6 It is a schematic diagram of the functional modules of the device for image-guided speech generation provided by the embodiments of the present invention;

[0028] Figure 7Schematic diagram of the hardware structure of the computer device provided by the embodiment of the present invention. Detailed implementation manners

[0029] To make the objectives, technical solutions and effects of the present invention clearer and more definite, the present invention will be further described in detail below. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. The embodiments of the present invention will be introduced below with reference to the accompanying drawings.

[0030] The voice generation method based on image indication provided by the embodiment of the present invention can be applied to an application environment such as Figure 1 which includes a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0031] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only for example).

[0032] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0033] The server 105 may be a server providing various services, such as a background server that provides support for the content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (only for example). The background server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device. The server 105 may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services (″Virtual Private Server″, or simply referred to as ″VPS″). The server 105 may also be a server of a distributed system, or a server combined with a blockchain.

[0034] It should be noted that the method for generating speech based on image indication provided by the embodiments of the present application can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the device for generating speech based on image indication provided by the embodiments of the present invention can also be disposed in the first terminal device 101, the second terminal device 102, or the third terminal device 103. Alternatively, the method for generating speech based on image indication provided by the embodiments of the present invention can generally also be executed by the server 105. Correspondingly, the device for generating speech based on image indication provided by the embodiments of the present invention can generally be disposed in the server 105.

[0035] It should be understood that the numbers of the above terminal devices, networks, and servers are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0036] As Figure 2 shown, the method for generating speech based on image indication provided by the embodiments of the present invention specifically includes the following steps:

[0037] S201, obtain a prompt image and a target text of the speech to be generated.

[0038] In this embodiment, the prompt image is an image used to provide visual scene information required for speech synthesis, which can be manually input by the user, or can also be fixedly set by the system in a fixed application scenario. For example, if the prompt image is a forest scene in virtual reality, then elements such as trees, terrain, and vegetation in the image will affect the propagation and reverberation characteristics of sound, thereby providing reference information for subsequent speech synthesis. The target text is the text content that needs to be converted into speech, which can be text in any language, such as Chinese, English, etc. Specifically, the text of the speech to be generated can be obtained through various methods such as user input, file reading, and web crawling. For example, in a medical scenario, the target text may be "Please wait for the examination at the nurse's station", while in a financial scenario, it may be "Welcome to our bank. Please go to the counter to handle business."

[0039] In specific implementation, a user can upload a prompt image through the interface of the speech generation system. The image can be a high-resolution color photo or a rendering of a virtual scene, etc. For example, in a medical scenario, the prompt image can be an image of the interior of a hospital ward, showing the layout of the ward, the wall materials, medical equipment, etc.; in a financial scenario, the prompt image can be an image of a bank lobby, showing the size of the lobby, the decoration materials, the counter positions, etc. And the user also inputs a target text. The speech synthesis system receives and stores the text content and preliminarily verifies the image and the text, such as whether the image format is supported and whether the text conforms to the language specification, etc., to ensure that it meets the input requirements to guarantee the accuracy of speech generation. For example, in a medical scenario, the target text can be "Please wait for the examination at the nurse's station", and in a financial scenario, the target text can be "Welcome to our bank. Please go to the counter to handle business", etc. By obtaining the prompt image and the target text as the basic inputs for speech synthesis, the user can upload different images and input corresponding texts according to different scenario requirements, so as to flexibly adapt to the speech generation effects of various application scenarios.

[0040] S202. Perform image encoding processing on the prompt image to obtain an acoustic embedding feature matching the environment in the prompt image.

[0041] In this embodiment, image encoding processing is performed on the obtained prompt image to extract environment-related features from the image, such as the size, material, layout of the room, etc. These environmental features will all affect the propagation and reverberation characteristics of sound. Then, the features obtained from the prompt image are converted into an acoustic embedding feature matching the environment. This acoustic embedding feature is used to represent the acoustic characteristics of the environment corresponding to the image, such as the reverberation type of the room, etc., so as to provide a reverberation hint matching the target environment for the subsequent speech generation process, thereby enhancing the immersion and realism of the generated speech.

[0042] Exemplarily, in a medical scenario, if the prompt image is an image of the interior of an operating room, showing the layout of the operating room, the wall materials, equipment, etc., by performing image encoding processing on the prompt image and converting and extracting the acoustic features of the operating room, such as the weak reverberation caused by the metal walls and equipment, etc., the generated speech can carry the acoustic features of the operating room, making medical staff feel more realistic when listening to the voice instructions and reducing the possibility of misoperation. Or in a financial scenario, if the prompt image is an image of a bank VIP room, showing the layout of the VIP room, the decoration materials, furniture, etc., by performing image encoding processing on the prompt image and converting and extracting the acoustic features of the VIP room, such as the softer reverberation caused by the carpet and soft materials, the generated speech can carry the acoustic features of the VIP room, making customers feel more comfortable when listening to the voice prompts and being beneficial to improving customer satisfaction.

[0043] S203. Input the target text and the acoustic embedding features into a pre-trained speech generation model, and perform speech generation processing for environmental fusion of the target text with the acoustic embedding features as the environmental reverberation condition to generate corresponding target reverberant speech.

[0044] In this embodiment, the target text and the extracted acoustic embedding features are input into a pre-trained speech generation model. Specifically, existing TTS (Text-to-Speech) models such as the VITS or Flow model can be used. These models can convert text into speech and can adjust the style and characteristics of the speech through conditional input. Through the speech generation model, speech generation processing is performed by fusing the acoustic embedding features as the environmental reverberation condition with the features corresponding to the target text. When converting the target text into a speech signal, the environmental reverberation represented by the acoustic embedding features can be synchronously fused, so as to generate speech features with corresponding environmental reverberation characteristics and further generate target reverberant speech, making the generated target reverberant speech not only match the text content, but also have a reverberation effect matching the image environment, making the speech synthesis more natural and realistic, and being able to adaptively adjust the reverberation effect according to different scene prompts, thus providing a more immersive auditory experience.

[0045] For example, in a medical scenario, the target text is "Please wait for the examination at the nurse's station", and the acoustic embedding features extracted based on the prompt image are the acoustic features of the operating room. At this time, the generated speech has the reverberation effect of the operating room, making the patient feel more real and natural when listening to the voice prompt, which helps to reduce the patient's anxiety. In a financial scenario, the target text is "Welcome to our bank. Please select the type of business to handle", and the acoustic embedding features are the acoustic features of the bank hall. At this time, the generated speech will have the reverberation effect of the bank hall, allowing customers to hear a voice prompt effect similar to that of handling business on-site even when handling business remotely, improving the immersion of the generated speech.

[0046] S204. Perform reverberation recognition on the target reverberant speech through a pre-trained reverberation classifier to confirm the reverberation category of the target reverberant speech.

[0047] In this embodiment, after generating the target reverberant speech, the target reverberant speech is further subjected to reverberation recognition by a pre-trained reverberation classifier. The reverberation classifier may specifically include a number of convolutional blocks and Transformer blocks, which can extract speech features and perform classification. Among them, the convolutional blocks are used to extract local features of the target reverberant speech, and the Transformer blocks are used to extract global features of the target reverberant speech. After combining the two, they are concatenated, and finally the reverberation category is output. The category output by the reverberation classifier can be used to verify whether the reverberation effect of the generated speech meets the expectations, that is, to confirm whether the reverberation category of the target reverberant speech is consistent with the target environment, so as to ensure the quality and accuracy of speech synthesis.

[0048] In the above embodiment, the present invention discloses a method for generating speech based on image indication, which includes obtaining a prompt image and a target text of the speech to be generated; performing image encoding processing on the prompt image to obtain an acoustic embedding feature matching the environment in the prompt image; inputting the target text and the acoustic embedding feature into a pre-trained speech generation model, and performing environment fusion speech generation processing on the target text with the acoustic embedding feature as the environmental reverberation condition to generate a corresponding target reverberant speech; performing reverberation recognition on the target reverberant speech by a pre-trained reverberation classifier to confirm the reverberation category of the target reverberant speech. By using the prompt of the image modality to guide speech generation and embedding the scene reverberation perception into the speech synthesis process, the acoustic environment of the corresponding scene can be reflected during speech synthesis, so that the generated speech can not only match the text, but also match the visual scene reflected in the image, and adaptively adjust the reverberation effect according to different scenes, providing a more immersive and realistic speech generation effect.

[0049] In one embodiment, after step S204, the method further includes:

[0050] Receiving the recognition feedback information of the user on the reverberation category;

[0051] Counting the recognition feedback information within a specified time period according to a preset update strategy to obtain feedback statistical data;

[0052] Updating the speech generation model according to the feedback statistical data.

[0053] In this embodiment, the generation effect is further optimized by collecting the evaluation feedback information of the user on the generated speech. For example, a user interface is provided to receive the evaluation and feedback of the user on the reverberation effect of the generated speech. The specific recognition feedback information can be qualitative feedback, such as selecting "satisfied with the reverberation effect" or "not satisfied with the reverberation effect", or quantitative feedback, such as giving a specific score to the reverberation effect of the speech within a preset range, etc. The user submits the recognition feedback information through the interface, and after the background receives it, the recognition feedback information can be preliminarily processed, such as converting the qualitative feedback information into numerical data, etc., to provide data support for subsequent model updates.

[0054] Based on the collected recognition feedback information, data statistics are performed according to a preset update strategy. For example, the preset update strategy can be to count the user feedback once at time intervals such as daily or weekly, or to perform statistics after a certain number of feedbacks are collected. Based on the preset update strategy, when the update condition is met, the recognition feedback information within a specified time period is statistically analyzed to obtain feedback statistical data that can reflect the overall evaluation of the user on the speech synthesis reverberation effect. For example, the recognition feedback information within a specified time period can be statistically analyzed, and indicators such as the satisfaction ratio and average score are calculated as the feedback statistical data, so as to reflect the changes in user needs and enable the system to adjust in time to improve its adaptability.

[0055] After that, the speech generation model is adjusted and optimized according to the feedback statistical data. For example, the parameters of the model are adjusted, the structure of the model is improved, or the model is retrained, etc. Specifically, the direction that needs to be optimized can be determined by analyzing the feedback statistical data. For example, if the feedback statistical data shows that the user's satisfaction with the reverberation effect is low, the extraction method of the acoustic embedding features can be adjusted to optimize the accuracy of the environmental reverberation. By regularly updating the model according to the user feedback, the reverberation effect of the speech synthesis is continuously improved, better adapting to different scenario requirements and providing a more natural and realistic speech experience.

[0056] In one embodiment, as Figure 3 shown, step S202 includes:

[0057] S301. Perform image encoding on the prompt image, and extract corresponding environmental prompt features from the prompt image;

[0058] S302. Determine the corresponding scene reverberation type according to the environmental prompt features, and convert the environmental prompt features into corresponding acoustic embedding features.

[0059] In this embodiment, when obtaining the corresponding acoustic embedding features from the prompt image, the extracted image is first encoded. Specifically, it can be implemented through an image encoding module based on the CLIP (Contrastive Language-Image Pretraining) architecture. The CLIP model can map the image to a feature space, effectively extracting the semantic information in the image. Specifically, features that can reflect the environmental characteristics are extracted from the extracted image, such as the size of the room, the material of the walls, the layout of the space, etc. Then, based on the extracted environmental prompt features, the reverberation type that matches the environment represented by the image is determined. For example, a large and empty space may have a longer reverberation time, while a small and narrow space may have a shorter reverberation time. The environmental prompt features can be mapped to the corresponding scene reverberation type through a predefined mapping table or model, and the environmental prompt features are converted into acoustic embedding features through an adapter module. The adapter module can include two fully connected layers and a GeLU activation function. The converted acoustic embedding features can be used as the input condition for the speech generation model to accurately match the reverberation type of the environment represented by the image, thereby generating more natural speech.

[0060] For example, the extracted environmental prompt features are the size of the ward, the material of the walls, etc. According to the extracted features, the reverberation type of the ward is determined to be "short reverberation" because the ward space is small and the wall material is sound-absorbing. The environmental prompt features are converted into acoustic embedding features through the adapter module, which are then used to guide the generation of speech with the reverberation effect of the ward, improving the authenticity and naturalness of the generated speech.

[0061] In one embodiment, as Figure 4 shown, step S203 includes:

[0062] S401. Input the target text and the acoustic embedding features into a pre-trained speech generation model, perform text encoding on the target text, and generate text features with semantic information and prosodic information;

[0063] S402. Perform acoustic feature prediction on the text features, and adjust the prediction result according to the acoustic embedding features to generate speech features with the corresponding environmental reverberation effect;

[0064] S403. Perform decoding processing on the speech features to generate the target reverberant speech of the target text, and the environmental reverberation effect of the target reverberant speech matches the environment in the prompt image.

[0065] In this embodiment, when performing voice generation processing for environment integration, the voice generation model includes components such as a text encoder, an acoustic model, and a vocoder. After inputting the target text and the acoustic embedding features into the voice generation model, the text encoder in the voice generation model maps each character or word in the text to a high-dimensional space through an embedding layer. For example, the text is encoded through a recurrent neural network (RNN), Transformer, or other sequence models, etc., to generate text features containing semantic and prosodic information. Among them, semantic information refers to the meaning of the text, while prosodic information includes prosodic characteristics of speech such as intonation, speech rate, and pause, providing a reliable data basis for generating natural and fluent speech. Then, the acoustic model predicts acoustic features from the text features, converting the text features into acoustic features. Acoustic features describe the physical characteristics of speech, such as fundamental frequency, formants, energy, etc. At the same time, the predicted acoustic features are adjusted based on the input acoustic embedding features to add an environmental reverberation effect. For example, parameters such as reverberation time, reverberation intensity, and frequency response are adjusted, so that the generated voice features carry the environmental reverberation effect while containing the semantic and prosodic information of the target text, thereby enhancing the immersion of the voice. Finally, the voice features with the environmental reverberation effect are input into the vocoder for decoding processing. After converting the voice features into a voice waveform, the target reverberant voice of the target text can be generated. At this time, the generated voice has a reverberation effect matching the environment in the prompt image, achieving the purpose of adaptively adjusting the reverberation effect according to different scenarios.

[0066] In one embodiment, as Figure 5 shown, before step S202, the method further includes:

[0067] S501. Construct an image encoder, a voice generation model, and a reverberation classifier to be trained;

[0068] S502. Collect model training data, where the model training data includes sample texts, sample prompt images, sample reverberation types corresponding to the sample prompt images, and sample reverberant voices corresponding to the sample texts and sample prompt images;

[0069] S503. Perform model training for multi-task learning on the image encoder, voice generation model, and reverberation classifier to be trained according to the model training data, and obtain a trained image encoder, voice generation model, and reverberation classifier.

[0070] In this embodiment, in the training stage, first construct an image encoder, a speech generation model, and a reverberation classifier to be trained. Among them, a pre-trained CLIP model is used as the basic architecture of the image encoder, and an adapter module composed of two fully connected layers plus a GeLU activation function is added to extract acoustic embedding features; the speech generation model (TTS model) can be implemented through a VITS structure, a flow model structure, etc.; the reverberation classifier includes several convolutional blocks and Transformer blocks for extracting speech features and classifying them. Then, collect model training data. Specifically, the model training data includes sample text w, sample prompt image i, and sample reverberation type r corresponding to the sample prompt image c , and sample reverberation speech s corresponding to the sample text and sample prompt image r . Among them, the sample text w is the text data used for training, usually sentences or phrases related to the application scenario; the sample prompt image i is the image data used for training, usually an environmental image related to the application scenario; the sample reverberation type r c is the reverberation type corresponding to the sample prompt image, usually obtained through manual annotation or automatic detection. For example, the reverberation type in a ward may be "short reverberation", and the reverberation type in a bank lobby may be "medium reverberation", etc.; the sample reverberation speech s r is the speech data corresponding to the sample text and sample prompt image, usually recorded in the actual environment by a professional recording device, with the corresponding environmental reverberation effect. By collecting data of multiple modalities, including text, images, and speech, etc., rich data resources are provided for model training, enabling the model to better adapt to different application scenarios and improving the adaptability of speech generation

[0071] Based on the collected model training data of multiple modalities, perform multi-task learning model training on the constructed image encoder, speech generation model, and reverberation classifier to be trained. Multi-task learning is to train multiple related tasks simultaneously and share some parts of the model to improve the performance of the model. That is, in this embodiment, the collected model training data is used to jointly train the image encoder, speech generation model, and reverberation classifier. During the training process, the model learns how to extract acoustic embedding features from images, how to generate speech with environmental reverberation effects based on the acoustic embedding features, and how to identify the reverberation type in the speech, so as to obtain a trained image encoder, speech generation model, and reverberation classifier. Through joint training, the model can simultaneously complete the tasks of image encoding, speech generation, and reverberation classification, improving the performance and adaptability of the speech generation system

[0072] In one embodiment, step S503 includes:

[0073] Input the sample prompt image into the image encoder to extract sample acoustic embedding features corresponding to the image environment, and calculate a first loss value based on the difference between the sample acoustic embedding features and the sample reverberation type;

[0074] Input the sample acoustic embedding features and the sample text into the speech generation model to generate corresponding predicted reverberant speech, and calculate a second loss value based on the difference between the predicted reverberant speech and the sample reverberant speech;

[0075] Input the predicted reverberant speech into the reverberation classifier for reverberation recognition to obtain a predicted reverberation type, and calculate a third loss value based on the difference between the predicted reverberation type and the sample reverberation type;

[0076] Calculate a corresponding total loss value based on the first loss value, the second loss value, and the third loss value, and perform parameter adjustment based on the total loss value until a trained image encoder, speech generation model, and reverberation classifier are obtained when a preset convergence condition is met.

[0077] In this embodiment, during the joint training process, input the sample prompt image i into the image encoder to extract environmental prompt features and convert them into sample acoustic embedding features a e , obtain the sample acoustic embedding features a e , and after that, perform cross - entropy (CrossEntropy) on it and the sample reverberation type r c to calculate a first loss value based on the difference between the two. Then input the sample acoustic embedding features a e and the sample text w into the speech generation model to generate corresponding predicted reverberant speech s p , and calculate a second loss value based on the difference between the predicted reverberant speech s p and the sample reverberant speech s r . Specifically, the mean squared error (MSE) between the predicted reverberant speech s p and the sample reverberant speech s r can be used as the second loss value to measure the difference between the two. Then input the predicted reverberant speech s p into the reverberation classifier to obtain the predicted reverberation type r p of the predicted reverberant speech s p , and then perform cross - entropy on the predicted reverberation type r p and the sample reverberation type r c as the third loss value.

[0078] Calculate the corresponding total loss value according to the first loss value, second loss value, and third loss value obtained under different tasks. The total loss value reflects the overall performance of the model in three tasks: image encoding, speech generation, and reverberation classification. The specific total loss value can be expressed as Loss = λ1CrossEntropy(a e ,r c ) + λ2MSE(s p ,s r ) + λ3CrossEntropy(r p ,r c ), where λ1, λ2, and λ3 are hyperparameters, representing the weights of the first loss value, second loss value, and third loss value respectively. Update the parameters of the model using the backpropagation algorithm according to the total loss value, and minimize the total loss value until the preset convergence condition is met, such as the total loss value is less than the preset threshold or the preset number of training rounds is reached, etc., then the trained image encoder, speech generation model, and reverberation classifier are obtained. By jointly training multiple tasks, the scene reverberation perception is embedded into the process of speech synthesis, so that the synthesized speech can not only match the text, but also match the visual scene reflected in the image. Therefore, the trained speech synthesis system can adaptively adjust the reverberation effect according to different scenes, making the synthesis of controllable reverberant speech feasible.

[0079] In one embodiment, before step S503, the method further includes:

[0080] Pre-train the reverberation classifier to be trained according to the model training data, obtain the pre-trained reverberation classifier, and freeze the parameters of the pre-trained reverberation classifier.

[0081] In this embodiment, before the overall joint training of the model, the sample reverberant speech s r and the sample reverberation type r c in the model training data can be used to pre-train the reverberation classifier to be trained, that is, input the sample reverberant speech s r into the reverberation classifier to obtain the corresponding predicted reverberation type r p , and based on the predicted reverberation type r p and the sample reverberation type r cThe cross-entropy is used as the loss function in the pre-training stage to pre-train the echo classifier. After the pre-training is completed, the parameters of the pre-trained echo classifier are frozen, and then the other parts in the overall architecture, namely the image encoder and the speech generation model, are trained as a whole. Through pre-training, the echo classifier can learn the mapping relationship between speech features and echo types, so as to provide more accurate echo category labels in subsequent multi-task learning, improve the efficiency of multi-task learning, and freezing the parameters of the echo classifier can prevent it from being interfered by other tasks during multi-task learning, maintain the stability of classification performance, and improve the performance of the entire speech generation system.

[0082] It should be noted that there is not necessarily a certain order among the above steps. Those of ordinary skill in the art can understand according to the description of the embodiments of the present invention that in different embodiments, the above steps can have different execution orders, that is, they can be executed in parallel, or they can be exchanged and executed, etc.

[0083] Further referring to Figure 6 , as an implementation of the method shown above Figure 2 , the present invention provides an embodiment of a speech generation device based on image indication. This device embodiment corresponds to the method embodiment shown in Figure 2 , and this device can be specifically applied to various electronic devices.

[0084] As shown in Figure 6 , the speech generation device 60 based on image indication described in this embodiment includes:

[0085] An acquisition module 601, configured to acquire a prompt image and a target text of the speech to be generated;

[0086] An image processing module 602, configured to perform image encoding processing on the prompt image to obtain an acoustic embedding feature matching the environment in the prompt image;

[0087] A speech generation module 603, configured to input the target text and the acoustic embedding feature into a pre-trained speech generation model, and perform environment-fused speech generation processing on the target text with the acoustic embedding feature as the environmental echo condition to generate a corresponding target echo speech;

[0088] An echo classification module 604, configured to perform echo recognition on the target echo speech through a pre-trained echo classifier to confirm the echo category of the target echo speech.

[0089] The module referred to in the present invention refers to a series of computer program instruction segments that can complete specific functions. It is more suitable for describing the execution process of speech generation based on image indication than a program. For the specific implementation manners of each module, please refer to the corresponding method embodiments above, and details are not described herein again.

[0090] In one embodiment, the image processing module 602 includes:

[0091] An image encoding unit, configured to perform image encoding on the prompt image and extract corresponding environmental prompt features from the prompt image;

[0092] An acoustic conversion unit, configured to determine a corresponding scene reverberation type according to the environmental prompt features and convert the environmental prompt features into corresponding acoustic embedding features.

[0093] In one embodiment, the speech generation module 603 includes:

[0094] A text encoding unit, configured to input the target text and the acoustic embedding features into a pre-trained speech generation model, perform text encoding on the target text, and generate text features with semantic information and prosodic information;

[0095] An acoustic prediction unit, configured to perform acoustic feature prediction on the text features and adjust the prediction result according to the acoustic embedding features to generate speech features with corresponding environmental reverberation effects;

[0096] A speech decoding unit, configured to perform decoding processing on the speech features to generate the target reverberant speech of the target text, and the environmental reverberation effect of the target reverberant speech matches the environment in the prompt image.

[0097] In one embodiment, the apparatus 60 further includes:

[0098] A construction module, configured to construct an image encoder, a speech generation model, and a reverberation classifier to be trained;

[0099] A data acquisition module, configured to acquire model training data, where the model training data includes sample texts, sample prompt images, sample reverberation types corresponding to the sample prompt images, and sample reverberant speeches corresponding to the sample texts and sample prompt images;

[0100] A model training module, configured to perform multi-task learning model training on the image encoder, the speech generation model, and the reverberation classifier to be trained according to the model training data, and obtain a trained image encoder, a speech generation model, and a reverberation classifier.

[0101] In one embodiment, the model training module includes:

[0102] The first loss calculation unit is configured to input the sample prompt image into the image encoder, extract sample acoustic embedding features corresponding to the image environment, and calculate a first loss value according to the difference between the sample acoustic embedding features and the sample reverberation type;

[0103] The second loss calculation unit is configured to input the sample acoustic embedding features and the sample text into the speech generation model to generate corresponding predicted reverberant speech, and calculate a second loss value according to the difference between the predicted reverberant speech and the sample reverberant speech;

[0104] The third loss calculation unit is configured to input the predicted reverberant speech into the reverberation classifier for reverberation recognition to obtain a predicted reverberation type, and calculate a third loss value according to the difference between the predicted reverberation type and the sample reverberation type;

[0105] The parameter adjustment unit is configured to calculate a corresponding total loss value according to the first loss value, the second loss value, and the third loss value, and perform parameter adjustment according to the total loss value until a trained image encoder, speech generation model, and reverberation classifier are obtained when a preset convergence condition is satisfied.

[0106] In one embodiment, the apparatus 60 further includes:

[0107] The pre-training module is configured to pre-train the reverberation classifier to be trained according to the model training data to obtain a pre-trained reverberation classifier, and freeze the parameters of the pre-trained reverberation classifier.

[0108] In one embodiment, the apparatus 60 further includes:

[0109] The feedback receiving module is configured to receive the user's recognition feedback information on the reverberation category;

[0110] The statistics module is configured to statistically calculate the recognition feedback information within a specified time period according to a preset update strategy to obtain feedback statistical data;

[0111] The model update module is configured to update the speech generation model according to the feedback statistical data.

[0112] In the above embodiments, the present invention discloses a voice generation device based on image indication, which obtains a prompt image and a target text for which voice is to be generated; performs image encoding processing on the prompt image to obtain an acoustic embedding feature matching the environment in the prompt image; inputs the target text and the acoustic embedding feature into a pre-trained voice generation model, and performs voice generation processing with environmental fusion on the target text using the acoustic embedding feature as the environmental reverberation condition to generate a corresponding target reverberant voice; and performs reverberation recognition on the target reverberant voice through a pre-trained reverberation classifier to confirm the reverberation category of the target reverberant voice. By using the prompt of the image modality to guide voice generation and embedding scene reverberation perception into the process of speech synthesis, the acoustic environment of the corresponding scene can be reflected during speech synthesis, so that the generated voice can not only match the text, but also match the visual scene reflected in the image, and adaptively adjust the reverberation effect according to different scenes, providing a more immersive and realistic voice generation effect.

[0113] Another embodiment of the present invention provides a computer device, as Figure 7 shown, the computer device 70 includes:

[0114] One or more processors 701 and a memory 702, Figure 7 Taking one processor 701 as an example for introduction, the processor 701 and the memory 702 can be connected through a bus or other means, Figure 7 Taking the connection through the bus as an example.

[0115] The processor 701 is used to complete various control logics of the computer device 70. It can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISCMachine), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination of these components. Additionally, the processor 701 can also be any conventional processor, microprocessor, or state machine. The processor 701 can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP and / or any other such configuration.

[0116] The memory 702, being a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions corresponding to the image-based voice generation method in the embodiments of the present invention. The processor 701 executes various functional applications and data processing of the computer device 70 by running the non-volatile software programs, instructions, and units stored in the memory 702, that is, implements the image-based voice generation method in the above method embodiments.

[0117] The memory 702 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device 70, etc. In addition, the memory 702 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 702 optionally includes a memory remotely set relative to the processor 701, and these remote memories can be connected to the computer device 70 through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. One or more units are stored in the memory 702 and, when executed by one or more processors 701, perform the steps of the image-based voice generation method in any of the above method embodiments.

[0118] In the above embodiments, the present invention discloses a computer device, which obtains a prompt image and a target text for which voice is to be generated; performs image encoding processing on the prompt image to obtain an acoustic embedding feature matching the environment in the prompt image; inputs the target text and the acoustic embedding feature into a pre-trained voice generation model, and performs environment-fused voice generation processing on the target text with the acoustic embedding feature as the environmental reverberation condition to generate a corresponding target reverberant voice; and performs reverberation recognition on the target reverberant voice through a pre-trained reverberation classifier to confirm the reverberation category of the target reverberant voice. By using the prompt of the image modality to guide voice generation and embedding scene reverberation perception into the process of speech synthesis, it is possible to reflect the acoustic environment of the corresponding scene during speech synthesis, so that the generated voice can not only match the text but also match the visual scene reflected in the image, and adaptively adjust the reverberation effect according to different scenes, providing a more immersive and realistic voice generation effect.

[0119] The embodiments of the present invention provide a non-volatile computer-readable storage medium, and the computer-readable storage medium stores computer-executable instructions, which, when executed by one or more processors, perform the steps of the image-based voice generation method in any of the above method embodiments.

[0120] In the above embodiments, the present invention discloses a non-volatile computer-readable storage medium, which obtains a prompt image and a target text of the voice to be generated; performs image encoding processing on the prompt image to obtain an acoustic embedding feature matching the environment in the prompt image; inputs the target text and the acoustic embedding feature into a pre-trained voice generation model, and performs environment fusion voice generation processing on the target text with the acoustic embedding feature as the environmental reverberation condition to generate a corresponding target reverberant voice; and performs reverberation recognition on the target reverberant voice through a pre-trained reverberation classifier to confirm the reverberation category of the target reverberant voice. By using the prompt of the image modality to guide voice generation and embedding scene reverberation perception into the process of speech synthesis, the acoustic environment of the corresponding scene can be reflected during the speech synthesis process, so that the generated voice can not only match the text, but also match the visual scene reflected in the image, and adaptively adjust the reverberation effect according to different scenes, providing a more immersive and realistic voice generation effect.

[0121] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0122] The present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0123] In summary, in the method, apparatus, device, and medium for image-indicated speech generation disclosed in the present invention, the method includes: obtaining a prompt image and a target text for which speech is to be generated; performing image encoding processing on the prompt image to obtain an acoustic embedding feature matching the environment in the prompt image; inputting the target text and the acoustic embedding feature into a pre-trained speech generation model, and performing environment-fused speech generation processing on the target text with the acoustic embedding feature as the environmental reverberation condition to generate a corresponding target reverberant speech; and performing reverberation recognition on the target reverberant speech through a pre-trained reverberation classifier to confirm the reverberation category of the target reverberant speech. By using the prompt of the image modality to guide speech generation and embedding scene reverberation perception into the speech synthesis process, the acoustic environment of the corresponding scene can be reflected during speech synthesis, so that the generated speech can not only match the text, but also match the visual scene reflected in the image, adaptively adjust the reverberation effect according to different scenes, and provide a more immersive and realistic speech generation effect.

[0124] Of course, those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. The storage medium can be a memory, a magnetic disk, a floppy disk, a flash memory, an optical memory, etc.

[0125] It should be noted that if non-company software tools or components appear in the embodiments of the present application, they are only used for illustrative introduction and do not represent actual use. It should be understood that the application of the present invention is not limited to the above examples. Those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for generating speech based on image indication, characterized in that, Including: Obtain a prompt image and a target text for which speech is to be generated; Perform image encoding processing on the prompt image to obtain an acoustic embedding feature matching the environment in the prompt image; Input the target text and the acoustic embedding feature into a pre-trained speech generation model, and perform environment-fused speech generation processing on the target text with the acoustic embedding feature as the environmental reverberation condition to generate a corresponding target reverberant speech; Perform reverberation recognition on the target reverberant speech through a pre-trained reverberation classifier to confirm the reverberation category of the target reverberant speech.

2. The method for generating voice based on image indication according to claim 1, wherein The performing image encoding processing on the prompt image to obtain an acoustic embedding feature matching the environment in the prompt image includes: Perform image encoding on the prompt image, and extract corresponding environmental prompt features from the prompt image; Determine a corresponding scene reverberation type according to the environmental prompt features, and convert the environmental prompt features into corresponding acoustic embedding features.

3. The method for generating voice based on image indication according to claim 1, wherein The inputting the target text and the acoustic embedding feature into a pre-trained speech generation model, and performing environment-fused speech generation processing on the target text with the acoustic embedding feature as the environmental reverberation condition to generate a corresponding target reverberant speech includes: Input the target text and the acoustic embedding feature into a pre-trained speech generation model, perform text encoding on the target text, and generate a text feature with semantic information and prosodic information; Perform acoustic feature prediction on the text feature, and adjust the prediction result according to the acoustic embedding feature to generate a speech feature with a corresponding environmental reverberation effect; Perform decoding processing on the speech feature to generate the target reverberant speech of the target text, and the environmental reverberation effect of the target reverberant speech matches the environment in the prompt image.

4. The method for generating voice based on image indication according to claim 1, wherein Before the performing image encoding processing on the prompt image to obtain an acoustic embedding feature matching the environment in the prompt image, the method further includes: Construct an image encoder, a speech generation model, and a reverberation classifier to be trained; Collect model training data, where the model training data includes sample texts, sample prompt images, sample reverberation types corresponding to the sample prompt images, and sample reverberant speeches corresponding to the sample texts and sample prompt images; Perform model training of multi-task learning on the image encoder, speech generation model, and reverberation classifier to be trained according to the model training data to obtain a trained image encoder, speech generation model, and reverberation classifier.

5. The method for generating speech based on image indication according to claim 4, wherein The performing model training of multi-task learning on the image encoder, speech generation model, and reverberation classifier to be trained according to the model training data to obtain a trained image encoder, speech generation model, and reverberation classifier includes: Input the sample prompt image into the image encoder, extract a sample acoustic embedding feature corresponding to the image environment, and calculate a first loss value according to the difference between the sample acoustic embedding feature and the sample reverberation type; Input the sample acoustic embedding features and the sample text into the speech generation model to generate corresponding predicted reverberant speech, and calculate a second loss value according to the difference between the predicted reverberant speech and the sample reverberant speech; Input the predicted reverberant speech into the reverberation classifier for reverberation recognition to obtain a predicted reverberation type, and calculate a third loss value according to the difference between the predicted reverberation type and the sample reverberation type; Calculate a corresponding total loss value according to the first loss value, the second loss value and the third loss value, and perform parameter adjustment according to the total loss value until the trained image encoder, speech generation model and reverberation classifier are obtained when a preset convergence condition is satisfied.

6. The method for generating speech based on image indication according to claim 4, wherein Before the method performs multi-task learning model training on the image encoder, speech generation model and reverberation classifier to be trained according to the model training data, the method further includes: Pre-train the reverberation classifier to be trained according to the model training data to obtain a pre-trained reverberation classifier, and freeze the parameters of the pre-trained reverberation classifier.

7. The method for generating speech based on image indication according to any one of claims 1-6, characterized in that, After the reverberation category of the target reverberant speech is confirmed by performing reverberation recognition on the target reverberant speech through the pre-trained reverberation classifier, the method further includes: Receive the user's recognition feedback information on the reverberation category; Count the recognition feedback information within a specified time period according to a preset update policy to obtain feedback statistical data; Update the model of the speech generation model according to the feedback statistical data.

8. A voice generation device based on image indication, characterized in that including: An acquisition module, configured to acquire a prompt image and a target text of the speech to be generated; An image processing module, configured to perform image encoding processing on the prompt image to obtain acoustic embedding features matching the environment in the prompt image; A speech generation module, configured to input the target text and the acoustic embedding features into a pre-trained speech generation model, and perform environment-fused speech generation processing on the target text with the acoustic embedding features as the environmental reverberation condition to generate corresponding target reverberant speech; A reverberation classification module, configured to perform reverberation recognition on the target reverberant speech through a pre-trained reverberation classifier to confirm the reverberation category of the target reverberant speech.

9. A computer device, characterized in that, including at least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image indication-based speech generation method according to any one of claims 1-7.

10. A non-volatile computer-readable storage medium, characterized in that, The non-volatile computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by one or more processors, the one or more processors can execute the image indication-based speech generation method according to any one of claims 1-7.