Training and speech generation methods and apparatuses for speech generation model, electronic device, computer-readable storage medium, and computer program product

By unifying the visual and text modal representation spaces through multimodal coding and loss calculation, the problems of inaccurate and inconvenient speech-to-text generation in existing technologies are solved, achieving more efficient and accurate speech-to-text generation and improving the user experience.

WO2025236958A9PCT designated stage Publication Date: 2026-06-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2025-04-14
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

In existing technologies, the method of generating speech text by encoding the hand and lip movements of the target object is not accurate or convenient enough, resulting in a poor user experience.

Method used

By acquiring multimodal sample data, using decoders and multiple encoders to encode visual and textual modalities, calculating target loss and updating model parameters, the representation space of visual and textual modalities is unified, generating more accurate speech-text.

Benefits of technology

It improves the accuracy and efficiency of generating voice text from prompt images, provides a more convenient user interaction method, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025088811_04062026_PF_FP_ABST
    Figure CN2025088811_04062026_PF_FP_ABST
Patent Text Reader

Abstract

The present application provides training and speech generation methods and apparatuses for a speech generation model, an electronic device, a computer-readable storage medium, and a computer program product. The method comprises: obtaining a first speech generation model; obtaining sample data of a plurality of modalities; on the basis of a prompt image sequence and speech text, respectively calling a plurality of encoders to perform encoding, so as to obtain a multi-modal encoding vector sequence; on the basis of the multi-modal encoding vector sequence, calling a decoder to perform decoding, so as to obtain decoded text; determining a probability distribution for the decoded text and the sample data of the plurality of modalities, and determining a target loss on the basis of the probability distribution; and on the basis of the target loss, updating parameters of the decoder and at least one of the encoders, wherein the updated decoder and the plurality of updated encoders are configured to form a second speech generation model, and the second speech generation model is used to generate target speech text corresponding to a prompt image sequence of a target object.
Need to check novelty before this filing date? Find Prior Art

Description

Training methods for speech generation models, speech generation methods, devices, electronic devices, computer-readable storage media, and computer program products.

[0001] Cross-references to related applications

[0002] This application is based on and claims priority to Chinese Patent Application No. 2024106268722, filed on May 15, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to artificial intelligence technology, and more particularly to a training method for a speech generation model, a speech generation method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology

[0004] In interpersonal communication or scenarios requiring human-computer interaction, when the target cannot directly express their intentions through speech (e.g., due to limitations of the current environment or physiological barriers to vocalization), relevant technologies support users to express their intentions through body language and lip movements, or by inputting text on the terminal device. For example, related technologies can generate speech-text by encoding the target's hand and lip movements; however, the speech-text generated by these technologies is not accurate or convenient enough, resulting in a poor user experience. Summary of the Invention

[0005] In view of this, embodiments of this application provide a training method for a speech generation model, a speech generation method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the efficiency and accuracy of generating speech text from prompt images.

[0006] The technical solution of this application embodiment is implemented as follows:

[0007] This application provides a method for training a speech generation model, the method being executed by an electronic device, the method comprising:

[0008] Obtain a first speech generation model, wherein the first speech generation model includes a decoder and multiple encoders corresponding to multiple modalities respectively;

[0009] Acquire sample data for multiple modalities, wherein the sample data for multiple modalities includes a sequence of prompt images of sample objects and the corresponding speech text;

[0010] Based on the prompt image sequence and the voice text, the multiple encoders are called to encode the text, resulting in a multimodal encoded vector sequence;

[0011] The decoder is invoked based on the multimodal encoded vector sequence to perform decoding, thereby obtaining the decoded text;

[0012] Determine the probability distribution of the decoded text and the sample data of the multiple modalities, and determine the target loss based on the probability distribution;

[0013] The parameters of the decoder and at least one encoder are updated based on the target loss to obtain the updated decoder and the plurality of encoders, wherein the updated decoder and the plurality of encoders are used to form a second speech generation model, and the second speech generation model is used to generate target speech text corresponding to the prompt image sequence of the target object.

[0014] This application provides a speech generation method based on a speech generation model, wherein the speech generation model is the second speech generation model described above, and the method is executed by an electronic device. The method includes:

[0015] Obtain the sequence of prompt images for the target object;

[0016] Based on the prompt image sequence, the second speech generation model is invoked to generate target speech text corresponding to the prompt image sequence of the target object.

[0017] This application provides a training apparatus for a speech generation model, the apparatus comprising:

[0018] The acquisition module is configured to acquire a first speech generation model, wherein the first speech generation model includes a decoder and multiple encoders corresponding to multiple modalities; and acquire sample data of multiple modalities, wherein the sample data of multiple modalities includes a sequence of prompt image of sample objects and speech text corresponding to the prompt image sequence.

[0019] The encoding module is configured to call the multiple encoders to encode the prompt image sequence and the speech text respectively, thereby obtaining a multimodal encoded vector sequence;

[0020] The decoding module is configured to call the decoder based on the multimodal encoded vector sequence to obtain decoded text;

[0021] The determination module is configured to determine the probability distribution of the decoded text and the sample data of the multiple modalities, and determine the target loss based on the probability distribution;

[0022] The generation module is configured to update the parameters of the decoder and at least one encoder based on the target loss to obtain the updated decoder and the plurality of encoders, wherein the updated decoder and the plurality of encoders are used to form a second speech generation model, and the second speech generation model is used to generate target speech text corresponding to the prompt image sequence of the target object.

[0023] This application provides a speech generation apparatus for a speech generation model, wherein the speech generation model is the second speech generation model described above, and the apparatus includes:

[0024] The acquisition module is configured to acquire a sequence of prompt images for the target object;

[0025] The generation module is configured to call the second speech generation model based on the prompt image sequence to generate target speech text corresponding to the prompt image sequence of the target object.

[0026] This application provides an electronic device, the electronic device comprising:

[0027] Memory is used to store executable instructions for a computer;

[0028] When the processor executes the computer-executable instructions stored in the memory, it implements the training method of the speech generation model provided in the embodiments of this application, or the speech generation method of the speech generation model.

[0029] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which, when executed by a processor, implements the training method of the speech generation model provided in this application, or the speech generation method of the speech generation model.

[0030] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the training method of the speech generation model provided in this application, or the speech generation method of the speech generation model.

[0031] The embodiments of this application have the following beneficial effects:

[0032] By encoding sample data from multiple modalities, including the prompt image sequence (corresponding to the visual modality) and the speech text (corresponding to the text modality), the target loss is calculated by comparing the resulting multimodal encoded vector sequence (i.e., the encoded vector sequence of multiple modalities) with the probability distribution of the decoded text. This allows the target loss to reflect the differences between the representation spaces of multiple modalities in the first speech generation model. Based on the target loss, the parameters of the first speech generation model are updated through backpropagation, enabling the first speech generation model to gradually learn the intrinsic relationship between the information of the visual and text modalities, improving the model's performance and generalization ability. This results in the unification of the representation spaces of the visual and text modalities in the second speech generation model after training. This means that the second speech generation model can map data from different modalities into a common representation space, allowing data from different modalities to have similar semantic expressions in this space. This lays the foundation for accurate speech text generation in the future, ensuring that the target speech text output by the second speech generation model accurately matches the intent of the prompt image of the target object, thereby guaranteeing the accuracy of speech text generation from the prompt image. At the same time, compared with related technologies that require manual text input to generate voice text, it is more convenient, more efficient, and provides a better user experience. Attached Figure Description

[0033] Figure 1 is a schematic diagram of the architecture of the training system 100 for the speech generation model provided in an embodiment of this application;

[0034] Figure 2A is a schematic diagram of the structure of the electronic device 500-1 provided in an embodiment of this application;

[0035] Figure 2B is a schematic diagram of the structure of the electronic device 500-2 provided in an embodiment of this application;

[0036] Figure 3A is a flowchart illustrating the training method of the speech generation model provided in an embodiment of this application;

[0037] Figure 3B is a schematic diagram of the first process for obtaining a multimodal encoded vector sequence according to an embodiment of this application;

[0038] Figure 3C is a schematic diagram of the process for obtaining the visual encoding vector sequence provided in an embodiment of this application;

[0039] Figure 3D is a schematic diagram of the second process for obtaining a multimodal encoded vector sequence provided in an embodiment of this application;

[0040] Figure 3E is a schematic diagram of the third process for obtaining a multimodal encoded vector sequence provided in an embodiment of this application;

[0041] Figure 3F is a schematic diagram of the fourth process for obtaining the multimodal encoded vector sequence provided in an embodiment of this application;

[0042] Figure 3G is a schematic diagram of the process for aligning the length of the encoded vector sequence provided in an embodiment of this application;

[0043] Figure 3H is a schematic diagram of the first process for determining the probability distribution provided in an embodiment of this application;

[0044] Figure 3I is a schematic diagram of the first process for determining the target loss provided in an embodiment of this application;

[0045] Figure 3J is a schematic diagram of the second process for determining the probability distribution provided in an embodiment of this application;

[0046] Figure 3K is a schematic diagram of the second process for determining the target loss provided in an embodiment of this application;

[0047] Figure 3L is a schematic diagram of the third process for determining the probability distribution provided in an embodiment of this application;

[0048] Figure 3M is a schematic diagram of the third process for determining the target loss provided in an embodiment of this application;

[0049] Figure 3N is a schematic diagram of the first process for updating parameters provided in an embodiment of this application;

[0050] Figure 30 is a schematic diagram of the second process for updating parameters provided in an embodiment of this application;

[0051] Figure 3P is a schematic diagram of the third process for updating parameters provided in an embodiment of this application;

[0052] Figure 3Q is a schematic flowchart of the speech generation method of the speech generation model provided in the embodiment of this application;

[0053] Figure 4A is a schematic diagram of the first architecture of the speech generation model provided in an embodiment of this application;

[0054] Figure 4B is a schematic diagram of the second architecture of the speech generation model provided in the embodiment of this application;

[0055] Figure 4C is a schematic diagram of the third architecture of the speech generation model provided in the embodiments of this application;

[0056] Figure 5A is a schematic diagram of the structure of the visual encoder model provided in an embodiment of this application;

[0057] Figure 5B is a schematic diagram of the structure of the machine learning model provided in the embodiment of this application;

[0058] Figure 6 is a schematic diagram of the alignment of the encoded vector sequence provided in an embodiment of this application;

[0059] Figure 7 is a schematic diagram of the prompt voice encoding information provided in an embodiment of this application;

[0060] Figure 8 is a schematic diagram of the training data provided in an embodiment of this application;

[0061] Figure 9 is a visual coding model architecture diagram provided in an embodiment of this application.

[0062] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation

[0063] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0064] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0065] In the following description, the terms "first / second / third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0066] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functions of that module or unit.

[0067] Unless otherwise specified, "at least one" as used below refers to one or more cases, and "multiple" can refer to two or more cases.

[0068] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0069] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0070] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0071] 1) Speech generation model: includes a decoder and multiple encoders corresponding to multiple modalities. It is trained using sample data from multiple modalities based on the sample object to generate target speech text corresponding to the visual expression of the target object.

[0072] 2) Visual Modal: Through intuitive means such as hand gestures, lip movements, and body language, the recipient can perceive and understand the visual expression of the target object's facial expressions, eye contact, and body posture, thereby obtaining the target object's emotional state and intentions. For example, for a sequence of prompt images for the target object, each prompt image in the sequence is a static image showing the hand or lip posture of the sample object when performing a body or lip movement. Each prompt image identifies the posture and relative position of the target object's hands and lips when performing the key actions.

[0073] 3) Text modality: A textual expression used to convey the emotional state and intentions of the target audience, such as "The weather is nice today".

[0074] 4) Speech modality: The expression of speech signals used to convey the emotional state and intentions of the target object.

[0075] 5) Multimodal encoding vector sequence: Call the encoders corresponding to multiple modalities in the language generation model to encode the sample data of multiple modalities respectively, and use the resulting multimodal encoding vector sequence as the multimodal encoding vector sequence.

[0076] 6) Representation space: This refers to the vector space composed of feature vectors or feature representations, used to describe and represent different attributes, features, or information of data. In different modalities, the representation space embodies the features of different types of data and their representations in the vector space. For example, the representation space of the visual modality is a vector space composed of visual feature vectors, while the representation space of the text modality is a vector space composed of text feature vectors.

[0077] 7) Cued-speech (CS): A coding system used by people who cannot speak or whose speech is impaired to express spoken language.

[0078] 8) Sign Language Source Language (SL): A form of language that uses gestures, hand movements, and facial expressions to communicate. It uses gestures and hand movements to represent words, phrases, and sentences, through different shapes, positions, and movements of the fingers, palms, wrists, etc., combined with facial expressions and body movements.

[0079] 9) Impaired speech (IS): refers to speech signals whose quality has deteriorated or are difficult to understand due to various reasons. For example, when noise is mixed with a speech signal, it will lead to a decrease in speech quality, making it difficult to hear and understand the speech content clearly; distortion or damage to the speech signal during transmission, recording or processing will lead to spectral distortion, temporal distortion, distortion impact, etc., making the speech sound unnatural or distorted; people with speech disorders or articulation problems due to physiological or neurological reasons will have unclear pronunciation, abnormal speech rate, unstable pitch, etc., making their speech difficult or incomprehensible.

[0080] 10) Knowledge Distillation (KD): This is a model training technique used to transfer knowledge from a teacher model to a student model to assist the student model's learning and improve its performance. In knowledge distillation, the teacher model is typically a complex and accurate model with good performance. The student model, on the other hand, is a lightweight model, usually unable to achieve the complexity of the teacher model due to limitations in computing resources or model size. For example, in this embodiment, knowledge distillation is used to align the output of the visual encoder to the feature space of the text encoder's output, that is, the text encoder model is used as the teacher model, and the visual encoder model is used as the student model, with the text encoder guiding the visual encoder.

[0081] 11) Prompt Image Sequence: Multiple prompt images form a prompt image sequence. For each prompt image, the posture and position of the target object's limbs, lips, and hands are identified when the object makes a limb movement, lip movement, or hand movement.

[0082] 12) Keypoint sequence: In each prompt image of the prompt image sequence, the points that identify the posture and position of the limbs, lips or hands are taken as keypoints. The keypoints in each prompt image of the prompt image sequence are combined in the order in which the actions occur to obtain the keypoint sequence.

[0083] Related technologies generate speech text by encoding the hand and lip movements of the target object. However, the method of generating speech text by encoding is only applicable to users who understand the encoding method, and has a small scope of application. Moreover, the expression of the generated speech text is not accurate or natural compared to normal communication. The method of users manually inputting text to complete the interaction process is not convenient, has low efficiency, and has a poor user experience.

[0084] To address the aforementioned issues, embodiments of this application provide a training method, apparatus, electronic device, computer-readable storage medium, and computer program product for a speech generation model, which can improve the efficiency and accuracy of generating speech text from prompt images.

[0085] The following describes exemplary applications of the electronic devices provided in the embodiments of this application. These electronic devices can be implemented as various types of terminals such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals, or as servers. Exemplary applications of the electronic devices as terminals or servers will be described below.

[0086] Referring to Figure 1, which is a schematic diagram of the architecture of a speech generation model training system 100 provided in an embodiment of this application, in order to support the training application of a speech generation model, a terminal 400 (exemplarily showing a graphical interface 410) connects to a server 200 through a network 300. The network 300 can be a wide area network or a local area network, or a combination of both.

[0087] The method provided in this application can be applied to several different scenarios, which can improve the user's interactive experience, enhance the flexibility and convenience of human-computer interaction, and help to achieve more intelligent and personalized human-computer interaction.

[0088] 1. Assisted Communication: For people with hearing or speech impairments, sign language can be recognized to help them communicate with others. For example, terminal 400 acquires a sequence of sign language images from people with hearing or speech impairments and sends it to server 200. Server 200 identifies the position of fingers, the shape of gestures, and the trajectory of movements in the sign language image sequence, translates the sign language into text or speech, and sends it back to terminal 400 for display to the target group. In this way, people with hearing or speech impairments can communicate with others in real time through gestures.

[0089] 2. Gesture Control: Body movements and gesture recognition can be applied to smart devices or wearable devices, allowing users to control device functions through gestures, such as adjusting volume or switching songs. For example, terminal 400 detects the user's body movements and sends them to server 200. Server 200 determines the volume increase (decrease) by recognizing the user's body movements, such as sliding the arm up (or down) or raising (lowering) the palm, and sends this information to terminal 400, which then increases (or decreases) the volume.

[0090] 3. Facial Expression Recognition: By recognizing facial expressions and lip movements, it's possible to determine a user's emotional state and intentions, thereby providing a better user experience and personalized services, such as automated sentiment analysis and emotional interaction in games. For example, in a game, terminal 400 can acquire facial expression change data from the player and send it to server 200. Server 200 edits the player's facial expression change data to obtain a sequence of prompt images, decodes the prompt image sequence to obtain the player's target intention data, and sends it to terminal 400. This allows terminal 400 to adjust the game difficulty, provide corresponding props, or change the storyline based on the target intention data, thus providing a more challenging and personalized gaming experience.

[0091] 4. Virtual Reality and Augmented Reality: By recognizing body movements and lip movements, users' actions and expressions can be transmitted in real time to virtual reality or augmented reality environments, achieving a more immersive interactive experience, such as the imitation of virtual character body movements. For example, terminal 400 acquires a sequence of user prompt images and sends it to terminal 200. Terminal 200 recognizes the user's body movements and lip movements to determine the user's target intention and sends it to terminal 400. Terminal 400 then maps the user's actions and expressions into a virtual character or virtual scene in real time according to the user's target intention. This allows actors or artists to control the actions and interactions of virtual characters through their own movements and expressions during performances, making the augmented reality experience more vivid.

[0092] 5. Intelligent Driving Assistance: By recognizing the driver's body language and lip movements, abnormal driver conditions can be detected promptly, improving driving safety. For example, terminal 400 (such as an in-vehicle camera) acquires a sequence of images of driver cues during driving (such as yawning, frequent blinking, hands off the steering wheel, etc.) and sends them to server 200. Server 200 determines whether the driver is fatigued or distracted by recognizing the sequence of cues. For example, if frequent blinking is detected as a sign of fatigue, an alarm message is generated and sent to terminal 400. Terminal 400 then reminds the driver to rest or take appropriate safety measures based on the alarm message, such as playing an alarm sound or providing rest instructions.

[0093] 6. Medical Rehabilitation and Assisted Treatment: Based on the patient's limb movements and lip movements, the system monitors the patient's rehabilitation progress, assisting doctors in diagnosis and treatment. For example, terminal 400 acquires image sequences of prompts (such as the amplitude and frequency of limb movements) during rehabilitation training and sends them to server 200. Server 200 analyzes the patient's rehabilitation status by recognizing the image sequences of these prompts. For instance, if it recognizes that the amplitude of the patient's limb movements is gradually increasing, indicating good rehabilitation results, it generates an assessment report and sends it to terminal 400. Terminal 400 displays the assessment report to both the doctor and the patient, allowing the doctor to adjust the rehabilitation plan based on the report, and the patient to understand their rehabilitation progress.

[0094] 7. Smart Home Control: In a smart home environment, users can control various devices in their homes, such as lights, curtains, and appliances, through body movements and lip movements, enhancing convenience. For example, terminal 400 (such as a smart camera) acquires a sequence of user prompts (such as waving an arm or opening a mouth) and sends it to server 200. Server 200 determines the user's intention by recognizing their body movements and lip movements. For instance, recognizing a user waving their arm to open the curtains generates a corresponding control command and sends it to terminal 400. Terminal 400 then operates the smart home devices according to the control command, such as opening the curtains or adjusting the light brightness.

[0095] Terminal 400 sends sample data of multiple modalities, including the prompt image sequence of the sample object and the corresponding speech text, as well as the prompt image sequence of the target object, to server 200 via network 300. Server 200 calls multiple encoders corresponding to the multiple modalities in the first speech generation model to encode the received sample data of multiple modalities of the sample object, obtaining a multimodal encoded vector sequence, and calls the decoder in the first speech generation model to decode the multimodal encoded vector sequence to obtain decoded text. Then, it determines the probability distribution of the decoded text and the sample data of multiple modalities, and determines the target loss based on the probability distribution. Based on the target loss, it updates the parameters of the decoder and at least one encoder to obtain a second speech generation model. Finally, server 200 generates target speech text corresponding to the prompt image sequence of the target object by calling the second speech generation model, and sends the target speech text to terminal 400 via network 300 for display through graphical interface 410. At the same time, the target speech text is processed by the speech signal generator of the second speech generation model to generate corresponding speech information, which is then played on the terminal.

[0096] In some embodiments, the terminal 400 calls multiple encoders corresponding to multiple modalities in the first speech generation model to encode the sample data of the received sample object in multiple modalities to obtain a multimodal encoded vector sequence, and calls the decoder in the first speech generation model to decode the multimodal encoded vector sequence to obtain decoded text; then, it determines the probability distribution of the decoded text and the sample data of multiple modalities, and determines the target loss based on the probability distribution; it updates the parameters of the decoder and at least one encoder based on the target loss to obtain a second speech generation model; finally, it generates target speech text corresponding to the prompt image sequence of the target object by calling the second speech generation model, and displays the target speech text through the graphical interface 410. At the same time, the target speech text passes through the speech signal generator of the second speech generation model to generate corresponding speech information, which is then broadcast on the terminal.

[0097] In some embodiments, server 200 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals and servers can be connected directly or indirectly via wired or wireless communication, which is not limited in this embodiment.

[0098] The embodiments of this application can be implemented using artificial intelligence (AI) technology. AI is the theory, methods, techniques, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.

[0099] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0100] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP deals with natural language—the language people use in daily life—and is closely related to linguistics research; it also involves crucial techniques for model training in computer science, mathematics, and artificial intelligence. Pre-trained models evolved from Large Language Models (LLMs) in NLP. After fine-tuning, large language models can be widely applied to downstream tasks. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.

[0101] The electronic device implementing the training method of the speech generation model provided in this application embodiment can be the terminal 400 or server 200 in FIG1. ​​Referring to FIG2A, FIG2A is a schematic diagram of the structure of the electronic device 500-1 provided in this application embodiment. The electronic device 500-1 shown in FIG2A includes: at least one processor 510-1, at least one network interface 520-1, a user interface 530-1, and a memory 550-1. The various components in the electronic device 500-1 are coupled together through a bus system 540-1. It is understood that the bus system 540-1 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 540-1 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, all buses are labeled as bus system 540-1 in FIG2A.

[0102] The processor 510-1 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0103] User interface 530-1 includes one or more output devices 531-1 that enable the presentation of media content, including at least one of one or more speakers and one or more visual displays. User interface 530-1 also includes one or more input devices 532-1, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0104] The memory 550-1 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 550-1 may optionally include one or more storage devices physically located away from the processor 510-1.

[0105] The memory 550-1 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550-1 described in this application embodiment is intended to include any suitable type of memory.

[0106] In some embodiments, memory 550-1 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0107] Operating system 551-1 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, for implementing various basic business functions and handling hardware-based tasks;

[0108] The network communication module 552-1 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520-1, such as Bluetooth, WiFi, and Universal Serial Bus (USB).

[0109] Presentation module 553-1 is used to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531-1 (e.g., a display screen, a speaker, etc.) associated with user interface 530-1;

[0110] The input processing module 554-1 is used to detect and translate one or more user inputs or interactions from one or more input devices 532-1.

[0111] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2A shows a training apparatus 555-1 for a language understanding model stored in memory 550-1. This apparatus can be software in the form of programs and plugins, including the following software modules: an acquisition module 5551-1, an encoding module 5552-1, a decoding module 5553-1, a determination module 5554-1, and a generation module 5555-1. These modules are logically related and can therefore be arbitrarily combined or further split according to their implemented functions. The functions of each module will be described below.

[0112] The electronic device implementing the speech generation method of the speech generation model provided in this application embodiment can be the terminal 400 or server 200 in FIG1. ​​Referring to FIG2B, FIG2B is a schematic diagram of the structure of the electronic device 500-2 provided in this application embodiment. The electronic device 500-2 shown in FIG2B includes: at least one processor 510-2, at least one network interface 520-2, user interface 530-2, and memory 550-2. The various components in the electronic device 500-2 are coupled together through a bus system 540-2. It is understood that the bus system 540-2 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 540-2 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, all buses are labeled as bus system 540-2 in FIG2B.

[0113] The processor 510-2 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor.

[0114] User interface 530-2 includes one or more output devices 531-2 that enable the presentation of media content, including at least one of one or more speakers and one or more visual displays. User interface 530-2 also includes one or more input devices 532-2, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0115] The memory 550-2 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 550-2 may optionally include one or more storage devices physically located away from the processor 510-2.

[0116] The memory 550-2 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550-2 described in this application embodiment is intended to include any suitable type of memory.

[0117] In some embodiments, the memory 550-2 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0118] Operating system 551-2 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic business functions and handle hardware-based tasks.

[0119] The network communication module 552-2 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520-2, such as Bluetooth, WiFi, and Universal Serial Bus (USB).

[0120] Presentation module 553-2 is used to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531-2 (e.g., a display screen, a speaker, etc.) associated with user interface 530-2;

[0121] The input processing module 554-2 is used to detect and translate one or more user inputs or interactions from one or more input devices 532-2.

[0122] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2B shows a training apparatus 555-2 for a language understanding model stored in memory 520-2, which can be software in the form of programs and plug-ins, including the following software modules: an acquisition module 5551-2 and a generation module 5552-2. These modules are logically related and can therefore be arbitrarily combined or further split according to the functions they implement. The functions of each module will be described below.

[0123] In some embodiments, the terminal or server can implement the language understanding model training method provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.

[0124] The following describes the training method of the speech generation model provided in the embodiments of this application. As mentioned above, the electronic device implementing the training method of the speech generation model in the embodiments of this application can be a terminal or a server, or a combination of both. Therefore, the executing entity of each step will not be described again below.

[0125] Referring to Figure 3A, which is a flowchart illustrating the training method of the speech generation model provided in the embodiments of this application, the steps shown in Figure 3A will be explained in conjunction with the steps shown in Figure 3A.

[0126] In step 101, a first speech generation model is obtained, wherein the first speech generation model includes a decoder and multiple encoders corresponding to multiple modalities.

[0127] In some embodiments, the first speech generation model may be a first speech generation model pre-trained locally, or it may be a first speech generation model pre-trained by other devices obtained from other devices, including a decoder and multiple encoders corresponding to multiple modalities.

[0128] As an example, see Figure 4A, which is a first architecture diagram of the speech generation model provided in the embodiments of this application. Figure 4A shows a text decoder, a visual encoder corresponding to the visual modality, and a text encoder corresponding to the text modality.

[0129] In step 102, sample data of multiple modalities are acquired, wherein the sample data of multiple modalities includes a sequence of prompt images of sample objects and the corresponding speech text.

[0130] In some embodiments, sample data for multiple modalities of the sample object includes visual modal sample data, such as a sequence of cue images of the sample object, and text modal sample data, such as speech text corresponding to the cue image sequence. Each cue image in the cue image sequence is a static image showing the hand or lip posture of the sample object when it makes a body movement or lip movement. These static images can capture the visual features of the sample object in a specific scene, providing the model with rich visual information and helping the model better understand the behavior and state of the sample object.

[0131] Here, visual modal sample data refers to data describing sample objects through visual information. Through intuitive methods such as hand gestures, lip movements, and body language, the recipient can perceive and understand the visual expression of the target object's facial expressions, eye contact, and body posture, thereby acquiring the target object's emotional state and intentions. Textual modal sample data refers to data describing sample objects in text form. This textual data can be an explanation or description of the image content or other information related to the image, providing semantic supplementation to the model. By combining visual and textual modal data, the model can more comprehensively understand the characteristics and behaviors of the sample objects.

[0132] As an example, sample data of the visual modality of the sample object can be collected by means of camera shooting, image databases or network sources; audio text related to the sample data of the visual modality of the sample object can be recorded by recording equipment, or audio text matching the sample data of the visual modality of the sample object can be extracted from existing audio databases or corpora.

[0133] In step 103, multiple encoders are invoked to encode the prompt image sequence and the speech text respectively, resulting in a multimodal encoded vector sequence.

[0134] In some embodiments, referring to FIG3B, FIG3B is a schematic diagram of the first process for obtaining a multimodal encoded vector sequence provided by the embodiments of this application. When the multiple modalities include visual modalities and text modalities, step 103 in FIG3A can be implemented by steps 1031A to 1033A in FIG3B, which will be described in detail below.

[0135] In step 1031A, the visual encoder is invoked to encode the prompt image sequence to obtain the visual encoding vector sequence of the visual modality, wherein the visual encoder is the encoder corresponding to the visual modality among multiple encoders.

[0136] In some embodiments, referring to FIG3C, FIG3C is a flowchart of obtaining a visual encoding vector sequence provided by an embodiment of the present application. Step 1031A in FIG3B can be implemented by steps 10311A to 10314A in FIG3C, as described in detail below.

[0137] In step 10311A, multiple prompt images in the prompt image sequence are identified to obtain the location of key points in each prompt image.

[0138] In some embodiments, a deep learning model can be used to perform human pose estimation (HPE) on multiple prompt images in a prompt image sequence to identify the locations of key points on the human body, which are then used as the locations of key points in each prompt image. Here, human pose estimation is a computer vision task aimed at detecting and locating key points (such as the head, shoulders, elbows, knees, etc.) of the human body from images or videos. The location information of these key points can be used to estimate the pose of the entire human body.

[0139] As an example, a deep learning model used to identify multiple prompt images in a prompt image sequence can be OpenPose, DeepPose, PoseNet, or DensePose.

[0140] In step 10312A, a key point sequence is generated based on the location of key points in each prompt image.

[0141] As an example, the location of key points in each prompt image can be the location of lips, fingers, etc. Based on the locations of key points in all prompt images in the corresponding prompt image sequence of the video, a key point sequence can be generated. Referring to Figure 9, based on the prompt image sequence shown in Figure 9, the points in each prompt image that mark the locations of hands and lips are key points. Based on the locations of key points in each prompt image, the key point sequence shown in Figure 9 is obtained.

[0142] In step 10313A, the embedding vector of the keypoint sequence is obtained.

[0143] In some embodiments, a machine learning model can be used to perform pose embedding processing on each keypoint in the keypoint sequence to obtain the embedding vector of each keypoint in the keypoint sequence. The embedding vectors are then concatenated to obtain the embedding vector of the corresponding keypoint sequence.

[0144] As an example, the pose embedding of keypoint sequences can be performed using a recurrent neural network (RNN) or a convolutional neural network (CNN) to obtain the embedding vector of the keypoint sequence.

[0145] In step 10314A, the visual encoder is invoked to encode the embedding vector to obtain the visual encoding vector of the visual modality.

[0146] In some embodiments, a pre-trained machine learning model can be used as a visual encoder to encode the embedding vector to obtain a visual encoding vector of the visual modality.

[0147] As an example, the visual encoder model can be a convolutional neural network model, as shown in Figure 5A, which is a schematic diagram of the structure of the visual encoder model provided in an embodiment of this application. The convolutional neural network shown in Figure 5A includes an input layer, a convolutional layer, a pooling layer, and an output layer (fully connected layer and softmax activation function layer). After the embedding vector of the keypoint sequence of the visual modality is input into the convolutional neural network, the embedding vector is sequentially processed by multiple convolutional layers to obtain the feature map corresponding to the embedding vector. Then, the feature map is sequentially processed by multiple max pooling layers to obtain the feature vector of the visual modality. Finally, the visual encoding vector sequence of the visual modality is output by the output layer.

[0148] This application embodiment converts human actions in a video into a keypoint sequence of a visual modality, extracting key information from the prompt image in the form of a keypoint sequence. These keypoints can capture information such as limb movements or lip movements of the human body in the image, thus preserving the key information in the image in the form of a keypoint sequence. This clearly describes changes in human posture and movement. Compared to directly processing the original image or video data, the keypoint sequence greatly reduces data complexity. Through the above method, the most critical information in the image is extracted, redundant data is reduced, computational efficiency is improved, and subsequent analysis and recognition of the visual modality prompt image features are facilitated, effectively utilizing the characteristic information in the prompt image.

[0149] Please refer to Figure 3B for further explanation following step 1031A above.

[0150] In step 1032A, a text encoder is invoked to encode multiple morphemes in the speech text to obtain a text encoding vector sequence of the text modality. The text encoding vector sequence includes text encoding vectors corresponding to multiple morphemes, and the text encoder is the encoder corresponding to the text modality among multiple encoders.

[0151] In some embodiments, the speech text is segmented into multiple morphemes (e.g., words, phrases, or sentences) by language segmentation. The text encoder can be a machine learning model that encodes the multiple morphemes in the speech text separately, converts the morphemes into text encoding vectors, and combines the text encoding vectors corresponding to the morphemes in the speech text to obtain a text encoding vector sequence of the text modality.

[0152] As an example, machine learning models used to encode speech text can be recurrent neural networks, convolutional neural networks, Transformers, Bidirectional Encoder Representations from Transformer (BERT), etc.

[0153] Referring to Figure 5B, which is a schematic diagram of the machine learning model provided in this embodiment, Figure 5B illustrates the input embedding layer, positional embedding, self-attention mechanism, feed-forward neural network, and output layer. First, the input embedding layer converts the morphemes of the speech text into fixed-dimensional text embedding vectors, typically using word embedding techniques to map morphemes to vector representations in a continuous space. Second, the positional embedding layer embeds the positional information of each morpheme into the feature vector, obtaining the positional embedding vector. Then, through the self-attention mechanism, the importance of each morpheme in the context is modeled, and semantic relationships are learned. The feed-forward neural network layer performs nonlinear transformations and mappings on the positional embedding vector of each morpheme, thereby extracting higher-level semantic and contextual information. Finally, the entire text encoding vector sequence is generated based on the positional embedding vectors of each morpheme.

[0154] In step 1033A, the visual encoding vector sequence and the text encoding vector sequence are used as the first multimodal encoding vector sequence, wherein the first multimodal encoding vector sequence is used for decoding by the decoder.

[0155] In some embodiments, when the multiple modalities include a visual modality and a text modality, the multimodal encoding vector sequence includes a visual encoding vector sequence corresponding to the visual modality and a text encoding vector sequence corresponding to the text modality.

[0156] This application embodiment encodes multiple morphemes in the speech text into a text encoding vector sequence of the text modality, so that the encoding vector of each morpheme can capture its semantic information and contextual relationship, thereby better expressing the meaning of the entire text and improving the accuracy of subsequent speech generation.

[0157] In some embodiments, referring to FIG3D, FIG3D is a second flowchart of obtaining a multimodal encoded vector sequence provided by an embodiment of the present application. When the multiple modalities also include a speech modality, the sample data of the multiple modalities also include the speech signal features of the speech modality. Step 103 in FIG3A can be implemented by steps 1031B to 1032B in FIG3D, which will be described in detail below.

[0158] In step 1031B, the speech encoder is invoked to encode the speech signal features to obtain the speech encoding vector sequence of the speech mode, wherein the speech encoder is the encoder corresponding to the speech mode among multiple encoders.

[0159] As an example, see Figure 4B, which is a second architecture diagram of the speech generation model provided in the embodiments of this application. Figure 4B shows a text decoder, a visual encoder corresponding to the visual modality, a text encoder corresponding to the text modality, and a speech encoder corresponding to the speech modality.

[0160] In some embodiments, the speech encoder is invoked to encode the features of the speech signal to obtain a speech encoding vector sequence of the speech modality. This can be achieved as follows: First, the original speech signal is preprocessed, such as noise removal (eliminating background noise and retaining a clean speech signal), filtering (removing noise within a specific frequency range or enhancing the speech signal within a specific frequency range), and volume normalization. Second, the preprocessed speech signal is divided into short time frames, with a commonly used frame length of 20-40 milliseconds, typically using 50% or 75% overlap. Then, the speech signal of each frame is multiplied by a window function, such as a Hamming window or a rectangular window, to reduce boundary artifacts between frames. A Fast Fourier Transform (FFT) is applied to each frame to convert the time-domain signal into a frequency-domain representation. Finally, the speech encoding vector is extracted from the spectrum and concatenated to obtain the speech encoding vector sequence.

[0161] As an example, a speech encoder model used to encode features of a speech signal can be a Mel-filterbank extractor.

[0162] In step 1032B, the visual encoding vector sequence, the text encoding vector sequence, and the speech encoding vector sequence are used as the second multimodal encoding vector sequence, wherein the second multimodal encoding vector sequence is used to replace the first multimodal encoding vector sequence for decoding by the decoder.

[0163] In some embodiments, when the multiple modalities include a visual modality, a text modality, and a speech modality, the multimodal coding vector sequence includes a visual coding vector sequence corresponding to the visual modality, a text coding vector sequence corresponding to the text modality, and a speech coding vector sequence corresponding to the speech modality.

[0164] By introducing the encoded vector sequence of the speech modality, the speech generation model can more comprehensively capture the correlations between multimodal information. The addition of speech signal features enables the model to better understand the semantic alignment between speech content and visual and textual information. Fusing the speech encoded vector sequence with encoded vector sequences from other modalities to form a unified multimodal representation allows the speech generation model to process multimodal data more efficiently. Especially in tasks requiring simultaneous processing of multiple modalities, this significantly improves the model's performance in multimodal tasks, enhances its semantic understanding ability, improves generation quality, and increases the robustness and adaptability of the speech generation model.

[0165] In some embodiments, referring to FIG3E, FIG3E is a schematic diagram of the third process for obtaining a multimodal coding vector sequence provided by the embodiments of this application. When the multiple modalities also include a speech modality, step 103 in FIG3A can also be implemented by steps 1031C to 1034C in FIG3E, which will be described in detail below.

[0166] In step 1031C, multiple encoders are invoked to encode the prompt image sequence and the speech text respectively, to obtain the visual encoding vector sequence of the visual modality, the text encoding vector sequence of the text modality, and the speech encoding vector sequence of the speech modality.

[0167] In some embodiments, as described above, when multiple modalities include a visual modality, a text modality, and a speech modality, a visual encoder can be invoked to encode the sample data of the visual modality to obtain a visual encoding vector sequence of the visual modality; a text encoder can be invoked to encode the sample data of the text modality to obtain a text encoding vector sequence of the text modality; and a speech encoder can be invoked to encode the sample data of the speech modality to obtain a speech encoding vector sequence of the speech modality.

[0168] In some embodiments, the sample data of multiple modalities may further include speech signal features of the speech modality; see Figure 3F, which is a schematic diagram of the fourth process for obtaining the multimodal coding vector sequence provided by the embodiments of this application. Step 1031C in Figure 3E can be implemented by steps 10311C to 10313C in Figure 3F, as described in detail below.

[0169] In step 10311C, the visual encoder is invoked to encode the prompt image sequence to obtain the visual encoding vector sequence of the visual modality, wherein the visual encoder is the encoder corresponding to the visual modality among multiple encoders.

[0170] In some embodiments, step 10311C above can be achieved by performing the following process:

[0171] First, the multiple prompt images in the prompt image sequence are identified separately to obtain the location of key points in each prompt image.

[0172] In some embodiments, a deep learning model can be used to perform human pose estimation (HPE) on multiple prompt images in a prompt image sequence to identify the locations of key points on the human body, which are then used as the locations of key points in each prompt image. Here, human pose estimation is a computer vision task aimed at detecting and locating key points (such as the head, shoulders, elbows, knees, etc.) of the human body from images or videos. The location information of these key points can be used to estimate the pose of the entire human body.

[0173] As an example, a deep learning model used to identify multiple prompt images in a sequence of prompt images can be an open pose estimation, deep pose estimation, pose net, or dense pose.

[0174] Secondly, a key point sequence is generated based on the location of key points in each prompt image.

[0175] As an example, the location of key points in each prompt image can be the location of lips, fingers, etc. Based on the locations of key points in all prompt images in the corresponding prompt image sequence of the video, a key point sequence can be generated. Referring to Figure 9, based on the prompt image sequence shown in Figure 9, the points in each prompt image that mark the locations of hands and lips are key points. Based on the locations of key points in each prompt image, the key point sequence shown in Figure 9 is obtained.

[0176] Then, obtain the embedding vector of the keypoint sequence.

[0177] In some embodiments, a machine learning model can be used to perform pose embedding processing on each keypoint in the keypoint sequence to obtain the embedding vector of each keypoint in the keypoint sequence. The embedding vectors are then concatenated to obtain the embedding vector of the corresponding keypoint sequence.

[0178] As an example, the keypoint sequence can be embedded in pose using a recurrent neural network or a convolutional neural network to obtain the embedding vector of the keypoint sequence.

[0179] Finally, the visual encoder is invoked to encode the embedding vector, resulting in the visual encoded vector of the visual modality.

[0180] In some embodiments, a pre-trained machine learning model can be used as a visual encoder to encode the embedding vector to obtain a visual encoding vector of the visual modality.

[0181] As an example, the visual encoder model can be a convolutional neural network model, as shown in Figure 5A, which is a schematic diagram of the structure of the visual encoder model provided in an embodiment of this application. The convolutional neural network shown in Figure 5A includes an input layer, a convolutional layer, a pooling layer, and an output layer. After the embedding vector of the keypoint sequence of the visual modality is input into the convolutional neural network, the embedding vector is sequentially processed by multiple convolutional layers to obtain the feature map corresponding to the embedding vector. Then, the feature map is sequentially processed by multiple max pooling layers to obtain the feature vector of the visual modality. Finally, the output layer outputs the visual encoding vector sequence of the visual modality.

[0182] This application embodiment converts human actions in a video into a keypoint sequence of a visual modality, extracting key information from the prompt image in the form of a keypoint sequence. These keypoints can capture information such as limb movements or lip movements of the human body in the image, thus preserving the key information in the image in the form of a keypoint sequence. This clearly describes changes in human posture and movement. Compared to directly processing the original image or video data, the keypoint sequence greatly reduces data complexity. Through the above method, the most critical information in the image is extracted, redundant data is reduced, computational efficiency is improved, and subsequent analysis and recognition of the visual modality prompt image features are facilitated, effectively utilizing the characteristic information in the prompt image.

[0183] In step 10312C, a text encoder is invoked to encode multiple morphemes in the speech text to obtain a text encoding vector sequence of the text modality. The text encoding vector sequence includes text encoding vectors corresponding to multiple morphemes, and the text encoder is the encoder corresponding to the text modality among multiple encoders.

[0184] In some embodiments, step 10312C above can be implemented by performing the following process: splitting the speech text into multiple morphemes (e.g., words, phrases, or sentences) through language segmentation. The text encoder can be a machine learning model, which encodes the multiple morphemes in the speech text separately, converts the morphemes into text encoding vectors, and combines the text encoding vectors corresponding to the morphemes in the speech text to obtain a text encoding vector sequence of the text modality.

[0185] As an example, machine learning models used for encoding speech text can be recurrent neural networks, convolutional neural networks, Transformers, Transformer-based bidirectional encoder representations, etc.

[0186] Referring to Figure 5B, which is a schematic diagram of the machine learning model provided in this embodiment, Figure 5B illustrates the input embedding layer, position embedding, self-attention mechanism, feedforward neural network, and output layer. First, the input embedding layer converts the morphemes of the speech text into fixed-dimensional text embedding vectors, typically using word embedding techniques to map morphemes to vector representations in a continuous space. Second, the position embedding layer embeds the position information of each morpheme into the feature vector, obtaining the position embedding vector. Then, through the self-attention mechanism, the importance of each morpheme in the context is modeled, and semantic relationships are learned. The feedforward neural network layer performs nonlinear transformations and mappings on the position embedding vector of each morpheme, thereby extracting higher-level semantic and contextual information. Finally, the entire text encoding vector sequence is generated based on the position embedding vectors of each morpheme.

[0187] In step 10313C, the speech encoder is invoked to encode the speech signal features to obtain the speech encoding vector sequence of the speech modality, wherein the speech encoder is the encoder corresponding to the speech modality among multiple encoders.

[0188] In some embodiments, the speech encoder is invoked to encode the features of the speech signal to obtain a speech coding vector sequence of the speech modality. This can be achieved as follows: First, the original speech signal is preprocessed, such as noise removal, filtering, and volume normalization. Second, the preprocessed speech signal is divided into short frames, with a commonly used frame length of 20-40 milliseconds and typically 50% or 75% overlap. Then, the speech signal of each frame is multiplied by a window function, such as a Hamming window or a rectangular window, to reduce boundary artifacts between frames. A fast Fourier transform is applied to each frame to convert the time-domain signal into a frequency-domain representation. Finally, the speech coding vector is extracted from the spectrum and concatenated to obtain the speech coding vector sequence.

[0189] Please refer to Figure 3E for a continuation of step 1031C from the previous section.

[0190] In step 1032C, the visual encoding vector sequence and the speech encoding vector sequence are concatenated to obtain a concatenated encoding vector sequence.

[0191] In some embodiments, the feature vectors in the visual coding vector sequence and the feature vectors in the speech coding vector sequence can be concatenated to obtain a concatenated coding vector. The concatenated coding vectors corresponding to each position can be combined to obtain a concatenated coding vector sequence.

[0192] As an example, if the visual encoding vectors in the visual encoding vector sequence output by the visual encoder are The speech encoding vectors in the speech encoding vector sequence output by the speech encoder are Where f represents the length of the encoded vector sequence, and D represents the dimension of the encoded vectors in the encoded vector sequence, then the visual encoded vector sequence and the speech encoded vector sequence are concatenated to obtain the concatenated encoded vector.

[0193] In step 1033C, the concatenated encoded vector sequence is subjected to feature mapping to obtain the fused encoded vector sequence.

[0194] As an example, in the above example, the concatenated encoded vector in the concatenated encoded vector sequence has a dimension of 2D. By performing feature mapping on the concatenated encoded vector sequence, the dimension of the concatenated encoded vector is transformed from 2D to D, resulting in the fused encoded vector. The fused coding vectors at each position are combined to obtain a fused coding vector sequence.

[0195] In step 1034C, the fused encoding vector sequence and the text encoding vector sequence are used as the third multimodal encoding vector sequence, which is used to replace the second multimodal encoding vector sequence for decoding by the decoder.

[0196] As an example, refer to Figure 4C, which is a schematic diagram of the third architecture of the speech generation model provided in the embodiment of this application. Figure 4C shows a text decoder, a visual encoder corresponding to the visual modality, a text encoder corresponding to the text modality, and a speech encoder corresponding to the speech modality. As shown in Figure 4C, after feature concatenation mapping of the visual encoding vector sequence output by the visual encoder and the speech encoding vector sequence output by the speech encoder, a fused encoding vector sequence is obtained, which is then combined with the output of the text encoder as the third multimodal encoding vector sequence.

[0197] This application embodiment calls the encoders corresponding to multiple modalities in the first speech generation model to encode sample data of multiple different modalities, obtaining a multimodal encoding vector sequence. Furthermore, it provides three different modal combination encoding methods, enabling the first speech generation model to fully learn the features of different modalities of visual, text, and speech modes. This improves the speech generation model's ability to perform speech generation tasks and increases the accuracy of speech generation. At the same time, the combination of different modalities ensures that the execution of speech generation tasks can adapt to different application scenarios, thereby enhancing the user experience.

[0198] Please refer to Figure 3A for further explanation following step 103 above.

[0199] In step 104, the decoder is invoked based on the multimodal encoded vector sequence to perform decoding, thereby obtaining the decoded text.

[0200] In some embodiments, the first speech generation model is a pre-trained speech recognition model, and after pre-training, the representation space of the text modality and the representation space of the speech modality of the first speech generation model have been unified to the same representation space.

[0201] Here, representation space refers to the vector space composed of feature vectors or feature representations, used to describe and represent different attributes, features, or information of data. In different modalities, the representation space embodies the features of different types of data and their representations in the vector space. For example, the representation space of the visual modality is a vector space composed of visual feature vectors, while the representation space of the text modality is a vector space composed of text feature vectors.

[0202] Referring to Figure 3G, which is a flowchart illustrating the alignment of the length of the encoded vector sequence provided in an embodiment of this application, before executing step 104 of Figure 3A, the length of the multimodal encoded vector sequence can be aligned through steps 201 to 202 of Figure 3G, as described in detail below.

[0203] In step 201, the visual encoding vector sequence and the text encoding vector sequence are mapped to obtain the vector sequence representation of the visual encoding vector sequence and the text encoding vector sequence in the representation space.

[0204] In some embodiments, a linear mapping function is used to map the visual encoded vector sequence and the text encoded vector sequence to obtain a vector sequence representation of the visual encoded vector sequence and the text encoded vector sequence in the same representation space.

[0205] As an example, refer to Figure 6, which is a schematic diagram of the alignment of the encoding vector sequence provided in the embodiment of this application. The visual encoding vector sequence is represented as “E1 E2 E3 E3 E4 E5 E5 E6” in the representation space, and the text encoding vector sequence is represented as “T1 T2 T3 T4 T5 T6” in the representation space.

[0206] In step 202, the length of the vector sequence representation of the visual encoding vector sequence is adjusted to be the same as the length of the vector sequence representation of the text encoding vector sequence.

[0207] As an example, in the example above, the length of the vector sequence representation of the visual encoding vector sequence is 8, and the length of the vector sequence representation of the text encoding vector sequence is 6. Since there is redundant information "E3" and "E5" in the vector sequence representation of the visual encoding vector sequence, the redundant information in the vector sequence representation of the visual encoding vector sequence is deleted, and the vector sequence representation of the visual encoding vector sequence is obtained as "E1 E2 E3 E4 E5 E6". At this time, the length of the vector sequence representation of the visual encoding vector sequence is the same as the length of the vector sequence representation of the text encoding vector sequence.

[0208] The training process of the first speech generation model in this embodiment is based on a pre-trained speech recognition model. Since the representation space of the text modality and the representation space of the speech modality have been unified into the same representation space, the conversion function from prompt image sequence to speech text can be realized on the basis of the original function of the pre-trained speech recognition model through fine-tuning. At the same time, multiple functions of the pre-trained speech recognition model itself can be reused, such as recognizing the prompt image sequence of the target object, translating the text, and generating speech, avoiding repeated development and improving development efficiency.

[0209] Aligning the feature sequence lengths of two modalities makes them more consistent in the feature space, which is beneficial for speech generation models to perform speech generation tasks. This process can improve the overall efficiency and performance of speech generation models when processing multimodal data.

[0210] Please refer to Figure 3A for further explanation following step 104 above.

[0211] In step 105, the probability distribution of the decoded text and sample data of multiple modalities is determined, and the target loss is determined based on the probability distribution.

[0212] In some embodiments, determining the probability distribution of the decoded text and sample data of multiple modalities in step 105 can be achieved by: determining the probability distribution of the encoded vector sequence of sample data of at least two modalities relative to preset conditions, wherein different modalities correspond to different preset conditions; determining the probability distribution of the speech text relative to sample data of at least one modality; and determining the probability distribution of the decoded text relative to sample data of at least one modality.

[0213] Here, the probability distribution of the encoded vector sequences of sample data from at least two modalities relative to preset conditions is determined, where different modalities correspond to different preset conditions, which can specifically depend on the data involved in the model's encoding structure; the probability distribution of the speech text relative to sample data from at least one modality is determined, that is, the conditional probability distribution of the speech text conditioned on sample data from at least one modality is determined; the probability distribution of the decoded text relative to sample data from at least one modality is determined, that is, the conditional probability distribution of the decoded text conditioned on sample data from at least one modality is determined. Here, different modalities correspond to different conditions, which can specifically depend on the model's encoding structure.

[0214] For example, the speech generation model architecture shown in Figure 4A illustrates a text decoder, a visual encoder corresponding to the visual modality, and a text encoder corresponding to the text modality. The speech generation model in Figure 4A only includes two modalities: visual and text. Therefore, the probability distribution of the decoded text and sample data from multiple modalities is determined based on the visual and text modalities, and the target loss is determined based on the probability distribution. For the architecture in Figure 4A, the preset conditions corresponding to the text modality can be the decoded text and the speech text; the preset conditions corresponding to the visual modality can be the decoded text and a sequence of prompt image sequences; the condition corresponding to the speech text is a sequence of prompt image sequences; and the condition corresponding to the decoded text is a sequence of prompt image sequences.

[0215] For example, the speech generation model architecture shown in Figure 4B illustrates a text decoder, a visual encoder corresponding to the visual modality, a text encoder corresponding to the text modality, and a speech encoder corresponding to the speech modality. Compared to Figure 4A, the speech generation model in Figure 4B further incorporates language modalities for training. Therefore, based on the speech and text modalities, the probability distribution of the decoded text and sample data from multiple modalities is determined, and then the target loss is determined based on the probability distribution. For the architecture in Figure 4B, the preset conditions corresponding to the text modality can be decoded text and speech text; the preset conditions corresponding to the speech modality can be decoded text and speech signal features; the condition corresponding to the speech text is speech signal features; and the condition corresponding to the decoded text is speech signal features.

[0216] For example, in the speech generation model architecture shown in Figure 4C, the visual encoding vector sequence output by the visual encoder and the speech encoding vector sequence output by the speech encoder are concatenated and mapped to obtain a fused encoding vector sequence, which is then combined with the output of the text encoder as a third multimodal encoding vector sequence. Therefore, the probability distribution of the decoded text and sample data from multiple modalities is determined based on the visual modality, speech modality, and text modality, and the target loss is determined based on the probability distribution. For the architecture in Figure 4C, the preset conditions corresponding to the text modality can be the decoded text and the speech text; the preset conditions corresponding to the fused encoding vector sequence can be the decoded text, the prompt image sequence, and the speech signal features; the conditions corresponding to the speech text are the prompt image sequence and the speech signal features; and the conditions corresponding to the decoded text are the prompt image sequence and the speech signal features.

[0217] This application's embodiments determine the probability distribution of encoded vector sequences of sample data from at least two modalities relative to preset conditions. By setting different preset conditions for different modalities and calculating the probability distribution of their encoded vector sequences, a deeper understanding of the characteristics and intrinsic relationships of each modal's data can be achieved. For example, in cross-modal image-text retrieval, by constructing semantic spaces for different modalities and learning the sample semantic distribution, the semantics of samples belonging to different modalities can be more accurately aligned. Different modalities correspond to different preset conditions, enabling the model to better adapt to the characteristics of various modal data, thereby improving the model's generalization ability in multimodal tasks. By calculating the probability distribution of speech text and modal sample data, the relationship between speech and modalities such as vision and text can be modeled more accurately, thereby improving the accuracy of speech recognition and generation. By calculating the probability distribution of decoded text and modal sample data, the consistency between decoded text and input modal data can be ensured, improving decoding quality. For example, in text generation tasks, by modeling conditional probability distributions, text that better conforms to context and semantic requirements can be generated. Through the above methods, more intelligent and flexible multimodal interaction can be achieved, improving the convenience of interaction and user experience.

[0218] In some embodiments, determining the target loss based on the probability distribution in step 105 can be achieved in the following ways:

[0219] For encoded vector sequences of sample data from at least two modalities, the difference in probability distribution of encoded vector sequences of different modalities is determined based on the probability distribution of encoded vector sequences of sample data from at least two modalities relative to preset conditions, and used as the sub-target loss of encoded vector sequences; based on the probability distribution of speech text relative to sample data from at least one modality, the sub-target loss of speech text negatively correlated with the probability distribution is determined; based on the probability distribution of decoded text relative to sample data from at least one modality, the sub-target loss of decoded text negatively correlated with the probability distribution is determined.

[0220] For example, in the architecture shown in Figure 4A, the difference in probability distribution between the encoded vector sequences of visual and text modal samples relative to preset conditions is determined, and used as the sub-target loss of the encoded vector sequences; the sub-target loss of speech text negatively correlated with the probability distribution of sample data of visual and text modal samples is determined; the sub-target loss of speech text negatively correlated with the probability distribution of sample data of decoded text relative to visual modal samples is determined; the sub-target loss of speech text negatively correlated with the probability distribution of sample data of decoded text is determined; the sub-target loss of the encoded vector sequences, the sub-target loss of speech text, and the sub-target loss of decoded text are fused to obtain the target loss.

[0221] For example, in the architecture shown in Figure 4B, the difference in probability distribution between the encoded vector sequences of the speech modality and the text modality is determined based on the probability distribution of the encoded vector sequences relative to preset conditions, and is used as the sub-target loss of the encoded vector sequences; the sub-target loss of the speech text is determined based on the probability distribution of the speech text relative to the sample data of the speech modality; the sub-target loss of the speech text is determined based on the probability distribution of the decoded text relative to the sample data of the speech modality; the sub-target loss of the encoded vector sequences, the sub-target loss of the speech text, and the sub-target loss of the decoded text are fused to obtain the target loss.

[0222] For example, in the architecture shown in Figure 4C, based on the probability distribution of the encoded vector sequences of sample data from the visual, speech, and text modal modes relative to preset conditions, the difference in probability distribution between the fused encoded vector sequence and the encoded vector sequence of the text modal is determined as the sub-target loss of the encoded vector sequence; based on the probability distribution of the speech text relative to the sample data of the visual and speech modal modes, the sub-target loss of the speech text negatively correlated with the probability distribution is determined; based on the probability distribution of the decoded text relative to the sample data of the visual and speech modal modes, the sub-target loss of the speech text negatively correlated with the probability distribution is determined; the sub-target loss of the encoded vector sequence, the sub-target loss of the speech text, and the sub-target loss of the decoded text are fused to obtain the target loss.

[0223] The following section will provide a detailed explanation of the three different architectures of the speech generation model shown in Figures 4A to 4C.

[0224] Regarding the architecture of the speech generation model shown in Figure 4A, refer to Figure 3H. Figure 3H is a schematic diagram of the first process for determining the probability distribution provided in the embodiment of this application. The "determining the probability distribution of the decoded text and the sample data of multiple modalities" in step 105 of Figure 3A can be achieved by determining the probability distribution of the encoded vector sequence of the sample data of the visual modality and the text modality relative to the preset conditions, determining the probability distribution of the speech text relative to the prompt image sequence, and determining the probability distribution of the decoded text relative to the prompt image sequence, as described above for the encoded vector sequence of the sample data of the visual modality and the text modality. The following is a detailed explanation in conjunction with steps 1051A to 1053A of Figure 3H.

[0225] In step 1051A, a first probability distribution of the text encoded vector sequence relative to the decoded text and the speech text is determined, and a second probability distribution of the visual encoded vector sequence relative to the decoded text and the prompt image sequence is determined.

[0226] In some embodiments, step 1051A above can be implemented as follows: A pre-trained first probability distribution model is invoked to determine a first probability distribution of the text encoded vector sequence relative to the decoded text and the speech text. The first probability distribution model is trained as follows: First sample data is acquired, including sample decoded text and sample speech text, and a reference probability distribution of the sample text encoded vector sequence relative to the sample decoded text and the sample speech text; the sample decoded text, sample speech text, and sample text encoded vector sequence are used as input to the initialized first probability distribution model, and the predicted probability distribution of the text encoded vector sequence relative to the sample decoded text and the sample speech text is output; a loss value is determined based on the difference between the predicted probability distribution and the reference probability distribution; and the parameters of the first probability distribution model are updated according to the loss value.

[0227] By calculating the probability distribution of the text encoding vector sequence relative to the decoded text and speech text, semantic information from different modalities can be more accurately aligned. This alignment method ensures that text, speech, and visual information maintain consistency and relevance during generation. The probability distribution model can capture the complex dependencies between the text encoding vector and the decoded and speech texts, thereby improving the model's ability to understand the semantics of multimodal data.

[0228] As an example, the first probability distribution of the text encoded vector sequence relative to the decoded text and the speech text can be represented as follows: The first probability distribution model can be a Conditional Random Field (CRF) or a Hidden Markov Model (HMM).

[0229] As an example, the difference operation can be used to calculate the difference between the predicted probability distribution of the text-encoded vector sequence relative to the decoded text and the reference probability distribution, and the square or absolute value of the difference can be taken as the loss value; or the exponential operation can be used to calculate the cross-entropy between the predicted probability distribution of the text-encoded vector sequence relative to the decoded text and the reference probability distribution, and take it as the loss value. For example, if the predicted probability distribution of the text-encoded vector sequence relative to the decoded text and the reference probability distribution is p1, and the reference probability distribution is p2, then the difference between the predicted probability distribution of the text-encoded vector sequence relative to the decoded text and the reference probability distribution is expressed as p1-p2, and the loss value can be expressed as |p1-p2|.

[0230] In some embodiments, a second probability distribution of the visual encoded vector sequence relative to the decoded text and prompt image sequences can be determined using a pre-trained second probability distribution model; that is, the conditional probability distribution of the visual encoded vector sequence relative to the decoded text and prompt image sequences. The second probability distribution model can be trained as follows: First, a second sample dataset is obtained, including the decoded text, prompt image sequences, and a reference probability distribution of the visual encoded vector sequence relative to the decoded text and prompt image sequences. Then, the decoded text and prompt image sequences from the sample data are used as input to the initialized second probability distribution model, outputting the predicted probability distribution of the visual encoded vector sequence relative to the decoded text and prompt image sequences. Finally, a loss value is determined based on the difference between the predicted probability distribution of the visual encoded vector sequence relative to the decoded text and prompt image sequences and the reference probability distribution, and the parameters of the second probability distribution model are updated based on the loss value using a backpropagation algorithm.

[0231] As an example, the second probability distribution of the visual encoded vector sequence relative to the decoded text and prompt image sequences can be represented as follows: The second probability distribution model can be a conditional random field or a hidden Markov model.

[0232] As an example, the difference operation can be used to calculate the difference between the predicted probability distribution of the visual encoded vector sequence relative to the decoded text and prompt image sequence and the reference probability distribution, and the square or absolute value of the difference can be taken as the loss value; or the exponential operation can be used to calculate the cross-entropy between the predicted probability distribution of the visual encoded vector sequence relative to the decoded text and prompt image sequence and the reference probability distribution, and take it as the loss value.

[0233] In step 1052A, a third probability distribution of the speech text relative to the prompt image sequence is determined.

[0234] In some embodiments, a pre-trained third probability distribution model can be used to determine the third probability distribution of the speech text relative to the prompt image sequence, i.e., the conditional probability distribution of the speech text relative to the prompt image sequence. The third probability distribution model can be trained as follows: First, a third sample dataset is obtained, including the prompt image sequence and a reference probability distribution of the speech text relative to the prompt image sequence; then, the prompt image sequence from the sample data is used as input to the initialized third probability distribution model, outputting the predicted probability distribution of the speech text relative to the prompt image sequence; finally, a loss value is determined based on the difference between the predicted probability distribution of the speech text relative to the prompt image sequence and the reference probability distribution, and the parameters of the third probability distribution model are updated based on the loss value using a backpropagation algorithm.

[0235] As an example, the third probability distribution of the voice text relative to the prompt image sequence can be represented as p(x text The third probability distribution model (|xcued-speech) can be a Conditional Random Field (CRF) or a Hidden Markov Model (HMM). Difference operations can be used to calculate the difference between the predicted probability distribution of the speech text relative to the prompt image sequence and the reference probability distribution, and the square or absolute value of the difference is taken as the loss value. For example, if the predicted probability distribution of the speech text relative to the prompt image sequence is p3 and the reference probability distribution is p4, then the difference between the predicted probability distribution and the reference probability distribution is represented as p3-p4, and the loss value can be represented as |p3-p4|. Alternatively, exponential operations can be used to calculate the cross-entropy between the predicted probability distribution of the speech text relative to the prompt image sequence and the reference probability distribution, as the loss value.

[0236] As an example, the speech text includes three parts, namely A, B, and C. The reference probability distribution of the speech text relative to the prompt image sequence is represented as p1 = [0.5, 0.3, 0.2], and the predicted probability distribution is represented as p2 = [0.4, 0.4, 0.2]. The formula for cross-entropy is expressed as formula (1):

[0237] Where H(p1,p2) is the cross-entropy between the predicted probability distribution of the speech text relative to the prompt image sequence and the reference probability distribution, i.e., the loss value, n is the number of speech texts, and p... 1,i p is the probability of the i-th speech text relative to the prompt image sequence in the reference probability distribution. 2,i It is the probability of the i-th speech text relative to the prompt image sequence in the predicted probability distribution.

[0238] The loss value H(p1,p2) = -(0.5log(0.4) + 0.3log(0.4) + 0.2log(0.2)) = 0.6388.

[0239] In step 1053A, when the target language of the speech generation task of the first speech generation model is different from the source language, a fourth probability distribution of the decoded text relative to the prompt image sequence is determined.

[0240] In some embodiments, when the target language of the speech generation task of the first speech generation model is different from the source language, for example, the target language of the speech generation task of the first speech generation model is English and the source language is Chinese, then it is necessary to first translate the Chinese source language into English.

[0241] In some embodiments, a pre-trained fourth probability distribution model can be used to determine the fourth probability distribution of the decoded text relative to the prompt image sequence, i.e., the conditional probability distribution of the decoded text relative to the prompt image sequence. The fourth probability distribution model can be trained as follows: First, a fourth sample dataset is obtained, including the prompt image sequence, the source speech text, and a reference probability distribution of the decoded text relative to the prompt image sequence; then, the prompt image sequence and the source speech text in the sample data are used as input to the initialized fourth probability distribution model, outputting the predicted probability distribution of the decoded text relative to the prompt image sequence; finally, a loss value is determined based on the difference between the predicted probability distribution of the decoded text relative to the prompt image sequence and the reference probability distribution, and the parameters of the fourth probability distribution model are updated based on the loss value using a backpropagation algorithm.

[0242] As an example, the fourth probability distribution of the decoded text relative to the prompt image sequence can be represented as follows: The fourth probability distribution model can be a conditional random field or a hidden Markov model. Difference operations can be used to calculate the difference between the predicted probability distribution of the decoded text relative to the prompt image sequence and the reference probability distribution, and the square or absolute value of the difference can be taken as the loss value; or exponential operations can be used to calculate the cross-entropy between the predicted probability distribution of the decoded text relative to the prompt image sequence and the reference probability distribution, and take this as the loss value.

[0243] This application's embodiments ensure consistency and alignment of information across different modalities (text, speech, and vision) during generation by calculating a first probability distribution between the text-encoded vector sequence and the decoded text and speech text, and a second probability distribution between the visual-encoded vector sequence and the decoded text and cue image sequence. Determining a third probability distribution of the speech text relative to the cue image sequence guides the speech generation model to generate speech text that better matches the cue images. When the target language differs from the source language, calculating a fourth probability distribution of the decoded text relative to the cue image sequence helps optimize cross-language speech generation, enabling the speech generation model to better handle cross-language speech synthesis tasks and generate speech that better conforms to the characteristics and semantics of the target language. By considering the probability distributions of multiple modalities, the model can better adapt to different types of input data, including text, speech, and visual information. This multimodal processing approach allows the model to more flexibly generate high-quality output when facing complex input scenarios.

[0244] In some embodiments, referring to Figure 3I, which is a schematic diagram of the first process for determining target loss provided by an embodiment of this application, the "determining target loss based on probability distribution" in step 105 of Figure 3A can be achieved by determining the difference in probability distribution between the encoded vector sequences of visual modality and text modality relative to preset conditions, as described above, and using this difference as the sub-target loss of the encoded vector sequence. Furthermore, the sub-target loss of speech text negatively correlated with the probability distribution is determined based on the probability distribution of speech text relative to visual modality sample data, and the sub-target loss of speech text negatively correlated with the probability distribution is determined based on the probability distribution of decoded text relative to visual modality sample data. This is implemented below in conjunction with steps 1054A to 1057A of Figure 3I, and will be described in detail below.

[0245] In step 1054A, the difference between the first probability distribution and the second probability distribution is determined as the first sub-target loss.

[0246] In some embodiments, the difference between a first probability distribution and a second probability distribution can be determined by divergence calculation. Examples include KL divergence and cross entropy.

[0247] As an example, the loss of the first sub-target can be determined using formula (2):

[0248] in, This is the first probability distribution. For the second probability distribution, y textIt is the decoded text corresponding to the target language, xcued-speech and x text These are the source language prompt image sequence and the corresponding source audio text, θcs E These are the parameters of the visual encoder.

[0249] In step 1055A, the second sub-target loss that is negatively correlated with the third probability distribution is determined.

[0250] As an example, the second sub-target loss that is negatively correlated with the third probability distribution can be determined by the following formula (3).

[0251] Where, θt D These are the parameters of the decoder.

[0252] In step 1056A, the loss of the third sub-target that is negatively correlated with the fourth probability distribution is determined.

[0253] As an example, the loss of the third sub-target that is negatively correlated with the fourth probability distribution can be determined by the following formula (4).

[0254] In step 1057A, the first sub-target loss, the second sub-target loss, and the third sub-target loss are fused to obtain the first target loss.

[0255] As an example, the first target loss is obtained by weighted summing the first sub-target loss, the second sub-target loss, and the third sub-target loss using the formula (5) shown below.

[0256] Here, α is the weighting parameter of the loss for the third sub-target.

[0257] This application's embodiments comprehensively consider information from three modalities: text, speech, and vision, fully leveraging the complementarity among them. Text provides semantic information, speech provides prosody and emotional information, and visual information provides an intuitive description of the scene and context. In this way, the model can generate more natural and accurate speech. Through explicit probability distribution calculation and loss function design, the model can more efficiently learn the relationships between text, speech, and visual information, thereby converging faster during training, helping to reduce training time and computational resource consumption, and improving model development efficiency.

[0258] The following description refers to the architecture of the speech generation modality shown in Figure 4B. Referring to Figure 3J, which is a schematic diagram of the second process for determining the probability distribution provided in the embodiment of this application, the "determining the probability distribution of the decoded text and the sample data of multiple modalities" in step 105 of Figure 3A can also be achieved by determining the probability distribution of the encoded vector sequence of the sample data of the speech modality and the text modality relative to the preset conditions, determining the probability distribution of the speech text relative to the speech signal features, and determining the probability distribution of the decoded text relative to the speech signal features, based on the encoded vector sequence of the sample data of the speech modality and the text modality described above. This is implemented below with reference to steps 1051B to 1053B of Figure 3J, and will be explained in detail below.

[0259] In step 1051B, a fifth probability distribution of the text encoding vector sequence relative to the decoded text and the speech text, and a sixth probability distribution of the speech encoding vector sequence relative to the decoded text and the speech signal features are determined.

[0260] In some embodiments, a fifth probability distribution of the text encoded vector sequence relative to the decoded text and the speech text can be determined using a pre-trained fifth probability distribution model; that is, the conditional probability distribution of the text encoded vector sequence relative to the decoded text and the speech text. Here, the fifth probability distribution is used merely to distinguish it from the first probability distribution mentioned above.

[0261] In some embodiments, a pre-trained sixth probability distribution model can be used to determine the sixth probability distribution of the speech encoded vector sequence relative to the decoded text and speech signal features, i.e., the conditional probability distribution of the speech encoded vector sequence relative to the decoded text and speech signal features. The sixth probability distribution model can be trained as follows: First, a sixth sample dataset is obtained, including decoded text, speech signal features, and a reference probability distribution of the speech encoded vector sequence relative to the decoded text and speech signal features; then, the decoded text and speech signal features in the sample data are used as input to the initialized sixth probability distribution model, outputting the predicted probability distribution of the speech encoded vector sequence relative to the decoded text and speech signal features; finally, a loss value is determined based on the difference between the predicted probability distribution of the speech encoded vector sequence relative to the decoded text and speech signal features and the reference probability distribution, and the parameters of the sixth probability distribution model are updated based on the loss value using a backpropagation algorithm.

[0262] As an example, the sixth probability distribution of the speech encoded vector sequence relative to the features of the decoded text and speech signal can be represented as follows: The sixth probability distribution model can be a conditional random field or a hidden Markov model.

[0263] As an example, the difference operation can be used to calculate the difference between the predicted probability distribution and the reference probability distribution of the speech encoded vector sequence relative to the features of the decoded text and speech signal, and the square or absolute value of the difference can be taken as the loss value; or the exponential operation can be used to calculate the cross-entropy between the predicted probability distribution and the reference probability distribution of the speech encoded vector sequence relative to the features of the decoded text and speech signal, and take it as the loss value.

[0264] In step 1052B, the seventh probability distribution of the speech text relative to the speech signal features is determined.

[0265] In some embodiments, the seventh probability distribution of speech text relative to speech signal features, i.e., the conditional probability distribution of speech text relative to speech signal features, can be determined by a pre-trained seventh probability distribution model. The seventh probability distribution model can be trained as follows: First, a seventh sample dataset is obtained, which includes speech signal features and a reference probability distribution of speech text relative to the speech signal features; then, the speech signal features in the sample data are used as input to the initialized seventh probability distribution model, and the predicted probability distribution of speech text relative to the speech signal features is output; finally, a loss value is determined based on the difference between the predicted probability distribution of speech text relative to the speech signal features and the reference probability distribution, and the parameters of the seventh probability distribution model are updated based on the loss value using a backpropagation algorithm.

[0266] As an example, the seventh probability distribution of speech text relative to the features of the speech signal can be represented as follows: The seventh probability distribution model can be a conditional random field or a hidden Markov model. Difference operations can be used to calculate the difference between the predicted probability distribution of the speech text relative to the speech signal features and the reference probability distribution, and the square or absolute value of the difference can be taken as the loss value; or exponential operations can be used to calculate the cross-entropy between the predicted probability distribution of the speech text relative to the speech signal features and the reference probability distribution, and take this as the loss value.

[0267] In step 1053B, when the target language of the speech generation task of the first speech generation model is different from the source language, the eighth probability distribution of the decoded text relative to the speech signal features is determined.

[0268] In some embodiments, when the target language of the speech generation task of the first speech generation model is different from the source language, the eighth probability distribution of the decoded text relative to the speech signal features, i.e., the conditional probability distribution of the decoded text relative to the speech signal features, can be determined by a pre-trained eighth probability distribution model. The eighth probability distribution model can be trained as follows: First, an eighth sample dataset is obtained, which includes speech signal features, source speech text, and a reference probability distribution of the decoded text relative to the speech signal features; then, the speech signal features and source speech text in the sample data are used as input to the initialized eighth probability distribution model, and the predicted probability distribution of the decoded text relative to the speech signal features is output; finally, a loss value is determined based on the difference between the predicted probability distribution of the decoded text relative to the speech signal features and the reference probability distribution, and the parameters of the eighth probability distribution model are updated based on the loss value using a backpropagation algorithm.

[0269] As an example, the eighth probability distribution of the decoded text relative to the features of the speech signal can be represented as follows: The eighth probability distribution model can be a conditional random field or a hidden Markov model.

[0270] As an example, the difference operation can be used to calculate the difference between the predicted probability distribution of the decoded text relative to the speech signal features and the reference probability distribution, and the square or absolute value of the difference can be taken as the loss value; or the exponential operation can be used to calculate the cross-entropy between the predicted probability distribution of the decoded text relative to the speech signal features and the reference probability distribution, and take it as the loss value.

[0271] By using a pre-trained probability distribution model, the conditional probability relationship between decoded text and speech signal features can be modeled more accurately. This allows the model to capture the complex dependencies between speech signal features and decoded text, thereby improving the accuracy of cross-language speech generation. As a result, the speech generation model can generate text content that better meets the user's needs based on the user's speech signal features, providing personalized services.

[0272] In some embodiments, referring to Figure 3K, which is a schematic diagram of the second process for determining target loss provided by the embodiments of this application, the "determining target loss based on probability distribution" in step 105 of Figure 3A can also be achieved by determining the difference in probability distribution between the encoded vector sequences of the speech modality and the text modality based on the probability distribution of the encoded vector sequences of the sample data of the speech modality and the text modality relative to the preset conditions, as described above, and using this as the sub-target loss of the encoded vector sequence; determining the sub-target loss of the speech text negatively correlated with the probability distribution based on the probability distribution of the sample data of the speech text relative to the speech modality; and determining the sub-target loss of the speech text negatively correlated with the probability distribution based on the probability distribution of the sample data of the decoded text relative to the speech modality. This is implemented in conjunction with steps 1054B to 1057B of Figure 3K, and will be described in detail below.

[0273] In step 1054B, the difference between the fifth probability distribution and the sixth probability distribution is determined as the fourth sub-target loss.

[0274] In some embodiments, the difference between the fifth and sixth probability distributions can be determined by divergence calculations. Examples include KL divergence and cross-entropy.

[0275] As an example, the loss of the fourth sub-target can be determined using formula (6):

[0276] Where ximpaired-speech is the speech signal feature in the sample data, θ isE These are parameters of the speech encoder, y text It is the decoded text corresponding to the target language, x text It is the source speech text. This represents the fifth probability distribution of the text encoding vector sequence relative to the decoded text and the speech text. This is the sixth probability distribution of the speech encoding vector sequence relative to the features of the decoded text and speech signal.

[0277] In step 1055B, the loss of the fifth sub-target that is negatively correlated with the seventh probability distribution is determined.

[0278] As an example, the loss of the fifth sub-target that is negatively correlated with the seventh probability distribution can be determined by the following formula (7).

[0279] In step 1056B, the loss of the sixth sub-target that is negatively correlated with the eighth probability distribution is determined.

[0280] As an example, the loss of the sixth sub-target that is negatively correlated with the eighth probability distribution can be determined by the following formula (8).

[0281] In step 1057B, the fourth sub-target loss, the fifth sub-target loss, and the sixth sub-target loss are fused to obtain the second target loss.

[0282] As an example, the second target loss is obtained by weighted summing of the fourth sub-target loss, the fifth sub-target loss, and the sixth sub-target loss using the formula (9) shown below.

[0283] Here, α is the weighting parameter of the loss for the sixth sub-target.

[0284] This application's embodiments calculate the differences in probability distributions between different modalities (such as the difference between the fifth and sixth probability distributions) and use this as a sub-objective loss for optimization, thereby generating more natural and context-aware speech content. Through explicit probability distribution calculation and loss function design, the model can consider multiple factors simultaneously during training, avoiding overfitting issues that may result from single-objective optimization. The speech generation model can learn the relationships between different modal information more efficiently, thus converging faster during training. This efficient training method reduces training time and computational resource consumption. In summary, by calculating the probability distributions between different modalities and optimizing the objective loss based on these distributions, the performance of multimodal models in terms of consistency, generation quality, robustness, and generalization ability can be significantly improved, making them perform better in complex multimodal tasks.

[0285] The following description refers to the architecture of the speech generation modality shown in Figure 4C. Referring to Figure 3L, which is a schematic diagram of the third process for determining the probability distribution provided in the embodiment of this application, the "determining the probability distribution of the decoded text and the sample data of multiple modalities" in step 105 of Figure 3A can also be achieved by determining the probability distribution of the encoded vector sequences of the sample data of the visual modality, speech modality, and text modality relative to preset conditions through the encoded vector sequences of the sample data of the visual modality, speech modality, and text modality described above; determining the probability distribution of the speech text relative to the prompt image sequence and speech signal features; and determining the probability distribution of the decoded text relative to the prompt image sequence and speech signal features. This is implemented below with reference to steps 1051C to 1053C of Figure 3L, and will be explained in detail below.

[0286] In step 1051C, the ninth probability distribution of the text encoded vector sequence relative to the decoded text and the speech text is determined, as well as the tenth probability distribution of the fused encoded vector sequence relative to the decoded text, the prompt image sequence, and the speech signal features.

[0287] In some embodiments, the ninth probability distribution of the text encoded vector sequence relative to the decoded text and the speech text can be determined by a pre-trained ninth probability distribution model; that is, the conditional probability distribution of the text encoded vector sequence relative to the decoded text and the speech text. Here, the ninth probability distribution is used only to distinguish it from the first probability distribution described above. The training process of the ninth probability distribution model is the same as the training process of the first probability distribution model described above.

[0288] In some embodiments, the tenth probability distribution of the fused encoded vector sequence relative to the decoded text, prompt image sequence, and speech signal features can be determined using a pre-trained tenth probability distribution model; that is, the conditional probability distribution of the fused encoded vector sequence relative to the decoded text, prompt image sequence, and speech signal features. The tenth probability distribution model can be trained as follows: First, a tenth sample dataset is obtained, including decoded text, prompt image sequence, speech signal features, and a reference probability distribution of the fused encoded vector sequence relative to the decoded text, prompt image sequence, and speech signal features. Then, the decoded text, prompt image sequence, and speech signal features in the sample data are used as input to the initialized tenth probability distribution model, outputting the predicted probability distribution of the fused encoded vector sequence relative to the decoded text, prompt image sequence, and speech signal features. Finally, a loss value is determined based on the difference between the predicted probability distribution of the fused encoded vector sequence relative to the decoded text, prompt image sequence, and speech signal features and the reference probability distribution, and the parameters of the tenth probability distribution model are updated based on the loss value using a backpropagation algorithm.

[0289] As an example, the tenth probability distribution of the fused encoded vector sequence relative to the features of the decoded text, the prompt image sequence, and the speech signal can be represented as follows: The tenth probability distribution model can be a conditional random field or a hidden Markov model.

[0290] As an example, the difference operation can be used to calculate the difference between the predicted probability distribution and the reference probability distribution of the fused encoded vector sequence relative to the decoded text, prompt image sequence, and speech signal features, and take the square or absolute value of the difference as the loss value; or the exponential operation can be used to calculate the cross-entropy between the predicted probability distribution and the reference probability distribution of the fused encoded vector sequence relative to the decoded text, prompt image sequence, and speech signal features, and take it as the loss value.

[0291] In step 1052C, the eleventh probability distribution of the speech text relative to the prompt image sequence and speech signal features is determined.

[0292] In some embodiments, the eleventh probability distribution of the speech text relative to the prompt image sequence and speech signal features can be determined by a pre-trained eleventh probability distribution model, i.e., the conditional probability distribution of the speech text relative to the prompt image sequence and speech signal features. The eleventh probability distribution model can be trained as follows: First, an eleventh sample dataset is obtained, which includes the prompt image sequence, speech signal features, and a reference probability distribution of the speech text relative to the prompt image sequence and speech signal features; then, the prompt image sequence and speech signal features in the sample data are used as input to the initialized eleventh probability distribution model, and the predicted probability distribution of the speech text relative to the prompt image sequence and speech signal features is output; finally, a loss value is determined based on the difference between the predicted probability distribution of the speech text relative to the prompt image sequence and speech signal features and the reference probability distribution, and the parameters of the eleventh probability distribution model are updated based on the loss value using the backpropagation algorithm.

[0293] As an example, the eleventh probability distribution of the speech text relative to the prompt image sequence and speech signal features can be represented as p(x text |xcued-speech,ximpaired-speech), the eleventh probability distribution model can be a conditional random field or a hidden Markov model.

[0294] As an example, interpolation can be used to calculate the difference between the predicted probability distribution of the speech text relative to the prompt image sequence and speech signal features and the reference probability distribution, and the square or absolute value of the difference can be taken as the loss value; or exponential operation can be used to calculate the cross-entropy between the predicted probability distribution of the speech text relative to the prompt image sequence and speech signal features and the reference probability distribution, and take this as the loss value. For example, if the predicted probability distribution of the speech text relative to the prompt image sequence and speech signal features is p5, and the reference probability distribution is p6, then the difference between the predicted probability distribution of the speech text relative to the prompt image sequence and speech signal features and the reference probability distribution is represented as p5-p6, and the loss value can be represented as |p5-p6|.

[0295] In step 1053C, when the target language of the speech generation task of the first speech generation model is different from the source language, the twelfth probability distribution of the decoded text relative to the prompt image sequence and speech signal features is determined.

[0296] In some embodiments, when the target language of the speech generation task of the first speech generation model is different from the source language, the twelfth probability distribution of the decoded text relative to the prompt image sequence and speech signal features can be determined by a pre-trained twelfth probability distribution model, i.e., the conditional probability distribution of the decoded text relative to the prompt image sequence and speech signal features. The twelfth probability distribution model can be trained as follows: First, a twelfth sample dataset is obtained, which includes the prompt image sequence and speech signal features, the source speech text, and the reference probability distribution of the decoded text relative to the prompt image sequence and speech signal features; then, the fused features of the prompt image sequence and speech signal features in the sample data, as well as the source speech text, are used as input to the initialized twelfth probability distribution model, and the predicted probability distribution of the decoded text relative to the prompt image sequence and speech signal features is output; finally, the loss value is determined based on the difference between the predicted probability distribution of the decoded text relative to the prompt image sequence and speech signal features and the reference probability distribution, and the parameters of the twelfth probability distribution model are updated based on the loss value using the backpropagation algorithm.

[0297] As an example, the twelfth probability distribution of the decoded text relative to the features of the prompt image sequence and the speech signal can be represented as follows: The twelfth probability distribution model can be a conditional random field or a hidden Markov model.

[0298] As an example, the difference operation can be used to calculate the difference between the predicted probability distribution of the decoded text relative to the features of the prompt image sequence and the reference probability distribution, and the square or absolute value of the difference can be taken as the loss value; or the exponential operation can be used to calculate the cross-entropy between the predicted probability distribution of the decoded text relative to the features of the prompt image sequence and the reference probability distribution, and take it as the loss value.

[0299] In some embodiments, referring to Figure 3M, which is a schematic diagram of the third process for determining target loss provided by the embodiments of this application, the "determining target loss based on probability distribution" in step 105 of Figure 3A can also be achieved by determining the difference in probability distribution between the fused encoding vector sequence and the text encoding vector sequence based on the probability distribution of the encoded vector sequence of sample data based on visual modality, speech modality, and text modality relative to preset conditions, as described above, and using it as the sub-target loss of the encoded vector sequence; determining the sub-target loss of speech text negatively correlated with the probability distribution based on the probability distribution of speech text relative to the sample data of visual modality and speech modality; and determining the sub-target loss of speech text negatively correlated with the probability distribution based on the probability distribution of decoded text relative to the sample data of visual modality and speech modality. This is implemented in conjunction with steps 1054C to 1057C of Figure 3M, and will be described in detail below.

[0300] In step 1054C, the difference between the ninth probability distribution and the tenth probability distribution is determined as the seventh sub-target loss.

[0301] In some embodiments, the difference between the ninth and tenth probability distributions can be determined by divergence calculations. Examples include KL divergence and cross-entropy.

[0302] As an example, the difference between the ninth and tenth probability distributions can be determined by formula (10) as the loss of the seventh sub-target.

[0303] In step 1055C, the loss of the eighth sub-target that is negatively correlated with the eleventh probability distribution is determined.

[0304] As an example, the loss of the eighth sub-target that is negatively correlated with the eleventh probability distribution can be determined by the following formula (11).

[0305] In step 1056C, the loss of the ninth sub-target that is negatively correlated with the twelfth probability distribution is determined.

[0306] As an example, the loss of the ninth sub-target that is negatively correlated with the twelfth probability distribution can be determined by the following formula (12).

[0307] In step 1057C, the loss of the seventh sub-target, the loss of the eighth sub-target, and the loss of the ninth sub-target are fused to obtain the loss of the third target.

[0308] As an example, the loss of the seventh sub-target, the loss of the eighth sub-target, and the loss of the ninth sub-target are weighted and summed using the formula (13) shown below to obtain the loss of the third target.

[0309] Here, α is the weighting parameter of the loss for the ninth sub-target.

[0310] Please refer to Figure 3A for further explanation following step 105 above.

[0311] In step 106, the parameters of the decoder and at least one encoder are updated based on the target loss to obtain an updated decoder and multiple encoders. The updated decoder and multiple encoders are used to form a second speech generation model, which is used to generate target speech text corresponding to the prompt image sequence of the target object.

[0312] In some embodiments, referring to FIG3N, FIG3N is a first flowchart of the update parameters provided in the embodiments of this application. The "update the parameters of the decoder and at least one encoder based on the target loss" in step 106 of FIG3A can be implemented through steps 1061A to 1062A of FIG3N, which will be described in detail below.

[0313] In step 1061A, the parameters of the visual encoder are updated based on the first sub-target loss.

[0314] As an example, as shown in formula (2) above, the parameters θ of the visual encoder are updated based on the loss of the first sub-target. csE .

[0315] In step 1062A, the parameters of the visual encoder and decoder are updated based on the second sub-target loss and the third sub-target loss.

[0316] As an example, when the target language of the first speech generation model's speech generation task is the same as the source language, the parameters θ of the visual encoder are updated based on the second sub-target loss, as shown in formula (3) above. csE and the decoder parameter θ tD When the target language of the first speech generation model's speech generation task is different from the source language, the parameters θ of the visual encoder are updated based on the third sub-target loss, as shown in formula (4) above. csE and the decoder parameter θ tD .

[0317] The first sub-target loss reflects the difference between the text-encoded vector sequence and the visual-encoded vector sequence. By updating the parameters of the visual encoder based on this loss, it can be ensured that the vectors generated by the visual encoder are better aligned with the text-encoded vectors, thereby improving the consistency between multimodal information. The second sub-target loss is related to the matching degree between the speech text and the prompt image sequence, and the third sub-target loss is related to the negative correlation between the speech text and the prompt image sequence. By combining these two losses to update the parameters of the visual encoder, the visual encoder can better capture the features of the prompt image sequence and generate encoded vectors that better match the speech text and decoded text, further improving the fusion effect of visual information with other modal information. This update method enables the decoder to generate speech text that is more consistent with the semantics of the prompt image, while avoiding the generation of content that does not match the prompt image, thereby improving the quality and accuracy of speech generation. By updating parameters for different sub-target losses, the model can better coordinate the relationship between text, speech, and visual information. The update of the visual encoder enables visual information to be better aligned with text and speech information, while the update of the decoder ensures that the generated speech text can better reflect the content of visual information. This collaborative optimization method can significantly improve the effect of multimodal information fusion and generate more natural and consistent output.

[0318] In some embodiments, referring to FIG30, FIG30 is a second flowchart of the update parameters provided in the embodiments of this application. The "update the parameters of the decoder and at least one encoder based on the target loss" in step 106 of FIG3A can be implemented through steps 1061B to 1062B of FIG30, which will be described in detail below.

[0319] In step 1061B, the parameters of the speech encoder are updated based on the fourth sub-target loss.

[0320] As an example, as shown in formula (6) above, the parameters θ of the speech encoder are updated based on the fourth sub-target loss. isE .

[0321] In step 1062B, the parameters of the speech encoder and decoder are updated based on the fifth sub-target loss and the sixth sub-target loss.

[0322] As an example, when the target language of the first speech generation model's speech generation task is the same as the source language, the parameters θ of the speech encoder are updated based on the fifth sub-target loss, as shown in formula (7) above. isE and the decoder parameter θ tD When the target language of the first speech generation model's speech generation task is different from the source language, the parameters θ of the speech encoder are updated based on the sixth sub-target loss, as shown in formula (8) above. isE and the decoder parameter θ tD.

[0323] In some embodiments, referring to FIG3P, FIG3P is a schematic diagram of the third process of updating parameters provided in the embodiments of this application. The "updating the parameters of the decoder and at least one encoder based on the target loss" in step 106 of FIG3A can be implemented through steps 1061C to 1062C of FIG3P, which will be described in detail below.

[0324] In step 1061C, the parameters of the visual encoder and the speech encoder are updated based on the seventh sub-target loss.

[0325] As an example, as shown in formula (10) above, the parameters θ of the visual encoder are updated based on the loss of the seventh sub-target. csE and the parameters θ of the speech encoder isE .

[0326] In step 1062C, the parameters of the visual encoder, speech encoder, and decoder are updated based on the eighth sub-target loss and the ninth sub-target loss.

[0327] As an example, when the target language of the first speech generation model's speech generation task is the same as the source language, the parameters θ of the visual encoder are updated based on the eighth sub-target loss, as shown in formula (11) above. csE The parameters θ of the voice encoder isE and the decoder parameter θ tD When the target language of the first speech generation model's speech generation task is different from the source language, as shown in formula (12) above, the parameters θ of the visual encoder are updated based on the ninth sub-target loss. csE The parameters θ of the voice encoder isE and the decoder parameter θ tD .

[0328] In some embodiments, the first speech generation model further includes a speech signal generator, wherein the speech signal generator is used to form a second speech generation model with the decoder and the updated plurality of encoders, the second speech generation model being used to generate speech signals.

[0329] As an example, continuing to refer to Figure 4A, the speech signal generator shown in 4A includes a text-to-unit module and a unit-to-speech module, which together with the decoder and multiple updated encoders form a second speech generation model, which is used to generate speech signals.

[0330] By encoding sample data from multiple modalities, including the prompt image sequence (corresponding to the visual modality) and the speech text (corresponding to the text modality), the target loss is calculated by comparing the resulting multimodal encoded vector sequence (i.e., the encoded vector sequence of multiple modalities) with the probability distribution of the decoded text. This allows the target loss to reflect the differences between the representation spaces of multiple modalities in the first speech generation model. Based on the target loss, the parameters of the first speech generation model are updated through backpropagation, enabling the first speech generation model to gradually learn the intrinsic relationship between the information of the visual and text modalities, improving the model's performance and generalization ability. This results in the unification of the representation spaces of the visual and text modalities in the second speech generation model after training. This means that the second speech generation model can map data from different modalities into a common representation space, allowing data from different modalities to have similar semantic expressions in this space. This lays the foundation for accurate speech text generation in the future, ensuring that the target speech text output by the second speech generation model accurately matches the intent of the prompt image of the target object, thereby guaranteeing the accuracy of speech text generation from the prompt image. At the same time, compared with related technologies that require manual text input to generate voice text, it is more convenient, more efficient, and provides a better user experience.

[0331] The following describes the speech generation method of the speech generation model provided in the embodiments of this application. As mentioned above, the electronic device implementing the speech generation method of the speech generation model in the embodiments of this application can be a terminal or a server, or a combination of both. Therefore, the executing entity of each step will not be described again below.

[0332] Referring to Figure 3Q, which is a schematic flowchart of the speech generation method of the speech generation model provided in the embodiment of this application, wherein the speech generation model is the second speech generation model described above, and the following will be explained in conjunction with steps 301 to 302 shown in Figure 3Q.

[0333] In step 301, the prompt image sequence of the target object is obtained.

[0334] In some embodiments, each prompt image in the prompt image sequence displays a static image of the hand or lip posture of the sample object at a certain video frame moment when the sample object makes a limb or lip movement.

[0335] As an example, a sequence of prompt images of the target object can be obtained by using methods such as camera shooting, image databases, or online sources.

[0336] In step 302, the second speech generation model is invoked based on the prompt image sequence to generate target speech text corresponding to the prompt image sequence of the target object.

[0337] In some embodiments, the second speech generation model described above can be invoked based on the prompt image sequence to generate target speech text corresponding to the prompt image sequence of the target object.

[0338] In some embodiments, the target speech text generated based on the speech generation model can be further used to generate the corresponding target speech through the speech signal generator in the speech generation model.

[0339] This application embodiment trains the first speech generation model in three parallel ways based on sample data from multiple modalities of the sample object: 1) Encoding sample data from the visual and text modalities, and calling the text decoder to decode the encoding results of the visual encoder and text encoder respectively, in order to update the parameters of the visual encoder; 2) Encoding sample data from the visual, text, and speech modalities, and calling the text decoder to decode the encoding results of the visual encoder, text encoder, and speech encoder respectively, in order to update the parameters of the speech encoder; 3) Encoding sample data from the visual, text, and speech modalities, fusing the encoding results corresponding to the visual and text modalities, and then calling the text decoder to decode the fused encoding result and the encoding result of the text encoder respectively, in order to update the parameters of the visual encoder and speech encoder. This allows the speech generation model to fully learn the features of different modalities of the visual, text, and speech modalities, which can improve the ability of the speech generation model to perform speech generation tasks and improve the accuracy of speech generation. At the same time, the combination of different modalities can ensure that the execution of speech generation tasks can adapt to different application scenarios and improve the user experience.

[0340] The following will describe an exemplary application of the embodiments of this application in a scenario where fluent speech is provided to target objects with hearing or speech impairments.

[0341] The system collects the visual representation of a target object with hearing or speech impairments from the prompting speech (equivalent to the prompting image sequence mentioned above) via a terminal. After uploading this visual representation to the server, the server uses a trained second speech generation model to recognize the received visual representation, generate target speech text corresponding to the visual representation, and send it to the terminal. The terminal then outputs the speech text corresponding to the target speech text. Alternatively, the server can recognize the visual representation of the target object, generate target speech text corresponding to the visual representation, generate the corresponding speech text, and then send the speech to the terminal for playback. The training process of the speech generation model is explained in detail below.

[0342] For people with hearing or speech impairments, lip reading or sign language are the main means of communication. However, lip reading makes it difficult to distinguish sounds with similar lip shapes, such as [u] and [y]. Sign language requires a long learning period before one can communicate with others. Related technologies use a prompting voice system to encode finger shapes and hand positions, combined with lip reading, to provide users with hearing or speech impairments with a clear visual representation of all phonemes in spoken language.

[0343] Referring to Figure 7, which is a schematic diagram of the prompting speech encoding information provided in an embodiment of this application, as shown in Figure 7, five different hand positions are used to encode the vowel groups in Mandarin, and eight different gestures are used to encode the consonant groups in Mandarin. Through prompting speech, people with hearing impairments can distinguish pronunciations that are indistinguishable when lip-reading by combining hand information. However, the prompting speech system encodes the visual expression provided by the prompting speech user based on the user's hand and lip movements. On the one hand, it cannot output fluent, natural, and universal speech expressions; on the other hand, it is not applicable to non-prompting speech users.

[0344] The related speech engine technology can clone a person's voice with just 15 seconds of speech sample, supporting cross-language communication. Users can choose a personalized, human voice instead of a synthesized one with a distinctly mechanical feel, and the voice maintains consistency across various languages. This technology can support Augmentative & Alternative Communication (AAC) devices, providing personalized voices across multiple languages ​​for users who cannot speak, thus helping those with speech impairments. Examples include providing therapeutic applications for users with speech-related disorders and educational enhancement services for users with learning needs. However, the core of this speech engine technology is a speech synthesis and cloning method that requires users to input text within the application before it can be converted into speech. This communication method cannot support streaming, real-time expression.

[0345] The method provided in this application encodes the visual expression from the prompting speech and maps it to a representation space that is consistent with the mapping space of the text modality and speech modality in the speech recognition translation synthesis engine (equivalent to the speech recognition model above). Based on the prompting speech of the target object with hearing impairment or speech impairment, it can directly generate the corresponding fluent speech, providing the target object with a customizable, natural and fluent voice that can be used across multiple languages, thereby helping the target object to communicate with others without barriers.

[0346] Referring to Figure 8, which is a schematic diagram of the training data provided in an embodiment of this application, the training data includes prompts for the visual modality, phonetic symbols and audio for the speech modality, and text for the text modality. To construct a latent representation space for the visual modality shared with the speech and text modalities, the training data must first be preprocessed and encoded.

[0347] Referring to Figure 9, which is a visual coding model architecture diagram provided in the embodiment of this application, as shown in Figure 9, firstly, the video of the target object is preprocessed to extract the key points of each frame of the video to obtain the key point sequence. The visual modality keypoints (CS keypoints) include the coordinate positions of a total of J keypoints such as lips, fingers and hands.

[0348] Next, the keypoint sequence of each frame of the video is input into the visual encoder. The keypoint information input into the visual encoder (CS Encoder) is represented as follows: Where F represents the total number of video frames, and C represents the coordinate dimension of the keypoints. For example, when the keypoint coordinates are 2D, C=2, and when the keypoint coordinates are 3D, C=3. Then, the keypoint sequence is embedded in pose and encoded using a Transformer block to obtain a visual encoded vector sequence.

[0349] The State-of-the-Art (SOTA) pose detection baseline model can be used as the pre-trained baseline model for the visual encoder. Based on the pre-trained visual encoder obtained by training with high-resource human pose data, the parameters of the pre-trained visual encoder model can be fine-tuned using low-resource prompt speech training data (such as the training data shown in Figure 8).

[0350] To reduce video redundancy while preserving the diversity and integrity of semantic information, related technologies employ clustering pruning techniques. The k-nearest neighbor (DPC-KNN) peak density clustering algorithm is used to cluster the token sequences of the input F frames based on similarity. Based on the criteria of higher density relative to nearest neighbors and greater distance relative to other high-density tokens, the higher-scoring f frames are used as cluster centers, i.e., the key tokens retained after pruning. f is much smaller than the original sequence length F (f < 0.05). <F)。

[0351] Unlike clustering and pruning techniques in related technologies, this application embodiment aligns the visual encoding vector sequence obtained after encoding by the visual encoder to the output feature space of the text encoder when fine-tuning the pre-trained visual encoding model, thereby ensuring the semantic integrity of the former. This includes the following two aspects of processing, see Figure 4A, which is a schematic diagram of the first architecture of the speech generation model (CSeamless) provided in this application embodiment.

[0352] First, sequence length alignment is performed. The visual encoded vector sequence (sequence length F) obtained by the visual encoder is aligned with the output sequence length f of the text encoder using a length adapter module. The length adapter can be a Transformer-based length adapter from SEAMLESSM4T V2.

[0353] By aligning the lengths of the feature vector sequences of the visual modality and the text modality, the length of the visual encoding vector sequence is compressed from F to f (f << F). At this point, the output of the visual encoder is... Here, D represents the dimension of the features input from the text encoder to the text decoder, which improves the processing efficiency of the visual encoding model.

[0354] Then, feature space alignment is performed. Using knowledge distillation as shown in Equation (2) above as the objective function, the outputs of the visual encoder and the text encoder are aligned in feature space to extract knowledge from powerful large-scale text-to-text machine translation models (such as SEAMLESSM4T-NLLB) to guide the visual expression of cued speech to text (and text translation), that is, the task from visual modality to text modality, thereby ensuring the semantic integrity while compressing the output sequence of the visual encoder.

[0355] Among them, xcued-speech and x text These are the cued-speech in the source language and the corresponding source text, y text This refers to the text corresponding to the target language. In this embodiment, the source language is Chinese, and the target languages ​​include both Chinese and English. When the target language is English, the source text x is processed... text Perform a text translation from Chinese to English to obtain pseudo-labels y. text When the target language is the same as the source language, it is equivalent to a cued-speech to text recognition task (CS2T).

[0356] By fixing the pre-trained text encoder, the parameters θ of the visual encoder (CS encoder) are adjusted. csE The parameters θ of the text decoder shown in Figure 4A tD Perform joint fine-tuning, adjusting the loss function for the corresponding recognition task. Loss function for translation task As an additional fine-tuning training objective function, the loss function for the recognition task Loss function for translation task As shown in formulas (3) and (4) above.

[0357] In summary, the loss function for the overall fine-tuning training of the first architecture of the speech generation model provided in this application embodiment is shown in formula (5) above, where α is a weight parameter determined by the proportion of the fine-tuning training set in the full training set.

[0358] As an example, the text encoder and text decoder shown in Figure 4A can adopt the SEAMLESSM4T-NLLB text encoder & decoder model based on Transformer and pre-trained weights, and fix the weights of the text encoder to fine-tune the parameters of the text decoder through the above joint training; the text to unit module can adopt the non-autoregressive (NAR) T2U model and weights in SEAMLESSM4T V2; the unit to speech module can adopt the unit vocoder model based on High-Fidelity Generative Adversarial Network (HiFi-GAN) in SEAMLESSM4T V2.

[0359] In the practical application of voice prompt systems, the prompts from hearing-impaired individuals, while expressing hand and lip movements, also include impaired voice or unvoiced sound. The degree of impairment varies from person to person; some have mild impairment and can produce somewhat discernible sounds, while others have severe impairment and produce sounds that are difficult to distinguish. In practice, this impaired voice can help improve the accuracy and intelligibility of communication. Therefore, the training of the voice generation model can be combined with the training of the impaired voice coding model.

[0360] Referring to Figure 4B, which is a schematic diagram of the second architecture of the speech generation model provided in this application embodiment, as shown in Figure 4B, compared with the training architecture of the speech generation model shown in Figure 4A, an impaired-speech encoder (IS Encoder) is added. The speech encoder in SEAMLESSM4T V2 can be used as the initial speech encoder, and the length of the speech encoding vector sequence of the speech modality is aligned with the length of the text vector sequence of the text modality through the length adapter of the speech modality. Then, the speech encoder is fine-tuned using a low-resource impaired speech dataset (referring to the inability of the speaker to produce normal, clear, and fluent speech due to physiological or medical reasons).

[0361] As shown in Figure 4B, an 80-dimensional spectral feature is extracted from the damaged speech using a Mel filter bank extractor. Where F' represents the total number of speech frames, which serves as the input to the speech encoder. Using knowledge distillation as shown in formula (6) above as the objective function, the outputs of the speech encoder and the text encoder are aligned in feature space.

[0362] Among them, ximpaired-speech and x text These are the impaired-speech in the source language and the corresponding source text, y text This refers to the text corresponding to the target language. When the target language is English, it is obtained by analyzing the source text x. text Perform a text translation from Chinese to English to obtain pseudo-labels y. text When the target language is the same as the source language (Chinese), it is equivalent to the impaired-speech to text recognition task (IS2T).

[0363] By fixing the pre-trained text encoder, the parameters θ of the speech encoder (IS encoder) are adjusted. isE The parameters θ of the text decoder shown in Figure 4B tD Perform joint fine-tuning and identify the loss function and translation loss function As an additional fine-tuning training objective function, the loss function for identification. and translation loss function As shown in formulas (7) and (8) above.

[0364] In summary, the loss function for the overall fine-tuning training of the second architecture of the speech generation model provided in this application embodiment is shown in formula (9) above.

[0365] In practical applications, there is a strong correlation between visual modal prompts and impaired speech in the speech modality. Hearing-impaired individuals can improve the accuracy and intelligibility of their communication by jointly expressing these prompts and speech. Therefore, the visual encoder and speech encoder can be jointly trained during the training of the speech generation model.

[0366] Referring to Figure 4C, Figure 4C is a schematic diagram of the third architecture of the speech generation model provided in this application embodiment. As shown in Figure 4C, the damaged speech can be processed by the speech encoder to output... and the output of the visual encoder The Concatenation Projector module first performs concatenation processing to obtain... Then perform mapping processing to obtain Then it is aligned with the feature space of the text encoder output. The objective function of knowledge distillation corresponding to the alignment of feature vector sequences is shown in formula (10) above.

[0367] In formula (10), xcued-speech, ximpaired-speech, and x text These are the cued-speech and impaired-speech in the source language and their corresponding text, respectively, while ytext corresponds to the text in the target language. When the target language is English, the source text x is processed... text Perform a text translation from Chinese to English to obtain pseudo-labels y. text When the target language is the same as the source language (Chinese), it is equivalent to a text recognition task that combines cued-speech and impaired-speech (CS_IS2T).

[0368] By fixing the pre-trained text encoder, the parameters θ of the visual encoder (CS encoder) are adjusted. csE The parameters θ of the speech encoder (IS Encoder) isE The parameters θ of the text decoder shown in Figure 4C tD Perform joint fine-tuning and increase the recognition loss function. and translation loss function As an additional fine-tuning training objective function, the loss function for identification. and translation loss function The above formulas (11) and (12) are shown respectively.

[0369] In summary, the loss function for the overall fine-tuning training of the third architecture of the speech generation model provided in this application embodiment is shown in formula (13) above.

[0370] The training methods described above can directly generate fluent speech based on prompts from people who are unable to speak or whose speech is impaired. It supports real-time streaming speech output, as well as the generation of natural and fluent voices in multiple languages. Users can also choose their own voice timbre, which can help more people with hearing loss or speech impairment to integrate into the digital society better and faster.

[0371] The following description continues to illustrate the exemplary structure of the speech generation model training device 555-1 provided in the embodiments of this application as a software module. In some embodiments, as shown in FIG2A, the software module in the speech generation model training device 555-1 stored in the memory 550-1 may include:

[0372] The acquisition module 5551-1 is configured to acquire a first speech generation model, wherein the first speech generation model includes a decoder and multiple encoders corresponding to multiple modalities; and acquire sample data of multiple modalities, wherein the sample data of multiple modalities includes a sequence of prompt image of sample objects and speech text corresponding to the prompt image sequence.

[0373] The encoding module 5552-1 is configured to call multiple encoders to encode the prompt image sequence and the speech text respectively, so as to obtain a multimodal encoded vector sequence.

[0374] The decoding module 5553-1 is configured to call the decoder based on the multimodal encoded vector sequence to obtain the decoded text.

[0375] The determination module 5554-1 is configured to determine the probability distribution of the decoded text and sample data of multiple modalities, and to determine the target loss based on the probability distribution.

[0376] The generation module 5555-1 is configured to update the parameters of the decoder and at least one encoder based on the target loss to obtain the updated decoder and multiple encoders. The updated decoder and multiple encoders are used to form a second speech generation model, which is used to generate target speech text corresponding to the prompt image sequence of the target object.

[0377] In some embodiments, when the multiple modalities include a visual modality and a text modality, the encoding module 5552-1 is further configured to call a visual encoder to encode the prompt image sequence to obtain a visual encoding vector sequence for the visual modality, wherein the visual encoder is the encoder corresponding to the visual modality among the multiple encoders; call a text encoder to encode multiple morphemes in the speech text to obtain a text encoding vector sequence for the text modality, wherein the text encoding vector sequence includes text encoding vectors corresponding to multiple morphemes, and the text encoder is the encoder corresponding to the text modality among the multiple encoders; and use the visual encoding vector sequence and the text encoding vector sequence as a first multimodal encoding vector sequence, wherein the first multimodal encoding vector sequence is used for decoding by the decoder.

[0378] In some embodiments, the encoding module 5552-1 is further configured to identify multiple prompt images in the prompt image sequence to obtain the position of key points in each prompt image; generate a key point sequence based on the position of key points in each prompt image; obtain the embedding vector of the key point sequence; and call the visual encoder to encode the embedding vector to obtain the visual encoding vector of the visual modality.

[0379] In some embodiments, the determining module 5554-1 is further configured to determine the probability distribution of the encoded vector sequence of sample data of at least two modalities relative to preset conditions, wherein different modalities correspond to different preset conditions; determine the probability distribution of speech text relative to sample data of at least one modality; and determine the probability distribution of decoded text relative to sample data of at least one modality.

[0380] In some embodiments, the determining module 5554-1 is further configured to: determine a first probability distribution of the text encoded vector sequence relative to the decoded text and the speech text, and a second probability distribution of the visual encoded vector sequence relative to the decoded text and the prompt image sequence; determine a third probability distribution of the speech text relative to the prompt image sequence; if the target language of the speech generation task of the first speech generation model is different from the source language, determine a fourth probability distribution of the decoded text relative to the prompt image sequence; determine the difference between the first probability distribution and the second probability distribution as a first sub-target loss; determine a second sub-target loss negatively correlated with the third probability distribution; determine a third sub-target loss negatively correlated with the fourth probability distribution; and fuse the first sub-target loss, the second sub-target loss, and the third sub-target loss to obtain a first target loss.

[0381] In some embodiments, the determining module 5554-1 is further configured to invoke a pre-trained first probability distribution model to determine a first probability distribution of the text encoded vector sequence relative to the decoded text and the speech text, wherein the first probability distribution model is trained by: acquiring first sample data, wherein the first sample data includes sample decoded text and sample speech text, and a reference probability distribution of the sample text encoded vector sequence relative to the sample decoded text and the sample speech text; using the sample decoded text, the sample speech text, and the sample text encoded vector sequence as input to the initialized first probability distribution model, and outputting a predicted probability distribution of the text encoded vector sequence relative to the sample decoded text and the sample speech text; determining a loss value based on the difference between the predicted probability distribution and the reference probability distribution; and updating the parameters of the first probability distribution model according to the loss value.

[0382] In some embodiments, the generation module 5555-1 is further configured to update the parameters of the visual encoder based on the first sub-target loss; and to update the parameters of the visual encoder and decoder based on the second sub-target loss and the third sub-target loss.

[0383] In some embodiments, the first speech generation model is a pre-trained speech recognition model, and after pre-training, the representation space of the text modality and the representation space of the speech modality of the first speech generation model have been unified to the same representation space; the decoding module 5553-1 is further configured to map the visual encoded vector sequence and the text encoded vector sequence to obtain the vector sequence representation of the visual encoded vector sequence and the text encoded vector sequence in the representation space; and adjust the length of the vector sequence representation of the visual encoded vector sequence to be the same as the length of the vector sequence representation of the text encoded vector sequence.

[0384] In some embodiments, when the multiple modalities also include a speech modality, the sample data of the multiple modalities also include speech signal features of the speech modality; the encoding module 5552-1 is further configured to call a speech encoder to encode the speech signal features to obtain a speech encoding vector sequence of the speech modality, wherein the speech encoder is the encoder corresponding to the speech modality among multiple encoders; the visual encoding vector sequence, the text encoding vector sequence and the speech encoding vector sequence are used as a second multimodal encoding vector sequence, wherein the second multimodal encoding vector sequence is used to replace the first multimodal encoding vector sequence for decoding by the decoder.

[0385] In some embodiments, the determining module 5554-1 is further configured to: determine a fifth probability distribution of the text encoding vector sequence relative to the decoded text and the speech text, and a sixth probability distribution of the speech encoding vector sequence relative to the decoded text and the speech signal features; determine a seventh probability distribution of the speech text relative to the speech signal features; when the target language of the speech generation task of the first speech generation model is different from the source language, determine an eighth probability distribution of the decoded text relative to the speech signal features; determine the difference between the fifth probability distribution and the sixth probability distribution as a fourth sub-target loss; determine a fifth sub-target loss negatively correlated with the seventh probability distribution; determine a sixth sub-target loss negatively correlated with the eighth probability distribution; and fuse the fourth sub-target loss, the fifth sub-target loss, and the sixth sub-target loss to obtain a second target loss.

[0386] In some embodiments, the generation module 5555-1 is further configured to update the parameters of the speech encoder based on the fourth sub-target loss; and to update the parameters of the speech encoder and decoder based on the fifth sub-target loss and the sixth sub-target loss.

[0387] In some embodiments, when the multiple modalities also include a speech modality, the encoding module 5552-1 is further configured to call multiple encoders to encode based on the prompt image sequence and the speech text, respectively, to obtain a visual encoding vector sequence for the visual modality, a text encoding vector sequence for the text modality, and a speech encoding vector sequence for the speech modality; to concatenate the visual encoding vector sequence and the speech encoding vector sequence to obtain a concatenated encoding vector sequence; to perform feature mapping on the concatenated encoding vector sequence to obtain a fused encoding vector sequence; and to use the fused encoding vector sequence and the text encoding vector sequence as a third multimodal encoding vector sequence, wherein the third multimodal encoding vector sequence is used to replace the second multimodal encoding vector sequence for decoding by the decoder.

[0388] In some embodiments, the sample data of multiple modalities further includes speech signal features of the speech modality. The encoding module 5552-1 is further configured to call a visual encoder to encode the prompt image sequence to obtain a visual encoding vector sequence of the visual modality, wherein the visual encoder is the encoder corresponding to the visual modality among multiple encoders; call a text encoder to encode multiple morphemes in the speech text to obtain a text encoding vector sequence of the text modality, wherein the text encoding vector sequence includes text encoding vectors corresponding to multiple morphemes, and the text encoder is the encoder corresponding to the text modality among multiple encoders; and call a speech encoder to encode the speech signal features to obtain a speech encoding vector sequence of the speech modality, wherein the speech encoder is the encoder corresponding to the speech modality among multiple encoders.

[0389] In some embodiments, the determining module 5554-1 is further configured to: determine a ninth probability distribution of the text encoded vector sequence relative to the decoded text and the speech text, and a tenth probability distribution of the fused encoded vector sequence relative to the decoded text, the prompt image sequence, and the speech signal features; determine an eleventh probability distribution of the speech text relative to the prompt image sequence and the speech signal features; when the target language of the speech generation task of the first speech generation model is different from the source language, determine a twelfth probability distribution of the decoded text relative to the prompt image sequence and the speech signal features; determine the difference between the ninth and tenth probability distributions as the seventh sub-target loss; determine an eighth sub-target loss negatively correlated with the eleventh probability distribution; determine a ninth sub-target loss negatively correlated with the twelfth probability distribution; and fuse the seventh, eighth, and ninth sub-target losses to obtain a third target loss.

[0390] In some embodiments, the generation module 5555-1 is further configured to update the parameters of the visual encoder and the speech encoder based on the seventh sub-target loss; and to update the parameters of the visual encoder, the speech encoder, and the decoder based on the eighth sub-target loss and the ninth sub-target loss.

[0391] In some embodiments, the first speech generation model further includes a speech signal generator, wherein the speech signal generator is used to form a second speech generation model with the decoder and the updated plurality of encoders, the second speech generation model being used to generate speech signals.

[0392] The following description continues to illustrate the exemplary structure of the speech generation device 555-2 for the speech generation model provided in this application embodiment as a software module. In some embodiments, as shown in FIG2B, the software module in the speech generation model training device 555-2 stored in the memory 550-2 may include:

[0393] The acquisition module 5551-2 is configured to acquire the prompt image sequence of the target object.

[0394] The generation module 5552-2 is configured to call the second speech generation model based on the prompt image sequence to generate target speech text corresponding to the prompt image sequence of the target object.

[0395] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the training method of the speech generation model described above, or to perform the speech generation method of the speech generation model described above.

[0396] This application provides a computer-readable storage medium storing computer-executable instructions or computer programs. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the training method of the speech generation model provided in this application, as shown in FIG3A, or execute the speech generation method of the speech generation model described above, as shown in FIG3Q.

[0397] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0398] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0399] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0400] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0401] In summary, this application embodiment uses sample data from multiple modalities, including the prompt image sequence of the target object (corresponding to the visual modality) and the speech text (corresponding to the text modality), and calls multiple encoders to encode the data. The resulting multimodal encoding vector sequence (i.e., the encoding vector sequence of multiple modalities) is compared with the probability distribution of the decoded text to calculate the target loss. This allows the target loss to reflect the differences between the representation spaces of multiple modalities in the first speech generation model. Based on the target loss, the parameters of the first speech generation model are updated through backpropagation, enabling the first speech generation model to gradually learn the intrinsic relationship between the information of the visual modality and the text modality. This improves the model's performance and generalization ability, resulting in a unified representation space for the visual modality and the text modality in the second speech generation model obtained after training. This means that the second speech generation model can map data from different modalities into a common representation space, allowing data from different modalities to have similar semantic expressions in this space. This lays the foundation for accurate speech text generation in the future, ensuring that the target speech text output by the second speech generation model accurately matches the intent of the prompt image of the target object, thereby guaranteeing the accuracy of speech text generation from the prompt image. Meanwhile, compared to related technologies that require manual text input to generate speech text, this method is more convenient, efficient, and provides a better user experience. The training process for the first speech generation model is based on a pre-trained speech recognition model. With the representation spaces of the text modality and the speech modality already unified into the same representation space, fine-tuning allows for the conversion of prompt image sequences into speech text based on the original functions of the pre-trained speech recognition model. Furthermore, multiple functions of the pre-trained speech recognition model can be reused, such as recognizing prompt image sequences of target objects, translating text, and generating speech, avoiding redundant development and improving development efficiency.

[0402] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A method for training a speech generation model, the method being executed by an electronic device, the method comprising: Obtain a first speech generation model, wherein the first speech generation model includes a decoder and multiple encoders corresponding to multiple modalities respectively; Acquire sample data for multiple modalities, wherein the sample data for multiple modalities includes a sequence of prompt images of sample objects and the corresponding speech text; Based on the prompt image sequence and the voice text, the multiple encoders are called to encode the text, resulting in a multimodal encoded vector sequence; The decoder is invoked based on the multimodal encoded vector sequence to perform decoding, thereby obtaining the decoded text; Determine the probability distribution of the decoded text and the sample data of the multiple modalities, and determine the target loss based on the probability distribution; The parameters of the decoder and at least one encoder are updated based on the target loss to obtain the updated decoder and the plurality of encoders, wherein the updated decoder and the plurality of encoders are used to form a second speech generation model, and the second speech generation model is used to generate target speech text corresponding to the prompt image sequence of the target object.

2. The method according to claim 1, wherein, When the multiple modalities include a visual modality and a text modality, the process of calling the multiple encoders to encode based on the prompt image and the spoken text respectively, to obtain a multimodal encoded vector sequence, includes: The visual encoder is invoked to encode the prompt image sequence to obtain a visual encoding vector sequence of the visual modality, wherein the visual encoder is the encoder corresponding to the visual modality among the plurality of encoders; The text encoder is invoked to encode multiple morphemes in the speech text to obtain a text encoding vector sequence of the text modality, wherein the text encoding vector sequence includes text encoding vectors corresponding to the multiple morphemes respectively, and the text encoder is the encoder corresponding to the text modality among the multiple encoders; The visual encoding vector sequence and the text encoding vector sequence are used as a first multimodal encoding vector sequence, wherein the first multimodal encoding vector sequence is used for decoding by the decoder.

3. The method according to claim 2, wherein, The step of calling the visual encoder to encode the prompt image sequence to obtain a visual encoding vector sequence of the visual modality includes: The positions of key points in each prompt image are obtained by identifying multiple prompt images in the prompt image sequence. Based on the position of the key points in each of the prompt images, a key point sequence is generated; Obtain the embedding vector of the key point sequence; The visual encoder is invoked to encode the embedding vector to obtain a visual encoding vector sequence of the visual modality.

4. The method according to claim 2 or 3, wherein, Determining the probability distribution of the decoded text and the sample data of the multiple modalities includes: Determine the probability distribution of the encoded vector sequence of sample data for at least two modalities relative to preset conditions, wherein different modalities correspond to different preset conditions; Determine the probability distribution of the speech text relative to sample data of at least one of the modalities; Determine the probability distribution of the decoded text relative to sample data of at least one of the modalities.

5. The method according to any one of claims 2 to 4, wherein, Determining the probability distribution of the decoded text and the sample data of the multiple modalities includes: Determine a first probability distribution of the text encoded vector sequence relative to the decoded text and the speech text, and a second probability distribution of the visual encoded vector sequence relative to the decoded text and the prompt image sequence; Determine the third probability distribution of the spoken text relative to the prompt image sequence; When the target language of the speech generation task of the first speech generation model is different from the source language, a fourth probability distribution of the decoded text relative to the prompt image sequence is determined. Determining the target loss based on the probability distribution includes: The difference between the first probability distribution and the second probability distribution is determined as the first sub-target loss; Determine the second sub-target loss that is negatively correlated with the third probability distribution; Determine the loss of the third sub-target that is negatively correlated with the fourth probability distribution; The first sub-target loss, the second sub-target loss, and the third sub-target loss are fused together to obtain the first target loss.

6. The method according to claim 5, wherein, Determining the first probability distribution of the text encoding vector sequence relative to the decoded text and the speech text includes: A pre-trained first probability distribution model is invoked to determine a first probability distribution of the text encoded vector sequence relative to the decoded text and the speech text, wherein the first probability distribution model is trained in the following manner: Acquire first sample data, wherein the first sample data includes sample decoded text and sample speech text, and a reference probability distribution of the sample text encoding vector sequence relative to the sample decoded text and the sample speech text; The sample decoded text, the sample speech text, and the sample text encoded vector sequence are used as inputs to an initialized first probability distribution model, and the predicted probability distribution of the text encoded vector sequence relative to the sample decoded text and the sample speech text is output. The loss value is determined based on the difference between the predicted probability distribution and the reference probability distribution; Update the parameters of the first probability distribution model based on the loss value.

7. The method according to claim 5, wherein, The step of updating the parameters of the decoder and at least one of the encoders based on the target loss includes: The parameters of the visual encoder are updated based on the first sub-target loss; The parameters of the visual encoder and the decoder are updated based on the second sub-target loss and the third sub-target loss.

8. The method according to any one of claims 2 to 4, wherein, The first speech generation model is a pre-trained speech recognition model, and after the pre-training, the representation space of the text modality and the representation space of the speech modality of the first speech generation model have been unified to the same representation space. Before calling the decoder based on the multimodal encoded vector sequence to obtain the decoded text, the method further includes: The visual encoding vector sequence and the text encoding vector sequence are mapped to obtain the vector sequence representations of the visual encoding vector sequence and the text encoding vector sequence in the representation space; The length of the vector sequence representation of the visual encoding vector sequence is adjusted to be the same as the length of the vector sequence representation of the text encoding vector sequence.

9. The method according to any one of claims 2 to 4, wherein, When the plurality of modalities also includes a speech modality, the sample data of the plurality of modalities also includes the speech signal features of the speech modality; The step of encoding the prompt image sequence and the speech text by calling the multiple encoders respectively to obtain a multimodal encoded vector sequence further includes: The speech signal features are encoded by calling a speech encoder to obtain a speech encoding vector sequence of the speech modality, wherein the speech encoder is the encoder corresponding to the speech modality among the plurality of encoders; The visual encoding vector sequence, the text encoding vector sequence, and the speech encoding vector sequence are used as a second multimodal encoding vector sequence, wherein the second multimodal encoding vector sequence is used to replace the first multimodal encoding vector sequence for decoding by the decoder.

10. The method according to claim 9, wherein, Determining the probability distribution of the decoded text and the sample data of the multiple modalities includes: Determine a fifth probability distribution of the text encoding vector sequence relative to the decoded text and the speech text, and a sixth probability distribution of the speech encoding vector sequence relative to the decoded text and the speech signal features; Determine the seventh probability distribution of the speech text relative to the features of the speech signal; When the target language of the speech generation task of the first speech generation model is different from the source language, the eighth probability distribution of the decoded text relative to the speech signal features is determined. Determining the target loss based on the probability distribution includes: The difference between the fifth probability distribution and the sixth probability distribution is determined as the fourth sub-target loss; Determine the loss of the fifth sub-target that is negatively correlated with the seventh probability distribution; Determine the loss of the sixth sub-target that is negatively correlated with the eighth probability distribution; The fourth sub-target loss, the fifth sub-target loss, and the sixth sub-target loss are fused together to obtain the second target loss.

11. The method according to claim 10, wherein, The step of updating the parameters of the decoder and at least one of the encoders based on the target loss includes: The parameters of the speech encoder are updated based on the fourth sub-target loss; The parameters of the speech encoder and the decoder are updated based on the fifth sub-target loss and the sixth sub-target loss.

12. The method according to any one of claims 1 to 11, wherein, When the multiple modalities also include a speech modality, the process of encoding the multiple encoders based on the prompt image sequence and the speech text to obtain a multimodal encoded vector sequence includes: Based on the prompt image sequence and the speech text, the multiple encoders are called to encode the visual encoding vector sequence of the visual modality, the text encoding vector sequence of the text modality, and the speech encoding vector sequence of the speech modality. The visual encoding vector sequence and the speech encoding vector sequence are concatenated to obtain a concatenated encoding vector sequence. The concatenated encoded vector sequence is then subjected to feature mapping to obtain a fused encoded vector sequence; The fused encoded vector sequence and the text encoded vector sequence are used as a third multimodal encoded vector sequence, wherein the third multimodal encoded vector sequence is used to replace the second multimodal encoded vector sequence for decoding by the decoder.

13. The method according to claim 12, wherein, The sample data for the multiple modalities also includes speech signal features of the speech modality; The process of encoding the visual modality (visual encoding vector sequence), the text modality (text encoding vector sequence), and the speech modality (speech encoding vector sequence) by calling the multiple encoders respectively, to obtain the speech encoding vector sequence, includes: The visual encoder is invoked to encode the prompt image sequence to obtain the visual encoding vector sequence of the visual modality, wherein the visual encoder is the encoder corresponding to the visual modality among the plurality of encoders; The text encoder is invoked to encode multiple morphemes in the speech text to obtain a text encoding vector sequence of the text modality, wherein the text encoding vector sequence includes text encoding vectors corresponding to the multiple morphemes, and the text encoder is the encoder corresponding to the text modality among the multiple encoders; The speech signal features are encoded by calling a speech encoder to obtain a speech encoding vector sequence of the speech modality, wherein the speech encoder is the encoder corresponding to the speech modality among the plurality of encoders.

14. The method according to claim 12, wherein, Determining the probability distribution of the decoded text and the sample data of the multiple modalities includes: Determine the ninth probability distribution of the text encoding vector sequence relative to the decoded text and the speech text, and the tenth probability distribution of the fused encoding vector sequence relative to the decoded text, the prompt image sequence, and the speech signal features; Determine the eleventh probability distribution of the speech text relative to the prompt image sequence and the speech signal features; When the target language of the speech generation task of the first speech generation model is different from the source language, the twelfth probability distribution of the decoded text relative to the prompt image sequence and the speech signal features is determined. Determining the target loss based on the probability distribution includes: The difference between the ninth probability distribution and the tenth probability distribution is determined as the seventh sub-target loss; Determine the loss of the eighth sub-target that is negatively correlated with the eleventh probability distribution; Determine the loss of the ninth sub-target that is negatively correlated with the twelfth probability distribution; The loss of the seventh sub-target, the loss of the eighth sub-target, and the loss of the ninth sub-target are fused together to obtain the third target loss.

15. The method according to claim 14, wherein, The step of updating the parameters of the decoder and at least one of the encoders based on the target loss includes: The parameters of the visual encoder and the speech encoder are updated based on the seventh sub-target loss; The parameters of the visual encoder, the speech encoder, and the decoder are updated based on the eighth sub-target loss and the ninth sub-target loss.

16. The method according to any one of claims 1 to 14, wherein, The first speech generation model further includes a speech signal generator, wherein the speech signal generator is used to form a second speech generation model with the decoder and the updated plurality of encoders, and the second speech generation model is used to generate speech signals.

17. A speech generation method for a speech generation model, the method being executed by an electronic device, the speech generation model being the second speech generation model according to any one of claims 1 to 16, the method comprising: Obtain the sequence of prompt images for the target object; Based on the prompt image sequence, the second speech generation model is invoked to generate target speech text corresponding to the prompt image sequence of the target object.

18. A training apparatus for a speech generation model, the apparatus comprising: The acquisition module is configured to acquire a first speech generation model, wherein the first speech generation model includes a decoder and multiple encoders corresponding to multiple modalities; and acquire sample data of multiple modalities, wherein the sample data of multiple modalities includes a sequence of prompt image of sample objects and speech text corresponding to the prompt image sequence. The encoding module is configured to call the multiple encoders to encode the prompt image sequence and the speech text respectively, thereby obtaining a multimodal encoded vector sequence; The decoding module is configured to call the decoder based on the multimodal encoded vector sequence to obtain decoded text; The determination module is configured to determine the probability distribution of the decoded text and the sample data of the multiple modalities, and determine the target loss based on the probability distribution; The generation module is configured to update the parameters of the decoder and at least one encoder based on the target loss to obtain the updated decoder and the plurality of encoders, wherein the updated decoder and the plurality of encoders are used to form a second speech generation model, and the second speech generation model is used to generate target speech text corresponding to the prompt image sequence of the target object.

19. A speech generation apparatus, the apparatus comprising: The acquisition module is configured to acquire a sequence of prompt images for the target object; The generation module is configured to call a speech generation model based on the prompt image to generate target speech text corresponding to the prompt image sequence of the target object, wherein the speech generation model is the second speech generation model as described in any one of claims 1 to 16.

20. An electronic device, the electronic device comprising: Memory is used to store executable instructions for a computer; The processor, when executing computer-executable instructions stored in the memory, implements the training method of the speech generation model according to any one of claims 1 to 16, or the speech generation method of the speech generation model according to claim 17.

21. A computer-readable storage medium storing computer-executable instructions or a computer program, wherein the computer-executable instructions or the computer program, when executed by a processor, implement the training method of the speech generation model according to any one of claims 1 to 16, or the speech generation method of the speech generation model according to claim 17.

22. A computer program product comprising computer-executable instructions or a computer program, wherein the computer-executable instructions or the computer program, when executed by a processor, implement the training method of the speech generation model according to any one of claims 1 to 16, or the speech generation method of the speech generation model according to claim 17.