Speech generation model training method, speech generation method and device, electronic equipment, computer readable storage medium and computer program product

By employing multimodal technology, the accuracy and efficiency of generating speech text from prompt images have been improved, solving the problems of inaccuracy and inconvenience in existing speech text generation technologies and providing a more efficient interaction method.

CN120977314APending Publication Date: 2025-11-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410626872.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-15
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In existing technologies, the method of generating speech text by encoding the hand and lip movements of the target object is not accurate or convenient enough, resulting in a poor user experience, especially for scenarios where it is impossible to express intentions directly with speech.

Method used

A multimodal speech generation model is adopted. Sample data from multiple modalities are acquired, encoded, and decoded. The model is trained using a decoder and multiple encoders. By combining sample data from visual and text modalities, the target loss is calculated and the model parameters are updated, thereby unifying the representation spaces of visual and text modalities and generating more accurate speech and text.

Benefits of technology

It improves the accuracy and efficiency of generating voice text from prompt images, provides a more convenient interaction method, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977314A_ABST
    Figure CN120977314A_ABST
Patent Text Reader

Abstract

The invention provides a training method of a voice generation model, a voice generation method and device, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the steps of obtaining a first voice generation model; acquiring sample data of a plurality of modals; based on the prompt word image sequence and the voice text, a plurality of encoders are called for encoding, and a multi-modal encoding vector sequence is obtained; calling a decoder for decoding based on the multi-modal coding vector sequence to obtain a decoded text; determining probability distribution of the decoded text and the sample data of the multiple modals, and determining target loss based on the probability distribution; and parameters of the decoder and the at least one encoder are updated based on the target loss, the updated decoder and the plurality of encoders are used for forming a second voice generation model, and the second voice generation model is used for generating a target voice text corresponding to the prompt word image sequence of the target object. According to the invention, the efficiency and accuracy of generating the voice text from the prompt word image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to artificial intelligence technology, and more particularly to a training method for a speech generation model, a speech generation method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] In interpersonal communication or scenarios requiring human-computer interaction, when the target cannot directly express their intentions through speech (e.g., due to limitations of the current environment or physiological barriers to vocalization), relevant technologies support users in expressing their intentions through body language and lip movements, or by inputting text on the terminal device. These technologies generate speech-text by encoding the target's hand and lip movements; however, the resulting speech-text is often inaccurate and inconvenient, leading to a poor user experience. Summary of the Invention

[0003] This application provides a training method for a speech generation model, a speech generation method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the efficiency and accuracy of generating speech text from prompt images.

[0004] The technical solution of this application embodiment is implemented as follows:

[0005] This application provides a method for training a speech generation model, the method comprising:

[0006] Obtain a first speech generation model, wherein the first speech generation model includes a decoder and multiple encoders corresponding to multiple modalities respectively;

[0007] Acquire sample data for multiple modalities, wherein the sample data for multiple modalities includes a sequence of prompt images of sample objects and the corresponding speech text;

[0008] Based on the prompt image sequence and the voice text, the multiple encoders are called to encode the text, resulting in a multimodal encoded vector sequence;

[0009] The decoder is invoked based on the multimodal encoded vector sequence to perform decoding, thereby obtaining the decoded text;

[0010] Determine the probability distribution of the decoded text and the sample data of the multiple modalities, and determine the target loss based on the probability distribution;

[0011] The parameters of the decoder and at least one encoder are updated based on the target loss to obtain the updated decoder and the plurality of encoders, wherein the updated decoder and the plurality of encoders are used to form a second speech generation model, and the second speech generation model is used to generate target speech text corresponding to the prompt image sequence of the target object.

[0012] This application provides a speech generation method for a speech generation model, wherein the speech generation model is the second speech generation model described above, and the method includes:

[0013] Obtain the sequence of prompt images for the target object;

[0014] Based on the prompt image sequence, the second speech generation model is invoked to generate target speech text corresponding to the prompt image sequence of the target object.

[0015] This application provides a training device for a speech generation model, the device comprising:

[0016] An acquisition module is used to acquire a first speech generation model, wherein the first speech generation model includes a decoder and multiple encoders corresponding to multiple modalities; and to acquire sample data of multiple modalities, wherein the sample data of multiple modalities includes a sequence of prompt image of sample objects and speech text corresponding to the prompt image sequence.

[0017] The encoding module is used to call the multiple encoders to encode the prompt image sequence and the voice text respectively, so as to obtain a multimodal encoded vector sequence;

[0018] The decoding module is used to call the decoder based on the multimodal encoded vector sequence to decode the text and obtain the decoded text.

[0019] The determination module is used to determine the probability distribution of the decoded text and the sample data of the multiple modalities, and to determine the target loss based on the probability distribution;

[0020] A generation module is used to update the parameters of the decoder and at least one encoder based on the target loss to obtain the updated decoder and the plurality of encoders, wherein the updated decoder and the plurality of encoders are used to form a second speech generation model, and the second speech generation model is used to generate target speech text corresponding to the prompt image sequence of the target object.

[0021] This application provides a speech generation apparatus for a speech generation model, wherein the speech generation model is the second speech generation model described above, and the apparatus includes:

[0022] The acquisition module is used to acquire the sequence of prompt images for the target object;

[0023] The generation module is used to call the second speech generation model based on the prompt image sequence to generate target speech text corresponding to the prompt image sequence of the target object.

[0024] This application provides an electronic device, the electronic device comprising:

[0025] Memory is used to store executable instructions for a computer;

[0026] The processor, when executing computer-executable instructions stored in the memory, implements the training method of the speech generation model provided in the embodiments of this application, or the speech generation method of the speech generation model.

[0027] This application provides a computer-readable storage medium storing a computer program or computer-executable instructions, which, when executed by a processor, implements the training method of the speech generation model provided in this application, or the speech generation method of the speech generation model.

[0028] This application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, they implement the training method of the speech generation model provided in this application, or the speech generation method of the speech generation model.

[0029] The embodiments of this application have the following beneficial effects:

[0030] By encoding sample data from multiple modalities, including prompt image sequences (corresponding to the visual modality) and speech text (corresponding to the text modality), the target loss is calculated by comparing the resulting multimodal encoded vector sequence (i.e., the encoded vector sequence of multiple modalities) with the probability distribution of the decoded text. This target loss reflects the differences between the representation spaces of multiple modalities in the first speech generation model. Based on the target loss, the parameters of the first speech generation model are updated through backpropagation, resulting in a unified representation space for the visual and text modalities in the trained second speech generation model. This ensures that the target speech text output by the second speech generation model accurately matches the intent of the prompt image, thus guaranteeing the accuracy of speech text generation from the prompt image. Furthermore, compared to related technologies that require manual text input to generate speech text, this method is more convenient, efficient, and provides a better user experience. Attached Figure Description

[0031] Figure 1 This is a schematic diagram of the architecture of the training system 100 for the speech generation model provided in this application embodiment;

[0032] Figure 2A This is a schematic diagram of the structure of the electronic device 500-1 provided in the embodiments of this application;

[0033] Figure 2B This is a schematic diagram of the structure of the electronic device 500-2 provided in the embodiments of this application;

[0034] Figure 3A This is a flowchart illustrating the training method of the speech generation model provided in the embodiments of this application;

[0035] Figure 3B This is a schematic diagram of the first process for obtaining a multimodal encoded vector sequence provided in an embodiment of this application;

[0036] Figure 3C This is a schematic diagram of the process for obtaining a visual encoding vector sequence provided in an embodiment of this application;

[0037] Figure 3D This is a schematic diagram of the second process for obtaining a multimodal encoded vector sequence provided in an embodiment of this application;

[0038] Figure 3E This is a schematic diagram of the third process for obtaining a multimodal encoded vector sequence provided in an embodiment of this application;

[0039] Figure 3F This is a schematic diagram of the fourth process for obtaining a multimodal encoded vector sequence provided in an embodiment of this application;

[0040] Figure 3G This is a schematic diagram of the process for aligning the length of the encoded vector sequence provided in an embodiment of this application;

[0041] Figure 3H This is a schematic diagram of the first process for determining the probability distribution provided in an embodiment of this application;

[0042] Figure 3I This is a schematic diagram of the first process for determining the target loss provided in an embodiment of this application;

[0043] Figure 3J This is a schematic diagram of the second process for determining the probability distribution provided in an embodiment of this application;

[0044] Figure 3K This is a schematic diagram of the second process for determining the target loss provided in an embodiment of this application;

[0045] Figure 3L This is a schematic diagram of the third process for determining the probability distribution provided in an embodiment of this application;

[0046] Figure 3M This is a schematic diagram of the third process for determining the target loss provided in an embodiment of this application;

[0047] Figure 3N This is a schematic diagram of the first process for updating parameters provided in an embodiment of this application;

[0048] Figure 3O This is a schematic diagram of the second process for updating parameters provided in the embodiments of this application;

[0049] Figure 3P This is a schematic diagram of the third process for updating parameters provided in the embodiments of this application;

[0050] Figure 3Q This is a flowchart illustrating the speech generation method of the speech generation model provided in this application embodiment;

[0051] Figure 4A This is a schematic diagram of the first architecture of the speech generation model provided in the embodiments of this application;

[0052] Figure 4B This is a schematic diagram of the second architecture of the speech generation model provided in the embodiments of this application;

[0053] Figure 4C This is a schematic diagram of the third architecture of the speech generation model provided in the embodiments of this application;

[0054] Figure 5A This is a schematic diagram of the structure of the visual encoder model provided in the embodiments of this application;

[0055] Figure 5B This is a schematic diagram of the structure of the machine learning model provided in the embodiments of this application;

[0056] Figure 6 This is a schematic diagram of the alignment of encoded vector sequences provided in an embodiment of this application;

[0057] Figure 7 This is a schematic diagram of the prompt voice encoding information provided in the embodiments of this application;

[0058] Figure 8 This is a schematic diagram of the training data provided in the embodiments of this application;

[0059] Figure 9 This is a visual coding model architecture diagram provided in the embodiments of this application.

[0060] It should be noted that the terms "first" and "second" mentioned above are only used to distinguish between different options and do not represent the degree of superiority or inferiority of the options or their priority in the implementation process. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0062] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0063] In the following description, the terms "first / second / third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0064] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functions of that module or unit.

[0065] Unless otherwise specified, "at least one" as used below refers to one or more cases, and "multiple" can refer to two or more cases.

[0066] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0067] In the implementation of this application, the collection and processing of relevant data should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.

[0068] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.

[0069] 1) Speech generation model: includes a decoder and multiple encoders corresponding to multiple modalities. It is trained using sample data from multiple modalities based on the sample object to generate target speech text corresponding to the visual expression of the target object.

[0070] 2) Visual Modal: Through intuitive means such as hand gestures, lip movements, and body language, the recipient can perceive and understand the visual expression of the target object's facial expressions, eye contact, and body posture, thereby obtaining the target object's emotional state and intentions. For example, for a sequence of prompt images for the target object, each prompt image in the sequence is a static image showing the hand or lip posture of the sample object when performing a body or lip movement. Each prompt image identifies the posture and relative position of the target object's hands and lips when performing the key actions.

[0071] 3) Text modality: A textual expression used to convey the emotional state and intentions of the target audience, such as "The weather is nice today".

[0072] 4) Speech modality: The expression of speech signals used to convey the emotional state and intentions of the target object.

[0073] 5) Multimodal encoding vector sequence: Call the encoders corresponding to multiple modalities in the language generation model to encode the sample data of multiple modalities respectively, and use the resulting multimodal encoding vector sequence as the multimodal encoding vector sequence.

[0074] 6) Representation space: This refers to the vector space composed of feature vectors or feature representations, used to describe and represent different attributes, features, or information of data. In different modalities, the representation space embodies the features of different types of data and their representations in the vector space. For example, the representation space of the visual modality is a vector space composed of visual feature vectors, while the representation space of the text modality is a vector space composed of text feature vectors.

[0075] 7) Cued-speech (CS), a coding system used by people who are unable to speak or whose speech is impaired to express spoken language.

[0076] 8) Sign Language Source Language (SL) is a form of communication that uses gestures, hand movements, and facial expressions to communicate. It uses gestures and hand movements to represent words, phrases, and sentences, employing different shapes, positions, and movements of the fingers, palms, wrists, etc., in conjunction with facial expressions and body movements.

[0077] 9) Impaired speech (IS) refers to speech signals whose quality has deteriorated or become difficult to understand due to various reasons. For example, when noise is mixed with a speech signal, it will lead to a decrease in speech quality, making it difficult to hear and understand the speech content clearly; distortion or damage to the speech signal during transmission, recording or processing will lead to spectral distortion, temporal distortion, distortion impact, etc., making the speech sound unnatural or distorted; people with speech disorders or articulation problems due to physiological or neurological reasons will have unclear pronunciation, abnormal speech rate, unstable pitch, etc., making their speech difficult or incomprehensible.

[0078] 10) Knowledge distillation (KD) is a model training technique used to transfer knowledge from a teacher model to a student model to assist the student model's learning and improve its performance. In knowledge distillation, the teacher model is typically a complex and accurate model with good performance. The student model, on the other hand, is a lightweight model, usually unable to achieve the complexity of the teacher model due to limitations in computing resources or model size. For example, in the embodiments of this application, knowledge distillation is used to align the output of the visual encoder to the feature space of the text encoder's output, that is, the text encoder model is used as the teacher model, and the visual encoder model is used as the student model, with the text encoder guiding the visual encoder.

[0079] 11) Prompt Image Sequence: Multiple prompt images form a prompt image sequence. For each prompt image, the posture and position of the target object's limbs, lips, and hands are identified when the object makes a limb movement, lip movement, or hand movement.

[0080] 12) Keypoint sequence: In each prompt image of the prompt image sequence, the points that identify the posture and position of the limbs, lips or hands are taken as keypoints. The keypoints in each prompt image of the prompt image sequence are combined in the order in which the actions occur to obtain the keypoint sequence.

[0081] Related technologies generate speech text by encoding the hand and lip movements of the target object. However, the method of generating speech text by encoding is only applicable to users who understand the encoding method, and has a small scope of application. Moreover, the expression of the generated speech text is not accurate or natural compared to normal communication. The method of users manually inputting text to complete the interaction process is not convenient, has low efficiency, and has a poor user experience.

[0082] To address the aforementioned issues, embodiments of this application provide a training method, apparatus, electronic device, computer-readable storage medium, and computer program product for a speech generation model, which can improve the efficiency and accuracy of generating speech text from prompt images.

[0083] The following describes exemplary applications of the electronic devices provided in the embodiments of this application. These electronic devices can be implemented as various types of terminals such as laptops, tablets, desktop computers, set-top boxes, smartphones, smart speakers, smartwatches, smart TVs, and in-vehicle terminals, or as servers. Exemplary applications of the electronic devices as terminals or servers will be described below.

[0084] See Figure 1 , Figure 1 This is a schematic diagram of the architecture of the speech generation model training system 100 provided in the embodiments of this application. In order to support the training application of a speech generation model, the terminal 400 (exemplarily showing the graphical interface 410) connects to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0085] The method provided in this application can be applied to several different scenarios, which can improve the user's interactive experience, enhance the flexibility and convenience of human-computer interaction, and help to achieve more intelligent and personalized human-computer interaction.

[0086] 1. Assisted Communication: For people with hearing or speech impairments, sign language can be translated into text or speech by recognizing body movements and lip readings. For example, by recognizing finger positions, gesture shapes, and movement trajectories, sign language can be translated into text or speech. This allows people with hearing or speech limitations to communicate with others in real time through gestures.

[0087] 2. Gesture Control: Body movements and gesture recognition can be applied to smart devices or wearable devices, allowing users to control device functions through gestures, such as adjusting volume and switching songs. For example, the volume can be increased by swiping up the arm or raising the palm, and decreased by swiping down the arm or lowering the palm.

[0088] 3. Facial Expression Recognition: By recognizing facial expressions and lip movements, it's possible to determine a user's emotional state and intentions, thereby providing a better user experience and personalized services. Examples include automated sentiment analysis and emotional interactions within games. For instance, in games, the difficulty can be adjusted, corresponding items provided, or the storyline altered based on changes in the player's facial expressions, resulting in a more challenging and personalized gaming experience.

[0089] 4. Virtual Reality and Augmented Reality: By recognizing body movements and lip movements, users' actions and expressions can be transmitted in real time to virtual reality or augmented reality environments, achieving a more immersive interactive experience, such as the imitation of virtual character movements. For example, by recognizing users' body movements and lip movements, their actions and expressions can be mapped onto virtual characters or virtual scenes in real time, allowing actors or artists to control the actions and interactions of virtual characters through their own movements and expressions during performances, making the augmented reality experience more vivid.

[0090] Terminal 400 sends sample data of multiple modalities, including the prompt image sequence of the sample object and the corresponding speech text, as well as the prompt image sequence of the target object, to server 200 via network 300. Server 200 calls multiple encoders corresponding to the multiple modalities in the first speech generation model to encode the received sample data of multiple modalities of the sample object, obtaining a multimodal encoded vector sequence, and calls the decoder in the first speech generation model to decode the multimodal encoded vector sequence to obtain decoded text. Then, it determines the probability distribution of the decoded text and the sample data of multiple modalities, and determines the target loss based on the probability distribution. Based on the target loss, it updates the parameters of the decoder and at least one encoder to obtain a second speech generation model. Finally, server 200 generates target speech text corresponding to the prompt image sequence of the target object by calling the second speech generation model, and sends the target speech text to terminal 400 via network 300 for display through graphical interface 410. At the same time, the target speech text is processed by the speech signal generator of the second speech generation model to generate corresponding speech information, which is then played on the terminal.

[0091] In some embodiments, the terminal 400 calls multiple encoders corresponding to multiple modalities in the first speech generation model to encode the sample data of the received sample object in multiple modalities to obtain a multimodal encoded vector sequence, and calls the decoder in the first speech generation model to decode the multimodal encoded vector sequence to obtain decoded text; then, it determines the probability distribution of the decoded text and the sample data of multiple modalities, and determines the target loss based on the probability distribution; it updates the parameters of the decoder and at least one encoder based on the target loss to obtain a second speech generation model; finally, it generates target speech text corresponding to the prompt image sequence of the target object by calling the second speech generation model, and displays the target speech text through the graphical interface 410. At the same time, the target speech text passes through the speech signal generator of the second speech generation model to generate corresponding speech information, which is then broadcast on the terminal.

[0092] In some embodiments, server 200 may be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Terminals and servers can be connected directly or indirectly via wired or wireless communication, which is not limited in this embodiment.

[0093] The embodiments of this application can be implemented using artificial intelligence (AI) technology. AI is the theory, methods, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can react in a way similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.

[0094] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0095] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP deals with natural language—the language people use in daily life—and is closely related to linguistics research; it also involves crucial techniques for model training in computer science, mathematics, and artificial intelligence. Pre-trained models evolved from Large Language Models (LLMs) in NLP. After fine-tuning, large language models can be widely applied to downstream tasks. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.

[0096] The electronic device that implements the training method of the speech generation model provided in the embodiments of this application may be Figure 1 Terminal 400 or server 200. See also Figure 2A , Figure 2A This is a schematic diagram of the structure of the electronic device 500-1 provided in the embodiments of this application. Figure 2A The illustrated electronic device 500-1 includes at least one processor 510-1, at least one network interface 520-1, a user interface 530-1, and a memory 550-1. The various components in the electronic device 500-1 are coupled together via a bus system 540-1. It is understood that the bus system 540-1 is used to implement communication between these components. In addition to a data bus, the bus system 540-1 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 2A The general designated all buses as Bus System 540-1.

[0097] The processor 510-1 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0098] User interface 530-1 includes one or more output devices 531-1 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 530-1 also includes one or more input devices 532-1, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0099] The memory 550-1 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 550-1 may optionally include one or more storage devices physically located away from the processor 510-1.

[0100] The memory 550-1 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550-1 described in this application embodiment is intended to include any suitable type of memory.

[0101] In some embodiments, memory 550-1 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0102] Operating system 551-1 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, for implementing various basic business functions and handling hardware-based tasks;

[0103] The network communication module 552-1 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520-1, such as Bluetooth, WiFi, and Universal Serial Bus (USB).

[0104] Presentation module 553-1 is used to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531-1 (e.g., a display screen, a speaker, etc.) associated with user interface 530-1;

[0105] The input processing module 554-1 is used to detect and translate one or more user inputs or interactions from one or more input devices 532-1.

[0106] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2A A training device 555-1 for a language understanding model stored in memory 550-1 is shown. This device can be software in the form of programs or plug-ins, and includes the following software modules: an acquisition module 5551-1, an encoding module 5552-1, a decoding module 5553-1, a determination module 5554-1, and a generation module 5555-1. These modules are logically connected and can therefore be arbitrarily combined or further split according to their implemented functions. The functions of each module will be described below.

[0107] The electronic device that implements the speech generation method of the speech generation model provided in the embodiments of this application may be Figure 1 Terminal 400 or server 200. See also Figure 2B , Figure 2B This is a schematic diagram of the structure of the electronic device 500-2 provided in the embodiments of this application. Figure 2BThe illustrated electronic device 500-2 includes at least one processor 510-2, at least one network interface 520-2, a user interface 530-2, and a memory 550-2. The various components in the electronic device 500-2 are coupled together via a bus system 540-2. It is understood that the bus system 540-2 is used to implement communication between these components. In addition to a data bus, the bus system 540-2 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 2B The general designated all buses as Bus System 540-2.

[0108] The processor 510-2 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor.

[0109] User interface 530-2 includes one or more output devices 531-2 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 530-2 also includes one or more input devices 532-2, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0110] The memory 550-2 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 550-2 may optionally include one or more storage devices physically located away from the processor 510-2.

[0111] The memory 550-2 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550-2 described in this application embodiment is intended to include any suitable type of memory.

[0112] In some embodiments, the memory 550-2 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.

[0113] Operating system 551-2 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic business functions and handle hardware-based tasks.

[0114] The network communication module 552-2 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520-2, such as Bluetooth, WiFi, and Universal Serial Bus (USB).

[0115] Presentation module 553-2 is used to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531-2 (e.g., a display screen, a speaker, etc.) associated with user interface 530-2;

[0116] The input processing module 554-2 is used to detect and translate one or more user inputs or interactions from one or more input devices 532-2.

[0117] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2B A training device 555-2 for a language understanding model stored in memory 520-2 is shown. This device can be software in the form of programs and plug-ins, and includes the following software modules: an acquisition module 5551-2 and a generation module 5552-2. These modules are logically linked and can therefore be arbitrarily combined or further split according to their implemented functions. The functions of each module will be described below.

[0118] In some embodiments, the terminal or server can implement the language understanding model training method provided in this application by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.

[0119] The following describes the training method of the speech generation model provided in the embodiments of this application. As mentioned above, the electronic device implementing the training method of the speech generation model in the embodiments of this application can be a terminal or a server, or a combination of both. Therefore, the executing entity of each step will not be described again below.

[0120] See Figure 3A , Figure 3A This is a flowchart illustrating the training method of the speech generation model provided in this application embodiment, which will be combined with... Figure 3A The steps shown are explained.

[0121] In step 101, a first speech generation model is obtained, wherein the first speech generation model includes a decoder and multiple encoders corresponding to multiple modalities.

[0122] In some embodiments, the first speech generation model may be a first speech generation model pre-trained locally, or it may be a first speech generation model pre-trained by other devices obtained from other devices, including a decoder and multiple encoders corresponding to multiple modalities.

[0123] As an example, see Figure 4A , Figure 4A This is a schematic diagram of the first architecture of the speech generation model provided in the embodiments of this application. Figure 4A A text decoder, a visual encoder corresponding to the visual modality, and a text encoder corresponding to the text modality are shown.

[0124] In step 102, sample data of multiple modalities are acquired, wherein the sample data of multiple modalities includes a sequence of prompt images of sample objects and the corresponding speech text.

[0125] In some embodiments, the sample data for multiple modalities of the sample object includes visual modal sample data, such as a sequence of prompt images of the sample object, and text modal sample data, such as speech text corresponding to the prompt image sequence. Each prompt image in the prompt image sequence is a static image showing the hand or lip posture of the sample object when it makes a body movement or lip movement.

[0126] As an example, sample data of the visual modality of the sample object can be collected by means of camera shooting, image databases or network sources; audio text related to the sample data of the visual modality of the sample object can be recorded by recording equipment, or audio text matching the sample data of the visual modality of the sample object can be extracted from existing audio databases or corpora.

[0127] In step 103, multiple encoders are invoked to encode the prompt image sequence and the speech text respectively, resulting in a multimodal encoded vector sequence.

[0128] In some embodiments, see Figure 3B , Figure 3B This is a schematic diagram of the first process for obtaining a multimodal encoded vector sequence according to an embodiment of this application. When the multiple modalities include visual modalities and text modalities, Figure 3A Step 103 can be achieved through Figure 3B Steps 1031A to 1033A are implemented, and the details are explained below.

[0129] In step 1031A, the visual encoder is invoked to encode the prompt image sequence to obtain the visual encoding vector sequence of the visual modality, wherein the visual encoder is the encoder corresponding to the visual modality among multiple encoders.

[0130] In some embodiments, see Figure 3C , Figure 3C This is a schematic diagram of the process for obtaining a visual encoding vector sequence provided in an embodiment of this application. Figure 3B Step 1031A can be achieved through Figure 3C Steps 10311A to 10314A are implemented, and the details are explained below.

[0131] In step 10311A, multiple prompt images in the prompt image sequence are identified to obtain the location of key points in each prompt image.

[0132] In some embodiments, a deep learning model can be used to estimate the human pose of each of the multiple prompt images in the prompt image sequence, and identify the key point locations of the human body as the locations of the key points in each prompt image.

[0133] As an example, a deep learning model used to identify multiple prompt images in a prompt image sequence can be OpenPose, DeepPose, PoseNet, or DensePose.

[0134] In step 10312A, a key point sequence is generated based on the location of key points in each prompt image.

[0135] As an example, the location of key points in each prompt image could be the location of lips, fingers, etc. Based on the locations of key points in all prompt images within the corresponding video prompt image sequence, a key point sequence can be generated. See also Figure 9 ,based on Figure 9In the sequence of cue images shown, the points marking the positions of the hand and lips in each cue image are keypoints. Based on the positions of the keypoints in each cue image, the following is obtained: Figure 9 The key point sequence shown.

[0136] In step 10313A, the embedding vector of the keypoint sequence is obtained.

[0137] In some embodiments, a machine learning model can be used to perform pose embedding processing on each keypoint in the keypoint sequence to obtain the embedding vector of each keypoint in the keypoint sequence. The embedding vectors are then concatenated to obtain the embedding vector of the corresponding keypoint sequence.

[0138] As an example, the pose embedding of keypoint sequences can be performed using a recurrent neural network (RNN) or a convolutional neural network (CNN) to obtain the embedding vector of the keypoint sequence.

[0139] In step 10314A, the visual encoder is invoked to encode the embedding vector to obtain the visual encoding vector of the visual modality.

[0140] In some embodiments, a pre-trained machine learning model can be used as a visual encoder to encode the embedding vector to obtain a visual encoding vector of the visual modality.

[0141] As an example, the model of a visual encoder can be a convolutional neural network model, see [link to relevant documentation]. Figure 5A , Figure 5A This is a schematic diagram of the structure of the visual encoder model provided in the embodiments of this application. Figure 5A The convolutional neural network shown includes an input layer, a convolutional layer, a pooling layer, and an output layer (fully connected layer and softmax layer). After the embedding vector of the keypoint sequence of the visual modality is input into the convolutional neural network, the embedding vector is processed by multiple convolutional layers to obtain the feature map corresponding to the embedding vector. Then, the feature map is processed by multiple max pooling layers to obtain the feature vector of the visual modality. Finally, the output layer outputs the visual encoding vector sequence of the visual modality.

[0142] This application embodiment converts human actions in a video into a key point sequence of a visual modality, extracting key information from the prompt image in the form of a key point sequence. This facilitates subsequent analysis and recognition of the visual modality prompt image features, and effectively utilizes the characteristic information in the prompt image.

[0143] See also Figure 3B The following explanation will continue from step 1031A above.

[0144] In step 1032A, a text encoder is invoked to encode multiple morphemes in the speech text to obtain a text encoding vector sequence of the text modality. The text encoding vector sequence includes text encoding vectors corresponding to multiple morphemes, and the text encoder is the encoder corresponding to the text modality among multiple encoders.

[0145] In some embodiments, the speech text is segmented into multiple morphemes (e.g., words, phrases, or sentences) by language segmentation. The text encoder can be a machine learning model that encodes the multiple morphemes in the speech text separately, converts the morphemes into text encoding vectors, and combines the text encoding vectors corresponding to the morphemes in the speech text to obtain a text encoding vector sequence of the text modality.

[0146] As an example, machine learning models used to encode speech text can be recurrent neural networks, convolutional neural networks, Transformers, Bidirectional Encoder Representations from Transformer (BERT), etc.

[0147] See Figure 5B , Figure 5B This is a schematic diagram of the structure of the machine learning model provided in the embodiments of this application. Figure 5B The diagram illustrates the input embedding layer, positional embedding, self-attention mechanism, feed-forward neural network, and output layer. First, the input embedding layer converts the morphemes of the speech text into fixed-dimensional text embedding vectors, typically using word embedding techniques to map morphemes to vector representations in a continuous space. Second, the positional embedding layer embeds the positional information of each morpheme into the feature vector, resulting in positional embedding vectors. Then, the self-attention mechanism models the importance of each morpheme in the context and learns semantic relationships. The feed-forward neural network layer performs nonlinear transformations and mappings on the positional embedding vectors of each morpheme, thereby extracting higher-level semantic and contextual information. Finally, the entire text encoding vector sequence is generated based on the positional embedding vectors of each morpheme.

[0148] In step 1033A, the visual encoding vector sequence and the text encoding vector sequence are used as the first multimodal encoding vector sequence, wherein the first multimodal encoding vector sequence is used for decoding by the decoder.

[0149] In some embodiments, when the multiple modalities include a visual modality and a text modality, the multimodal encoding vector sequence includes a visual encoding vector sequence corresponding to the visual modality and a text encoding vector sequence corresponding to the text modality.

[0150] This application embodiment encodes multiple morphemes in the speech text into a text encoding vector sequence of the text modality, so that the encoding vector of each morpheme can capture its semantic information and contextual relationship, thereby better expressing the meaning of the entire text and improving the accuracy of subsequent speech generation.

[0151] In some embodiments, see Figure 3D , Figure 3D This is a schematic diagram of the second process for obtaining a multimodal encoded vector sequence provided in an embodiment of this application. When the multiple modalities also include a speech modality, the sample data of the multiple modalities also include the speech signal features of the speech modality. Figure 3A Step 103 can be achieved through Figure 3D Steps 1031B to 1032B are implemented, and the details are explained below.

[0152] In step 1031B, the speech encoder is invoked to encode the speech signal features to obtain the speech encoding vector sequence of the speech mode, wherein the speech encoder is the encoder corresponding to the speech mode among multiple encoders.

[0153] As an example, see Figure 4B , Figure 4B This is a schematic diagram of the second architecture of the speech generation model provided in the embodiments of this application. Figure 4B The diagram shows a text decoder, a visual encoder corresponding to the visual modality, a text encoder corresponding to the text modality, and a speech encoder corresponding to the speech modality.

[0154] In some embodiments, the speech encoder is invoked to encode the features of the speech signal to obtain a speech coding vector sequence of the speech modality. This can be achieved as follows: First, the original speech signal is preprocessed, such as noise removal, filtering, and volume normalization. Second, the preprocessed speech signal is divided into short frames, with a commonly used frame length of 20-40 milliseconds and typically 50% or 75% overlap. Then, the speech signal of each frame is multiplied by a window function, such as a Hamming window or a rectangular window, to reduce boundary artifacts between frames. A Fast Fourier Transform (FFT) is applied to each frame to convert the time-domain signal into a frequency-domain representation. Finally, the speech coding vector is extracted from the spectrum and concatenated to obtain the speech coding vector sequence.

[0155] As an example, a speech encoder model used to encode features of a speech signal can be a Mel-filterbank extractor.

[0156] In step 1032B, the visual coding vector sequence, the text coding vector sequence, and the speech coding vector sequence are used as the second multimodal coding vector sequence, wherein the second multimodal coding vector sequence is used to replace the first multimodal coding vector sequence for decoding by the decoder.

[0157] In some embodiments, when the multiple modalities include a visual modality, a text modality, and a speech modality, the multimodal coding vector sequence includes a visual coding vector sequence corresponding to the visual modality, a text coding vector sequence corresponding to the text modality, and a speech coding vector sequence corresponding to the speech modality.

[0158] In some embodiments, see Figure 3E , Figure 3E This is a schematic diagram of the third process for obtaining a multimodal encoded vector sequence provided in an embodiment of this application. When the multiple modalities also include a speech modality, Figure 3A Step 103 can also be done through Figure 3E Steps 1031C to 1034C are implemented, and the details are explained below.

[0159] In step 1031C, multiple encoders are invoked to encode the prompt image sequence and the speech text respectively, to obtain the visual encoding vector sequence of the visual modality, the text encoding vector sequence of the text modality, and the speech encoding vector sequence of the speech modality.

[0160] In some embodiments, as described above, when multiple modalities include a visual modality, a text modality, and a speech modality, a visual encoder can be invoked to encode the sample data of the visual modality to obtain a visual encoding vector sequence of the visual modality; a text encoder can be invoked to encode the sample data of the text modality to obtain a text encoding vector sequence of the text modality; and a speech encoder can be invoked to encode the sample data of the speech modality to obtain a speech encoding vector sequence of the speech modality.

[0161] In some embodiments, the sample data for multiple modalities further includes speech signal features of the speech modality; see [link to documentation]. Figure 3F , Figure 3F This is a schematic diagram of the fourth process for obtaining a multimodal encoded vector sequence provided in an embodiment of this application. Figure 3E Step 1031C can be achieved through Figure 3F Steps 10311C to 10313C are implemented, and the details are explained below.

[0162] In step 10311C, the visual encoder is invoked to encode the prompt image sequence to obtain the visual encoding vector sequence of the visual modality, wherein the visual encoder is the encoder corresponding to the visual modality among multiple encoders.

[0163] In some embodiments, the process of calling a visual encoder to encode the visual modality prompt image sequence to obtain the visual modality visual encoding vector sequence is described in steps 10311A to 10314A above, and will not be repeated here.

[0164] In step 10312C, a text encoder is invoked to encode multiple morphemes in the speech text to obtain a text encoding vector sequence of the text modality. The text encoding vector sequence includes text encoding vectors corresponding to multiple morphemes, and the text encoder is the encoder corresponding to the text modality among multiple encoders.

[0165] In some embodiments, the process of calling a text encoder to encode multiple morphemes in the speech text to obtain a text encoding vector sequence of the text modality is described in step 1032A above, and will not be repeated here.

[0166] In step 10313C, the speech encoder is invoked to encode the speech signal features to obtain the speech encoding vector sequence of the speech modality, wherein the speech encoder is the encoder corresponding to the speech modality among multiple encoders.

[0167] In some embodiments, the process of calling a speech encoder to encode the features of the speech signal to obtain a speech encoding vector sequence of the speech modality is described in step 1031B above, and will not be repeated here.

[0168] See also Figure 3EThe following will be explained following step 1031C above.

[0169] In step 1032C, the visual encoding vector sequence and the speech encoding vector sequence are concatenated to obtain a concatenated encoding vector sequence.

[0170] In some embodiments, the feature vectors in the visual coding vector sequence and the feature vectors in the speech coding vector sequence can be concatenated to obtain a concatenated coding vector. The concatenated coding vectors corresponding to each position can be combined to obtain a concatenated coding vector sequence.

[0171] As an example, if the visual encoding vectors in the visual encoding vector sequence output by the visual encoder are The speech encoding vectors in the speech encoding vector sequence output by the speech encoder are Where f represents the length of the encoded vector sequence, and D represents the dimension of the encoded vectors in the encoded vector sequence, then the visual encoded vector sequence and the speech encoded vector sequence are concatenated to obtain the concatenated encoded vector.

[0172] In step 1033C, the concatenated encoded vector sequence is subjected to feature mapping to obtain the fused encoded vector sequence.

[0173] As an example, in the above example, the concatenated encoded vector in the concatenated encoded vector sequence has a dimension of 2D. By performing feature mapping on the concatenated encoded vector sequence, the dimension of the concatenated encoded vector is transformed from 2D to D, resulting in the fused encoded vector. The fused coding vectors at each position are combined to obtain a fused coding vector sequence.

[0174] In step 1034C, the fused encoding vector sequence and the text encoding vector sequence are used as the third multimodal encoding vector sequence, which is used to replace the second multimodal encoding vector sequence for decoding by the decoder.

[0175] As an example, see Figure 4C , Figure 4C This is a schematic diagram of the third architecture of the speech generation model provided in the embodiments of this application. Figure 4C The diagram illustrates a text decoder, a visual encoder corresponding to the visual modality, a text encoder corresponding to the text modality, and a speech encoder corresponding to the speech modality, as shown below. Figure 4C As shown, the visual encoding vector sequence output by the visual encoder and the speech encoding vector sequence output by the speech encoder are concatenated and mapped to obtain a fused encoding vector sequence, which is then combined with the output of the text encoder as a third multimodal encoding vector sequence.

[0176] This application embodiment calls the encoders corresponding to multiple modalities in the first speech generation model to encode sample data of multiple different modalities, obtaining a multimodal encoding vector sequence. Furthermore, it provides three different modal combination encoding methods, enabling the first speech generation model to fully learn the features of different modalities of visual, text, and speech modes. This improves the speech generation model's ability to perform speech generation tasks and increases the accuracy of speech generation. At the same time, the combination of different modalities ensures that the execution of speech generation tasks can adapt to different application scenarios, thereby enhancing the user experience.

[0177] See also Figure 3A The following will be an explanation following step 103 above.

[0178] In step 104, the decoder is invoked based on the multimodal encoded vector sequence to perform decoding, thereby obtaining the decoded text.

[0179] In some embodiments, the first speech generation model is a pre-trained speech recognition model, and after pre-training, the representation space of the text modality and the representation space of the speech modality of the first speech generation model have been unified to the same representation space; see also Figure 3G , Figure 3G This is a schematic diagram of the process for aligning the length of the encoded vector sequence provided in an embodiment of this application. During execution... Figure 3A Before step 104, it can be done through Figure 3G Steps 201 to 202 align the length of the multimodal encoded vector sequence, as detailed below.

[0180] In step 201, the visual encoding vector sequence and the text encoding vector sequence are mapped to obtain the vector sequence representation of the visual encoding vector sequence and the text encoding vector sequence in the representation space.

[0181] In some embodiments, a linear mapping function is used to map the visual encoded vector sequence and the text encoded vector sequence to obtain a vector sequence representation of the visual encoded vector sequence and the text encoded vector sequence in the same representation space.

[0182] As an example, see Figure 6 , Figure 6 This is a schematic diagram of the alignment of the encoding vector sequence provided in the embodiments of this application. The visual encoding vector sequence is represented as “E1 E2 E3 E3 E4 E5 E5 E6” in the representation space, and the text encoding vector sequence is represented as “T1 T2 T3 T4 T5 T6” in the representation space.

[0183] In step 202, the length of the vector sequence representation of the visual encoding vector sequence is adjusted to be the same as the length of the vector sequence representation of the text encoding vector sequence.

[0184] As an example, in the example above, the length of the vector sequence representation of the visual encoding vector sequence is 8, and the length of the vector sequence representation of the text encoding vector sequence is 6. Since there is redundant information "E3" and "E5" in the vector sequence representation of the visual encoding vector sequence, the redundant information in the vector sequence representation of the visual encoding vector sequence is deleted, and the vector sequence representation of the visual encoding vector sequence is obtained as "E1 E2 E3 E4 E5 E6". At this time, the length of the vector sequence representation of the visual encoding vector sequence is the same as the length of the vector sequence representation of the text encoding vector sequence.

[0185] The training process of the first speech generation model in this embodiment is based on a pre-trained speech recognition model. Since the representation space of the text modality and the representation space of the speech modality have been unified into the same representation space, the conversion function from prompt image sequence to speech text can be realized on the basis of the original function of the pre-trained speech recognition model through fine-tuning. At the same time, multiple functions of the pre-trained speech recognition model itself can be reused, such as recognizing the prompt image sequence of the target object, translating the text, and generating speech, avoiding repeated development and improving development efficiency.

[0186] Aligning the feature sequence lengths of two modalities makes them more consistent in the feature space, which is beneficial for speech generation models to perform speech generation tasks. This process can improve the overall efficiency and performance of speech generation models when processing multimodal data.

[0187] See also Figure 3A The following will be an explanation following step 104 above.

[0188] In step 105, the probability distribution of the decoded text and sample data of multiple modalities is determined, and the target loss is determined based on the probability distribution.

[0189] In some embodiments, determining the probability distribution of the decoded text and sample data of multiple modalities in step 105 can be achieved by: determining the probability distribution of the encoded vector sequences of sample data of at least two modalities relative to preset conditions, wherein different modalities correspond to different preset conditions, which may specifically depend on the data involved in the model's encoding structure; determining the probability distribution of the speech text relative to sample data of at least one modality, that is, the conditional probability distribution of the speech text conditioned on sample data of at least one modality; and determining the probability distribution of the decoded text relative to sample data of at least one modality, that is, the conditional probability distribution of the decoded text conditioned on sample data of at least one modality. Here, different modalities correspond to different conditions, which may specifically depend on the model's encoding structure.

[0190] For example, in Figure 4A The architecture of the speech generation model shown includes a text decoder, a visual encoder corresponding to the visual modality, and a text encoder corresponding to the text modality. Figure 4A The speech generation model only includes two modalities: visual and text. Therefore, the probability distribution of the decoded text and sample data from multiple modalities is determined based on the visual and text modalities, and the target loss is determined based on the probability distribution. For... Figure 4A The architecture allows for preset conditions for text modalities to be decoded text and speech text, and preset conditions for visual modalities to be decoded text and a sequence of prompt images; speech text is accompanied by a sequence of prompt images; and decoded text is accompanied by a sequence of prompt images.

[0191] For example, in Figure 4B The architecture of the speech generation model shown includes a text decoder, a visual encoder corresponding to the visual modality, a text encoder corresponding to the text modality, and a speech encoder corresponding to the speech modality. Figure 4B The speech generation model compared to Figure 4A Furthermore, training incorporates language modalities, thus determining the probability distribution of the decoded text and sample data from multiple modalities based on the speech and text modalities, and subsequently determining the target loss based on the probability distribution. This is specifically for... Figure 4B In this architecture, the preset conditions for the text modality can be decoded text and speech text, and the preset conditions for the speech modality can be decoded text and speech signal features; the condition for speech text is speech signal features; and the condition for decoded text is speech signal features.

[0192] For example, in Figure 4C The illustrated speech generation model architecture performs feature concatenation mapping on the visual encoded vector sequence output by the visual encoder and the speech encoded vector sequence output by the speech encoder to obtain a fused encoded vector sequence, which is then combined with the output of the text encoder as a third multimodal encoded vector sequence. Therefore, based on the visual modality, speech modality, and text modality, the probability distribution of the decoded text and sample data from multiple modalities is determined, and the target loss is then determined based on the probability distribution. For... Figure 4C The architecture allows for the following preset conditions: for text modalities, the preset conditions can be decoded text and speech text; for fused encoded vector sequences, the preset conditions can be decoded text, prompt image sequences, and speech signal features; for speech text, the preset conditions are prompt image sequences and speech signal features; and for decoded text, the preset conditions are prompt image sequences and speech signal features.

[0193] In some embodiments, determining the target loss based on the probability distribution in step 105 can be achieved in the following ways:

[0194] For encoded vector sequences of sample data from at least two modalities, the difference in probability distribution of encoded vector sequences of different modalities is determined based on the probability distribution of encoded vector sequences of sample data from at least two modalities relative to preset conditions, and used as the sub-target loss of encoded vector sequences; based on the probability distribution of speech text relative to sample data from at least one modality, the sub-target loss of speech text negatively correlated with the probability distribution is determined; based on the probability distribution of decoded text relative to sample data from at least one modality, the sub-target loss of decoded text negatively correlated with the probability distribution is determined.

[0195] For example, in Figure 4A In the architecture shown, the difference in probability distribution between the encoded vector sequences of visual and text modal samples relative to preset conditions is determined, and used as the sub-target loss of the encoded vector sequences. Based on the probability distribution of speech text relative to visual modal sample data, the sub-target loss of speech text negatively correlated with the probability distribution is determined. Based on the probability distribution of decoded text relative to visual modal sample data, the sub-target loss of speech text negatively correlated with the probability distribution is determined. The sub-target loss of the encoded vector sequences, the sub-target loss of speech text, and the sub-target loss of decoded text are fused to obtain the target loss.

[0196] For example, in Figure 4B In the architecture shown, the difference in probability distribution between the encoded vector sequences of the speech modality and the text modality is determined based on the probability distribution of the encoded vector sequences relative to preset conditions, and is used as the sub-target loss of the encoded vector sequences. Based on the probability distribution of the speech text relative to the sample data of the speech modality, the sub-target loss of the speech text that is negatively correlated with the probability distribution is determined. Based on the probability distribution of the decoded text relative to the sample data of the speech modality, the sub-target loss of the speech text that is negatively correlated with the probability distribution is determined. The sub-target loss of the encoded vector sequences, the sub-target loss of the speech text, and the sub-target loss of the decoded text are fused to obtain the target loss.

[0197] For example, in Figure 4CIn the architecture shown, the difference in probability distribution between the encoded vector sequences of visual, speech, and text modal sample data and the encoded vector sequence of the text modal is determined based on the probability distribution of these sequences relative to preset conditions, and is used as the sub-target loss of the encoded vector sequences. Based on the probability distribution of speech text relative to the sample data of visual and speech modal data, the sub-target loss of speech text negatively correlated with the probability distribution is determined. Based on the probability distribution of decoded text relative to the sample data of visual and speech modal data, the sub-target loss of speech text negatively correlated with the probability distribution is determined. The sub-target losses of the encoded vector sequences, speech text, and decoded text are fused to obtain the target loss.

[0198] Below, in conjunction with Figures 4A to 4C The three different architectures of the speech generation model are shown and explained in detail.

[0199] against Figure 4A The architecture of the speech generation model shown is available in [reference]. Figure 3H , Figure 3H This is a schematic diagram of the first process for determining the probability distribution provided in an embodiment of this application. Figure 3A Step 105, "determining the probability distribution of the decoded text and sample data of multiple modalities," can be achieved by determining the probability distribution of the encoded vector sequences of the visual and text modal sample data relative to preset conditions, the probability distribution of the speech text relative to the prompt image sequence, and the probability distribution of the decoded text relative to the prompt image sequence, as described above, using the encoded vector sequences of the visual and text modal sample data. The following will combine... Figure 3H Detailed explanation of steps 1051A to 1053A.

[0200] In step 1051A, a first probability distribution of the text encoded vector sequence relative to the decoded text and the speech text is determined, and a second probability distribution of the visual encoded vector sequence relative to the decoded text and the prompt image sequence is determined.

[0201] In some embodiments, a first probability distribution of the text-encoded vector sequence relative to the decoded text and speech text can be determined by a pre-trained first probability distribution model; that is, the conditional probability distribution of the text-encoded vector sequence relative to the decoded text and speech text. The first probability distribution model can be trained as follows: First, a first sample dataset is obtained, including decoded text, speech text, and a reference probability distribution of the text-encoded vector sequence relative to the decoded text and speech text. Then, the decoded text and speech text in the sample data are used as input to the initialized first probability distribution model, outputting the predicted probability distribution of the text-encoded vector sequence relative to the decoded text and speech text. Finally, a loss value is determined based on the difference between the predicted probability distribution of the text-encoded vector sequence relative to the decoded text and speech text and the reference probability distribution, and the parameters of the first probability distribution model are updated based on the loss value using a backpropagation algorithm.

[0202] As an example, the first probability distribution of the text encoded vector sequence relative to the decoded text and the speech text can be represented as follows: The first probability distribution model can be a Conditional Random Field (CRF) or a Hidden Markov Model (HMM).

[0203] As an example, the difference operation can be used to calculate the difference between the predicted probability distribution of the text encoded vector sequence relative to the decoded text and the reference probability distribution, and the square or absolute value of the difference can be taken as the loss value; or the exponential operation can be used to calculate the cross-entropy between the predicted probability distribution of the text encoded vector sequence relative to the decoded text and the reference probability distribution, and take it as the loss value.

[0204] In some embodiments, a second probability distribution of the visual encoded vector sequence relative to the decoded text and prompt image sequences can be determined using a pre-trained second probability distribution model; that is, the conditional probability distribution of the visual encoded vector sequence relative to the decoded text and prompt image sequences. The second probability distribution model can be trained as follows: First, a second sample dataset is obtained, including the decoded text, prompt image sequences, and a reference probability distribution of the visual encoded vector sequence relative to the decoded text and prompt image sequences. Then, the decoded text and prompt image sequences from the sample data are used as input to the initialized second probability distribution model, outputting the predicted probability distribution of the visual encoded vector sequence relative to the decoded text and prompt image sequences. Finally, a loss value is determined based on the difference between the predicted probability distribution of the visual encoded vector sequence relative to the decoded text and prompt image sequences and the reference probability distribution, and the parameters of the second probability distribution model are updated based on the loss value using a backpropagation algorithm.

[0205] As an example, the second probability distribution of the visual encoded vector sequence relative to the decoded text and prompt image sequences can be represented as follows: The second probability distribution model can be a conditional random field or a hidden Markov model.

[0206] As an example, the difference operation can be used to calculate the difference between the predicted probability distribution of the visual encoded vector sequence relative to the decoded text and prompt image sequence and the reference probability distribution, and the square or absolute value of the difference can be taken as the loss value; or the exponential operation can be used to calculate the cross-entropy between the predicted probability distribution of the visual encoded vector sequence relative to the decoded text and prompt image sequence and the reference probability distribution, and take it as the loss value.

[0207] In step 1052A, a third probability distribution of the speech text relative to the prompt image sequence is determined.

[0208] In some embodiments, a pre-trained third probability distribution model can be used to determine the third probability distribution of the speech text relative to the prompt image sequence, i.e., the conditional probability distribution of the speech text relative to the prompt image sequence. The third probability distribution model can be trained as follows: First, a third sample dataset is obtained, including the prompt image sequence and a reference probability distribution of the speech text relative to the prompt image sequence; then, the prompt image sequence from the sample data is used as input to the initialized third probability distribution model, outputting the predicted probability distribution of the speech text relative to the prompt image sequence; finally, a loss value is determined based on the difference between the predicted probability distribution of the speech text relative to the prompt image sequence and the reference probability distribution, and the parameters of the third probability distribution model are updated based on the loss value using a backpropagation algorithm.

[0209] As an example, the third probability distribution of the voice text relative to the prompt image sequence can be represented as p(x text |x cued-speech The third probability distribution model can be a conditional random field or a hidden Markov model. Difference operations can be used to calculate the difference between the predicted probability distribution of the speech text relative to the prompt image sequence and the reference probability distribution, and the square or absolute value of the difference can be taken as the loss value; or exponential operations can be used to calculate the cross-entropy between the predicted probability distribution of the speech text relative to the prompt image sequence and the reference probability distribution, and take this as the loss value.

[0210] In step 1053A, when the target language of the speech generation task of the first speech generation model is different from the source language, a fourth probability distribution of the decoded text relative to the prompt image sequence is determined.

[0211] In some embodiments, when the target language of the speech generation task of the first speech generation model is different from the source language, for example, the target language of the speech generation task of the first speech generation model is English and the source language is Chinese, then it is necessary to first translate the Chinese source language into English.

[0212] In some embodiments, a pre-trained fourth probability distribution model can be used to determine the fourth probability distribution of the decoded text relative to the prompt image sequence, i.e., the conditional probability distribution of the decoded text relative to the prompt image sequence. The fourth probability distribution model can be trained as follows: First, a fourth sample dataset is obtained, including the prompt image sequence, the source speech text, and a reference probability distribution of the decoded text relative to the prompt image sequence; then, the prompt image sequence and the source speech text in the sample data are used as input to the initialized fourth probability distribution model, outputting the predicted probability distribution of the decoded text relative to the prompt image sequence; finally, a loss value is determined based on the difference between the predicted probability distribution of the decoded text relative to the prompt image sequence and the reference probability distribution, and the parameters of the fourth probability distribution model are updated based on the loss value using a backpropagation algorithm.

[0213] As an example, the fourth probability distribution of the decoded text relative to the prompt image sequence can be represented as follows: The fourth probability distribution model can be a conditional random field or a hidden Markov model. Difference operations can be used to calculate the difference between the predicted probability distribution of the decoded text relative to the prompt image sequence and the reference probability distribution, and the square or absolute value of the difference can be taken as the loss value; or exponential operations can be used to calculate the cross-entropy between the predicted probability distribution of the decoded text relative to the prompt image sequence and the reference probability distribution, and take this as the loss value.

[0214] In some embodiments, see Figure 3I , Figure 3I This is a schematic diagram of the first process for determining the target loss provided in an embodiment of this application. Figure 3A Step 105, "determining the target loss based on probability distribution," can be achieved by determining the difference in probability distribution between the encoded vector sequences of the visual and text modal data relative to preset conditions, as described above, and using this difference as the sub-target loss of the encoded vector sequence. This is further achieved by determining the sub-target loss of the speech text negatively correlated with the probability distribution of the speech text relative to the visual modal sample data, and by determining the sub-target loss of the speech text negatively correlated with the probability distribution of the decoded text relative to the visual modal sample data. The following will combine... Figure 3I Steps 1054A to 1057A are implemented, and the details are explained below.

[0215] In step 1054A, the difference between the first probability distribution and the second probability distribution is determined as the first sub-target loss.

[0216] In some embodiments, the difference between a first probability distribution and a second probability distribution can be determined by divergence calculation. Examples include KL divergence and cross entropy.

[0217] As an example, the loss of the first sub-target can be determined using formula (1):

[0218]

[0219] in, The first probability distribution, For the second probability distribution, y text It is the decoded text corresponding to the target language, x cued-speech and x text These are the source language prompt image sequence and the corresponding source language audio text, θ csE These are the parameters of the visual encoder.

[0220] In step 1055A, the second sub-target loss that is negatively correlated with the third probability distribution is determined.

[0221] As an example, the second sub-target loss that is negatively correlated with the third probability distribution can be determined by the following formula (2).

[0222]

[0223] Where, θ tD These are the parameters of the decoder.

[0224] In step 1056A, the loss of the third sub-target that is negatively correlated with the fourth probability distribution is determined.

[0225] As an example, the loss of the third sub-target that is negatively correlated with the fourth probability distribution can be determined by the following formula (3).

[0226]

[0227] In step 1057A, the first sub-target loss, the second sub-target loss, and the third sub-target loss are fused to obtain the first target loss.

[0228] As an example, the first target loss is obtained by weighted summing the first sub-target loss, the second sub-target loss, and the third sub-target loss using the formula (4) shown below.

[0229]

[0230] Here, α is the weighting parameter of the loss for the third sub-target.

[0231] The following is combined with Figure 4B The architecture of the shown speech generation modality is illustrated. See [link / reference] Figure 3J , Figure 3J This is a schematic diagram of the second process for determining the probability distribution provided in an embodiment of this application. Figure 3A Step 105, "determining the probability distribution of the decoded text and sample data of multiple modalities," can also be achieved by determining the probability distribution of the encoded vector sequences of the sample data for the speech and text modalities relative to preset conditions, the probability distribution of the speech text relative to the speech signal features, and the probability distribution of the decoded text relative to the speech signal features, as described above, using the encoded vector sequences of the sample data for the speech and text modalities. The following will combine... Figure 3J Steps 1051B to 1053B are implemented, and the details are explained below.

[0232] In step 1051B, a fifth probability distribution of the text encoding vector sequence relative to the decoded text and the speech text, and a sixth probability distribution of the speech encoding vector sequence relative to the decoded text and the speech signal features are determined.

[0233] In some embodiments, the fifth probability distribution of the text encoded vector sequence relative to the decoded text and the speech text can be determined by a pre-trained fifth probability distribution model; that is, the conditional probability distribution of the text encoded vector sequence relative to the decoded text and the speech text. Here, the fifth probability distribution is only used to distinguish it from the first probability distribution mentioned above. The training process of the fifth probability distribution model is the same as the training process of the first probability distribution model described above, and will not be repeated here.

[0234] In some embodiments, a pre-trained sixth probability distribution model can be used to determine the sixth probability distribution of the speech encoded vector sequence relative to the decoded text and speech signal features, i.e., the conditional probability distribution of the speech encoded vector sequence relative to the decoded text and speech signal features. The sixth probability distribution model can be trained as follows: First, a sixth sample dataset is obtained, including decoded text, speech signal features, and a reference probability distribution of the speech encoded vector sequence relative to the decoded text and speech signal features; then, the decoded text and speech signal features in the sample data are used as input to the initialized sixth probability distribution model, outputting the predicted probability distribution of the speech encoded vector sequence relative to the decoded text and speech signal features; finally, a loss value is determined based on the difference between the predicted probability distribution of the speech encoded vector sequence relative to the decoded text and speech signal features and the reference probability distribution, and the parameters of the sixth probability distribution model are updated based on the loss value using a backpropagation algorithm.

[0235] As an example, the sixth probability distribution of the speech encoded vector sequence relative to the features of the decoded text and speech signal can be represented as follows: The sixth probability distribution model can be a conditional random field or a hidden Markov model.

[0236] As an example, the difference operation can be used to calculate the difference between the predicted probability distribution and the reference probability distribution of the speech encoded vector sequence relative to the features of the decoded text and speech signal, and the square or absolute value of the difference can be taken as the loss value; or the exponential operation can be used to calculate the cross-entropy between the predicted probability distribution and the reference probability distribution of the speech encoded vector sequence relative to the features of the decoded text and speech signal, and take it as the loss value.

[0237] In step 1052B, the seventh probability distribution of the speech text relative to the speech signal features is determined.

[0238] In some embodiments, the seventh probability distribution of speech text relative to speech signal features, i.e., the conditional probability distribution of speech text relative to speech signal features, can be determined by a pre-trained seventh probability distribution model. The seventh probability distribution model can be trained as follows: First, a seventh sample dataset is obtained, which includes speech signal features and a reference probability distribution of speech text relative to the speech signal features; then, the speech signal features in the sample data are used as input to the initialized seventh probability distribution model, and the predicted probability distribution of speech text relative to the speech signal features is output; finally, a loss value is determined based on the difference between the predicted probability distribution of speech text relative to the speech signal features and the reference probability distribution, and the parameters of the seventh probability distribution model are updated based on the loss value using a backpropagation algorithm.

[0239] As an example, the seventh probability distribution of speech text relative to the features of the speech signal can be represented as p(x text |x impaired-sDeech The seventh probability distribution model can be a conditional random field or a hidden Markov model. Difference operations can be used to calculate the difference between the predicted probability distribution of the speech text relative to the speech signal features and the reference probability distribution, and the square or absolute value of the difference can be taken as the loss value; or exponential operations can be used to calculate the cross-entropy between the predicted probability distribution of the speech text relative to the speech signal features and the reference probability distribution, and take this as the loss value.

[0240] In step 1053B, when the target language of the speech generation task of the first speech generation model is different from the source language, the eighth probability distribution of the decoded text relative to the speech signal features is determined.

[0241] In some embodiments, when the target language of the speech generation task of the first speech generation model is different from the source language, the eighth probability distribution of the decoded text relative to the speech signal features, i.e., the conditional probability distribution of the decoded text relative to the speech signal features, can be determined by a pre-trained eighth probability distribution model. The eighth probability distribution model can be trained as follows: First, an eighth sample dataset is obtained, which includes speech signal features, source speech text, and a reference probability distribution of the decoded text relative to the speech signal features; then, the speech signal features and source speech text in the sample data are used as input to the initialized eighth probability distribution model, and the predicted probability distribution of the decoded text relative to the speech signal features is output; finally, a loss value is determined based on the difference between the predicted probability distribution of the decoded text relative to the speech signal features and the reference probability distribution, and the parameters of the eighth probability distribution model are updated based on the loss value using a backpropagation algorithm.

[0242] As an example, the eighth probability distribution of the decoded text relative to the features of the speech signal can be represented as follows: The eighth probability distribution model can be a conditional random field or a hidden Markov model.

[0243] As an example, the difference operation can be used to calculate the difference between the predicted probability distribution of the decoded text relative to the speech signal features and the reference probability distribution, and the square or absolute value of the difference can be taken as the loss value; or the exponential operation can be used to calculate the cross-entropy between the predicted probability distribution of the decoded text relative to the speech signal features and the reference probability distribution, and take it as the loss value.

[0244] In some embodiments, see Figure 3K , Figure 3K This is a schematic diagram of the second process for determining the target loss provided in an embodiment of this application. Figure 3A Step 105, "determining the target loss based on probability distribution," can also be achieved by using the probability distribution of the encoded vector sequences of the speech modality and text modality sample data relative to preset conditions, as described above, to determine the difference in probability distribution between the encoded vector sequences of the speech modality and text modality, which serves as the sub-target loss of the encoded vector sequence; determining the sub-target loss of speech text negatively correlated with the probability distribution based on the probability distribution of the speech text relative to the speech modality sample data; and determining the sub-target loss of speech text negatively correlated with the probability distribution based on the probability distribution of the decoded text relative to the speech modality sample data. This will be discussed in the following sections. Figure 3K Steps 1054B to 1057B are implemented, and the details are explained below.

[0245] In step 1054B, the difference between the fifth probability distribution and the sixth probability distribution is determined as the fourth sub-target loss.

[0246] In some embodiments, the difference between the fifth and sixth probability distributions can be determined by divergence calculations. Examples include KL divergence and cross-entropy.

[0247] As an example, the loss of the fourth sub-target can be determined using formula (5):

[0248]

[0249] Where, x impaired-speech It is the speech signal feature in the sample data, θ isE These are parameters of the speech encoder. The meanings of the other parameters are explained above and will not be repeated here.

[0250] In step 1055B, the loss of the fifth sub-target that is negatively correlated with the seventh probability distribution is determined.

[0251] As an example, the loss of the fifth sub-target that is negatively correlated with the seventh probability distribution can be determined by the following formula (6).

[0252]

[0253] In step 1056B, the loss of the sixth sub-target that is negatively correlated with the eighth probability distribution is determined.

[0254] As an example, the loss of the sixth sub-target that is negatively correlated with the eighth probability distribution can be determined by the following formula (7).

[0255]

[0256] In step 1057B, the fourth sub-target loss, the fifth sub-target loss, and the sixth sub-target loss are fused to obtain the second target loss.

[0257] As an example, the second target loss is obtained by weighted summing the losses of the fourth, fifth, and sixth sub-targets using the formula (4) shown below.

[0258]

[0259] Here, α is the weighting parameter of the loss for the sixth sub-target.

[0260] The following is combined with Figure 4C The architecture of the shown speech generation modality is illustrated. See [link / reference] Figure 3L , Figure 3L This is a schematic diagram of the third process for determining the probability distribution provided in the embodiments of this application. Figure 3AStep 105, "determining the probability distribution of the decoded text and sample data of multiple modalities," can also be achieved by determining the probability distribution of the encoded vector sequences of the sample data of the visual modality, speech modality, and text modality relative to preset conditions using the encoded vector sequences of the sample data for the visual modality, speech modality, and text modality, as described above; determining the probability distribution of the speech text relative to the prompt image sequence and speech signal features; and determining the probability distribution of the decoded text relative to the prompt image sequence and speech signal features. The following will combine... Figure 3L Steps 1051C to 1053C are implemented, and the details are explained below.

[0261] In step 1051C, the ninth probability distribution of the text encoded vector sequence relative to the decoded text and the speech text is determined, as well as the tenth probability distribution of the fused encoded vector sequence relative to the decoded text, the prompt image sequence, and the speech signal features.

[0262] In some embodiments, the ninth probability distribution of the text encoded vector sequence relative to the decoded text and the speech text can be determined by a pre-trained ninth probability distribution model; that is, the conditional probability distribution of the text encoded vector sequence relative to the decoded text and the speech text. Here, the ninth probability distribution is only used to distinguish it from the first probability distribution mentioned above. The training process of the ninth probability distribution model is the same as the training process of the first probability distribution model described above, and will not be repeated here.

[0263] In some embodiments, the tenth probability distribution of the fused encoded vector sequence relative to the decoded text, prompt image sequence, and speech signal features can be determined using a pre-trained tenth probability distribution model; that is, the conditional probability distribution of the fused encoded vector sequence relative to the decoded text, prompt image sequence, and speech signal features. The tenth probability distribution model can be trained as follows: First, a tenth sample dataset is obtained, including decoded text, prompt image sequence, speech signal features, and a reference probability distribution of the fused encoded vector sequence relative to the decoded text, prompt image sequence, and speech signal features. Then, the decoded text, prompt image sequence, and speech signal features in the sample data are used as input to the initialized tenth probability distribution model, outputting the predicted probability distribution of the fused encoded vector sequence relative to the decoded text, prompt image sequence, and speech signal features. Finally, a loss value is determined based on the difference between the predicted probability distribution of the fused encoded vector sequence relative to the decoded text, prompt image sequence, and speech signal features and the reference probability distribution, and the parameters of the tenth probability distribution model are updated based on the loss value using a backpropagation algorithm.

[0264] As an example, the tenth probability distribution of the fused encoded vector sequence relative to the features of the decoded text, the prompt image sequence, and the speech signal can be represented as follows: The tenth probability distribution model can be a conditional random field or a hidden Markov model.

[0265] As an example, the difference operation can be used to calculate the difference between the predicted probability distribution and the reference probability distribution of the fused encoded vector sequence relative to the decoded text, prompt image sequence, and speech signal features, and take the square or absolute value of the difference as the loss value; or the exponential operation can be used to calculate the cross-entropy between the predicted probability distribution and the reference probability distribution of the fused encoded vector sequence relative to the decoded text, prompt image sequence, and speech signal features, and take it as the loss value.

[0266] In step 1052C, the eleventh probability distribution of the speech text relative to the prompt image sequence and speech signal features is determined.

[0267] In some embodiments, the eleventh probability distribution of the speech text relative to the prompt image sequence and speech signal features can be determined by a pre-trained eleventh probability distribution model, i.e., the conditional probability distribution of the speech text relative to the prompt image sequence and speech signal features. The eleventh probability distribution model can be trained as follows: First, an eleventh sample dataset is obtained, which includes the prompt image sequence, speech signal features, and a reference probability distribution of the speech text relative to the prompt image sequence and speech signal features; then, the prompt image sequence and speech signal features in the sample data are used as input to the initialized eleventh probability distribution model, and the predicted probability distribution of the speech text relative to the prompt image sequence and speech signal features is output; finally, a loss value is determined based on the difference between the predicted probability distribution of the speech text relative to the prompt image sequence and speech signal features and the reference probability distribution, and the parameters of the eleventh probability distribution model are updated based on the loss value using the backpropagation algorithm.

[0268] As an example, the eleventh probability distribution of the speech text relative to the prompt image sequence and speech signal features can be represented as p(x text |x cued-speech ,x impaired-speech The eleventh probability distribution model can be a conditional random field or a hidden Markov model.

[0269] As an example, the difference operation can be used to calculate the difference between the predicted probability distribution of the speech text relative to the prompt image sequence and the speech signal features and the reference probability distribution, and the square or absolute value of the difference can be taken as the loss value; or the exponential operation can be used to calculate the cross-entropy between the predicted probability distribution of the speech text relative to the prompt image sequence and the speech signal features and the reference probability distribution, and take it as the loss value.

[0270] In step 1053C, when the target language of the speech generation task of the first speech generation model is different from the source language, the twelfth probability distribution of the decoded text relative to the prompt image sequence and speech signal features is determined.

[0271] In some embodiments, when the target language of the speech generation task of the first speech generation model is different from the source language, the twelfth probability distribution of the decoded text relative to the prompt image sequence and speech signal features can be determined by a pre-trained twelfth probability distribution model, i.e., the conditional probability distribution of the decoded text relative to the prompt image sequence and speech signal features. The twelfth probability distribution model can be trained as follows: First, a twelfth sample dataset is obtained, which includes the prompt image sequence and speech signal features, the source speech text, and the reference probability distribution of the decoded text relative to the prompt image sequence and speech signal features; then, the fused features of the prompt image sequence and speech signal features in the sample data, as well as the source speech text, are used as input to the initialized twelfth probability distribution model, and the predicted probability distribution of the decoded text relative to the prompt image sequence and speech signal features is output; finally, the loss value is determined based on the difference between the predicted probability distribution of the decoded text relative to the prompt image sequence and speech signal features and the reference probability distribution, and the parameters of the twelfth probability distribution model are updated based on the loss value using the backpropagation algorithm.

[0272] As an example, the twelfth probability distribution of the decoded text relative to the features of the prompt image sequence and the speech signal can be represented as follows: The twelfth probability distribution model can be a conditional random field or a hidden Markov model.

[0273] As an example, the difference operation can be used to calculate the difference between the predicted probability distribution of the decoded text relative to the features of the prompt image sequence and the reference probability distribution, and the square or absolute value of the difference can be taken as the loss value; or the exponential operation can be used to calculate the cross-entropy between the predicted probability distribution of the decoded text relative to the features of the prompt image sequence and the reference probability distribution, and take it as the loss value.

[0274] In some embodiments, see Figure 3M , Figure 3M This is a schematic diagram of the third process for determining the target loss provided in an embodiment of this application. Figure 3AStep 105, "determining the target loss based on probability distribution," can also be achieved by using the probability distribution of the encoded vector sequences of the visual, speech, and text modal sample data relative to preset conditions, as described above, to determine the difference in probability distribution between the fused encoded vector sequence and the text modal encoded vector sequence, which serves as the sub-target loss of the encoded vector sequence; determining the sub-target loss of speech text negatively correlated with the probability distribution based on the probability distribution of speech text relative to the visual and speech modal sample data; and determining the sub-target loss of speech text negatively correlated with the probability distribution based on the probability distribution of decoded text relative to the visual and speech modal sample data. This will be discussed in the following sections. Figure 3M Steps 1054C to 1057C are implemented, and the details are explained below.

[0275] In step 1054C, the difference between the ninth probability distribution and the tenth probability distribution is determined as the loss for the seventh sub-target.

[0276] In some embodiments, the difference between the ninth and tenth probability distributions can be determined by divergence calculations. Examples include KL divergence and cross-entropy.

[0277] As an example, the difference between the ninth and tenth probability distributions can be determined by formula (9) as the loss for the seventh sub-target.

[0278]

[0279] In step 1055C, the loss of the eighth sub-target that is negatively correlated with the eleventh probability distribution is determined.

[0280] As an example, the loss of the eighth sub-target that is negatively correlated with the eleventh probability distribution can be determined by the following formula (10).

[0281]

[0282] In step 1056C, the loss of the ninth sub-target that is negatively correlated with the twelfth probability distribution is determined.

[0283] As an example, the loss of the ninth sub-target that is negatively correlated with the twelfth probability distribution can be determined by the following formula (11).

[0284]

[0285] In step 1057C, the loss of the seventh sub-target, the loss of the eighth sub-target, and the loss of the ninth sub-target are fused to obtain the loss of the third target.

[0286] As an example, the loss of the seventh sub-target, the loss of the eighth sub-target, and the loss of the ninth sub-target are weighted and summed using the formula (12) shown below to obtain the loss of the third target.

[0287]

[0288] Here, α is the weighting parameter of the loss for the ninth sub-target.

[0289] See also Figure 3A The following will be an explanation following step 105 above.

[0290] In step 106, the parameters of the decoder and at least one encoder are updated based on the target loss to obtain an updated decoder and multiple encoders. The updated decoder and multiple encoders are used to form a second speech generation model, which is used to generate target speech text corresponding to the prompt image sequence of the target object.

[0291] In some embodiments, see Figure 3N , Figure 3N This is a schematic diagram of the first process for updating parameters provided in an embodiment of this application. Figure 3A Step 106, "updating the parameters of the decoder and at least one encoder based on the target loss," can be achieved through... Figure 3N Steps 1061A to 1062A are implemented, and the details are explained below.

[0292] In step 1061A, the parameters of the visual encoder are updated based on the first sub-target loss.

[0293] As an example, as shown in formula (1) above, the parameters θ of the visual encoder are updated based on the loss of the first sub-target. csE .

[0294] In step 1062A, the parameters of the visual encoder and decoder are updated based on the second sub-target loss and the third sub-target loss.

[0295] As an example, when the target language of the first speech generation model's speech generation task is the same as the source language, the parameters θ of the visual encoder are updated based on the second sub-target loss, as shown in formula (2) above. csE and the decoder parameter θ tD When the target language of the first speech generation model's speech generation task is different from the source language, the parameters θ of the visual encoder are updated based on the third sub-target loss, as shown in formula (3) above. csE and the decoder parameter θ tD .

[0296] In some embodiments, see Figure 3O , Figure 3OThis is a schematic diagram of the second process for updating parameters provided in an embodiment of this application. Figure 3A Step 106, "updating the parameters of the decoder and at least one encoder based on the target loss," can be achieved through... Figure 3O Steps 1061B to 1062B are implemented, and the details are explained below.

[0297] In step 1061B, the parameters of the speech encoder are updated based on the fourth sub-target loss.

[0298] As an example, as shown in formula (5) above, the parameters θ of the speech encoder are updated based on the fourth sub-target loss. isE .

[0299] In step 1062B, the parameters of the speech encoder and decoder are updated based on the fifth sub-target loss and the sixth sub-target loss.

[0300] As an example, when the target language of the first speech generation model's speech generation task is the same as the source language, the parameters θ of the speech encoder are updated based on the fifth sub-target loss, as shown in formula (6) above. isE and the decoder parameter θ tD When the target language of the first speech generation model's speech generation task is different from the source language, the parameters θ of the speech encoder are updated based on the sixth sub-target loss, as shown in formula (7) above. isE and the decoder parameter θ tD .

[0301] In some embodiments, see Figure 3P , Figure 3P This is a schematic diagram of the third process for updating parameters provided in the embodiments of this application. Figure 3A Step 106, "updating the parameters of the decoder and at least one encoder based on the target loss," can be achieved through... Figure 3P Steps 1061C to 1062C are implemented, and the details are explained below.

[0302] In step 1061C, the parameters of the visual encoder and the speech encoder are updated based on the seventh sub-target loss.

[0303] As an example, as shown in formula (9) above, the parameters θ of the visual encoder are updated based on the loss of the seventh sub-target. csE and the parameters θ of the speech encoder isE .

[0304] In step 1062C, the parameters of the visual encoder, speech encoder, and decoder are updated based on the eighth sub-target loss and the ninth sub-target loss.

[0305] As an example, when the target language of the first speech generation model's speech generation task is the same as the source language, the parameters θ of the visual encoder are updated based on the eighth sub-target loss, as shown in formula (10) above. csE The parameters θ of the voice encoder isE and the decoder parameter θ tD When the target language of the first speech generation model's speech generation task is different from the source language, the parameters θ of the visual encoder are updated based on the ninth sub-target loss, as shown in formula (11) above. csE The parameters θ of the voice encoder isE and the decoder parameter θ tD .

[0306] In some embodiments, the first speech generation model further includes a speech signal generator, wherein the speech signal generator is used to form a second speech generation model with the decoder and the updated plurality of encoders, the second speech generation model being used to generate speech signals.

[0307] As an example, see further. Figure 4A The speech signal generator shown in 4A includes a text-to-unit module and a unit-to-speech module, which together with a decoder and multiple updated encoders form a second speech generation model for generating speech signals.

[0308] The following describes the speech generation method of the speech generation model provided in the embodiments of this application. As mentioned above, the electronic device implementing the speech generation method of the speech generation model in the embodiments of this application can be a terminal or a server, or a combination of both. Therefore, the executing entity of each step will not be described again below.

[0309] See Figure 3Q , Figure 3Q This is a flowchart illustrating the speech generation method of the speech generation model provided in this application embodiment. The speech generation model is the second speech generation model described above. Below, it will be combined with... Figure 3Q Steps 301 to 302 shown will be explained.

[0310] In step 301, the prompt image sequence of the target object is obtained.

[0311] In some embodiments, each prompt image in the prompt image sequence displays a static image of the hand or lip posture of the sample object at a certain moment in a video frame when the sample object makes a limb or lip movement.

[0312] As an example, a sequence of prompt images of the target object can be obtained by using methods such as camera shooting, image databases, or online sources.

[0313] In step 302, the second speech generation model is invoked based on the prompt image sequence to generate target speech text corresponding to the prompt image sequence of the target object.

[0314] In some embodiments, the second speech generation model described above can be invoked based on the prompt image sequence to generate target speech text corresponding to the prompt image sequence of the target object.

[0315] In some embodiments, the target speech text generated based on the speech generation model can be further used to generate the corresponding target speech through the speech signal generator in the speech generation model.

[0316] This application embodiment trains the first speech generation model in three parallel ways based on sample data from multiple modalities of the sample object: 1) Encoding sample data from the visual and text modalities, and calling the text decoder to decode the encoding results of the visual encoder and text encoder respectively, in order to update the parameters of the visual encoder; 2) Encoding sample data from the visual, text, and speech modalities, and calling the text decoder to decode the encoding results of the visual encoder, text encoder, and speech encoder respectively, in order to update the parameters of the speech encoder; 3) Encoding sample data from the visual, text, and speech modalities, fusing the encoding results corresponding to the visual and text modalities, and then calling the text decoder to decode the fused encoding result and the encoding result of the text encoder respectively, in order to update the parameters of the visual encoder and speech encoder. This allows the speech generation model to fully learn the features of different modalities of the visual, text, and speech modalities, which can improve the ability of the speech generation model to perform speech generation tasks and improve the accuracy of speech generation. At the same time, the combination of different modalities can ensure that the execution of speech generation tasks can adapt to different application scenarios and improve the user experience.

[0317] The following will describe an exemplary application of the embodiments of this application in a scenario where fluent speech is provided to target objects with hearing or speech impairments.

[0318] The system collects the visual representation of a target object with hearing or speech impairments from the prompting speech (equivalent to the prompting image sequence mentioned above) via a terminal. After uploading this visual representation to the server, the server uses a trained second speech generation model to recognize the received visual representation, generate target speech text corresponding to the visual representation, and send it to the terminal. The terminal then outputs the speech text corresponding to the target speech text. Alternatively, the server can recognize the visual representation of the target object, generate target speech text corresponding to the visual representation, generate the corresponding speech text, and then send the speech to the terminal for playback. The training process of the speech generation model is explained in detail below.

[0319] For people with hearing or speech impairments, lip reading or sign language are the main means of communication. However, lip reading makes it difficult to distinguish sounds with similar lip shapes, such as [u] and [y]. Sign language requires a long learning period before one can communicate with others. Related technologies use a prompting voice system to encode finger shapes and hand positions, combined with lip reading, to provide users with hearing or speech impairments with a clear visual representation of all phonemes in spoken language.

[0320] See Figure 7 , Figure 7 This is a schematic diagram of the prompt voice encoding information provided in the embodiments of this application, such as... Figure 7 As shown, five different hand positions are used to encode vowel groups in Mandarin, and eight different hand gestures are used to encode consonant groups. With prompts, people with hearing impairments can distinguish sounds that are indistinguishable when lip-reading by combining hand information. However, the prompting system encodes the visual expression provided by the prompting user based on their hand and lip movements. On the one hand, it cannot output fluent, natural, and universal speech expressions; on the other hand, it is not suitable for non-prompting users.

[0321] The related speech engine technology can clone a person's voice with just 15 seconds of speech sample, supporting cross-language communication. Users can choose a personalized, human voice instead of a synthesized one with a noticeable mechanical feel, and the voice maintains consistency across various languages. This can support Augmentative & Alternative Communication (AAC) devices, providing personalized voices across multiple languages ​​for users who cannot speak, thus helping those with speech impairments. For example, it can provide therapeutic applications for users with speech-related disorders and educational enhancement services for users with learning needs. However, the core of this speech engine technology is a speech synthesis and cloning method that requires users to input text within the application before it can be converted into speech. This communication method cannot support streaming, real-time expression.

[0322] The method provided in this application encodes the visual expression from the prompting speech and maps it to a representation space that is consistent with the mapping space of the text modality and speech modality in the speech recognition translation synthesis engine (equivalent to the speech recognition model above). Based on the prompting speech of the target object with hearing impairment or speech impairment, it can directly generate the corresponding fluent speech, providing the target object with a customizable, natural and fluent voice that can be used across multiple languages, thereby helping the target object to communicate with others without barriers.

[0323] See Figure 8 , Figure 8 This is a schematic diagram of the training data provided in the embodiments of this application, such as... Figure 8 As shown, the training data includes cues from the visual modality, phonetic symbols and audio from the speech modality, and text from the text modality. To construct a latent representation space for the visual modality that is shared with the speech and text modalities, the training data must first be preprocessed and encoded.

[0324] See Figure 9 , Figure 9 This is a diagram of the visual coding model architecture provided in the embodiments of this application, such as... Figure 9 As shown, firstly, the video of the target object is preprocessed to extract the key points of each frame of the video, resulting in a key point sequence. The visual modality keypoints (CS keypoints) identify the coordinate positions of a total of J keypoints, including the lips, fingers, and hands.

[0325] Next, the keypoint sequence of each frame of the video is input into the visual encoder. The keypoint information input into the visual encoder (CSEncoder) is represented as follows: Where F represents the total number of video frames, and C represents the coordinate dimension of the keypoints. For example, when the keypoint coordinates are 2D, C=2, and when the keypoint coordinates are 3D, C=3. Then, the keypoint sequence is embedded in pose and encoded using a Transformer block to obtain a visual encoded vector sequence.

[0326] The State-of-the-Art (SOTA) pose detection baseline model can be used as the pre-trained baseline model for the visual encoder. Based on the pre-trained visual encoder trained using high-resource human pose data, low-resource prompt speech training data (e.g., ...) can be used. Figure 8 The training data shown is used to fine-tune the parameters of the pre-trained visual encoder model.

[0327] To reduce redundant information in the video while preserving the diversity and integrity of semantic information, related technologies employ clustering pruning techniques. The k-nearest neighbor peak density clustering algorithm (DPC-KNN) is used to cluster the token sequences of the input F frames according to similarity. Based on the criteria of higher density relative to nearest neighbors and greater distance relative to other high-density tokens, the f frames with higher scores are used as cluster centers, i.e., the key tokens retained after pruning, where f is much smaller than the original sequence length F (f << F).

[0328] Unlike clustering and pruning techniques in related technologies, this application's embodiments, when fine-tuning the pre-trained visual encoding model, align the visual encoding vector sequence obtained after encoding by the visual encoder to the output feature space of the text encoder, thereby ensuring the semantic integrity of the former. This includes the following two aspects of processing, see [link to relevant documentation]. Figure 4A , Figure 4A This is a schematic diagram of the first architecture of the speech generation model (CSeamless) provided in the embodiments of this application.

[0329] First, sequence length alignment is performed. The visual encoded vector sequence (sequence length F) obtained by the visual encoder is aligned with the output sequence length f of the text encoder using a length adaptation module. The length adapter can be a Transformer-based length adapter from SEAMLESSM4T V2.

[0330] By aligning the lengths of the feature vector sequences of the visual modality and the text modality, the length of the visual encoding vector sequence is compressed from F to f (f << F). At this point, the output of the visual encoder is... Here, D represents the dimension of the features input from the text encoder to the text decoder, which improves the processing efficiency of the visual encoding model.

[0331] Then, feature space alignment is performed. Using knowledge distillation as shown in Equation (1) above as the objective function, the outputs of the visual encoder and the text encoder are aligned in feature space to extract knowledge from powerful large-scale text-to-text machine translation models (such as SEAMLESSM4T-NLLB) to guide the visual expression of cued speech to text (and text translation), i.e., the task from visual modality to text modality, thereby ensuring the semantic integrity while compressing the output sequence of the visual encoder.

[0332] Where, x cued-speech and x text These are the cued-speech in the source language and the corresponding source text, y text This refers to the text corresponding to the target language. In this embodiment, the source language is Chinese, and the target languages ​​include both Chinese and English. When the target language is English, the source text x is processed... text Perform a text translation from Chinese to English to obtain pseudo-labels y. text When the target language is the same as the source language, it is equivalent to a cued-speech to text recognition task (CS2T).

[0333] By fixing the pre-trained text encoder, the parameters θ of the visual encoder (CS encoder) are adjusted. csE and Figure 4A The parameters θ of the text decoder are shown. tD Perform joint fine-tuning, adjusting the loss function for the corresponding recognition task. Loss function for translation task As an additional fine-tuning training objective function, the loss function for the recognition task Loss function for translation task As shown in formulas (2) and (3) above.

[0334] In summary, the loss function for the overall fine-tuning training of the first architecture of the speech generation model provided in this application embodiment is shown in the above formula (4), where α is a weight parameter determined by the proportion of the fine-tuning training set in the full training set.

[0335] As an example, Figure 4A The text encoder and text decoder shown can adopt the SEAMLESSM4T-NLLB text encoder & decoder model based on Transformer and pre-trained weights, and fix the weights of the text encoder and fine-tune the parameters of the text decoder through the above joint training; the text to unit module can adopt the non-autoregressive (NAR) T2U model and weights in SEAMLESSM4TV2; the unit to speech module can adopt the unit vocoder model based on High-Fidelity Generative Adversarial Network (HiFi-GAN) in SEAMLESSM4TV2.

[0336] In the practical application of voice prompt systems, the prompts from hearing-impaired individuals, while expressing hand and lip movements, also include impaired voice or unvoiced sound. The degree of impairment varies from person to person; some have mild impairment and can produce somewhat discernible sounds, while others have severe impairment and produce sounds that are difficult to distinguish. In practice, this impaired voice can help improve the accuracy and intelligibility of communication. Therefore, the training of the voice generation model can be combined with the training of the impaired voice coding model.

[0337] See Figure 4B , Figure 4BThis is a schematic diagram of the second architecture of the speech generation model provided in the embodiments of this application. For example... Figure 4B As shown, compared to Figure 4A The training architecture of the speech generation model shown adds an impaired-speech encoder (IS Encoder). The speech encoder in SEAMLESSM4T V2 can be used as the initial speech encoder, and the speech vector sequence of the speech modality is length-aligned with the text vector sequence of the text modality through a length adapter. Then, the speech encoder is fine-tuned using a low-resource impaired speech dataset (referring to speakers unable to produce normal, clear, and fluent speech due to physiological or medical reasons).

[0338] like Figure 4B As shown, an 80-dimensional spectral feature was extracted from damaged speech using a Mel filter bank extractor. Where F' represents the total number of speech frames, which serves as the input to the speech encoder. Using knowledge distillation as shown in formula (5) above as the objective function, the outputs of the speech encoder and the text encoder are aligned in feature space.

[0339] Where, x impaired-speech and x text These are the impaired-speech in the source language and the corresponding source text, y text This refers to the text corresponding to the target language. When the target language is English, it is obtained by analyzing the source text x. text Perform a text translation from Chinese to English to obtain pseudo-labels y. text When the target language is the same as the source language (Chinese), it is equivalent to the impaired-speech to text recognition task (IS2T).

[0340] By fixing the pre-trained text encoder, the parameters θ of the speech encoder (IS encoder) are adjusted. isE and Figure 4B The parameters θ of the text decoder are shown. tD Perform joint fine-tuning and identify the loss function and translation loss function As an additional fine-tuning training objective function, the loss function for identification. and translation loss function As shown in formulas (6) and (7) above.

[0341] In summary, the loss function for the overall fine-tuning training of the second architecture of the speech generation model provided in this application embodiment is shown in formula (8) above.

[0342] In practical applications, there is a strong correlation between visual modal prompts and impaired speech in the speech modality. Hearing-impaired individuals can improve the accuracy and intelligibility of their communication by jointly expressing these prompts and speech. Therefore, the visual encoder and speech encoder can be jointly trained during the training of the speech generation model.

[0343] See Figure 4C , Figure 4C This is a schematic diagram of the third architecture of the speech generation model provided in the embodiments of this application. For example... Figure 4C As shown, damaged speech can be output via a speech encoder. and the output of the visual encoder The concatenation projector module first performs concatenation processing to obtain... Then perform mapping processing to obtain Then it is aligned with the feature space of the text encoder output. The objective function of knowledge distillation corresponding to the alignment of feature vector sequences is shown in Equation (9) above.

[0344] x in formula (9) cued-speech ,x impaired-speech and x text These are the source language's cued-speech, impaired-speech, and corresponding text, y text The text corresponds to the target language. When the target language is English, it is obtained by analyzing the source text x. text Perform a text translation from Chinese to English to obtain pseudo-labels y. text When the target language is the same as the source language (Chinese), it is equivalent to a text recognition task that combines cued-speech and impaired-speech (CS_IS2T).

[0345] By fixing the pre-trained text encoder, the parameters θ of the visual encoder (CS encoder) are adjusted. csE The parameters θ of the speech encoder (IS Encoder) isE and Figure 4C The parameters θ of the text decoder are shown. tD Perform joint fine-tuning and increase the recognition loss function. and translation loss function As an additional fine-tuning training objective function, the loss function for identification. and translation loss function The above formulas (10) and (11) are shown respectively.

[0346] In summary, the loss function for the overall fine-tuning training of the third architecture of the speech generation model provided in this application embodiment is shown in formula (12) above.

[0347] The training methods described above can directly generate fluent speech based on prompts from people who are unable to speak or whose speech is impaired. It supports real-time streaming speech output, as well as the generation of natural and fluent voices in multiple languages. Users can also choose their own voice timbre, which can help more people with hearing loss or speech impairment to integrate into the digital society better and faster.

[0348] The following description continues to illustrate the exemplary structure of the training device 555-1 for the speech generation model provided in this application embodiment as a software module. In some embodiments, such as... Figure 2A As shown, the software modules in the training device 555-1 for the speech generation model stored in memory 550-1 may include:

[0349] The acquisition module 5551-1 is used to acquire a first speech generation model, wherein the first speech generation model includes a decoder and multiple encoders corresponding to multiple modalities; and to acquire sample data of multiple modalities, wherein the sample data of multiple modalities includes a sequence of prompt image of sample objects and speech text corresponding to the prompt image sequence.

[0350] The encoding module 5552-1 is used to call multiple encoders to encode the prompt image sequence and the voice text respectively, so as to obtain a multimodal encoded vector sequence.

[0351] The decoding module 5553-1 is used to call the decoder based on the multimodal encoded vector sequence to obtain the decoded text.

[0352] The determination module 5554-1 is used to determine the probability distribution of the decoded text and sample data of multiple modalities, and to determine the target loss based on the probability distribution.

[0353] The generation module 5555-1 is used to update the parameters of the decoder and at least one encoder based on the target loss to obtain the updated decoder and multiple encoders. The updated decoder and multiple encoders are used to form a second speech generation model, which is used to generate target speech text corresponding to the prompt image sequence of the target object.

[0354] In some embodiments, when the multiple modalities include a visual modality and a text modality, the encoding module 5552-1 is further configured to call a visual encoder to encode the prompt image sequence to obtain a visual encoding vector sequence for the visual modality, wherein the visual encoder is the encoder corresponding to the visual modality among the multiple encoders; call a text encoder to encode multiple morphemes in the speech text to obtain a text encoding vector sequence for the text modality, wherein the text encoding vector sequence includes text encoding vectors corresponding to multiple morphemes, and the text encoder is the encoder corresponding to the text modality among the multiple encoders; and use the visual encoding vector sequence and the text encoding vector sequence as a first multimodal encoding vector sequence, wherein the first multimodal encoding vector sequence is used for decoding by the decoder.

[0355] In some embodiments, the encoding module 5552-1 is further configured to identify multiple prompt images in the prompt image sequence respectively to obtain the position of key points in each prompt image; generate a key point sequence based on the position of key points in each prompt image; obtain the embedding vector of the key point sequence; and call the visual encoder to encode the embedding vector to obtain the visual encoding vector of the visual modality.

[0356] In some embodiments, the determining module 5554-1 is further configured to: determine a first probability distribution of the text encoded vector sequence relative to the decoded text and the speech text, and a second probability distribution of the visual encoded vector sequence relative to the decoded text and the prompt image sequence; determine a third probability distribution of the speech text relative to the prompt image sequence; determine a fourth probability distribution of the decoded text relative to the prompt image sequence when the target language of the speech generation task of the first speech generation model is different from the source language; determine the difference between the first probability distribution and the second probability distribution as a first sub-target loss; determine a second sub-target loss negatively correlated with the third probability distribution; determine a third sub-target loss negatively correlated with the fourth probability distribution; and fuse the first sub-target loss, the second sub-target loss, and the third sub-target loss to obtain a first target loss.

[0357] In some embodiments, the generation module 5555-1 is further configured to update the parameters of the visual encoder based on the first sub-target loss; and update the parameters of the visual encoder and decoder based on the second sub-target loss and the third sub-target loss.

[0358] In some embodiments, the first speech generation model is a pre-trained speech recognition model, and after pre-training, the representation space of the text modality and the representation space of the speech modality of the first speech generation model have been unified to the same representation space; the decoding module 5553-1 is further configured to map the visual encoded vector sequence and the text encoded vector sequence to obtain the vector sequence representation of the visual encoded vector sequence and the text encoded vector sequence in the representation space; and adjust the length of the vector sequence representation of the visual encoded vector sequence to be the same as the length of the vector sequence representation of the text encoded vector sequence.

[0359] In some embodiments, when the multiple modalities also include a speech modality, the sample data of the multiple modalities also include speech signal features of the speech modality; the encoding module 5552-1 is further configured to call the speech encoder to encode the speech signal features to obtain a speech encoding vector sequence of the speech modality, wherein the speech encoder is the encoder corresponding to the speech modality among the multiple encoders; the visual encoding vector sequence, the text encoding vector sequence and the speech encoding vector sequence are used as a second multimodal encoding vector sequence, wherein the second multimodal encoding vector sequence is used to replace the first multimodal encoding vector sequence for decoding by the decoder.

[0360] In some embodiments, the determining module 5554-1 is further configured to: determine a fifth probability distribution of the text encoding vector sequence relative to the decoded text and the speech text, and a sixth probability distribution of the speech encoding vector sequence relative to the decoded text and the speech signal features; determine a seventh probability distribution of the speech text relative to the speech signal features; determine an eighth probability distribution of the decoded text relative to the speech signal features when the target language of the speech generation task of the first speech generation model is different from the source language; determine the difference between the fifth probability distribution and the sixth probability distribution as a fourth sub-target loss; determine a fifth sub-target loss negatively correlated with the seventh probability distribution; determine a sixth sub-target loss negatively correlated with the eighth probability distribution; and fuse the fourth sub-target loss, the fifth sub-target loss, and the sixth sub-target loss to obtain a second target loss.

[0361] In some embodiments, the generation module 5555-1 is further configured to update the parameters of the speech encoder based on the fourth sub-target loss; and update the parameters of the speech encoder and decoder based on the fifth sub-target loss and the sixth sub-target loss.

[0362] In some embodiments, when the multiple modalities also include a speech modality, the encoding module 5552-1 is further configured to call multiple encoders to encode based on the prompt image sequence and the speech text respectively, to obtain a visual encoding vector sequence for the visual modality, a text encoding vector sequence for the text modality, and a speech encoding vector sequence for the speech modality; to concatenate the visual encoding vector sequence and the speech encoding vector sequence to obtain a concatenated encoding vector sequence; to perform feature mapping on the concatenated encoding vector sequence to obtain a fused encoding vector sequence; and to use the fused encoding vector sequence and the text encoding vector sequence as a third multimodal encoding vector sequence, wherein the third multimodal encoding vector sequence is used to replace the second multimodal encoding vector sequence for decoding by the decoder.

[0363] In some embodiments, the sample data of multiple modalities further includes speech signal features of the speech modality. The encoding module 5552-1 is also used to call the visual encoder to encode the prompt image sequence to obtain a visual encoding vector sequence of the visual modality, wherein the visual encoder is the encoder corresponding to the visual modality among multiple encoders; call the text encoder to encode multiple morphemes in the speech text to obtain a text encoding vector sequence of the text modality, wherein the text encoding vector sequence includes text encoding vectors corresponding to multiple morphemes, and the text encoder is the encoder corresponding to the text modality among multiple encoders; and call the speech encoder to encode the speech signal features to obtain a speech encoding vector sequence of the speech modality, wherein the speech encoder is the encoder corresponding to the speech modality among multiple encoders.

[0364] In some embodiments, the determining module 5554-1 is further configured to: determine the ninth probability distribution of the text encoded vector sequence relative to the decoded text and the speech text, and the tenth probability distribution of the fused encoded vector sequence relative to the decoded text, the prompt image sequence, and the speech signal features; determine the eleventh probability distribution of the speech text relative to the prompt image sequence and the speech signal features; when the target language of the speech generation task of the first speech generation model is different from the source language, determine the twelfth probability distribution of the decoded text relative to the prompt image sequence and the speech signal features; determine the difference between the ninth and tenth probability distributions as the seventh sub-target loss; determine the eighth sub-target loss negatively correlated with the eleventh probability distribution; determine the ninth sub-target loss negatively correlated with the twelfth probability distribution; and fuse the seventh, eighth, and ninth sub-target losses to obtain the third target loss.

[0365] In some embodiments, the generation module 5555-1 is further configured to update the parameters of the visual encoder and the speech encoder based on the seventh sub-target loss; and update the parameters of the visual encoder, the speech encoder, and the decoder based on the eighth sub-target loss and the ninth sub-target loss.

[0366] In some embodiments, the first speech generation model further includes a speech signal generator, wherein the speech signal generator is used to form a second speech generation model with the decoder and the updated plurality of encoders, the second speech generation model being used to generate speech signals.

[0367] The following description continues to illustrate the exemplary structure of the speech generation device 555-2 of the speech generation model provided in this application embodiment as a software module. In some embodiments, such as... Figure 2B As shown, the software modules in the training device 555-2 for the speech generation model stored in memory 550-2 may include:

[0368] The acquisition module 5551-2 is used to acquire the image sequence of prompts for the target object.

[0369] The generation module 5552-2 is used to generate target speech text corresponding to the prompt image sequence of the target object by calling the second speech generation model based on the prompt image sequence.

[0370] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the training method of the speech generation model described above, or to perform the speech generation method of the speech generation model described above.

[0371] This application provides a computer-readable storage medium storing computer-executable instructions or a computer program. When the computer-executable instructions or the computer program are executed by a processor, the processor will execute the training method of the speech generation model provided in this application. Figure 3A The training method of the speech generation model shown, or the speech generation method of the speech generation model described above in this application, such as... Figure 3Q The speech generation method of the speech generation model is shown.

[0372] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.

[0373] In some embodiments, computer-executable instructions may take the form of programs, software, software modules, scripts, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as stand-alone programs or as modules, components, subroutines, or other units suitable for use in a computing environment.

[0374] As an example, computer-executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple co-located files (e.g., files that store one or more modules, subroutines, or code sections).

[0375] As an example, computer-executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.

[0376] In summary, this application embodiment uses sample data from multiple modalities, including the target object's prompt image sequence (corresponding to the visual modality) and the speech text (corresponding to the text modality), and calls multiple encoders to encode the data. The resulting multimodal encoded vector sequence (i.e., the encoded vector sequence of multiple modalities) is compared with the probability distribution of the decoded text to calculate the target loss. This allows the target loss to reflect the differences between the representation spaces of multiple modalities in the first speech generation model. Based on the target loss, the parameters of the first speech generation model are updated through backpropagation, resulting in a unified representation space for the visual and text modalities in the trained second speech generation model. This ensures that the target speech text output by the second speech generation model accurately matches the intent of the target object's prompt image, thereby guaranteeing the accuracy of speech text generation from the prompt image. Furthermore, compared to related technologies that require manual text input to generate speech text, this method is more convenient, efficient, and provides a better user experience. The training process of the first speech generation model is carried out on the basis of the pre-trained speech recognition model. With the representation space of the text modality and the representation space of the speech modality already unified into the same representation space, fine-tuning can realize the function of converting prompt image sequences into speech text on the basis of the original functions of the pre-trained speech recognition model. At the same time, multiple functions of the pre-trained speech recognition model itself can be reused, such as recognizing prompt image sequences of target objects, translating text, and generating speech, avoiding redundant development and improving development efficiency.

[0377] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.

Claims

1. A training method for a speech generation model, characterized in that, The method includes: Obtain a first speech generation model, wherein the first speech generation model includes a decoder and multiple encoders corresponding to multiple modalities respectively; Acquire sample data for multiple modalities, wherein the sample data for multiple modalities includes a sequence of prompt images of sample objects and the corresponding speech text; Based on the prompt image sequence and the voice text, the multiple encoders are called to encode the text, resulting in a multimodal encoded vector sequence; The decoder is invoked based on the multimodal encoded vector sequence to perform decoding, thereby obtaining the decoded text; Determine the probability distribution of the decoded text and the sample data of the multiple modalities, and determine the target loss based on the probability distribution; The parameters of the decoder and at least one encoder are updated based on the target loss to obtain the updated decoder and the plurality of encoders, wherein the updated decoder and the plurality of encoders are used to form a second speech generation model, and the second speech generation model is used to generate target speech text corresponding to the prompt image sequence of the target object.

2. The method according to claim 1, characterized in that, When the multiple modalities include a visual modality and a text modality, the process of calling the multiple encoders to encode based on the prompt image and the spoken text respectively, to obtain a multimodal encoded vector sequence, includes: The visual encoder is invoked to encode the prompt image sequence to obtain a visual encoding vector sequence of the visual modality, wherein the visual encoder is the encoder corresponding to the visual modality among the plurality of encoders; The text encoder is invoked to encode multiple morphemes in the speech text to obtain a text encoding vector sequence of the text modality, wherein the text encoding vector sequence includes text encoding vectors corresponding to the multiple morphemes, and the text encoder is the encoder corresponding to the text modality among the multiple encoders; The visual encoding vector sequence and the text encoding vector sequence are used as a first multimodal encoding vector sequence, wherein the first multimodal encoding vector sequence is used for decoding by the decoder.

3. The method according to claim 2, characterized in that, The step of calling the visual encoder to encode the prompt image sequence to obtain a visual encoding vector sequence of the visual modality includes: The positions of key points in each prompt image are obtained by identifying multiple prompt images in the prompt image sequence. Based on the position of the key points in each of the prompt images, a key point sequence is generated; Obtain the embedding vector of the key point sequence; The visual encoder is invoked to encode the embedding vector to obtain a visual encoding vector sequence of the visual modality.

4. The method according to claim 2 or 3, characterized in that, Determining the probability distribution of the decoded text and the sample data of the multiple modalities includes: Determine a first probability distribution of the text encoded vector sequence relative to the decoded text and the speech text, and a second probability distribution of the visual encoded vector sequence relative to the decoded text and the prompt image sequence; Determine the third probability distribution of the spoken text relative to the prompt image sequence; When the target language of the speech generation task of the first speech generation model is different from the source language, a fourth probability distribution of the decoded text relative to the prompt image sequence is determined. Determining the target loss based on the probability distribution includes: The difference between the first probability distribution and the second probability distribution is determined as the first sub-target loss; Determine the second sub-target loss that is negatively correlated with the third probability distribution; Determine the loss of the third sub-target that is negatively correlated with the fourth probability distribution; The first sub-target loss, the second sub-target loss, and the third sub-target loss are fused together to obtain the first target loss.

5. The method according to claim 4, characterized in that, The step of updating the parameters of the decoder and at least one of the encoders based on the target loss includes: The parameters of the visual encoder are updated based on the first sub-target loss; The parameters of the visual encoder and the decoder are updated based on the second sub-target loss and the third sub-target loss.

6. The method according to claim 2 or 3, characterized in that, The first speech generation model is a pre-trained speech recognition model, and after the pre-training, the representation space of the text modality and the representation space of the speech modality of the first speech generation model have been unified to the same representation space. Before calling the decoder based on the multimodal encoded vector sequence to obtain the decoded text, the method further includes: The visual encoding vector sequence and the text encoding vector sequence are mapped to obtain the vector sequence representations of the visual encoding vector sequence and the text encoding vector sequence in the representation space; The length of the vector sequence representation of the visual encoding vector sequence is adjusted to be the same as the length of the vector sequence representation of the text encoding vector sequence.

7. The method according to claim 2 or 3, characterized in that, When the plurality of modalities also includes a speech modality, the sample data of the plurality of modalities also includes the speech signal features of the speech modality; The step of encoding the prompt image sequence and the speech text by calling the multiple encoders respectively to obtain a multimodal encoded vector sequence further includes: The speech signal features are encoded by calling a speech encoder to obtain a speech encoding vector sequence of the speech modality, wherein the speech encoder is the encoder corresponding to the speech modality among the plurality of encoders; The visual encoding vector sequence, the text encoding vector sequence, and the speech encoding vector sequence are used as a second multimodal encoding vector sequence, wherein the second multimodal encoding vector sequence is used to replace the first multimodal encoding vector sequence for decoding by the decoder.

8. The method according to claim 7, characterized in that, Determining the probability distribution of the decoded text and the sample data of the multiple modalities includes: Determine a fifth probability distribution of the text encoding vector sequence relative to the decoded text and the speech text, and a sixth probability distribution of the speech encoding vector sequence relative to the decoded text and the speech signal features; Determine the seventh probability distribution of the speech text relative to the features of the speech signal; When the target language of the speech generation task of the first speech generation model is different from the source language, the eighth probability distribution of the decoded text relative to the speech signal features is determined. Determining the target loss based on the probability distribution includes: The difference between the fifth probability distribution and the sixth probability distribution is determined as the fourth sub-target loss; Determine the loss of the fifth sub-target that is negatively correlated with the seventh probability distribution; Determine the loss of the sixth sub-target that is negatively correlated with the eighth probability distribution; The fourth sub-target loss, the fifth sub-target loss, and the sixth sub-target loss are fused together to obtain the second target loss.

9. The method according to claim 8, characterized in that, The step of updating the parameters of the decoder and at least one of the encoders based on the target loss includes: The parameters of the speech encoder are updated based on the fourth sub-target loss; The parameters of the speech encoder and the decoder are updated based on the fifth sub-target loss and the sixth sub-target loss.

10. The method according to any one of claims 1 to 9, characterized in that, When the multiple modalities also include a speech modality, the process of encoding the multiple encoders based on the prompt image sequence and the speech text to obtain a multimodal encoded vector sequence includes: Based on the prompt image sequence and the speech text, the multiple encoders are called to encode the visual encoding vector sequence of the visual modality, the text encoding vector sequence of the text modality, and the speech encoding vector sequence of the speech modality. The visual encoding vector sequence and the speech encoding vector sequence are concatenated to obtain a concatenated encoding vector sequence. The concatenated encoded vector sequence is then subjected to feature mapping to obtain a fused encoded vector sequence; The fused encoded vector sequence and the text encoded vector sequence are used as a third multimodal encoded vector sequence, wherein the third multimodal encoded vector sequence is used to replace the second multimodal encoded vector sequence for decoding by the decoder.

11. The method according to claim 10, characterized in that, The sample data for the multiple modalities also includes speech signal features of the speech modality; The process of encoding the visual modality (visual encoding vector sequence), the text modality (text encoding vector sequence), and the speech modality (speech encoding vector sequence) by calling the multiple encoders respectively, to obtain the speech encoding vector sequence, includes: The visual encoder is invoked to encode the prompt image sequence to obtain the visual encoding vector sequence of the visual modality, wherein the visual encoder is the encoder corresponding to the visual modality among the plurality of encoders; The text encoder is invoked to encode multiple morphemes in the speech text to obtain a text encoding vector sequence of the text modality, wherein the text encoding vector sequence includes text encoding vectors corresponding to the multiple morphemes, and the text encoder is the encoder corresponding to the text modality among the multiple encoders; The speech signal features are encoded by calling a speech encoder to obtain a speech encoding vector sequence of the speech modality, wherein the speech encoder is the encoder corresponding to the speech modality among the plurality of encoders.

12. The method according to claim 10, characterized in that, Determining the probability distribution of the decoded text and the sample data of the multiple modalities includes: Determine the ninth probability distribution of the text encoding vector sequence relative to the decoded text and the speech text, and the tenth probability distribution of the fused encoding vector sequence relative to the decoded text, the prompt image sequence, and the speech signal features; Determine the eleventh probability distribution of the speech text relative to the prompt image sequence and the speech signal features; When the target language of the speech generation task of the first speech generation model is different from the source language, the twelfth probability distribution of the decoded text relative to the prompt image sequence and the speech signal features is determined. Determining the target loss based on the probability distribution includes: The difference between the ninth probability distribution and the tenth probability distribution is determined as the seventh sub-target loss; Determine the loss of the eighth sub-target that is negatively correlated with the eleventh probability distribution; Determine the loss of the ninth sub-target that is negatively correlated with the twelfth probability distribution; The loss of the seventh sub-target, the loss of the eighth sub-target, and the loss of the ninth sub-target are fused together to obtain the third target loss.

13. The method according to claim 12, characterized in that, The step of updating the parameters of the decoder and at least one of the encoders based on the target loss includes: The parameters of the visual encoder and the speech encoder are updated based on the seventh sub-target loss; The parameters of the visual encoder, the speech encoder, and the decoder are updated based on the eighth sub-target loss and the ninth sub-target loss.

14. The method according to any one of claims 1 to 12, characterized in that, The first speech generation model further includes a speech signal generator, wherein the speech signal generator is used to form a second speech generation model with the decoder and the updated plurality of encoders, and the second speech generation model is used to generate speech signals.

15. A speech generation method for a speech generation model, characterized in that, The speech generation model is the second speech generation model according to any one of claims 1 to 14, and the method includes: Obtain the sequence of prompt images for the target object; Based on the prompt image sequence, the second speech generation model is invoked to generate target speech text corresponding to the prompt image sequence of the target object.

16. A training device for a speech generation model, characterized in that, The device includes: An acquisition module is used to acquire a first speech generation model, wherein the first speech generation model includes a decoder and multiple encoders corresponding to multiple modalities; and to acquire sample data of multiple modalities, wherein the sample data of multiple modalities includes a sequence of prompt image of sample objects and speech text corresponding to the prompt image sequence. The encoding module is used to call the multiple encoders to encode the prompt image sequence and the voice text respectively, so as to obtain a multimodal encoded vector sequence; The decoding module is used to call the decoder based on the multimodal encoded vector sequence to decode the text and obtain the decoded text. The determination module is used to determine the probability distribution of the decoded text and the sample data of the multiple modalities, and to determine the target loss based on the probability distribution; A generation module is used to update the parameters of the decoder and at least one encoder based on the target loss to obtain the updated decoder and the plurality of encoders, wherein the updated decoder and the plurality of encoders are used to form a second speech generation model, and the second speech generation model is used to generate target speech text corresponding to the prompt image sequence of the target object.

17. A speech generation device, characterized in that, The device includes: The acquisition module is used to acquire the sequence of prompt images for the target object; The generation module is used to call a speech generation model based on the prompt image to generate target speech text corresponding to the prompt image sequence of the target object, wherein the speech generation model is the second speech generation model as described in any one of claims 1 to 14.

18. An electronic device, characterized in that, The electronic device includes: Memory is used to store executable instructions for a computer; The processor, when executing computer-executable instructions stored in the memory, implements the training method of the speech generation model according to any one of claims 1 to 14, or the speech generation method of the speech generation model according to claim 15.

19. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the training method of the speech generation model according to any one of claims 1 to 14, or the speech generation method of the speech generation model according to claim 15.

20. A computer program product comprising computer-executable instructions or a computer program, characterized in that, When the computer-executable instructions or computer program are executed by a processor, they implement the training method of the speech generation model according to any one of claims 1 to 14, or the speech generation method of the speech generation model according to claim 15.