Audio and video data generation methods, training methods, and intelligent agents based on large models

CN121306089BActive Publication Date: 2026-08-11BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-08-11

AI Technical Summary

Benefits of technology

[0010] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform a method provided according to an embodiment of this disclosure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121306089B_ABST
    Figure CN121306089B_ABST
Patent Text Reader

Abstract

This disclosure provides a method, training method, and intelligent agent for generating audio and video data based on a large model, relating to the field of artificial intelligence technology, particularly to the fields of multimedia data generation, intelligent customer service, video production, animation production, metaverse, and digital human live streaming. The method includes: receiving speech text and audio and video prompt data related to a specified object, wherein the audio and video prompt data represents the pronunciation content of the specified object and the corresponding facial expression changes; using a large model to semantically fuse the speech text and audio and video prompt data to obtain semantic fusion features, wherein the semantic fusion features represent the pronunciation prosody of the specified object in response to the speech text, and the facial visual semantics aligned with the pronunciation prosody sequence; and generating target audio and video data based on the semantic fusion features, wherein the target audio and video data represents the target pronunciation content of the specified object in response to the speech text, and the corresponding visual facial expression content of the specified object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to the fields of multimedia data generation, intelligent customer service, video production, animation production, metaverse, and digital human live streaming. Background Technology

[0002] With the rapid development of artificial intelligence technology, large models can be used to drive digital human characters, such as cartoon characters, to interact with users. For example, by outputting voice responses and video data representing the actions of cartoon characters through the output of large models, users can watch the cartoon characters' actions in the video through the display interface of their terminal devices and receive voice responses played by audio playback devices. Users can also input their desired content based on the voice and actions of the cartoon characters to achieve an immersive interactive experience. Summary of the Invention

[0003] This disclosure provides a method for generating audio and video data based on a large model, a training method, an intelligent agent, a device, an electronic device, and a storage medium.

[0004] According to one aspect of this disclosure, a method for generating audio and video data based on a large model is provided, comprising: receiving speech text and audio and video prompt data related to a specified object, wherein the audio and video prompt data represents the pronunciation content of the specified object and the facial expression changes corresponding to the pronunciation content; using a large model to perform semantic fusion of the speech text and the audio and video prompt data to obtain semantic fusion features, wherein the semantic fusion features represent the pronunciation prosody of the specified object in response to the speech text and the facial expression visual semantics aligned with the pronunciation prosody time sequence; and generating target audio and video data based on the semantic fusion features, wherein the target audio and video data represents the target pronunciation content of the specified object in response to the speech text and the on-screen facial expression content of the specified object corresponding to the target pronunciation content.

[0005] According to another aspect of this disclosure, a method for training a large model is provided, comprising: receiving sample speech text, sample audio-visual cue data related to sample objects, and label semantic fusion features, wherein the label semantic fusion features characterize the pronunciation prosody of the sample objects in relation to the sample speech text in the sample audio-visual cue data, and the facial visual semantics aligned with the pronunciation prosody sequence; using a large model to perform semantic fusion on the sample speech text and the sample audio-visual cue data to obtain sample semantic fusion features; and training the large model based on the difference between the sample semantic fusion features and the label semantic fusion features to obtain a trained large model.

[0006] According to another aspect of this disclosure, an audio-visual data generation apparatus based on a large model is provided, comprising: a first receiving module, configured to receive speech text and audio-visual prompt data related to a specified object, wherein the audio-visual prompt data represents the pronunciation content of the specified object and the facial expression changes corresponding to the pronunciation content; a semantic fusion feature acquisition module, configured to perform semantic fusion on the speech text and the audio-visual prompt data using a large model to obtain semantic fusion features, wherein the semantic fusion features represent the pronunciation prosody of the specified object in response to the speech text, and the facial expression visual semantics aligned with the pronunciation prosody sequence; and a target audio-visual data generation module, configured to generate target audio-visual data based on the semantic fusion features, wherein the target audio-visual data represents the target pronunciation content of the specified object in response to the speech text, and the on-screen facial expression content of the specified object corresponding to the target pronunciation content.

[0007] According to another aspect of this disclosure, an apparatus for training a large model is provided, comprising: a second receiving module for receiving sample speech text, sample audio-visual cue data related to a sample object, and label semantic fusion features, wherein the label semantic fusion features characterize the pronunciation prosody of the sample object in relation to the sample speech text in the sample audio-visual cue data, and the facial visual semantics aligned with the pronunciation prosody sequence; a sample semantic fusion feature obtaining module for semantically fusing the sample speech text and the sample audio-visual cue data using the large model to obtain sample semantic fusion features; and a training module for training the large model based on the difference between the sample semantic fusion features and the label semantic fusion features to obtain a trained large model.

[0008] According to another aspect of this disclosure, an artificial intelligence agent is provided, comprising: an input module for receiving input information; a processing module for determining a target task based on the input information received by the input module, determining a large model based on the target task, and obtaining output information by calling the large model to execute the method provided in the embodiments of this disclosure; and an output module for outputting the output information obtained by the processing module.

[0009] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method provided according to an embodiment of this disclosure.

[0010] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform a method provided according to an embodiment of this disclosure.

[0011] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method provided according to embodiments of this disclosure.

[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0013] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0014] Figure 1 The illustration schematically shows an exemplary system architecture for applying a large-model-based audio and video data generation method and apparatus according to embodiments of the present disclosure.

[0015] Figure 2 A flowchart illustrating a method for generating audio and video data based on a large model according to an embodiment of the present disclosure is shown.

[0016] Figure 3 The illustration shows a schematic diagram of the principle of a large-model-based audio and video data generation method according to an embodiment of the present disclosure.

[0017] Figure 4 A flowchart illustrating a method for training a large model according to an embodiment of the present disclosure is shown schematically.

[0018] Figure 5 A schematic diagram illustrating a method for training a large model according to an embodiment of the present disclosure is shown.

[0019] Figure 6 A block diagram of a large-model-based audio and video data generation apparatus according to an embodiment of the present disclosure is shown schematically.

[0020] Figure 7 A block diagram of an apparatus for training a large model according to an embodiment of the present disclosure is shown schematically.

[0021] Figure 8 A schematic block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is shown.

[0022] Figure 9 A method for generating audio and video data based on a large model, which can be used to implement embodiments of the present disclosure, is illustrated. A schematic block diagram of an example electronic device for training a large model is also shown. Detailed Implementation

[0023] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0024] In the technical solution disclosed herein, the acquisition, storage, and application of user personal information comply with the provisions of relevant laws and regulations, necessary confidentiality measures have been taken, and there is no violation of public order and good morals.

[0025] The inventors discovered that with the rapid development of artificial intelligence technology, it is possible to use speech synthesis and video generation technologies to drive digital humans displayed on a screen to perform facial expressions and simultaneously broadcast voice messages, enabling multi-dimensional interaction with users and promptly meeting their needs. For example, a customer service digital human can interact with real users to promptly meet their actual needs related to products, such as inquiries about product information and discount rules. However, during the interaction between the digital human and the user, the displayed audio and video data may be out of sync, with facial expressions lagging behind the audio content. This reduces the digital human's expressiveness and negatively impacts the user's interactive experience.

[0026] This disclosure provides a method, training method, agent, device, electronic device, and storage medium for generating audio and video data based on a large model. The method includes: receiving speech text and audio / video prompt data related to a specified object, wherein the audio / video prompt data represents the pronunciation content of the specified object and the corresponding facial expression changes; semantically fusing the speech text and audio / video prompt data using a large model to obtain semantic fusion features, wherein the semantic fusion features represent the pronunciation prosody of the specified object in response to the speech text, and the facial expression visual semantics aligned with the pronunciation prosody sequence; and generating target audio / video data based on the semantic fusion features, wherein the target audio / video data represents the target pronunciation content of the specified object in response to the speech text, and the corresponding visual facial expression content of the specified object.

[0027] According to embodiments of this disclosure, by utilizing a large model to perform semantic fusion on speech text and audio / video prompt data representing a specified object, the semantic fusion features can represent, based on temporal alignment, the facial visual semantics of the specified object broadcasting speech text in the temporal space, the prosodic semantics of the speech text, and the pronunciation content attributes of the text words in the speech text. This allows the semantic fusion features to represent the prosodic pronunciation of the specified object in response to speech text, as well as the facial visual semantics naturally aligned with the prosodic pronunciation, achieving deep fusion of multimodal semantic attributes such as speech text, facial visual semantics of the specified object, and prosodic pronunciation based on semantic fusion features. Furthermore, based on the multimodal semantic attributes represented by the semantic fusion features and naturally aligned with the temporal order, the target audio data in the generated target audio / video data and the video content in the target video data achieve natural temporal alignment. This enables the use of target audio / video data to represent the speech content and corresponding facial expression content of the specified object displayed under conditions of synchronized speech and video, thereby improving the data quality of audio / video data in interactive scenarios and ultimately enhancing the interactive experience and audio / video browsing experience of the target object.

[0028] Figure 1 The illustration schematically shows an exemplary system architecture for applying a large-model-based audio and video data generation method and apparatus according to embodiments of the present disclosure.

[0029] It is important to note that Figure 1 The examples shown are merely examples of system architectures applicable to embodiments of this disclosure, intended to help those skilled in the art understand the technical content of this disclosure. However, they do not imply that embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. For instance, in another embodiment, an exemplary system architecture applicable to the large-model-based audio and video data generation method and apparatus may include a terminal device. However, the terminal device can implement the large-model-based audio and video data generation method and apparatus provided by embodiments of this disclosure without interacting with a server.

[0030] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0031] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).

[0032] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0033] Server 105 can be a server that provides various services, such as a backend management server that supports the content browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0034] It should be noted that the large-model-based audio and video data generation method provided in this disclosure can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Correspondingly, the large-model-based audio and video data generation device provided in this disclosure can also be disposed in the first terminal device 101, the second terminal device 102, or the third terminal device 103.

[0035] Alternatively, the large-model-based audio and video data generation method provided in this disclosure can generally also be executed by server 105. Correspondingly, the large-model-based audio and video data generation apparatus provided in this disclosure can generally be located in server 105. The large-model-based audio and video data generation method provided in this disclosure can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the large-model-based audio and video data generation apparatus provided in this disclosure can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0036] For example, when a user is reading an ebook online, the first terminal device 101, the second terminal device 102, and the third terminal device 103 can acquire the target content in the ebook that the user is looking at, and then send the acquired target content to the server 105. The server 105 analyzes the target content to determine its feature information; predicts the content that the user is interested in based on the feature information; and extracts the content that the user is interested in. Alternatively, a server or server cluster that can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105 can analyze the target content and ultimately extract the content that the user is interested in.

[0037] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0038] Figure 2 A flowchart illustrating a method for generating audio and video data based on a large model according to an embodiment of the present disclosure is shown.

[0039] like Figure 2 As shown, the audio and video data generation method based on a large model includes operations S210~S230.

[0040] In operation S210, voice text and audio / video prompt data related to a specified object are received.

[0041] In operating S220, a large model is used to semantically fuse speech text and audio / video prompt data to obtain semantic fusion features.

[0042] In operation S230, target audio and video data are generated based on semantic fusion features.

[0043] According to embodiments of this disclosure, the speech text represents the text content that a specified object in the target audio / video data to be generated needs to express. Speech text can be, for example, a speech script, a novel script, an advertising / marketing script, or an answer text used to respond to real-time input from a target object. Speech text can include multiple text characters. In some embodiments, a text character can be a string of syllables formed by spelling consonants and vowels. For example, if the speech text is Chinese text, the text character can be a Chinese character in the speech text; if the speech text is English or French, the text character can be an English word or a French word; or, the text character can also represent a string of syllables composed of consonants and vowels in an English or French word.

[0044] According to embodiments of this disclosure, the designated object can be an authorized cartoon character, a real-life human figure, or any other image that can be displayed on the display interface. The designated object can be of any type, such as a two-dimensional cartoon character or a three-dimensional human figure. Embodiments of this disclosure do not limit the type of the designated object, as long as it can obtain relevant authorization and be displayed on the display interface.

[0045] According to embodiments of this disclosure, audio-visual prompt data represents the pronunciation content of a specified object, as well as the corresponding facial expression changes. The audio-visual prompt data associated with the specified object can be audio-visual data recording the speech and facial expressions of the specified object. For example, the audio-visual prompt data can be video content carrying audio data obtained by recording a live broadcast of the specified object using a camera with sound recording capabilities.

[0046] According to embodiments of this disclosure, the audio prompt data and video prompt data in the audio and video prompt data are time-synchronized, thereby enabling the time-synchronized characterization of the facial expression changes of a specified object during the pronunciation process, which are aligned with the pronunciation content.

[0047] According to embodiments of this disclosure, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions or even trillions of model parameters. Large models can include large language models (LLMs), visual large models, multimodal large models, and so on. The large model involved in the embodiments of this disclosure can be a multimodal large model with multimodal data processing capabilities, or it can be a large language model. This model can perform semantic encoding on audio and video prompt data, mapping the visual semantics and audio pronunciation prosodic semantics of the audio and video prompt data into an encoding space, thereby enabling the large language model to fuse the encodings of speech text and audio and video data. This achieves semantic fusion of speech text and audio and video prompt data using a large model.

[0048] According to embodiments of this disclosure, the target audio-visual data represents the target pronunciation content of a specified object in response to spoken text, and the facial expression content of the specified object corresponding to the target pronunciation content. For example, the target audio-visual data includes target video data and target speech data. The target speech data represents the target pronunciation content of the specified object pronouncing the spoken text, and the target video data drives the lip movements of the specified object in response to the spoken text, and drives the specified object to perform facial expression movements corresponding to the pronunciation content.

[0049] According to embodiments of this disclosure, semantic fusion of speech text and audio / video prompt data using a large model can deeply integrate textual semantics in speech text, audio semantics of the pronunciation content of a specified object represented by audio / video data, and facial visual semantics that are temporally aligned with the pronunciation content in audio / video data, based on the large-scale model parameters and complex model structure of the large model. This enables the semantic fusion features to naturally align the pronunciation content of the specified object with the changes in facial visual semantics in speech text in a natural temporal sequence.

[0050] In some embodiments, articulation prosody represents the pronunciation rhythm, tone, rate of speech, stress, and emotion of structured text content such as text characters, words or phrases composed of multiple text characters, and text sentences, which are related to pronunciation effect. Semantic fusion features can learn audio semantic attributes related to the pronunciation effect of a specified object in relation to structured text content by fusing audio and video cue data. This enables the semantic fusion features to represent the articulation prosody of the specified object when it pronounces the text according to speech. Furthermore, by learning the temporal alignment relationship between the facial visual semantics represented in the audio and video data and the articulation prosody, the semantic fusion features can temporally align the articulation prosody and facial visual semantics of the specified object in relation to speech text.

[0051] In some embodiments, the semantic fusion features include multiple sequentially arranged sub-features. These sub-features represent the pronunciation content of multiple text characters in the target audio / video data segment, where a specified object pronounces them according to pronunciation attributes related to pronunciation rhythm, such as timbre, speech rate, pitch, and accent, as represented by the audio / video cues. Simultaneously, the sub-features can also represent the visual semantics of the specified object performing facial expressions corresponding to the multiple text characters, according to the facial expression visual semantics represented by the audio / video cues. Therefore, based on these sequentially arranged sub-features, the temporal alignment relationship between the facial expression visual semantics and pronunciation content of specified objects in multiple target data segments of the target audio / video data can be represented, achieving precise temporal synchronization between the pronunciation content represented by the speech data and the facial expression content represented by the video data in the generated target audio / video data, thus improving the data quality of the target audio / video data.

[0052] According to embodiments of this disclosure, semantic fusion features characterize the pronunciation prosody of a specified object in relation to spoken text, as well as the facial visual semantics aligned with the temporal sequence of the pronunciation prosody. By learning the alignment relationship between the changes in the facial visual semantics of a specified object in audio and video data and the changes in the pronunciation prosody in the spoken content through a large model, the semantic fusion features can accurately align the facial visual semantics of the specified object, the pronunciation attribute semantics represented by the pronunciation prosody, and the pronunciation attributes of the words in the spoken text in a precise temporal sequence. This allows the semantic fusion features to more accurately characterize the changes in facial visual semantics generated by the specified object during the pronunciation of spoken text.

[0053] This enables semantic fusion features to achieve natural temporal alignment of textual semantics, facial expression visual semantics, and pronunciation prosody. Consequently, the target audio and video data can represent the process of a specified object broadcasting audio and video in sync with the text, reducing inconsistencies between the voice content and facial expression changes of the specified object in the audio and video data, improving the data quality of the target audio and video data, and enhancing the immersive experience of the target object when browsing the target audio and video data.

[0054] In some embodiments, the large model-based audio and video generation method further includes: receiving demand information input by the target object; using the large model to perform intent understanding on the demand information to obtain speech text that matches the demand intent of the target object.

[0055] According to embodiments of this disclosure, the designated object is a digital human displayed on the display interface of a terminal device. The target object can input required information through any type of input operation, such as voice input or touch screen output. The required information can be any data type, such as voice or text.

[0056] In one example, the demand information can be the question input by the target object. A large model is used to understand the intent of the question information, outputting the answer text as speech text that satisfies the target object's intent. By using the large model to perform semantic understanding on the answer text and audio / video prompts, semantic fusion features are obtained. Based on these features, target audio / video data is generated, enabling the target object to respond with speech according to the text content of the answer text. During the speech response, the target object is also driven to perform facial expressions based on the visual semantics of facial expressions represented by the audio / video data. This allows the target object to experience interaction under synchronized audio-visual conditions by browsing the target video data displayed on the smart terminal device's screen and simultaneously listening to the target speech data, thus enhancing the interactive experience.

[0057] In some embodiments, semantic fusion of speech text and audio / video prompt data using a large model may include: performing semantic fusion of speech text and audio / video prompt data using a large model based on control instructions determined based on received intent description information to obtain semantic fusion features.

[0058] According to embodiments of this disclosure, the intent description information represents the prosodic intent and facial expression visual intent for spoken text. The prosodic intent includes the intent to control the prosodic pronunciation of structured text content such as text characters, words or phrases composed of multiple text characters, and text sentences. For example, the intent description information can be based on natural language representations such as "speak loudly" or "speak quickly and loudly," or, for example, the prosodic intent can be the prosodic pronunciation of structured text representations such as "[speaking speed: fast], [facial expression: happy]," where "[]" can represent structured markers related to the prosodic intent or facial expression visual intent. The facial expression visual intent can represent the intent to control the facial expression actions of a specified object during the pronunciation of spoken text. The facial expression visual intent can be, for example, text or strings related to facial expressions or emotions such as "happy," "frustrated," or "sad."

[0059] The intention describes the phonetic prosody intention and facial expression visual intention of the information representation. It describes the phonetic prosody that the target speech data should have when broadcasting speech text to a specified object in the target audio and video data, and the facial expressions and actions that the specified object in the target video data needs to display corresponding to the speech content of the target speech data.

[0060] According to embodiments of this disclosure, semantic fusion features represent the pronunciation content of a specified object pronouncing speech text according to its prosodic intention, and the facial expression visual semantics of performing facial expressions according to its visual intention. The prosodic intention and visual expression intention, represented by the intention expression information, can be used as control instructions to control a large model to perform semantic fusion on speech text and audio-visual prompt data. This allows the large model to leverage its semantic understanding capabilities to deeply fuse the pronunciation attributes represented by the prosodic intention with the pronunciation patterns of characters in the speech text, and the unique pronunciation patterns of the specified object, such as timbre and speech rate, in the audio-visual prompt data. Simultaneously, the large model can also deeply fuse the visual expression semantics represented by the visual expression intention with the lip movements corresponding to the pronunciation patterns of characters in the speech text, and the unique visual expression semantics of the specified object represented by the audio-visual prompt data. This enables the semantic fusion features to represent the specified object pronouncing speech text according to its prosodic intention. The semantic fusion feature aligns the pronunciation content of the specified object with the pronunciation prosodic sequence of the pronunciation content. This allows for the generation of target audio-visual data based on the semantic fusion feature. This data enables the specified object to pronounce the speech text according to its prosodic intent, and simultaneously executes facial expressions according to the visual intent during pronunciation. This improves the flexibility of the speech playback and facial expressions of the specified object in the target audio-visual data, making it suitable for real-time interactive scenarios such as interactions with the target object. This enhances the data quality of the target audio-visual data and improves the interactive experience of the target object.

[0061] In some embodiments, the intent description information includes sub-description information corresponding to multiple text characters in the spoken text. The sub-description information may represent at least one of a phonetic prosodic intent and an expressive visual intent for the multiple text characters.

[0062] In some embodiments, semantic fusion of speech text and audio-visual prompt data using a large model based on control instructions determined based on received intent description information may include: using a large model to perform semantic fusion on multiple text characters in audio-visual prompt data and speech text based on the pronunciation prosodic intent and facial expression visual intent represented by sub-description information, to obtain sub-features corresponding to the multiple text characters.

[0063] According to embodiments of this disclosure, the semantic fusion feature includes a plurality of sequentially arranged sub-features. Each sub-feature represents the pronunciation content of a specified object pronouncing multiple text characters according to the phonetic prosodic intent corresponding to the sub-description information, and the facial expression visual semantics of performing facial expression actions according to the facial expression visual intent corresponding to the sub-description information, within a data segment of the target audio / video data. The data segment of the target audio / video data can be represented as a time-aligned target speech data segment and a target video data segment.

[0064] In one example, the intent description information could be: "Say it loudly and happily: 'We've graduated!' ... Say it softly and sadly: 'See you at school again next year,'" where "Say it loudly and happily" and "Say it softly and sadly" are two different sub-descriptions. These two different sub-descriptions correspond to "We've graduated!" and "See you at school again next year" in the spoken text, respectively. This demonstrates refined pronunciation rhythm and facial expression control.

[0065] According to embodiments of this disclosure, a large model is used to fuse the pronunciation prosody and facial expression visual intent of a specified object represented by audio and video data, based on the pronunciation prosody and facial expression visual intent represented by the sub-description information in the intent description information. This also fuses the pronunciation patterns and lip movements of multiple corresponding text characters. This allows the sub-features to more accurately represent the pronunciation prosody of a specified object loudly proclaiming "We've graduated!", and to more accurately represent the lip movements and facial expression changes of the specified object during the pronunciation of "We've graduated!" according to happy facial expression visual semantics. Furthermore, the sub-features can align the pronunciation prosody and facial expression visual changes of the specified object proclaiming "We've graduated!" in a temporal sequence. Therefore, based on the facial expression visual intent and pronunciation prosody intent represented by multiple sub-description information in the intent description information, the generated target audio and video data can represent the speech semantic change process of the specified object in the process of broadcasting different texts in the speech text, as well as the picture content representing the facial expression visual change corresponding to the pronunciation prosody change. This can improve the expressiveness of the specified object in the target audio and video data, and drive the specified object to make synchronous changes in pronunciation prosody and facial expression based on multiple sequentially arranged sub-features, thereby improving the quality and expressiveness of the target audio and video data and enhancing the naturalness of the specified object's expression.

[0066] In some embodiments, the intent description information includes at least one of the following: natural language text representing facial visual intent and articulatory prosodic intent based on the grammatical structure of natural language; and facial intent text and prosodic intent text associated with structured tags.

[0067] In one example, natural language text representing facial expression and prosodic intent based on the grammatical structure of natural language could be: "Speak the input speech text aloud happily." Alternatively, it could be: "The prosodic tone is fast and loud; the facial expression is a happy expression." This allows the large model's natural language understanding capabilities to interpret the prosodic and facial expression intents represented in the intent description information. This enables users to generate intent description information based on natural language descriptions, improving the ease of control over the facial expression content and prosodic tone of a specified object in target audio / video data.

[0068] In one example, the facial expression intent text and prosodic intent text associated with structured markup could be: "[Say it out loud] 'The weather is nice today' (pleasant expression)" where "[]" is a structured markup representing the prosodic intent, "Say it out loud" is the prosodic intent text, "()" is a structured markup representing the facial expression intent, and "pleasant expression" is the facial expression intent text. Structured markup can represent the facial expression intent and prosodic intent of the text word "The weather is nice today" in spoken text.

[0069] According to embodiments of this disclosure, by using structured markers in the intent description information to represent associated facial expression intent text and prosodic intent text, large models can more accurately fuse the pronunciation prosodic intent and facial visual intent with the pronunciation mode and facial visual semantics of the specified object represented by the audio cue data based on the structured markers. It also fully integrates the pronunciation content of the text in the speech text, so that the semantic fusion features can accurately simulate the process of the specified object's pronunciation and facial expression changes in response to the speech text according to the facial visual intent and pronunciation prosodic intent. This improves the accuracy and naturalness of the target audio and video data in representing the specified object, and achieves temporal synchronization between the speech content and the facial expression content, thereby enhancing the viewing experience of the target audio and video data.

[0070] It should be noted that the text in the intent description information can represent at least one of facial expression visual intent and phonetic prosodic intent. For example, "happy" represents both the emotional expression of happiness and the phonetic prosodic intent of happiness. The embodiments of this disclosure do not limit the type of intent represented by the text in the intent description information.

[0071] In some embodiments, the intent description information is determined based on the following operations: using a large model to perform semantic understanding of the demand information of the target object to obtain the facial expression visual intent and phonetic prosodic intent related to the specified object; using the large model to perform a text generation task based on the demand information, facial expression visual intent and phonetic prosodic intent to obtain speech text and intent description information that match the demand intent represented by the demand information.

[0072] According to embodiments of this disclosure, a target object inputs requirement information by interacting with a designated object. For example, the target object can input the requirement text "How's the weather today?" by performing an interactive operation on a terminal device. It should be noted that the requirement information can include any type of data such as voice, text, images, and video. Embodiments of this disclosure do not limit the data type of the requirement information, as long as it can represent the target object's requirement.

[0073] According to embodiments of this disclosure, facial expression visual intent and phonetic prosodic intent are used to respond to the emotions represented by demand information. Utilizing a large model to perform semantic understanding of the demand information can include using the large model to understand the textual semantics, facial image semantics, and vocal emotion semantics of the demand information representation to determine the target object's emotion type or emotional demand intent. This enables the large model to generate speech text and intent description information that respond to the target object's emotion type or emotional demand intent based on the emotion type or emotional demand intent.

[0074] In one example, the demand information consists of a facial image representing the target object's facial expression and the target object's input demand speech. By utilizing a multimodal large model to perform semantic understanding on the facial image and demand speech, the target object's current emotion type is determined to be "sad." The large model can then use the target object's emotion type as a cue to process the input demand speech and generate a soothing speech text, "Please don't be sad...", along with the intent description information, "Speak calmly, with a gentle smile." Thus, the large model can fuse the visual and prosodic intent of the facial expression cues with the prosodic intent of the intent description information to represent the specified object's facial expression and prosodic meaning, and integrate the pronunciation content of multiple words in the speech text. This allows the target audio-visual data to drive the specified object to read the speech text with a smooth pronunciation rhythm and speed, simultaneously displaying the specified object's gentle smile during the speech text reading.

[0075] According to embodiments of this disclosure, by generating voice text and intent description information that respond to the emotional needs and intentions of the target object based on the target object's needs information, and by using a large model to fuse the audio and video prompt data and voice text of the specified object under the control command conditions determined by the intent description information, the target audio and video data displayed by driving the specified object during the interaction with the target object can more accurately respond to the target object's emotional needs, and provide a natural and smooth audio-visual synchronized interactive experience, satisfying the target object's multi-dimensional needs and intentions such as knowledge acquisition needs and emotional response needs, thereby improving the interactive experience.

[0076] In some embodiments, generating target audio and video data based on semantic fusion features may include: extracting features from semantic fusion features based on audio and video prompt data to obtain time-aligned target speech features and target visual features; and fusing target speech data determined based on target speech features and target video data determined based on target visual features to obtain target audio and video data.

[0077] According to embodiments of this disclosure, audio-visual prompt data and semantic fusion features can be fused using a Diffusion Model mechanism. This allows for fine-grained alignment and semantic enhancement of semantic fusion features representing multimodal semantics in the temporal space, utilizing the pronunciation prosody and facial expression visual semantics of a specified object represented by the audio-visual prompt data. Furthermore, it captures the dynamic mapping relationship between multimodal semantics and the specified object-related semantics corresponding to the audio-visual prompt data in a unified temporal dimension. This improves the accuracy and naturalness of the target speech features and target visual features in the unified temporal space, enhancing the pronunciation content and facial expression visual content corresponding to the speech text of the specified object. Finally, by fusing temporally aligned target audio data and target video data, the target audio-visual data enables the specified object to perform natural and fluent speech and facial expressions, improving the data quality and viewing experience of the target audio-visual data.

[0078] In some embodiments, feature extraction of semantic fusion features based on audio and video prompt data may include: denoising preset noise data based on audio and video prompt data and semantic fusion features to obtain time-aligned target speech features and target visual features.

[0079] In some embodiments, a diffusion model mechanism can be used to process audio-visual cue data and semantic fusion features to obtain temporally aligned target speech features and target visual features. For example, temporally aligned text semantics, pronunciation prosody, and facial expression visual semantics represented by semantic fusion features, as well as the pronunciation mode and facial expression visual semantics of the target object represented by audio-visual cue data, can be used as denoising conditions. A diffusion model network is then used to denoise the preset noise data according to the denoising conditions, thereby obtaining target speech features and target visual features. A trained audio-visual encoder is then used to process the temporally aligned target speech features and target visual features to obtain target audio data and target video data. Target audio data and target video data can represent speech content and video content of the same duration. Therefore, target audio-visual data can be obtained by concatenating target video data and target audio data according to a reference time.

[0080] In one example, a feature extraction network with a diffusion model mechanism can be constructed based on the Transformer architecture. This feature extraction network processes audio-visual cue data and semantic fusion features, performing time-step-level conditional modeling and feature reconstruction on multiple sub-features within the semantic fusion features to obtain time-step-aligned target speech and visual features. The cross-modal attention layer within the feature extraction network implements dynamic semantic interaction between speech and visual semantics in the semantic fusion features based on the audio-visual cue data. This ensures that the generated target speech and visual features maintain consistency and smoothness across facial expression and speech content, thereby improving the data quality of the target audio-visual data and enhancing the expressiveness and naturalness of the specified object within it.

[0081] Figure 3 The illustration shows a schematic diagram of the principle of a large-model-based audio and video data generation method according to an embodiment of the present disclosure.

[0082] like Figure 3 As shown, the large model can be constructed based on a Mixture of Experts (MoE) model structure. The feature extraction network is built based on a diffusion model mechanism, and the audio / video translation network is built based on a Variational Autoencoder. The speech text can be the streaming input text used to respond to the target object's current needs. The audio / video cue data represents the prosodic and facial visual semantics of the specified object. The intent description information, representing the facial visual intent and prosodic intent, is used to respond to the target object's current emotional needs.

[0083] A large model is used to process audio / video prompt data, speech text, and intent description information. This enables cross-modal fusion of the speech semantics (pronunciation attributes, prosody, etc.) represented by the audio / video prompt data, speech text, and intent description information, as well as the textual semantics (facial expression visual semantics and speech text), based on the large model. These are then mapped to the same high-dimensional latent space, allowing the semantic fusion feature 310 to form a multimodal semantic representation with temporal and semantic consistency in the same temporal space. This achieves the projection and normalization of semantic information from different modalities within the high-dimensional latent space, enabling the semantic fusion feature 310 to contain a joint latent representation of prosody, contextual text semantics, and facial expression visual semantics. For example, the first sub-feature in the semantic fusion feature 310 can represent the temporally aligned prosody and facial expression visual semantics in the first data segment of the target audio / video data.

[0084] Therefore, based on the semantic fusion features, the temporal synchronization of pronunciation prosody and facial expression visual semantics can reduce the problem of modal semantic delay or phase drift caused by cross-modal fusion of different semantic features. This enables the semantic fusion feature 310 to naturally coordinate the representation of the object for speech text, facial expression visual intent and pronunciation prosody intent, laying the foundation for generating high-precision target audio and video data with synchronized audio and video.

[0085] The feature extraction network can denoise the preset noise data based on audio-visual prompt data and semantic fusion features 310, enabling the network to extract acoustic and visual features from the semantic fusion features 310, which are represented in the same temporal and spatial domain, according to the denoising conditions, thus obtaining target speech features and target visual features. The feature extraction network can enhance the temporal synchronization accuracy of the audio and visual modal semantics represented in the semantic fusion features 30, and also achieve the accuracy of representing emotions through facial expressions at the latent space level, enabling the collaborative generation of pronunciation prosody such as tone and intonation matched with facial expressions and visual semantics. This allows multiple sub-features in the semantic fusion features 310 to stably represent the natural dynamic changes in pronunciation prosody and visual semantics, providing high-fidelity input features for subsequent feature extraction networks and audio-visual data conversion networks.

[0086] The audio-video data conversion network can process target speech features and target visual features to map the acoustic latent vectors and visual latent vectors represented by the target speech features and target visual features respectively to the real perceptual domain space, thereby obtaining target audio data represented by Mel-spectrum data and target video data represented by video frame sequences. Multiple video frames in the video frame sequence can represent the facial expressions and actions performed by the target object according to its visual intent. This achieves the restoration of the semantic fusion features 310 from the high-dimensional latent space into multimodal signals that the target object can perceive, realizing a synchronous generation effect of "what you hear is what you see".

[0087] In some embodiments, the audio / video prompt data includes first prompt data and second prompt data, with the second prompt data having higher data quality than the first prompt data. The data quality of the first and second prompt data can be determined based on indicators such as the clarity, data size, and number of facial expression types represented by the audio / video data.

[0088] For example, the amount of data in the second prompt data is greater than the amount of data in the first prompt data; or, for example, the clarity of the second prompt data is higher than the clarity of the first prompt data; or, for example, the number of facial expression types of the specified object represented by multiple video data in the second prompt data is greater than the number of facial expression types of the specified object represented by multiple video data in the first prompt data.

[0089] In one example, the semantic fusion features are determined by semantically fusing the first cue data and the speech text. Specifically, feature extraction of the semantic fusion features based on the audio / video cue data includes: feature extraction of the semantic fusion features based on the second cue data.

[0090] For example, a large model can be used to perform semantic understanding on speech text based on first cue data with lower data quality, outputting semantic fusion features. Based on second cue data, the large model can be used to extract features from the semantic fusion features, obtaining time-aligned target speech features and target visual features. Thus, the second cue data with higher data quality can be used as a noise reduction condition for the extension mechanism to further enhance the pronunciation prosody and facial expression visual semantics related to the specified object in the semantic fusion features, enabling the target speech features and target visual features to more accurately represent the expression and pronunciation prosody of the specified object. Therefore, in the process of executing the large model-based audio and video data generation method provided in this embodiment using a graphics processing chip or other computing unit, the amount of data that the computing unit needs to process during the generation of semantic fusion features is reduced, and the overall computational overhead of the method provided in this embodiment is reduced. This improves the generation efficiency of target audio and video data and the response efficiency to the interactive needs of the target object, making the method based on this embodiment adaptable to scenarios where virtual avatars interact with target objects in real-time, thus enhancing the interactive experience of the target object.

[0091] Figure 4 A flowchart illustrating a method for training a large model according to an embodiment of the present disclosure is shown schematically.

[0092] like Figure 4 As shown, the method for training the large model includes operations S410~S430.

[0093] During operation of S410, sample speech text, sample audio and video prompt data related to the sample object, and label semantic fusion features are received.

[0094] In operating the S420, a large model is used to semantically fuse sample speech text and sample audio and video prompt data to obtain sample semantic fusion features.

[0095] In operation S430, a large model is trained based on the difference between the sample semantic fusion features and the label semantic fusion features, resulting in a trained large model.

[0096] According to embodiments of this disclosure, the sample audio and video prompt data can be real data obtained by capturing video and voice data of the sample object during its speech expression.

[0097] According to embodiments of this disclosure, the tag semantic fusion features characterize the pronunciation prosody of sample objects in sample audio / video prompt data in relation to sample speech text, as well as the facial visual semantics aligned with the temporal sequence of the pronunciation prosody. For example, feature extraction can be performed on the tag speech data based on a trained encoder to obtain the tag semantic fusion features.

[0098] In some embodiments, training a large model based on the difference between sample semantic fusion features and label semantic fusion features may include using a loss function to process the sample semantic fusion features and label semantic fusion features to obtain a loss value. The model parameters of the large model are then adjusted based on the loss value until the loss function converges, resulting in a trained large model.

[0099] It should be noted that the technical terms involved in the method for training a large model provided in this disclosure, including but not limited to sample semantic fusion features, sample speech-text, sample intent description information, and large model, have the same or similar attributes as the technical terms involved in the audio and video data generation method based on a large model provided in this disclosure, including but not limited to semantic fusion features, speech-text, intent description information, and large model, and belong to the same type of technical features. The embodiments of this disclosure will not be repeated here.

[0100] The large model, feature extraction network, and audio / video data conversion network involved in the training method of the large model provided in this disclosure have the same or similar model structure as the large model, feature extraction network, and audio / video data conversion network involved in the speech data generation method based on the large model provided in this disclosure. Therefore, the embodiments of this disclosure will not be described again.

[0101] Figure 5 A schematic diagram illustrating a method for training a large model according to an embodiment of the present disclosure is shown.

[0102] like Figure 5 As shown, a trained encoder is used to process sample audio and video cue data to obtain label semantic fusion features. Label semantic fusion feature 502 integrates the timbre, pitch, and other pronunciation attributes and prosody of the sample objects represented by the sample audio and video data, while also integrating the visual semantics of the sample objects' facial expressions. Therefore, label semantic fusion feature 502 can be used as the ground truth for fusing multimodal semantics to train a large model.

[0103] The large model performs semantic fusion on sample audio / video prompts and sample speech / text to obtain sample semantic fusion features 501. A loss function is then applied to sample semantic fusion features 501 and label semantic fusion features 502 to obtain a loss value. The model parameters of the large model are adjusted using this loss value until the loss function converges, resulting in a trained large model.

[0104] In some embodiments, the label semantic fusion features are determined based on sample audio and video prompt data processed by a trained encoder. The trained encoder is determined based on the following operations: extracting features from reference audio and video data using the encoder to obtain pre-trained semantic fusion features; generating pre-trained audio and video data based on the pre-trained semantic fusion features; and training the encoder based on the differences between the pre-trained audio and video data and the reference audio and video data to obtain the trained encoder.

[0105] According to embodiments of this disclosure, the reference audio-visual prompt data is related to the sample object. For example, the reference audio-visual prompt data can be real data obtained by collecting voice and video data from the sample object. The reference audio-visual prompt data characterizes the pronunciation rhythm of the sample object in response to the reference voice text, as well as the facial visual semantics aligned with the timing of the pronunciation rhythm. The reference audio-visual data is the audio-visual data of the sample object broadcasting according to the reference voice text, and the reference audio-visual data characterizes the timing-aligned facial visual semantics and the reference voice content.

[0106] According to embodiments of this disclosure, the encoder can be constructed based on an attention network algorithm, or it can be constructed based on other types of neural network algorithms. The embodiments of this disclosure do not limit the specific algorithm type for constructing the speech encoder.

[0107] According to embodiments of this disclosure, pre-trained semantic fusion features can represent the pronunciation prosody of sample objects in reference audio-visual prompt data and the corresponding facial visual semantics. Features can be extracted from the pre-trained semantic fusion features using a feature extraction network to obtain pre-trained speech features and pre-trained visual features. The pre-trained speech features and pre-trained visual features are then processed using an audio-visual conversion network to obtain time-aligned pre-trained speech data and pre-trained video data. Finally, pre-trained audio-visual data is obtained by fusing the pre-trained speech data and pre-trained video data.

[0108] In some embodiments, a pre-training loss value can be obtained by processing pre-training audio / video data and reference audio / video data based on a loss function. The model parameters of the encoder, feature extraction network, and audio / video conversion network are then adjusted using this pre-training loss value until it converges, resulting in a trained encoder, a trained feature extraction network, and a trained audio / video conversion network. This enables the trained encoder to perform cross-modal semantic fusion on sample audio / video data and speech text, facilitating the training of a large model based on the difference between the sample semantic fusion features output by the large model and the label semantic fusion features output by the trained encoder.

[0109] In one example, during the training of the feature extraction network, temporal noise perturbation can be applied to the pre-trained semantic fusion features to simulate the continuity of natural speech and facial expression changes. Subsequently, the feature extraction network, based on a back-diffusion process using a diffusion model mechanism, gradually recovers the clear multimodal semantic representation of the pre-trained semantic fusion features, thereby capturing the coupling patterns between audio prosody changes, semantic transitions, and facial expression transitions in the temporal domain. This allows the temporally aligned pre-trained speech features and pre-trained visual features to accurately represent the facial expression changes of the sample object during pronunciation in the temporal domain. Therefore, the trained feature extraction network can be applied to the large-model-based audio and video data generation method provided in the embodiments of this disclosure to process the semantic fusion features using the trained feature extraction network to obtain temporally aligned target speech features and target visual features.

[0110] In some embodiments, generating pre-trained audio and video data based on pre-trained semantic fusion features includes: extracting features from the pre-trained semantic fusion features using reference audio and video prompt data as prompt information to obtain time-aligned pre-trained speech features and pre-trained visual features; and fusing the pre-trained speech data determined based on the pre-trained speech features and the pre-trained video data determined based on the pre-trained visual features to obtain pre-trained audio and video data.

[0111] In some embodiments, a large model is used to semantically fuse sample speech text and sample audio / video prompt data to obtain sample semantic fusion features, including:

[0112] Based on the control instructions determined by the received sample intent description information, a large model is used to perform semantic fusion on the sample speech text and sample audio-visual prompt data to obtain sample semantic fusion features. Among them, the sample intent description information represents the pronunciation prosody of the sample object in the sample audio-visual prompt data in relation to the pronunciation content of the sample speech text, as well as the facial expression visual semantics corresponding to the pronunciation content.

[0113] According to embodiments of this disclosure, by including control instructions determined based on sample intent description information in the information input to the large model during training, the trained large model can control the facial expression and prosody of the semantic fusion features according to the pronunciation prosodic intent and facial expression visual intent represented by the intent description information. This improves the ability of the target audio and video data to flexibly adapt to diverse needs to drive the interaction between the specified object and the target object, thereby enhancing the interactive experience of the target object.

[0114] Figure 6 A block diagram of a large-model-based audio and video data generation apparatus according to an embodiment of the present disclosure is shown schematically.

[0115] like Figure 6As shown, the audio and video data generation device 600 based on a large model includes: a first receiving module 610, a semantic fusion feature acquisition module 620, and a target audio and video data generation module 630.

[0116] The first receiving module 610 is used to receive voice text and audio-visual prompt data related to a specified object. The audio-visual prompt data represents the pronunciation content of the specified object and the facial expression changes corresponding to the pronunciation content.

[0117] The semantic fusion feature acquisition module 620 is used to perform semantic fusion of speech text and audio-visual prompt data using a large model to obtain semantic fusion features. The semantic fusion features characterize the pronunciation prosody of a specified object in relation to the speech text, as well as the facial visual semantics aligned with the pronunciation prosody time sequence.

[0118] The target audio and video data generation module 630 is used to generate target audio and video data based on semantic fusion features. The target audio and video data represents the target pronunciation content of a specified object in relation to the speech text, as well as the on-screen facial expressions of the specified object corresponding to the target pronunciation content.

[0119] According to embodiments of this disclosure, the semantic fusion feature acquisition module includes a first acquisition unit.

[0120] The first obtaining unit is used to perform semantic fusion on speech text and audio-visual prompt data using a large model based on the control instructions determined based on the received intent description information, to obtain semantic fusion features. The intent description information represents the pronunciation prosodic intent and facial expression visual intent of the speech text, and the semantic fusion features represent the pronunciation content of the specified object pronouncing the speech text according to the pronunciation prosodic intent, and the facial expression visual semantics of performing facial expression actions according to the facial expression visual intent.

[0121] According to embodiments of this disclosure, the intent description information includes sub-description information corresponding to multiple text characters in the speech text, and the semantic fusion features include multiple sub-features arranged in sequence; wherein, the first obtaining unit includes a sub-feature obtaining sub-unit.

[0122] The sub-feature acquisition sub-unit is used to perform semantic fusion on multiple text characters in audio-visual prompt data and speech text based on the pronunciation prosodic intent and facial expression visual intent represented by the sub-description information of the large model, and to obtain sub-features corresponding to multiple text characters. Among them, the sub-features represent the pronunciation content of the specified object in the data segment of the target audio-visual data, which pronounces multiple text characters according to the pronunciation prosodic intent corresponding to the sub-description information, and the facial expression visual semantics of performing facial expression actions according to the facial expression visual intent corresponding to the sub-description information.

[0123] According to embodiments of this disclosure, the intent description information includes at least one of the following: natural language text representing facial expression visual intent and phonetic prosodic intent based on the grammatical structure of natural language; and facial expression intent text and prosodic intent text associated with structured tags.

[0124] According to embodiments of this disclosure, the intent description information is determined based on the following operations: semantic understanding of the target object's demand information is performed using a large model to obtain facial expression visual intent and phonetic prosodic intent related to the specified object, wherein the target object inputs demand information by interacting with the specified object, and the facial expression visual intent and phonetic prosodic intent are used to respond to the emotions represented by the demand information; a text generation task is performed using a large model based on the demand information, facial expression visual intent, and phonetic prosodic intent to obtain speech text and intent description information that match the demand intent represented by the demand information.

[0125] According to embodiments of this disclosure, the audio and video data generation apparatus based on a large model further includes: a demand information receiving module and a voice and text acquisition module.

[0126] The requirement information receiving module is used to receive requirement information input by the target object.

[0127] The speech-text acquisition module is used to understand the intent of demand information using a large model and obtain speech-text that matches the demand intent of the target object.

[0128] According to embodiments of this disclosure, the target audio and video data generation module includes a feature extraction unit and a fusion unit.

[0129] The feature extraction unit is used to extract features from semantic fusion features based on audio and video prompt data, so as to obtain temporally aligned target speech features and target visual features.

[0130] The fusion unit is used to fuse target speech data determined based on target speech features and target video data determined based on target visual features to obtain target audio and video data.

[0131] According to embodiments of this disclosure, the feature extraction unit includes a noise reduction subunit.

[0132] The noise reduction subunit is used to reduce noise from preset noise data based on audio and video prompt data and semantic fusion features, so as to obtain time-aligned target speech features and target visual features.

[0133] According to embodiments of this disclosure, the audio / video prompt data includes first prompt data and second prompt data, the data quality of the second prompt data is higher than that of the first prompt data, and the semantic fusion feature is determined by semantic fusion of the first prompt data and the speech text; wherein, the feature extraction unit includes a feature extraction subunit.

[0134] The feature extraction subunit is used to extract features from semantic fusion features based on the second cue data.

[0135] Figure 7 A block diagram of an apparatus for training a large model according to an embodiment of the present disclosure is shown schematically.

[0136] like Figure 7 As shown, the apparatus 700 for training a large model includes: a second receiving module 710, a sample semantic fusion feature acquisition module 720, and a training module 730.

[0137] The second receiving module 710 is used to receive sample speech text, sample audio and video prompt data related to the sample object, and label semantic fusion features. The label semantic fusion features characterize the pronunciation prosody of the sample object in the sample audio and video prompt data in response to the sample speech text, as well as the facial expression visual semantics aligned with the pronunciation prosody time sequence.

[0138] The sample semantic fusion feature acquisition module 720 is used to perform semantic fusion of sample speech text and sample audio and video prompt data using a large model to obtain sample semantic fusion features.

[0139] Training module 730 is used to train a large model based on the difference between the sample semantic fusion features and the label semantic fusion features, and obtain the trained large model.

[0140] According to embodiments of this disclosure, the semantic fusion features of the tags are determined based on the sample audio and video prompt data processed by a trained encoder. The trained encoder is determined based on the following operations: using the encoder to extract features from the reference audio and video data to obtain pre-trained semantic fusion features, wherein the reference audio and video prompt data is related to the sample object; generating pre-trained audio and video data based on the pre-trained semantic fusion features; and training the encoder based on the differences between the pre-trained audio and video data and the reference audio and video data to obtain the trained encoder.

[0141] According to embodiments of this disclosure, generating pre-trained audio and video data based on pre-trained semantic fusion features includes: extracting features from the pre-trained semantic fusion features using reference audio and video prompt data as prompt information to obtain time-aligned pre-trained speech features and pre-trained visual features; and fusing the pre-trained speech data determined based on the pre-trained speech features and the pre-trained video data determined based on the pre-trained visual features to obtain pre-trained audio and video data.

[0142] According to embodiments of this disclosure, the sample semantic fusion feature acquisition module includes a second acquisition unit.

[0143] The second acquisition unit is used to perform semantic fusion on the sample speech text and sample audio-visual prompt data using a large model based on the control instructions determined based on the received sample intent description information, to obtain sample semantic fusion features. The sample intent description information represents the pronunciation prosody of the sample object in the sample audio-visual prompt data in relation to the pronunciation content of the sample speech text, as well as the facial expression visual semantics corresponding to the pronunciation content.

[0144] Figure 8 A schematic block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is shown.

[0145] In embodiments of this disclosure, such as Figure 8 As shown, the AI ​​agent 800 may include an input module 810, a processing module 820, and an output module 830.

[0146] Input module 810 is used to receive input information;

[0147] The processing module 820 is used to determine the target task based on the input information received by the input module, determine the large model based on the target task, and obtain output information by calling the large model to execute the audio and video data generation method based on the large model according to the embodiments of this disclosure, or by calling the large model to execute the training method of the large model according to the embodiments of this disclosure.

[0148] Output module 830 is used to output the output information obtained by the processing module.

[0149] According to embodiments of this disclosure, the input module 810 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment), and converting it into a format that the AI ​​agent 800 can understand and process. The input module 810 is the primary link for the AI ​​agent 800 to interact with the outside world, enabling the AI ​​agent 800 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.

[0150] In the example, the input module 810 can input the voice text, audio and video prompt data, sample voice text, sample audio and video prompt data, etc., as described above.

[0151] In the example, processing module 820 is the core support for the AI ​​agent 800's ability to handle complex tasks. Processing module 820 can execute the audio and video data generation method based on a large model described earlier, as well as the method for training the large model.

[0152] In the example, the performance of the processing module 820 is closely related to the large model on which the AI ​​agent 800 is based. To fully leverage the capabilities of the large model, the internal structure of the processing module 820 can be designed to be highly configurable and scalable to handle various types of tasks and requirements in real-world scenarios.

[0153] In the example, after the AI ​​agent 800 acquires the required voice, the processing module 820 can use a large model to process the voice text and audio-visual prompt data to obtain semantic fusion features, call the vertical semantic fusion features of the visual large model to obtain the target audio-visual data, and pass the target audio-visual data to the output module 830.

[0154] Understandably, while large language models possess excellent language understanding and generation capabilities, like humans, their ability to solve tasks is limited without the aid of any tools. However, once the AI ​​agent 800 is given the ability to invoke tools, it can perform tasks such as using a calculator to complete mathematical calculations, using Python to perform data analysis, and using a search engine to create weather forecasts.

[0155] In the example, output module 830 can output the target audio and video data or the trained large model described above.

[0156] The AI ​​agent 800 according to embodiments of this disclosure can simply and effectively improve the level of intelligence, as well as enhance flexibility and versatility.

[0157] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0158] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.

[0159] According to embodiments of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause a computer to perform the method described above.

[0160] According to an embodiment of this disclosure, a computer program product includes a computer program that, when executed by a processor, implements the method described above.

[0161] Figure 9A method for generating audio and video data based on a large model, which can be used to implement embodiments of the present disclosure, is illustrated. A schematic block diagram of an example electronic device for training the large model is also provided. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0162] like Figure 9 As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. The RAM 903 may also store various programs and data required for the operation of the electronic device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0163] Multiple components in electronic device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of displays, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows electronic device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0164] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as a method for generating audio and video data based on a large model. Methods for training large models. For example, in some embodiments, a method for generating audio and video data based on a large model. The method for training large models can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, the method for generating audio and video data based on a large model described above can be performed. One or more steps of the method for training large models. Alternatively, in other embodiments, computing unit 901 may be configured, by any other suitable means (e.g., by means of firmware), to perform a method for generating audio and video data based on a large model, or a method for training a large model.

[0165] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0166] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0167] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0168] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0169] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0170] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, distributed system servers, or servers incorporating blockchain technology.

[0171] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0172] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for generating audio and video data based on a large model, comprising: Receive voice text and audio-visual prompt data related to a specified object, wherein the audio-visual prompt data represents the pronunciation content of the specified object and the facial expression changes corresponding to the pronunciation content; Using a large model, the speech text and the audio-visual prompt data are semantically fused to obtain semantic fusion features. The semantic fusion features characterize the pronunciation prosody of the specified object in relation to the speech text, as well as the facial visual semantics aligned with the pronunciation prosody sequence. Target audio and video data is generated based on the semantic fusion features. The target audio and video data represents the target pronunciation content of the specified object in relation to the speech text, as well as the on-screen facial expressions of the specified object corresponding to the target pronunciation content.

2. The method according to claim 1, wherein, The step of semantically fusing the speech text and audio / video prompt data using a large model includes: Based on the control instructions determined by the received intent description information, the large model is used to perform semantic fusion on the speech text and the audio-visual prompt data to obtain the semantic fusion features. The intent description information represents the prosodic intent and facial expression visual intent of the speech text, and the semantic fusion features represent the pronunciation content of the specified object pronouncing the speech text according to the prosodic intent, and the facial expression visual semantics of performing facial expression actions according to the facial expression visual intent.

3. The method according to claim 2, wherein, The intent description information includes sub-description information corresponding to multiple text characters in the speech text, and the semantic fusion features include multiple sub-features arranged in sequence; The step of semantically fusing the speech text and audio / video prompt data using the large model based on the control command determined based on the received intent description information includes: Using the large model, based on the pronunciation prosody intent and facial expression visual intent represented by the sub-description information, semantic fusion is performed on the audio-visual prompt data and the multiple text characters in the speech text to obtain sub-features corresponding to the multiple text characters; Wherein, the sub-feature represents the pronunciation content of the specified object pronouncing the plurality of text words according to the pronunciation prosody intention corresponding to the sub-description information in the data segment of the target audio and video data, and the facial expression visual semantics of performing facial expression actions according to the facial expression visual intention corresponding to the sub-description information.

4. The method according to claim 2, wherein, The intent description information includes at least one of the following: Natural language text that represents expressive visual intentions and phonetic prosodic intentions based on the grammatical structure of natural language; Facial intent text and prosodic intent text associated with structured markup.

5. The method according to claim 2, wherein, The intent description information is determined based on the following operations: The large model is used to perform semantic understanding of the target object's demand information to obtain the facial expression visual intent and phonetic prosodic intent related to the specified object. The target object inputs the demand information by interacting with the specified object, and the facial expression visual intent and phonetic prosodic intent are used to respond to the emotions represented by the demand information. The large model is used to perform a text generation task based on the demand information, the facial expression visual intent, and the pronunciation prosody intent, to obtain the speech text and the intent description information that match the demand intent represented by the demand information.

6. The method according to claim 1 or 2, further comprising: Receive the requirement information input by the target object; By using a large model to understand the intent of the demand information, speech text that matches the demand intent of the target object is obtained.

7. The method according to claim 1, wherein, The step of generating target audio and video data based on the semantic fusion features includes: Based on the audio and video prompt data, feature extraction is performed on the semantic fusion features to obtain time-aligned target speech features and target visual features; The target speech data determined based on the target speech features and the target video data determined based on the target visual features are fused to obtain the target audio and video data.

8. The method according to claim 7, wherein, The step of extracting features from the semantic fusion features based on the audio and video prompt data includes: The preset noise data is denoised based on the audio and video prompt data and the semantic fusion features to obtain time-aligned target speech features and target visual features.

9. The method according to claim 7, wherein, The audio and video prompt data includes first prompt data and second prompt data. The data quality of the second prompt data is higher than that of the first prompt data. The semantic fusion feature is determined by semantic fusion of the first prompt data and the speech text. The step of extracting features from the semantic fusion features based on the audio and video prompt data includes: Based on the second prompt data, feature extraction is performed on the semantic fusion features.

10. A method for training large models, comprising: The system receives sample speech text, sample audio and video prompt data related to the sample object, and tag semantic fusion features. The tag semantic fusion features characterize the pronunciation prosody of the sample object in relation to the sample speech text in the sample audio and video prompt data, as well as the facial visual semantics aligned with the pronunciation prosody sequence. A large model is used to perform semantic fusion of the sample speech text and the sample audio and video prompt data to obtain sample semantic fusion features; The large model is trained based on the difference between the sample semantic fusion features and the label semantic fusion features to obtain the trained large model.

11. The method according to claim 10, wherein, The semantic fusion features of the tags are determined based on the sample audio and video prompt data processed by a trained encoder, and the trained encoder is determined based on the following operations: The reference audio and video prompt data is used to extract features from the reference audio and video prompt data to obtain pre-trained semantic fusion features, wherein the reference audio and video prompt data is related to the sample object; Pre-trained audio and video data are generated based on the pre-trained semantic fusion features; The encoder is trained based on the difference between the pre-trained audio / video data and the reference audio / video prompt data to obtain the trained encoder.

12. The method according to claim 11, wherein, The step of generating pre-trained audio and video data based on the pre-trained semantic fusion features includes: Based on the reference audio and video prompt data as prompt information, feature extraction is performed on the pre-trained semantic fusion features to obtain time-aligned pre-trained speech features and pre-trained visual features; and The pre-trained speech data determined based on the pre-trained speech features and the pre-trained video data determined based on the pre-trained visual features are fused to obtain the pre-trained audio-visual data.

13. The method according to claim 10, wherein, The method of using a large model to perform semantic fusion of the sample speech text and the sample audio-visual prompt data to obtain sample semantic fusion features includes: Based on the control instructions determined by the received sample intent description information, the large model is used to perform semantic fusion on the sample speech text and the sample audio-visual prompt data to obtain the sample semantic fusion features. The sample intent description information represents the pronunciation prosody of the sample object in the sample audio-visual prompt data in relation to the pronunciation content of the sample speech text, as well as the facial expression visual semantics corresponding to the pronunciation content.

14. An audio and video data generation device based on a large model, comprising: The first receiving module is used to receive voice text and audio-visual prompt data related to a specified object, wherein the audio-visual prompt data represents the pronunciation content of the specified object and the facial expression changes corresponding to the pronunciation content; The semantic fusion feature acquisition module is used to perform semantic fusion of the speech text and the audio-visual prompt data using a large model to obtain semantic fusion features. The semantic fusion features characterize the pronunciation prosody of the specified object in relation to the speech text, as well as the facial visual semantics aligned with the pronunciation prosody time sequence. The target audio and video data generation module is used to generate target audio and video data based on the semantic fusion features. The target audio and video data represents the target pronunciation content of the specified object in relation to the speech text, as well as the on-screen facial expressions of the specified object corresponding to the target pronunciation content.

15. An apparatus for training large models, comprising: The second receiving module is used to receive sample speech text, sample audio and video prompt data related to the sample object, and tag semantic fusion features. The tag semantic fusion features characterize the pronunciation prosody of the sample object in the sample audio and video prompt data in relation to the sample speech text, as well as the facial visual semantics aligned with the pronunciation prosody sequence. The sample semantic fusion feature acquisition module is used to perform semantic fusion of the sample speech text and the sample audio and video prompt data using a large model to obtain sample semantic fusion features. The training module is used to train the large model based on the difference between the sample semantic fusion features and the label semantic fusion features, so as to obtain the trained large model.

16. An intelligent agent of artificial intelligence, comprising: The input module is used to receive input information; The processing module is configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and execute the method of any one of claims 1 to 13 by calling the large model to obtain output information; An output module is used to output the output information obtained by the processing module.

17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 13.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 13.

19. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Video generation method and device, electronic equipment and storage medium

    CN120302120A

  • Digital human video generation method and device based on large model, intelligent agent, electronic equipment and storage medium

    CN120302122A