Digital human generation method and related device
By splitting the target text and distinguishing between static and dynamic text fragments, combining the pre-generated audio and video library and the real-time generation of audio and video packages, the problem of high computing overhead for digital human systems is solved, and efficient response and targeted needs are achieved.
Patent Information
- Application Number
- CN202311763610.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-06-20
AI Technical Summary
How to reduce the computing overhead of digital human systems to support instant interaction while ensuring response speed?
By splitting the target text, distinguishing between static and dynamic text clips, the static text clips match audio and video packages from the pre-generated audio and video library, while the dynamic text clip generates audio and video packages in real time, only generating audio and video packages when necessary to reduce computational overhead.
While reducing computing overhead, it meets the targeted needs of digital audio and video packages in different scenarios, ensuring response speed and playback continuity.
Smart Images

Figure CN120182437A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular, to a digital human generation method and related devices. Background Art
[0002] Digital humans, also known as virtual humans, virtual digital humans, virtual avatars, etc., refer to virtual characters that exist in the physical world and have a digital appearance created through technical means such as computer graphics, graphics rendering, deep learning, and speech synthesis. With the development of technologies such as computer vision, natural language processing, and speech recognition, digital humans are developing towards intelligence and diversification, and have the ability of emotional expression and communication. The applications of digital humans have evolved from the early pan-entertainment field to industries such as banking, healthcare, education, government affairs, and communication, serving as digital employees / virtual customer service / virtual lecturers of enterprises to provide services such as explanations, consultations, and trainings to customers. A digital human system generally includes a character image, a voice generation module, an animation generation module, an audio-video synthesis module, and an interaction module. According to different interaction modules, it can be divided into non-interactive digital humans, intelligent-driven digital humans, and real-person-driven digital humans. Different types of digital humans have different processing methods for voice and animation generation at the terminal.
[0003] Intelligent-driven digital humans are the most commonly used type of virtual digital humans, generally used in scenarios such as digital human employees, digital human customer service, and virtual lecturers. Users interact with digital humans in forms such as voice and text. The digital human uses artificial intelligence (AI) technologies such as speech recognition and semantic recognition in the intelligent system of the interaction module to identify the user's intention, and based on the intention, decides the subsequent output text and actions of the digital human, driving the voice generation module to generate the voice corresponding to the response text in real time. The animation generation module uses algorithms to render and generate matching character animations (including lip movements, expressions, body movements, etc.) in real time according to the response voice. The audio-video synthesis module generates the final video and pushes it to the user terminal device (such as a mobile phone / large screen / tablet, etc.) to enable interaction between the digital human and the user. Such interactions are generally real-time or quasi-real-time, and require the digital human to be able to quickly respond to the user's questions. However, in actual applications, to ensure immediate response to users, usually a large amount of computing power support is required. How to reduce the computing power as much as possible while ensuring the response speed is a technical problem that those skilled in the art are studying. Summary of the Invention
[0004] Embodiments of the present application disclose a digital human generation method and related devices, which can meet the targeted requirements for digital human audio-video packets in different scenarios while reducing the computing overhead as much as possible.
[0005] In a first aspect, embodiments of the present application provide a digital human generation method, which includes:
[0006] The target text is split into multiple text segments. Among them, the target text includes text in response to a user, the multiple text segments include at least one static text segment and at least one dynamic text segment, the dynamic text segment includes text content that changes with a preset factor, and the static text segment includes text content that does not change with the preset factor;
[0007] If the first text segment among the multiple text segments belongs to the dynamic text segment, a first audio-video package corresponding to the first text segment is generated, where the first audio-video package corresponding to the first text segment is used to play in the form of a digital human.
[0008] Using this method, based on the multiple text segments split from the target text, only dynamic text will specifically generate audio-video packages, which avoids the problem of large computational overhead caused by regenerating audio-video packages for all audio segments.
[0009] Combined with the first aspect, in a possible implementation manner of the first aspect, the method further includes:
[0010] If the first text segment among the multiple text segments belongs to the static text segment, a first audio-video package corresponding to the first text segment is matched from the audio-video library, where the audio-video library includes multiple pre-generated audio-video packages.
[0011] In this method, different methods are used to obtain the corresponding audio-video packages for the static text segments and dynamic text segments. For static text segments, their corresponding audio-video packages are directly searched from the audio-video library, while for dynamic text, audio-video packages are directly generated. This not only avoids the problem of large computational overhead caused by regenerating audio-video packages for all audio segments, but also avoids the problem that the audio-video packages cannot meet the different requirements of different scenarios when searching for audio-video packages from the audio-video library for all audio segments. That is, the above method can meet the targeted requirements of different scenarios for digital human audio-video packages while minimizing computational overhead.
[0012] Combined with the first aspect, or any of the above possible implementation manners of the first aspect, in a possible implementation manner of the first aspect, the method further includes: outputting the first audio-video package corresponding to the first text segment.
[0013] Combined with the first aspect, or any of the above possible implementation manners of the first aspect, in a possible implementation manner of the first aspect, the method further includes:
[0014] If the next second text segment of the first text segment in the time domain is a static text segment, match the second audio-video packet corresponding to the second text segment from the audio-video library, where the audio-video library includes a plurality of pre-generated audio-video packets;
[0015] Generate a first transition frame according to the last x frames of the video part in the first audio-video packet and the first y frames of the video part in the second audio-video packet, where the first transition frame is used for playing between the first audio-video packet and the second audio-video packet.
[0016] In this implementation manner, after obtaining the audio-video packet corresponding to the text segment, a smooth transition design is adopted between the audio-video packet corresponding to the current text segment and the audio-video packet corresponding to the next text segment, so that the playback continuity of the front and rear audio-video packets is better.
[0017] Combined with the first aspect, or any of the above possible implementation manners of the first aspect, in a possible implementation manner of the first aspect, the method further includes:
[0018] If the next second text segment of the first text segment in the time domain is a dynamic text segment, generate a first transition frame according to the last x frames of the video part in the first audio-video packet and the first y frames of the video part in the third audio-video packet, where the third audio-video packet includes the audio-video packet corresponding to the second text segment generated according to the second text segment, and the first transition frame is used for playing between the first audio-video packet and the third audio-video packet.
[0019] In this implementation manner, after obtaining the audio-video packet corresponding to the text segment, a smooth transition design is adopted between the audio-video packet corresponding to the current text segment and the audio-video packet corresponding to the next text segment, so that the playback continuity of the front and rear audio-video packets is better.
[0020] Combined with the first aspect, or any of the above possible implementation manners of the first aspect, in a possible implementation manner of the first aspect, it further includes:
[0021] Output the first transition frame.
[0022] Combined with the first aspect, or any of the above possible implementation manners of the first aspect, in a possible implementation manner of the first aspect, the method further includes:
[0023] If the next second text segment of the first text segment in the time domain is a dynamic text segment, generate a reference frame according to the action information of the limbs and / or face in the last z frames of the video part in the first audio-video packet, and the reference frame is used as the first frame when generating the video part in the third audio-video packet corresponding to the second text segment later.
[0024] In this implementation, after obtaining the audio-visual packets corresponding to the text segments, a smooth transition design is adopted between the audio-visual packets corresponding to the current text segment and the audio-visual packets corresponding to the next text segment, so that the playback continuity of the front and back audio-visual packets is better.
[0025] Combined with the first aspect, or any of the above possible implementation manners of the first aspect, in a possible implementation manner of the first aspect, the first text segment is any text segment except the last text segment among the multiple text segments.
[0026] Combined with the first aspect, or any of the above possible implementation manners of the first aspect, in a possible implementation manner of the first aspect, generating the first audio-visual packet corresponding to the first text segment includes:
[0027] Synthesizing an audio packet corresponding to the first text segment through a speech synthesis module;
[0028] Generating a video packet according to the first text segment and / or the audio packet;
[0029] Performing timestamp alignment on the audio packet and the video packet to obtain the first audio-visual packet.
[0030] Combined with the first aspect, or any of the above possible implementation manners of the first aspect, in a possible implementation manner of the first aspect, the video packet includes one or more of the lip movement parameters, expression parameters, and action parameters of the digital human.
[0031] Combined with the first aspect, or any of the above possible implementation manners of the first aspect, in a possible implementation manner of the first aspect, the preset factors include one or more of time, application scenario, user type, user object, business environment, business attribute, etc.
[0032] Combined with the first aspect, or any of the above possible implementation manners of the first aspect, in a possible implementation manner of the first aspect, it further includes: identifying the user intention and generating a target text in response to the intention.
[0033] In a second aspect, an embodiment of the present application provides a digital human generation device, and the device includes:
[0034] A splitting unit, configured to split a target text to obtain multiple text segments, where the target text includes text in response to a user, the multiple text segments include at least one static text segment and at least one dynamic text segment, the dynamic text segment includes text content that changes with a preset factor, and the static text segment includes text content that does not change with a preset factor;
[0035] A first generation unit, configured to generate a first audio-video packet corresponding to the first text segment when the first text segment among the multiple text segments belongs to the dynamic text segment, wherein the first audio-video packet corresponding to the first text segment is used to be played in the form of a digital human.
[0036] By adopting this method, among the multiple text segments split from the target text, only the dynamic text will specifically generate audio-video packets, which avoids the problem of large computational overhead caused by regenerating audio-video packets for all audio segments.
[0037] Combined with the second aspect, in a possible implementation manner of the second aspect, the device further includes:
[0038] A matching unit, configured to match a first audio-video packet corresponding to the first text segment from an audio-video library when the first text segment among the multiple text segments belongs to the static text segment, wherein the audio-video library includes a plurality of pre-generated audio-video packets.
[0039] In this method, different methods are used to obtain the corresponding audio-video packets for the static text segments and dynamic text segments therein. For the static text segments, the corresponding audio-video packets are directly searched from the audio-video library, while for the dynamic text, the audio-video packets are directly generated. This not only avoids the problem of large computational overhead caused by regenerating audio-video packets for all audio segments, but also avoids the problem that the audio-video packets cannot meet the different requirements of different scenarios due to searching for audio-video packets from the audio-video library for all audio segments. That is, the above method can meet the targeted requirements of different scenarios for digital human audio-video packets while reducing the computational overhead as much as possible.
[0040] Combined with the second aspect, or any of the above possible implementation manners of the second aspect, in a possible implementation manner of the second aspect, the device further includes:
[0041] An output unit, configured to output the first audio-video packet corresponding to the first text segment.
[0042] Combined with the second aspect, or any of the above possible implementation manners of the second aspect, in a possible implementation manner of the second aspect, the device further includes a second generation unit:
[0043] The matching unit is further configured to match a second audio-video packet corresponding to the second text segment from the audio-video library when the next second text segment of the first text segment in the time domain is a static text segment, wherein the audio-video library includes a plurality of pre-generated audio-video packets;
[0044] The second generation unit is configured to generate a first transition frame according to the last x frames of the video part in the first audio-visual packet and the first y frames of the video part in the second audio-visual packet, where the first transition frame is used for playing between the first audio-visual packet and the second audio-visual packet.
[0045] In this implementation, after obtaining the audio-visual packets corresponding to the text segments, a smooth transition design is adopted between the audio-visual packet corresponding to the current text segment and the audio-visual packet corresponding to the next text segment, so that the playback continuity of the front and back audio-visual packets is better.
[0046] Combined with the second aspect, or any of the above possible implementation manners of the second aspect, in a possible implementation manner of the second aspect, the apparatus further includes:
[0047] A third generation unit, configured to generate a first transition frame according to the last x frames of the video part in the first audio-visual packet and the first y frames of the video part in the third audio-visual packet when the second text segment in the time domain of the first text segment is a dynamic text segment, where the third audio-visual packet includes the audio-visual packet corresponding to the second text segment generated according to the second text segment, and the first transition frame is used for playing between the first audio-visual packet and the third audio-visual packet.
[0048] In this implementation, after obtaining the audio-visual packets corresponding to the text segments, a smooth transition design is adopted between the audio-visual packet corresponding to the current text segment and the audio-visual packet corresponding to the next text segment, so that the playback continuity of the front and back audio-visual packets is better.
[0049] Combined with the second aspect, or any of the above possible implementation manners of the second aspect, in a possible implementation manner of the second aspect, the output unit is further configured to: output the first transition frame.
[0050] Combined with the second aspect, or any of the above possible implementation manners of the second aspect, in a possible implementation manner of the second aspect, the apparatus further includes:
[0051] A fourth generation unit, configured to generate a reference frame according to the action information of the limbs and / or face in the last z frames of the video part in the first audio-visual packet when the second text segment in the time domain of the first text segment is a dynamic text segment, and the reference frame is used as the first frame when generating the video part in the third audio-visual packet corresponding to the second text segment later.
[0052] In this implementation, after obtaining the audio-visual packets corresponding to the text segments, a smooth transition design is adopted between the audio-visual packet corresponding to the current text segment and the audio-visual packet corresponding to the next text segment, so that the playback continuity of the front and back audio-visual packets is better.
[0053] In combination with the second aspect, or any of the above possible implementation manners of the second aspect, in a possible implementation manner of the second aspect, the first text segment is any text segment among the multiple text segments except the last text segment.
[0054] In combination with the second aspect, or any of the above possible implementation manners of the second aspect, in a possible implementation manner of the second aspect, in terms of generating the first audio-visual packet corresponding to the first text segment, the first generating unit is specifically configured to:
[0055] Synthesize an audio packet corresponding to the first text segment through a speech synthesis module;
[0056] Generate a video packet according to the first text segment and / or the audio packet;
[0057] Perform timestamp alignment on the audio packet and the video packet to obtain the first audio-visual packet.
[0058] In combination with the second aspect, or any of the above possible implementation manners of the second aspect, in a possible implementation manner of the second aspect, the video packet includes one or more of lip movement parameters, expression parameters, and action parameters of the digital human.
[0059] In combination with the second aspect, or any of the above possible implementation manners of the second aspect, in a possible implementation manner of the second aspect, the preset factors include one or more of time, application scenario, user type, user object, business environment, business attribute, etc.
[0060] In combination with the second aspect, or any of the above possible implementation manners of the second aspect, in a possible implementation manner of the second aspect, it further includes: an identification unit, configured to identify a user intention and generate a target text in response to the intention.
[0061] In a third aspect, an embodiment of the present application provides a digital human generation device, which includes a processor and a memory. Among them, the memory is used to store a computer program, and the processor is used to call the computer program to implement the method described in the first aspect or any of the possible implementation manners of the first aspect.
[0062] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is called by a processor, the method described in the first aspect or any of the possible implementation manners of the first aspect is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] The following introduces the drawings used in the embodiments of the present application.
[0064] Figure 1 It is a schematic diagram of the architecture of a digital human video processing system provided by an embodiment of the present application;
[0065] Figure 2 It is a schematic flowchart of a digital human generation method provided by an embodiment of the present application;
[0066] Figure 3 It is a schematic flowchart of a digital human generation method provided by an embodiment of the present application;
[0067] Figure 4 It is a schematic flowchart of a digital human generation method provided by an embodiment of the present application;
[0068] Figure 5 It is a schematic diagram of the structure of a digital human generation device provided by an embodiment of the present application;
[0069] Figure 6 It is a schematic diagram of the structure of a digital human generation device provided by an embodiment of the present application. Detailed implementation manners
[0070] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.
[0071] Please refer to Figure 1 , Figure 1 It is a schematic diagram of the architecture of a digital human video processing system provided by an embodiment of the present application. The digital human video processing system includes a service system 101, a digital human system 102, an audio and video playback system 103, and a terminal device 104, where:
[0072] Service system 101: Generates the text content to be broadcast by the digital human according to service requirements. For example, it obtains the intention information of the user and generates the text content in response to the intention information, that is, the target text. This service system can be deployed in the cloud or locally. When deployed in the cloud, the intention information of the user can be uploaded by the terminal device to the service system. When deployed locally, the service system can be the terminal device itself or other devices independent of the terminal device, which are equipped with sensors and can collect the intention information of the user.
[0073] Digital human system 102: Generates the audio and video file or audio and video stream of the digital human according to the target text. Specifically, it includes operations such as text splitting, pre-making the audio and video corresponding to the static text fragments to obtain an audio and video library and instantaneously searching for the audio and video package corresponding to a specific static text fragment from the audio and video library, instantaneously making the audio and video package corresponding to the dynamic text fragment, and audio and video file output.
[0074] The digital human system can be a software system or a hardware system with software deployed thereon. The hardware system can be a single device or a cluster composed of multiple devices. The digital human system can be deployed in the cloud or locally, which is not limited in this application.
[0075] Audio and video playback system 103: It is docked with the digital human system. The digital human system outputs audio and video files to the audio and video file system, and then this audio and video system interacts with the terminal device (or user equipment (UE)) to stream or play the digital human audio and video files on the terminal device. Optionally, there may be no audio and video playback system 103, and the digital human system 102 directly outputs the audio and video files to the terminal device.
[0076] Terminal device 104: When the user uses the business system, they watch the digital human audio and video files through the terminal device. The terminal device can be an accessory device or an associated device of the business system, such as a mobile phone, a PAD, a computer, etc., or the terminal device may be integrated with the business system. There can be many services provided by the business system, such as a banking business system, a telecommunications business system, etc., which are not limited here.
[0077] Please refer to Figure 2 , Figure 2 FIG. is a schematic flowchart of a digital human generation method provided by an embodiment of this application. This method can be implemented based on architectures including but not limited to those shown below, or can be implemented based on other architectures. This method includes but not limited to the following steps:
[0078] Step S201: The business system identifies the user's intention and generates a target text in response to the intention.
[0079] Specifically, the user's intention can be expressed in various ways. It can be expressed by voice, such as when the user says "Query phone bill", or when the user says "Help query the traffic conditions of the route home"; it can also be expressed by physical actions, such as when the user uses sign language, nods, shakes their head, blinks, etc.; it can also be expressed by the user's behavioral actions, such as when the user steps on a weighing scale or a height measuring ruler.
[0080] Correspondingly, information can be collected through sensors to obtain the user's intention. For example, the physical actions and behavioral actions of the user can be collected through an image sensor (such as a camera), and for another example, the user's voice can be collected through a microphone, and for another example, the user's selection operations and input operations (such as text) can be collected through a touch screen, etc. By analyzing the information collected by the sensors according to predefined rules, the user's intention can be determined.
[0081] After that, generate a target text that responds to the user's intention. For example, if it is determined that the user's intention is to "query the phone bill balance" based on the user's input of "query phone bill", then the corresponding target text can be "Hello, your cumulative consumption this month is 69.55 yuan, and your phone bill balance is 100.52 yuan. The specific information will be sent to your mobile phone later. Please check it."; Another example, if it is determined that the user's intention is to "agree to specifically understand the phone bill package A" through the user's body movement of nodding, then the target text can be the introduction text about package A; Another example, if the recognized user intention is to weigh oneself, then the response text can be information about the user's weight (such as "Your weight is 140KG"). It can be understood that the scenarios and text contents mentioned here by way of example are only for illustration, and actually there can be other scenarios and contents, which will not be elaborated one by one here.
[0082] There is no limitation here on how to specifically generate the response text. Multiple mapping relationships between user intentions and multiple texts can be established in advance. When the user intention is known, the text corresponding to the mapping relationship is used as the target text; A prediction model can also be constructed in advance. Inputting the user intention into the prediction model can obtain the corresponding target text; Corresponding algorithms or rules can also be configured in advance. Taking the user intention as the input and substituting it into the corresponding algorithms and rules can obtain the corresponding target text. Of course, there are other ways, which will not be elaborated one by one here.
[0083] Optionally, step S201 can also be completed by the digital human system.
[0084] Step S202: The digital human system splits the target text to obtain multiple text fragments.
[0085] Specifically, splitting rules can be configured in advance. For example, splitting can be performed according to punctuation marks, dynamic content identifiers (pre-set), etc. Another example is that a splitting algorithm can be configured in advance. Inputting the target text into the splitting algorithm can output multiple split text fragments. The splitting algorithm can be directly configured or obtained by training a large number of training samples into the corresponding model with the help of AI algorithms. Optionally, the splitting algorithm can use high-frequency common text fragments that conform to natural language rules as static content.
[0086] In the embodiments of the present application, after obtaining multiple text fragments, each text fragment is classified. The text fragments can include static text fragments and dynamic text fragments. Of course, other types of text fragments can also be included, which are specifically set according to needs. The dynamic text fragments include text contents that change with preset factors, and the static text fragments include text contents that do not change with preset factors. For example, the preset factors include one or more of factors such as time, application scenarios, user types, user objects, etc.
[0087] Optionally, the text type recognition model can be pre-configured. The text category recognition model can be trained with a large number of sample data, and each sample data includes a text segment and the type identifier of the text segment. Therefore, the finally trained text type recognition model can better recognize or predict the text type of each text segment. In the embodiments of the present application, during classification, the corresponding first label (such as 0) can be assigned to the static text segment, and the corresponding second label (such as 1) can be assigned to the dynamic text segment. Special symbols can also be added before and after the text segment. For example, special symbol @ can be added before and after the dynamic text segment to facilitate subsequent recognition and distinction.
[0088] Generally, the recognition result is that the multiple text segments include at least one static text segment and at least one dynamic text segment.
[0089] For ease of understanding, the following is an example. For example, the target text is "Hello, your cumulative consumption this month is 69.55 yuan, and the phone bill balance is 100.52 yuan. The specific information will be sent to your mobile phone later. Please check it.", and the text segments and text types obtained after splitting are shown in Table 1:
[0090] Table 1
[0091] Serial number Text segment Text type 1 Hello, your cumulative consumption this month is Static text segment 2 69.55 yuan Dynamic text segment 3 The phone bill balance is Static text segment 4 100.52 yuan Dynamic text segment 5 The specific information will be sent to your mobile phone later. Please check it Static text segment
[0092] As can be seen from Table 1, there may be static content before and after dynamic content, and there may also be dynamic content before and after static content.
[0093] In the embodiments of the present application, to determine the audio-video package corresponding to each text segment in the multiple text segments, that is, the content played in the form of a digital human. Different determination methods can be adopted for static text segments and dynamic text segments. For ease of understanding, the following takes the first text segment in the multiple text segments as an example for illustration. Among the multiple text segments, some text segments can be determined in the same way as the first text segment, or all text segments can be determined in the same way as the first text segment.
[0094] Step S203: If the first text segment in the multiple text segments belongs to the dynamic text segment, the digital human system generates the first audio-video package corresponding to the first text segment.
[0095] In an optional implementation solution, generating the first audio-video package may include:
[0096] Synthesize an audio packet corresponding to the first text segment through a speech synthesis module. The audio packet may include the audio of reading or explaining the first text segment, and of course also includes the timestamp information corresponding to the audio. The first text segment may be input into a corresponding algorithm or model, such as a first model, to output the audio packet, and the first model may be pre-trained.
[0097] Then, generate a video packet according to the first text segment and / or the audio packet. Here, a video for broadcasting the first text segment needs to be generated. For example, the video includes information such as the lip movement parameters, expression parameters, hand and body movements, digital human image, and digital human video background during the broadcast. Of course, it also includes the timestamp corresponding to the video. The digital human image can be pre-configured according to needs. For example, there are differences between male and female digital humans, or differences between old, young, and child digital humans, etc. Since the audio packet includes the audio for the first text segment, the video packet can be generated based on the first text segment or the audio packet. Since information such as the tone and pitch in the audio packet may affect the expression and body movements of the digital human, the video packet can also be generated according to the first text segment and the audio packet (such as the tone and pitch therein). The process of generating a video packet based on the audio packet is the process of obtaining a video by driving the lip movement, expression, etc. of the digital human through the voice. Optionally, the audio packet may include an audio stream, and the video packet may include a video stream.
[0098] After that, perform timestamp alignment on the audio packet and the video packet to obtain a first audio-video packet. Here, timestamp alignment can synchronize (or align) the voice in the audio packet with the lip movement, expression, body movements, etc. of the digital human in the video packet. Therefore, the finally obtained first audio-video packet is a video animation of the digital human broadcasting the above first text segment.
[0099] Optionally, if the first text segment is a dynamic text segment and it is not the starting text segment in the time dimension, then the first frame of the video part in the first audio-video can be generated according to the z-th frame closer to the end of the video part in the audio-video corresponding to the previous text segment of the first text segment. For example, refer to the lip movement, expression, body movements, etc. in the z-th frame to generate the first frame of the video part in the first audio-video (the first frame can be generated with reference to the generation principle in S209). This can ensure the continuity between the first audio-video corresponding to the first text segment and the audio-video corresponding to its previous text segment.
[0100] Optionally, if the first text segment is a dynamic text segment and it is the starting text segment in the time dimension, then the first frame of the video part in the first audio-video can be default, such as the lip movement, expression, body movements, etc. are default configured values.
[0101] Step S204: If the first text segment among the multiple text segments belongs to the static text segment, the digital human system matches the first audio-video package corresponding to the first text segment from the audio-video library.
[0102] Since static text usually does not change frequently, its appearance frequency is relatively high and the probability is relatively large. Therefore, the corresponding audio-video package can be pre-generated for static text, and an audio-video library can be constructed based on a large number of static texts and their corresponding audio-video packages for calling.
[0103] In the embodiment of the present application, when the first text segment is static text, there is no need to generate its corresponding audio-video package immediately. Instead, the audio-video package corresponding to the first text segment is directly searched from the existing audio-video library, and this process will be relatively fast.
[0104] When searching, in addition to using the first text segment, the required digital human image, the required digital human video background, etc. may also be used. For example, if the first text segment corresponds to two audio-video packages, one of which is broadcast by a male digital human and the other is broadcast by a female digital human, then when searching through the first text segment, the required digital human image needs to be combined to find the corresponding specific audio-video package. The role and usage method of the digital human video background are similar to those of the digital human image, and will not be elaborated here. Optionally, if the first text segment is the starting text segment, then when searching, the digital human image, digital human video background, etc. corresponding to the audio-video package can be selected according to certain rules, such as configuring the default; if the first text segment is not the starting text segment, then when searching, the digital human image, digital human video background, etc. in the audio-video package corresponding to the previous text segment can be used as one of the search keywords to ensure that the digital human image, digital human video background, etc. in the audio-video package corresponding to the first text segment searched are consistent with those in the audio-video package corresponding to the previous text segment.
[0105] In the embodiment of the present application, the audio-video library can exist in a database, or in memory, or in a cache, or in a corresponding file, and is not limited here.
[0106] Step S205: The digital human system outputs the first audio-video package corresponding to the first text segment.
[0107] In the embodiments of the present application, the first audio-video packet corresponding to the finally determined first text segment is used for playing in the form of a digital human. The output here includes the direct playing method, such as directly played by the above digital human system. The output here can also refer to sending to a terminal device for playing by the terminal device. This includes the subsequently generated first transition frame, which can be played by the digital human system or sent by the digital human system to the terminal device for playing by the terminal device.
[0108] In addition, the playing in the embodiments of the present application can be performed in the form of a data stream for immediate playing, or the audio-video packet corresponding to a text segment can be played after it is generated, or all the audio-video packets corresponding to all text segments can be generated and then played as a whole. The specific playing method is not limited here. Optionally, when playing in the form of a video stream, playing will be performed simultaneously during the acquisition, generation, or transmission of the audio-video packet. Of course, it may be necessary to wait briefly for the playing of some frames. For example, when the first transition frame mentioned later needs to be generated, although the first video frame in the audio-video corresponding to the dynamic text segment has been generated, it is necessary to wait for the first transition frame to be played before playing the first video frame. Of course, the waiting process generally will not be too long and can basically achieve immediate playing.
[0109] Optionally, in order to achieve smooth transition during the playing of the audio-video corresponding to different text segments, the following steps S206 - S207 (optional solution one, as shown in Figure 3 , Figure 4 ), or including step S208 (optional solution two, as shown in Figure 3 ), or including step S209 (optional solution three, as shown in Figure 4 ) can be included. The following will be introduced separately.
[0110] Step S206: If the next second text segment of the first text segment in the time domain is a static text segment, the digital human system matches the second audio-video packet corresponding to the second text segment from the audio-video library.
[0111] It can be understood that each text segment also corresponds to a timestamp. Therefore, the previous text segment and the next text segment of each text segment can be determined in the time domain, and the text type of each text segment was also marked when splitting into multiple text segments before. Therefore, it can be easily determined whether the next second text segment of the first text segment in the time domain is a static text segment or a dynamic text segment.
[0112] The method of matching the second audio-video packet corresponding to the second text segment from the audio-video library is the same as the method of matching the first audio-video packet corresponding to the first text segment before, and will not be elaborated here.
[0113] Step S207: The digital human system generates a first transition frame based on the last x frames of the video part in the first audio-video packet and the first y frames of the video part in the second audio-video packet.
[0114] Specifically, the last x frames and the first y frames can generate the first transition frame between the last x frames and the first y frames by means of frame interpolation. The first transition frame, including other subsequent transition frames, can be one frame or multiple frames. Generally speaking, if the content in the x frames and the y frames varies greatly, multiple first transition frames can be generated, and the subsequent playback effect will be better. Optionally, in addition to frame interpolation, other methods can also be used. For example, based on a preset generation process, using the last x frames and the first y frames as inputs, the first transition frame is obtained through this generation process; or, the last x frames and the first y frames are input into a trained transition frame prediction model to output the first transition frame. Here, the sizes of x and y can be preset as needed. Of course, there are other methods, which will not be exemplified one by one here.
[0115] In this case, the first transition frame is used for playback between the first audio-video packet and the second audio-video packet.
[0116] As mentioned before, the playback of the first transition frame can be performed by the above digital human system, or the digital human system can send it to the terminal device for the terminal device to play. The playback method can be in the form of a video stream or a non-video stream.
[0117] Step S208: If the next second text segment of the first text segment in the time domain is a dynamic text segment, the digital human system generates a first transition frame based on the last x frames of the video part in the first audio-video packet and the first y frames of the video part in the third audio-video packet.
[0118] Specifically, the method of generating the first transition frame can refer to the relevant description in step S207 above and will not be elaborated here.
[0119] Among them, the third audio-video packet includes the audio-video packet corresponding to the second text segment generated according to the second text segment. The method of generating the third audio-video packet according to the second text segment (when it is a dynamic text segment) is the same as the method of generating the first audio-video packet according to the first text segment (when it is a dynamic text segment), which will not be elaborated here.
[0120] In this case, the first transition frame is used for playback between the first audio-video packet and the second audio-video packet.
[0121] As mentioned above, the playback of the first transition frame can be performed by the above digital human system, or the digital human system can send it to the terminal device for playback by the terminal device. The playback method can be in the form of a video stream or a non-video stream.
[0122] Step S209: If the next second text segment of the first text segment in the time domain is a dynamic text segment, the digital human system generates a reference frame according to the action information of the limbs and / or face in the tail z-frame of the video part in the first audio-visual packet.
[0123] That is to say, when generating the reference frame, in addition to using the action information of the limbs (such as gestures) and / or face (such as mouth shapes, expressions) in the tail z-frame, other information may also be used, such as the digital human image, the digital human video background, and other information.
[0124] It should be noted that the reference frame is used as the first frame when generating the video part of the third audio-visual packet corresponding to the second text segment later. Therefore, the reference frame in step S209 can be generated when generating the third audio-visual packet according to the second text segment later, or can be generated before generating the third audio-visual packet according to the second text segment, which is equivalent to a preprocessing operation for generating the third audio-visual packet.
[0125] It should be noted that steps S206 - S209 are supplements to steps S201 - S205. Steps S206 - S209 mainly describe how to achieve smooth transition between the audio-visual packets corresponding to the first text segment and its subsequent text segment (i.e., the second text segment), so as to ensure the continuity and smoothness of the digital human playback. Therefore, it can be considered as an associated operation for generating the audio-visual packet of the first text segment. In the embodiments of the present application, the processing process of each of the above multiple text segments can be the same as that of the first text segment, including the steps of S201 - S205. Except for the last text segment in the time domain among the multiple text segments, each text segment can be the same as the first text segment, including the steps of S206 - S209. Of course, it is also possible not to consider smooth transition and directly play the audio-visual packet corresponding to the next text segment after the audio-visual packet corresponding to the previous text segment is played; of course, it is also possible to consider smooth transition but adopt a method other than S206 - S209.
[0126] In addition, in one implementation, the operations of generating corresponding audio-visual packets for different text segments among the above-mentioned multiple text segments can be performed synchronously; in another implementation, the operations of generating corresponding audio-visual packets for each text segment among the multiple text segments are sequentially performed in the order of time domain until the operations of generating corresponding audio-visual packets are completed for each text segment. In the embodiments of the present application, the specific method to be adopted can be selected and configured according to the characteristics of the specific application scenario, and no limitation is imposed herein.
[0127] In Figure 2 In the method shown, based on multiple text segments split from the target text, different methods are used to obtain corresponding audio-visual packets for static text segments and dynamic text segments among them. For static text segments, their corresponding audio-visual packets are directly searched from the audio-visual library, while for dynamic text, audio-visual packets are directly generated. This not only avoids the problem of large computational overhead caused by regenerating audio-visual packets for all audio segments, but also avoids the problem that the audio-visual packets cannot meet the different requirements of different scenarios when searching for audio-visual packets from the audio-visual library for all audio segments. That is, the above method can meet the targeted requirements of different scenarios for digital human audio-visual packets while minimizing the computational overhead as much as possible. Further, after obtaining the audio-visual packet corresponding to the current text segment, a smooth transition design is adopted between the audio-visual packet corresponding to the current text segment and the audio-visual packet corresponding to the next text segment, so that the playback continuity of the front and rear audio-visual packets is better.
[0128] The method of the embodiments of the present application is elaborated in detail above, and the device of the embodiments of the present application is provided below.
[0129] Please refer to Figure 5 , Figure 5 FIG. is a schematic structural diagram of a digital human generation device provided by an embodiment of the present application. The device can be the above-mentioned digital human system or a device or module in the digital human system. The device 50 may include a splitting unit 501 and a first generating unit 502. The detailed descriptions of each unit are as follows.
[0130] The splitting unit 501 is configured to split the target text to obtain multiple text segments, where the target text includes text in response to a user, the multiple text segments include at least one static text segment and at least one dynamic text segment, the dynamic text segment includes text content that changes with a preset factor, and the static text segment includes text content that does not change with a preset factor;
[0131] A first generation unit 502, configured to generate a first audio-video packet corresponding to the first text segment when the first text segment among the multiple text segments belongs to the dynamic text segment, where the first audio-video packet corresponding to the first text segment is used to be played in the form of a digital human.
[0132] By adopting this method, among the multiple text segments split from the target text, only the dynamic text will specifically generate audio-video packets, which avoids the problem of large computational overhead caused by regenerating audio-video packets for all audio segments.
[0133] In a possible implementation manner, the device further includes:
[0134] A matching unit, configured to match a first audio-video packet corresponding to the first text segment from an audio-video library when the first text segment among the multiple text segments belongs to the static text segment, where the audio-video library includes a plurality of pre-generated audio-video packets.
[0135] In this method, different methods are used to obtain corresponding audio-video packets for the static text segments and dynamic text segments therein. For the static text segments, their corresponding audio-video packets are directly searched from the audio-video library, while for the dynamic text, audio-video packets are directly generated. This not only avoids the problem of large computational overhead caused by regenerating audio-video packets for all audio segments, but also avoids the problem that the audio-video packets cannot meet the different requirements of different scenarios due to searching for audio-video packets from the audio-video library for all audio segments. That is, the above method can meet the targeted requirements of different scenarios for digital human audio-video packets while reducing the computational overhead as much as possible.
[0136] In a possible implementation manner, the device further includes:
[0137] An output unit, configured to output the first audio-video packet corresponding to the first text segment.
[0138] In a possible implementation manner, the device further includes a second generation unit:
[0139] The matching unit is further configured to match a second audio-video packet corresponding to the second text segment from the audio-video library when the next second text segment of the first text segment in the time domain is a static text segment, where the audio-video library includes a plurality of pre-generated audio-video packets;
[0140] The second generation unit is configured to generate a first transition frame according to the last x frames of the video part in the first audio-video packet and the first y frames of the video part in the second audio-video packet, where the first transition frame is used to be played between the first audio-video packet and the second audio-video packet.
[0141] In this implementation manner, after obtaining the audio - video packet corresponding to the text segment, a smooth transition design is adopted between the audio - video packet corresponding to the current text segment and the audio - video packet corresponding to the next text segment, so that the playback continuity of the front - and - back audio - video packets is better.
[0142] In a possible implementation manner, the device further includes:
[0143] A third generation unit, configured to, when the next second text segment of the first text segment in the time domain is a dynamic text segment, generate a first transition frame according to the last x frames of the video part in the first audio - video packet and the first y frames of the video part in the third audio - video packet, where the third audio - video packet includes the audio - video packet corresponding to the second text segment generated according to the second text segment, and the first transition frame is used for playback between the first audio - video packet and the third audio - video packet.
[0144] In this implementation manner, after obtaining the audio - video packet corresponding to the text segment, a smooth transition design is adopted between the audio - video packet corresponding to the current text segment and the audio - video packet corresponding to the next text segment, so that the playback continuity of the front - and - back audio - video packets is better.
[0145] In a possible implementation manner, the output unit is further configured to: output the first transition frame.
[0146] In a possible implementation manner, the device further includes:
[0147] A fourth generation unit, configured to, when the next second text segment of the first text segment in the time domain is a dynamic text segment, generate a reference frame according to the action information of the limbs and / or face in the last z frames of the video part in the first audio - video packet, and the reference frame is used as the first frame when generating the video part in the third audio - video packet corresponding to the second text segment later.
[0148] In this implementation manner, after obtaining the audio - video packet corresponding to the text segment, a smooth transition design is adopted between the audio - video packet corresponding to the current text segment and the audio - video packet corresponding to the next text segment, so that the playback continuity of the front - and - back audio - video packets is better.
[0149] In a possible implementation manner, the first text segment is any text segment except the last text segment among the multiple text segments.
[0150] In a possible implementation manner, in terms of generating the first audio - video packet corresponding to the first text segment, the first generation unit is specifically configured to:
[0151] Synthesize an audio packet corresponding to the first text segment through a speech synthesis module;
[0152] Generate a video packet according to the first text fragment and / or the audio packet;
[0153] Perform timestamp alignment on the audio packet and the video packet to obtain a first audio-visual packet.
[0154] In a possible implementation, the video packet includes one or more of the lip movement parameters, expression parameters, and action parameters of the digital human.
[0155] In a possible implementation, the preset factors include one or more of time, application scenario, user type, and user object.
[0156] In a possible implementation, it further includes: an identification unit for identifying the user intention and generating a target text in response to the intention.
[0157] It should be noted that the implementation of each unit can also correspond to the corresponding description of the method embodiment with reference to Figure 2 、 Figure 3 or Figure 4 as shown.
[0158] Please refer to Figure 6 , Figure 6 Figure 60 shows a digital human generation device 60 provided by an embodiment of the present application. The device 60 includes a processor 601, a memory 602, and a communication interface 603. The processor 601, the memory 602, and the communication interface 603 are interconnected through a bus.
[0159] The memory 602 includes but is not limited to a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), or a compact disc read-only memory (CD-ROM). The memory 602 is used for relevant computer programs and data. The communication interface 603 is used for receiving and sending data.
[0160] The processor 601 may be one or more central processing units (CPUs). When the processor 601 is a single CPU, the CPU may be a single-core CPU or a multi-core CPU.
[0161] The processor 601 in the device 60 is used to read the computer program code stored in the memory 602 and perform the following operations:
[0162] The target text is split into multiple text segments. Among them, the target text includes text in response to a user, the multiple text segments include at least one static text segment and at least one dynamic text segment, the dynamic text segment includes text content that changes with a preset factor, and the static text segment includes text content that does not change with the preset factor;
[0163] If the first text segment among the multiple text segments belongs to the dynamic text segment, a first audio-video package corresponding to the first text segment is generated, where the first audio-video package corresponding to the first text segment is used to be played in the form of a digital human.
[0164] By adopting this method, based on the multiple text segments split from the target text, only dynamic text will specifically generate audio-video packages, which avoids the problem of large computational overhead caused by regenerating audio-video packages for all audio segments.
[0165] In a possible implementation, the processor is further configured to:
[0166] If the first text segment among the multiple text segments belongs to the static text segment, a first audio-video package corresponding to the first text segment is matched from an audio-video library, where the audio-video library includes multiple pre-generated audio-video packages.
[0167] In this method, different methods are used to obtain corresponding audio-video packages for the static text segments and dynamic text segments therein. For static text segments, their corresponding audio-video packages are directly searched from the audio-video library, while for dynamic text, audio-video packages are directly generated. This not only avoids the problem of large computational overhead caused by regenerating audio-video packages for all audio segments, but also avoids the problem that the audio-video packages cannot meet the different requirements of different scenarios when searching for audio-video packages from the audio-video library for all audio segments. That is, the above method can meet the targeted requirements of digital human audio-video packages in different scenarios while minimizing computational overhead as much as possible.
[0168] In a possible implementation, the processor is further configured to output the first audio-video package corresponding to the first text segment by calling the communication interface 603.
[0169] In a possible implementation, the processor is further configured to:
[0170] If the next second text segment of the first text segment in the time domain is a static text segment, a second audio-video package corresponding to the second text segment is matched from the audio-video library, where the audio-video library includes multiple pre-generated audio-video packages;
[0171] Generate a first transition frame based on the last x frames of the video part in the first audio-video packet and the first y frames of the video part in the second audio-video packet, where the first transition frame is used for playing between the first audio-video packet and the second audio-video packet.
[0172] In this implementation, after obtaining the audio-video packet corresponding to the text segment, a smooth transition design is adopted between the audio-video packet corresponding to the current text segment and the audio-video packet corresponding to the next text segment, making the playback continuity of the front and back audio-video packets better.
[0173] In a possible implementation, the processor is further configured to:
[0174] If the next second text segment of the first text segment in the time domain is a dynamic text segment, generate a first transition frame based on the last x frames of the video part in the first audio-video packet and the first y frames of the video part in the third audio-video packet, where the third audio-video packet includes the audio-video packet corresponding to the second text segment generated according to the second text segment, and the first transition frame is used for playing between the first audio-video packet and the third audio-video packet.
[0175] In this implementation, after obtaining the audio-video packet corresponding to the text segment, a smooth transition design is adopted between the audio-video packet corresponding to the current text segment and the audio-video packet corresponding to the next text segment, making the playback continuity of the front and back audio-video packets better.
[0176] In a possible implementation, the processor is further configured to call the communication interface 603 to output the first transition frame.
[0177] In a possible implementation, the processor is further configured to:
[0178] If the next second text segment of the first text segment in the time domain is a dynamic text segment, generate a reference frame based on the action information of the limbs and / or face in the last z frames of the video part in the first audio-video packet, and the reference frame is used as the first frame when generating the video part in the third audio-video packet corresponding to the second text segment later.
[0179] In this implementation, after obtaining the audio-video packet corresponding to the text segment, a smooth transition design is adopted between the audio-video packet corresponding to the current text segment and the audio-video packet corresponding to the next text segment, making the playback continuity of the front and back audio-video packets better.
[0180] In a possible implementation, the first text segment is any text segment except the last text segment among the multiple text segments.
[0181] In a possible implementation, in terms of generating a first audio-video packet corresponding to the first text segment, the processor is specifically configured to:
[0182] Synthesize an audio packet corresponding to the first text segment through a speech synthesis module;
[0183] Generate a video packet according to the first text segment and / or the audio packet;
[0184] Perform timestamp alignment on the audio packet and the video packet to obtain a first audio-video packet.
[0185] In a possible implementation, the video packet includes one or more of lip movement parameters, expression parameters, and action parameters of a digital human.
[0186] In a possible implementation, the preset factors include one or more of time, application scenario, user type, and user object.
[0187] In a possible implementation, it further includes: identifying a user intention and generating a target text in response to the intention.
[0188] It should be noted that the implementation of each operation can also be correspondingly referred to the Figure 2 、 Figure 3 or Figure 4 corresponding description in the method embodiments shown.
[0189] The embodiment of the present application further provides a chip system, which includes at least one processor, a memory, and an interface circuit. The memory, the interface circuit, and the at least one processor are interconnected by lines. A computer program is stored in the at least one memory; when the computer program is executed by the processor, it implements the Figure 2 、 Figure 3 or Figure 4 method flow shown.
[0190] The embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When it runs on a processor, it implements the Figure 2 、 Figure 3 or Figure 4 method flow shown.
[0191] The embodiment of the present application further provides a computer program product. When the computer program product runs on a processor, it implements the Figure 2 、 Figure 3 or Figure 4 method flow shown.
[0192] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by hardware related to a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. The foregoing storage medium includes: various media such as ROM or random access memory RAM, magnetic disk, or optical disc that can store computer program code.
Claims
1. A method for generating a digital life, characterized in that, It includes: Split the target text to obtain multiple text segments. Among them, the target text includes text in response to a user, the multiple text segments include at least one static text segment and at least one dynamic text segment, the dynamic text segment includes text content that changes with a preset factor, and the static text segment includes text content that does not change with a preset factor; If the first text segment among the multiple text segments belongs to the dynamic text segment, generate a first audio-video package corresponding to the first text segment, where the first audio-video package corresponding to the first text segment is used to play in the form of a digital human.
2. The method according to claim 1, characterized in that, The method further includes: If the first text segment among the multiple text segments belongs to the static text segment, match a first audio-video package corresponding to the first text segment from an audio-video library, where the audio-video library includes multiple pre-generated audio-video packages.
3. The method according to claim 1 or 2, characterized in that, The method further includes: Output the first audio-video package corresponding to the first text segment.
4. The method according to any one of claims 1-3, characterized in that, The method further includes: If the next second text segment of the first text segment in the time domain is a static text segment, match a second audio-video package corresponding to the second text segment from an audio-video library, where the audio-video library includes multiple pre-generated audio-video packages; Generate a first transition frame according to the last x frames of the video part in the first audio-video package and the first y frames of the video part in the second audio-video package, where the first transition frame is used to play between the first audio-video package and the second audio-video package.
5. The method according to any one of claims 1-3, characterized in that, The method further includes: If the next second text segment of the first text segment in the time domain is a dynamic text segment, generate a first transition frame according to the last x frames of the video part in the first audio-video package and the first y frames of the video part in a third audio-video package, where the third audio-video package includes an audio-video package corresponding to the second text segment generated according to the second text segment, and the first transition frame is used to play between the first audio-video package and the third audio-video package.
6. The method according to claim 4 or 5, characterized in that, It further includes: Output the first transition frame.
7. The method according to any one of claims 1-3, characterized in that, The method further includes: If the next second text segment of the first text segment in the time domain is a dynamic text segment, generate a reference frame according to the action information of the limbs and / or face in the last z frames of the video part in the first audio-video package, and the reference frame is used as the first frame when generating the video part in the third audio-video package corresponding to the second text segment later.
8. The method according to any one of claims 1-7, characterized in that, The first text segment is any text segment except the last text segment among the multiple text segments.
9. The method according to any one of claims 1-7, characterized in that, The generation of the first audio-video package corresponding to the first text segment includes: Synthesize an audio package corresponding to the first text segment through a speech synthesis module; Generate a video package according to the first text segment and / or the audio package; Perform timestamp alignment on the audio package and the video package to obtain a first audio-video package.
10. The method according to claim 9, characterized in that, The video package includes one or more of the lip movement parameters, expression parameters, and action parameters of the digital human.
11. The method according to any one of claims 1-9, characterized in that, The preset factor includes one or more of time, application scenario, user type, user object, business environment, and business attribute.
12. The method according to any one of claims 1-11, characterized in that, It also includes: Identifying the user's intention and generating a target text in response to the intention.
13. A digital life generation device, characterized in that, Including: A splitting unit for splitting the target text to obtain a plurality of text segments. Among them, the target text includes text in response to the user, and the plurality of text segments include at least one static text segment and at least one dynamic text segment. The dynamic text segment includes text content that changes with a preset factor, and the static text segment includes text content that does not change with the preset factor; A first generating unit for generating a first audio-video packet corresponding to the first text segment when the first text segment among the plurality of text segments belongs to the dynamic text segment, where the first audio-video packet corresponding to the first text segment is used to be played in the form of a digital human.
14. The device according to claim 13, wherein, The device also includes: A matching unit for matching a first audio-video packet corresponding to the first text segment from an audio-video library when the first text segment among the plurality of text segments belongs to the static text segment, where the audio-video library includes a plurality of pre-generated audio-video packets.
15. The device according to claim 13 or 14, wherein, The device also includes: An output unit for outputting the first audio-video packet corresponding to the first text segment.
16. The device according to any one of claims 13-15, wherein, The device also includes a second generating unit: The matching unit is further used for matching a second audio-video packet corresponding to the second text segment from the audio-video library when the next second text segment of the first text segment in the time domain is a static text segment, where the audio-video library includes a plurality of pre-generated audio-video packets; The second generating unit is used to generate a first transition frame according to the last x frames of the video part in the first audio-video packet and the first y frames of the video part in the second audio-video packet, where the first transition frame is used to be played between the first audio-video packet and the second audio-video packet.
17. The device according to any one of claims 13-15, wherein, The device also includes: A third generating unit for generating a first transition frame according to the last x frames of the video part in the first audio-video packet and the first y frames of the video part in a third audio-video packet when the next second text segment of the first text segment in the time domain is a dynamic text segment. The third audio-video packet includes an audio-video packet corresponding to the second text segment generated according to the second text segment, and the first transition frame is used to be played between the first audio-video packet and the third audio-video packet.
18. The device according to claim 16 or 17, wherein, The output unit is further used for: Outputting the first transition frame.
19. The device according to any one of claims 13-15, wherein, The device also includes: A fourth generating unit for generating a reference frame according to the action information of the limbs and / or face in the last z frames of the video part in the first audio-video packet when the next second text segment of the first text segment in the time domain is a dynamic text segment. The reference frame is used as the first frame when generating the video part of the third audio-video packet corresponding to the second text segment later.
20. The device according to any one of claims 13-19, wherein, The first text segment is any text segment except the last text segment among the plurality of text segments.
21. The device according to any one of claims 13-19, wherein, Regarding generating the first audio-video packet corresponding to the first text segment, the first generating unit specifically is used for: Synthesize an audio packet corresponding to the first text segment through a speech synthesis module; Generate a video packet according to the first text segment and / or the audio packet; Perform timestamp alignment on the audio packet and the video packet to obtain a first audio-visual packet.
22. The device according to claim 21, wherein, The video packet includes one or more of the lip movement parameters, expression parameters, and action parameters of the digital human.
23. The device according to any one of claims 13-21, wherein, The preset factors include one or more of time, application scenario, user type, user object, business environment, and business attribute.
24. The device according to any one of claims 13-23, wherein, It further includes: An identification unit for identifying a user intention and generating a target text in response to the intention.
25. A digital life generation device, wherein, It includes a processor and a memory. Among them, the memory is used to store a computer program, and the processor is used to call the computer program to implement the method according to any one of claims 1-12.
26. A computer-readable storage medium, wherein, The computer-readable storage medium stores a computer program, and when the computer program is called by a processor, the method according to any one of claims 1-12 is implemented.