Digital-human generation method and related apparatus
By splitting the target text, different audio and video packet acquisition methods are adopted for static and dynamic text clips in digital human applications, which solves the problem of how to reduce computing power overhead while ensuring response speed, and realizes efficient digital audio and video packet generation and playback.
Patent Information
- Application Number
- PCT/CN2024/136676
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-12-04
- Publication Date
- 2025-06-26
AI Technical Summary
In digital human applications, in order to ensure instant response to users, large computing power support is usually required. How to minimize computing power overhead while ensuring response speed is a technical challenge.
By splitting the target text, static text clips and dynamic text clips are obtained. The audio and video packages are directly matched from the audio and video library for the static text clips, while the audio and video packages are generated for the dynamic text clips, avoiding the calculation overhead caused by full regeneration or search.
While minimizing the calculation overhead as much as possible, it meets the targeted demand for digital audio and video packages in different scenarios, and improves response speed and playback continuity.
Smart Images

Figure CN2024136676_26062025_PF_FP_ABST
Abstract
Description
A digital human generation method and related device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 20, 2023, with application number 202311763610.2, and invention name “A method for generating digital humans and related devices”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present invention relates to the field of artificial intelligence, and in particular to a digital human generation method and related devices. Background Art
[0003] Digital humans, also known as virtual humans, virtual digital humans, or avatars, refer to virtual characters with digital appearances that exist in the physical world, created through technologies such as computer graphics, graphics rendering, deep learning, and speech synthesis. With the advancement of technologies such as computer vision, natural language processing, and speech recognition, digital humans are becoming more intelligent and diverse, possessing the ability to express emotions and communicate. The application of digital humans has evolved from their early days in the entertainment field to include banking, healthcare, education, government affairs, and telecommunications. They serve as digital employees, virtual customer service representatives, and virtual instructors for businesses, providing explanations, consultations, and training to customers. A digital human system typically includes a character image, a speech generation module, an animation generation module, an audio and video synthesis module, and an interaction module. Based on the different interaction modules, digital humans can be categorized as non-interactive, intelligent-driven, and human-driven. Different types of digital humans handle speech and animation generation differently on the terminal.
[0004] Intelligent-driven digital humans are the most commonly used type of virtual digital human, typically used in scenarios such as digital human employees, digital human customer service, and virtual instructors. Users interact with the digital human through voice and text. The digital human recognizes the user's intent through artificial intelligence (AI) technologies such as speech recognition and semantic recognition in the interaction module's intelligent system. Based on this intent, the digital human determines the digital human's subsequent output text and actions. This drives the speech generation module to generate the corresponding speech response in real time. The animation generation module uses algorithms to render the corresponding character animation (including lip movements, facial expressions, and body movements) in real time based on the response speech. The audio and video synthesis module generates the final video and pushes it to the user's terminal device (such as a mobile phone, large screen, or tablet), enabling interaction between the digital human and the user. This type of interaction is generally real-time or near-real-time, requiring the digital human to respond quickly to user questions. However, in practical applications, ensuring immediate responses to users typically requires significant computing power. How to minimize computing power while ensuring responsiveness is a technical problem currently being explored by those skilled in the art. Summary of the Invention
[0005] The embodiments of the present application disclose a digital human generation method and related devices, which can meet the targeted needs of different scenarios for digital human audio and video packages while reducing computing overhead as much as possible.
[0006] In a first aspect, an embodiment of the present application provides a method for generating a digital human, the method comprising:
[0007] Splitting the target text to obtain a plurality of text segments, wherein the target text includes text responding to a user, and the plurality of text segments include at least one static text segment and at least one dynamic text segment, wherein the dynamic text segment includes text content that changes with a preset factor, and the static text segment includes text content that does not change with a preset factor;
[0008] If the first text segment among the multiple text segments belongs to the dynamic text segment, a first audio and video package corresponding to the first text segment is generated, wherein the first audio and video package corresponding to the first text segment is used to be played in the form of a digital human.
[0009] Using this method, based on the multiple text segments split from the target text, only the dynamic text will generate audio and video packages specifically, avoiding the problem of high computational overhead caused by regenerating audio and video packages for all audio segments.
[0010] In conjunction with the first aspect, in a possible implementation of the first aspect, the method further includes:
[0011] If the first text segment among the multiple text segments belongs to the static text segment, a first audio and video package corresponding to the first text segment is matched from an audio and video library, wherein the audio and video library includes multiple pre-generated audio and video packages.
[0012] In this method, different methods are used to obtain corresponding audio and video packages for static text segments and dynamic text segments. For static text segments, the corresponding audio and video packages are directly searched from the audio and video library, while for dynamic text, audio and video packages are directly generated. This avoids the problem of high computational overhead caused by regenerating audio and video packages for all audio segments, and also avoids the problem of searching for audio and video packages from the audio and video library for all audio segments, which results in the audio and video packages being unable to meet the different needs of different scenarios. That is, the above method can meet the targeted needs of digital human audio and video packages in different scenarios while reducing computational overhead as much as possible.
[0013] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in a possible implementation of the first aspect, the method further includes: outputting a first audio and video package corresponding to the first text segment.
[0014] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in a possible implementation of the first aspect, the method further includes:
[0015] If a second text segment that is next to the first text segment in the time domain is a static text segment, matching a second audio and video package corresponding to the second text segment from an audio and video library, wherein the audio and video library includes a plurality of pre-generated audio and video packages;
[0016] A first transition frame is generated according to the tail x frames of the video part in the first audio and video package and the head y frames of the video part in the second audio and video package, wherein the first transition frame is used to play between the first audio and video package and the second audio and video package.
[0017] In this implementation, after obtaining the audio and video package corresponding to the text segment, a smooth transition design is adopted between the audio and video package corresponding to the current text segment and the audio and video package corresponding to the next text segment, so that the playback continuity of the previous and next audio and video packages is better.
[0018] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in a possible implementation of the first aspect, the method further includes:
[0019] If the next second text segment of the first text segment in the time domain is a dynamic text segment, a first transition frame is generated based on the tail x frames of the video part in the first audio and video package and the head y frames of the video part in the third audio and video package, wherein the third audio and video package includes an audio and video package corresponding to the second text segment generated based on the second text segment, and the first transition frame is used to play between the first audio and video package and the third audio and video package.
[0020] In this implementation, after obtaining the audio and video package corresponding to the text segment, a smooth transition design is adopted between the audio and video package corresponding to the current text segment and the audio and video package corresponding to the next text segment, so that the playback continuity of the previous and next audio and video packages is better.
[0021] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in a possible implementation of the first aspect, the following further includes:
[0022] The first transition frame is output.
[0023] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in a possible implementation of the first aspect, the method further includes:
[0024] If the next second text segment of the first text segment in the time domain is a dynamic text segment, a reference frame is generated based on the limb and / or facial movement information in the tail z frame of the video part in the first audio and video package, and the reference frame is used as the first frame when generating the video part in the third audio and video package corresponding to the second text segment.
[0025] In this implementation, after obtaining the audio and video package corresponding to the text segment, a smooth transition design is adopted between the audio and video package corresponding to the current text segment and the audio and video package corresponding to the next text segment, so that the playback continuity of the previous and next audio and video packages is better.
[0026] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in a possible implementation of the first aspect, the first text segment is any text segment among the multiple text segments except the last text segment.
[0027] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in a possible implementation of the first aspect, generating a first audio and video package corresponding to the first text segment includes:
[0028] synthesizing an audio package corresponding to the first text segment by a speech synthesis module;
[0029] generating a video package based on the first text segment and / or the audio package;
[0030] Timestamps of the audio packet and the video packet are aligned to obtain a first audio and video packet.
[0031] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in a possible implementation of the first aspect, the video package includes one or more of the digital human's lip parameters, expression parameters, and action parameters.
[0032] In combination with the first aspect, or any of the above-mentioned possible implementations of the first aspect, in a possible implementation of the first aspect, the preset factors include one or more of time, application scenario, user type, user object, business environment, business attributes, etc.
[0033] In combination with the first aspect, or any of the foregoing possible implementations of the first aspect, in a possible implementation of the first aspect, the method further includes: identifying user intent and generating a target text that responds to the intent.
[0034] In a second aspect, an embodiment of the present application provides a digital human generation device, the device comprising:
[0035] a splitting unit, configured to split the target text into a plurality of text segments, wherein the target text includes text responding to a user, and the plurality of text segments include at least one static text segment and at least one dynamic text segment, wherein the dynamic text segment includes text content that changes with a preset factor, and the static text segment includes text content that does not change with a preset factor;
[0036] The first generating unit is used to generate a first audio and video package corresponding to the first text segment among the multiple text segments when the first text segment belongs to the dynamic text segment, wherein the first audio and video package corresponding to the first text segment is used to be played in the form of a digital human.
[0037] Using this method, based on the multiple text segments split from the target text, only the dynamic text will generate audio and video packages specifically, avoiding the problem of high computational overhead caused by regenerating audio and video packages for all audio segments.
[0038] In conjunction with the second aspect, in a possible implementation of the second aspect, the apparatus further includes:
[0039] A matching unit is used to match a first audio and video package corresponding to a first text segment among the multiple text segments from an audio and video library when the first text segment belongs to the static text segment, wherein the audio and video library includes multiple pre-generated audio and video packages.
[0040] In this method, different methods are used to obtain corresponding audio and video packages for static text segments and dynamic text segments. For static text segments, the corresponding audio and video packages are directly searched from the audio and video library, while for dynamic text, audio and video packages are directly generated. This avoids the problem of high computational overhead caused by regenerating audio and video packages for all audio segments, and also avoids the problem of searching for audio and video packages from the audio and video library for all audio segments, which results in the audio and video packages being unable to meet the different needs of different scenarios. That is, the above method can meet the targeted needs of digital human audio and video packages in different scenarios while reducing computational overhead as much as possible.
[0041] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in a possible implementation of the second aspect, the apparatus further includes:
[0042] An output unit is used to output a first audio and video package corresponding to the first text segment.
[0043] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in a possible implementation of the second aspect, the apparatus further includes a second generating unit:
[0044] The matching unit is further configured to, when a second text segment that is next to the first text segment in the time domain is a static text segment, match a second audio and video package corresponding to the second text segment from an audio and video library, wherein the audio and video library includes a plurality of pre-generated audio and video packages;
[0045] The second generation unit is used to generate a first transition frame based on the tail x frames of the video part in the first audio and video package and the head y frames of the video part in the second audio and video package, wherein the first transition frame is used to play between the first audio and video package and the second audio and video package.
[0046] In this implementation, after obtaining the audio and video package corresponding to the text segment, a smooth transition design is adopted between the audio and video package corresponding to the current text segment and the audio and video package corresponding to the next text segment, so that the playback continuity of the previous and next audio and video packages is better.
[0047] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in a possible implementation of the second aspect, the apparatus further includes:
[0048] The third generation unit is used to generate a first transition frame according to the tail x frames of the video part in the first audio and video package and the head y frames of the video part in the third audio and video package when the next second text segment of the first text segment in the time domain is a dynamic text segment, wherein the third audio and video package includes an audio and video package corresponding to the second text segment generated according to the second text segment, and the first transition frame is used to play between the first audio and video package and the third audio and video package.
[0049] In this implementation, after obtaining the audio and video package corresponding to the text segment, a smooth transition design is adopted between the audio and video package corresponding to the current text segment and the audio and video package corresponding to the next text segment, so that the playback continuity of the previous and next audio and video packages is better.
[0050] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in a possible implementation of the second aspect, the output unit is further configured to: output the first transition frame.
[0051] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in a possible implementation of the second aspect, the apparatus further includes:
[0052] The fourth generation unit is used to generate a reference frame based on the limb and / or facial movement information in the tail z frame of the video part in the first audio and video package when the next second text segment in the time domain of the first text segment is a dynamic text segment. The reference frame is used as the first frame when subsequently generating the video part in the third audio and video package corresponding to the second text segment.
[0053] In this implementation, after obtaining the audio and video package corresponding to the text segment, a smooth transition design is adopted between the audio and video package corresponding to the current text segment and the audio and video package corresponding to the next text segment, so that the playback continuity of the previous and next audio and video packages is better.
[0054] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in a possible implementation of the second aspect, the first text segment is any text segment among the multiple text segments except the last text segment.
[0055] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in a possible implementation of the second aspect, in terms of generating a first audio and video package corresponding to the first text segment, the first generation unit is specifically configured to:
[0056] synthesizing an audio package corresponding to the first text segment by a speech synthesis module;
[0057] generating a video package based on the first text segment and / or the audio package;
[0058] Timestamps of the audio packet and the video packet are aligned to obtain a first audio and video packet.
[0059] In combination with the second aspect, or any of the foregoing possible implementations of the second aspect, in a possible implementation of the second aspect, the video package includes one or more of the digital human's lip parameters, expression parameters, and action parameters.
[0060] In combination with the second aspect, or any of the above-mentioned possible implementations of the second aspect, in a possible implementation of the second aspect, the preset factors include one or more of time, application scenario, user type, user object, business environment, business attributes, etc.
[0061] In combination with the second aspect, or any of the above-mentioned possible implementations of the second aspect, in a possible implementation of the second aspect, it further includes: a recognition unit, which is used to recognize the user's intention and generate a target text that responds to the intention.
[0062] In a third aspect, an embodiment of the present application provides a digital human generation device, which includes a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to call the computer program to implement the method described in the first aspect or any possible implementation of the first aspect.
[0063] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is called by a processor, it implements the method described in the first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] The following is an introduction to the drawings used in the embodiments of this application.
[0065] FIG1 is a schematic diagram of the architecture of a digital human video processing system provided in an embodiment of the present application;
[0066] FIG2 is a schematic diagram of a flow chart of a method for generating a digital human according to an embodiment of the present application;
[0067] FIG3 is a flow chart of a method for generating a digital human according to an embodiment of the present application;
[0068] FIG4 is a flow chart of a method for generating a digital human according to an embodiment of the present application;
[0069] FIG5 is a schematic structural diagram of a digital human generation device provided in an embodiment of the present application;
[0070] FIG6 is a schematic structural diagram of a digital human generation device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0071] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.
[0072] Please refer to FIG1 , which is a schematic diagram of the architecture of a digital human video processing system provided in an embodiment of the present application. The digital human video processing system includes a business system 101, a digital human system 102, an audio and video playback system 103, and a terminal device 104, wherein:
[0073] Business system 101: Generates textual content broadcast by the digital human according to business needs. For example, it acquires user intent information and generates textual content responsive to that intent, i.e., the target text. This business system can be deployed in the cloud or locally. When deployed in the cloud, the user intent information can be uploaded from the terminal device to the business system. When deployed locally, the business system can be the terminal device itself, or another device independent of the terminal device, equipped with sensors to collect user intent information.
[0074] Digital Human System 102: Generates the digital human's audio and video files or audio and video streams based on the target text. This includes operations such as text segmentation, pre-creating audio and video corresponding to static text segments to obtain an audio and video library, instantly searching the library for audio and video packages corresponding to specific static text segments, instantly creating audio and video packages corresponding to dynamic text segments, and outputting the audio and video files.
[0075] The digital human system can be a software system or a hardware system with software deployed. The hardware system can be a single device or a cluster of multiple devices. The digital human system can be deployed in the cloud or locally, and this application does not limit this.
[0076] Audio and video playback system 103: This system interfaces with the Digihuman system and outputs audio and video files to the audio and video file system. This system then interacts with the terminal device (also known as the User Equipment (UE)) to stream the Digihuman audio and video files to the terminal device for playback. Optionally, the audio and video playback system 103 may not exist, and the Digihuman system 102 may directly output the audio and video files to the terminal device.
[0077] Terminal device 104: When using the service system, users view digital human audio and video files through the terminal device. This terminal device can be an auxiliary or associated device of the service system, such as a mobile phone, PAD, or computer, or it can be integrated with the service system. The service system can provide a variety of services, such as banking service systems and telecommunications service systems, which are not limited here.
[0078] Please refer to FIG2 , which is a flowchart of a method for generating a digital human provided by an embodiment of the present application. The method can be implemented based on the architecture shown in the figure, including but not limited to the following, or based on other architectures. The method includes but is not limited to the following steps:
[0079] Step S201: The business system identifies the user's intention and generates a target text that responds to the intention.
[0080] Specifically, user intention can be expressed in a variety of ways. It can be expressed through voice, such as the user saying "check the phone bill" or "help me check the traffic conditions on the way home"; it can also be expressed through body movements, such as the user expressing through sign language, nodding, shaking head, blinking, etc.; it can also be expressed through user behavior, such as the user stepping on a scale or a height measuring tape.
[0081] Accordingly, information collected by sensors can be used to obtain user intent. For example, an image sensor (such as a camera) can be used to collect user body movements and actions, a microphone can be used to collect user voice, and a touch screen can be used to collect user selection operations and input operations (such as text). The information collected by the sensors can be analyzed according to predefined rules to determine the user's intent.
[0082] Afterwards, a target text is generated to respond to the user's intention. For example, if the user's input "query phone bill" determines that the user's intention is "query phone bill balance", then the target text of the response can be "Hello, your cumulative consumption this month is 69.55 yuan, and the phone bill balance is 100.52 yuan. The specific information will be sent to your mobile phone later, please check it"; for example, if the user's body movement of nodding determines that the user's intention is "agree to learn more about phone bill package A", then the target text of the response can be an introduction text about package A; for example, if the identified user intention is to weigh oneself, then the response text can be information about the user's weight (such as "Your weight is 140KG"). It can be understood that the scenarios and text contents mentioned here are only examples. In fact, there can be other scenarios and contents, which will not be given one by one here.
[0083] The specific method for generating the response text is not limited here. You can pre-establish mapping relationships between multiple user intents and multiple text segments, and once the user intent is known, use the text corresponding to the mapping relationship as the target text. You can also pre-build a prediction model and input the user intent into the prediction model to obtain the corresponding target text. You can also pre-configure corresponding algorithms or rules and use the user intent as input to substitute into the corresponding algorithms and rules to obtain the corresponding target text. Of course, there are other methods, which are not listed here one by one.
[0084] Optionally, step S201 may also be completed by a digital human system.
[0085] Step S202: The Digital Human system splits the target text to obtain multiple text segments.
[0086] Specifically, splitting rules can be pre-configured, such as splitting according to punctuation marks, dynamic content identifiers (pre-set), etc.; another example is that a splitting algorithm can be pre-configured, and the target text can be input into the splitting algorithm to output multiple text fragments after splitting. The splitting algorithm can be directly configured, or it can be trained by inputting a large number of training samples into the corresponding model with the help of an AI algorithm. Optionally, the splitting algorithm can use high-frequency common copy fragments that conform to natural language rules as static content.
[0087] In the embodiment of the present application, after obtaining multiple text segments, each of the text segments is classified. The text segments may include static text segments and dynamic text segments. Of course, other types of text segments may also be included, which are specifically set as needed. The dynamic text segments include text content that changes with preset factors, and the static text segments include text content that does not change with preset factors. For example, the preset factors include one or more of factors such as time, application scenario, user type, and user object.
[0088] Optionally, a text type recognition model can be pre-configured. The text category recognition model can be obtained by training a large amount of sample data, each sample data including a text segment and a type identifier of the text segment. Therefore, the text type recognition model finally trained can better identify or predict the text type of each text segment. In an embodiment of the present application, when classifying, a corresponding first label (such as 0) can be marked on a static text segment, and a corresponding second label (such as 1) can be marked on a dynamic text segment. Special symbols can also be added before and after the text segment, such as adding a special symbol @ before and after the dynamic text segment to facilitate subsequent identification and distinction.
[0089] Typically, the recognition result is that the multiple text segments include at least one static text segment and at least one dynamic text segment.
[0090] For ease of understanding, let's take an example. For example, if the target text is "Hello, your monthly cumulative consumption is 69.55 yuan, and your phone balance is 100.52 yuan. Detailed information will be sent to your mobile phone later, please check." The text segments and text types obtained after segmentation are shown in Table 1:
[0091] Table 1
[0092] As can be seen from Table 1, there may be static content before and after dynamic content, and there may also be dynamic content before and after static content.
[0093] In the embodiment of the present application, it is necessary to determine the audio and video package corresponding to each of the multiple text segments, that is, the content played in the form of a digital human. Different determination methods can be used for static text segments and dynamic text segments. For ease of understanding, the following description is made using the first text segment among the multiple text segments as an example. Among the multiple text segments, some of the text segments can be determined in the same manner as the first text segment, or all of the text segments can be determined in the same manner as the first text segment.
[0094] Step S203: If the first text segment among the multiple text segments belongs to the dynamic text segment, the digital human system generates a first audio and video package corresponding to the first text segment.
[0095] In an optional implementation, generating the first audio and video package may include:
[0096] An audio package corresponding to the first text segment is synthesized by a speech synthesis module. The audio package may include audio of reading or explaining the first text segment, and of course, also includes timestamp information corresponding to the audio. The first text segment may be input into a corresponding algorithm or model, such as a first model, to output the audio package. The first model may be pre-trained.
[0097] Then, a video package is generated based on the first text segment and / or the audio package. Here, a video of the first text segment is generated. For example, the video includes information such as the digital human's lip parameters, facial expressions, hand and body movements, the digital human's image, the digital human's video background, and, of course, the corresponding timestamp of the video. The digital human image can be pre-configured as needed, for example, to distinguish between male and female digital humans, or between elderly, young, and child digital humans. Since the audio package includes audio for the first text segment, a video package can be generated based on the first text segment or the audio package. Since information such as tone and intonation in the audio package may affect the digital human's facial expressions and body movements, a video package can also be generated based on the first text segment and the audio package (such as tone and intonation). The process of generating a video package based on a voice package is the process of generating a video by driving the digital human's lip shape and facial expressions through voice. Optionally, the audio package can include an audio stream, and the video package can include a video stream.
[0098] Afterwards, the audio and video packets are time-stamp aligned to obtain a first audio-video packet. Time-stamp alignment allows the voice in the audio packet to be synchronized (or aligned) with the digital human figure, facial expression, and body movements in the video packet. The resulting first audio-video packet is therefore a video animation of the digital human reading the first text segment.
[0099] Optionally, if the first text segment is a dynamic text segment, and the first text segment is not the starting text segment in the time dimension, then the first frame of the video portion in the first audio and video can be generated based on the z-frame of the video portion in the audio and video corresponding to the previous text segment of the first text segment. For example, the first frame of the video portion in the first audio and video can be generated with reference to the lip shape, expression, body movement, etc. in the z-frame (the first frame can be generated by referring to the generation principle in S209). In this way, the continuity of the first audio and video corresponding to the first text segment and the audio and video corresponding to the previous text segment can be ensured.
[0100] Optionally, if the first text segment is a dynamic text segment and is a starting text segment in the time dimension, then the first frame of the video part in the first audio and video can be default, for example, the lip shape, facial expression, body movement, etc. are the default configuration values.
[0101] Step S204: If the first text segment among the multiple text segments belongs to the static text segment, the digital human system matches the first audio and video package corresponding to the first text segment from the audio and video library.
[0102] Since static text usually does not change frequently, it appears more frequently and has a higher probability. Therefore, corresponding audio and video packages can be generated in advance for static text, and an audio and video library can be built based on a large amount of static text and its corresponding audio and video packages for calling.
[0103] In an embodiment of the present application, when the first text segment is static text, there is no need to generate its corresponding audio and video package immediately. Instead, the audio and video package corresponding to the first text segment is directly searched from the existing audio and video library. This process is faster.
[0104] When searching, in addition to using the first text segment, you may also need the desired digital human image, the desired digital human video background, etc. For example, if the first text segment corresponds to two audio and video packages, one of which is presented by a male digital human and the other by a female digital human, then when searching through the first text segment, you need to combine the desired digital human image to find the corresponding specific audio and video package. The function and usage of the digital human video background are similar to those of the digital human image and will not be further described here. Optionally, if the first text segment is the starting text segment, then during the search, the digital human image, digital human video background, etc. corresponding to the audio and video package can be selected according to certain rules, such as configuring a default. If the first text segment is not the starting text segment, then during the search, the digital human image, digital human video background, etc. in the audio and video package corresponding to the previous text segment can be used as one of the search keywords to ensure that the digital human image, digital human video background, etc. in the audio and video package corresponding to the first text segment are consistent with the digital human image, digital human video background, etc. in the audio and video package corresponding to the previous text segment.
[0105] In the embodiment of the present application, the audio and video library may exist in a database, memory, cache, or corresponding file, which is not limited here.
[0106] Step S205: The Digital Human system outputs the first audio and video package corresponding to the first text segment.
[0107] In this embodiment of the present application, the first audio and video package corresponding to the finalized first text segment is used for playback in the form of a digital human. Output here includes direct playback, such as direct playback by the aforementioned digital human system. Output here can also refer to sending to a terminal device for playback by the terminal device. This includes the subsequently generated first transition frame, which can be played by the digital human system or sent from the digital human system to the terminal device for playback.
[0108] In addition, the playback in the embodiment of the present application can be performed instantly in the form of a data stream, or the audio and video package corresponding to a text segment can be played after the audio and video package is generated, or the entire audio and video package is played after all the audio and video packages corresponding to the text segment are generated. The specific playback method is not limited here. Optionally, when playing in the form of a video stream, the audio and video package will be played simultaneously during the process of obtaining, generating or transmitting the audio and video package. Of course, the playback of some frames may require a short wait. For example, when the first transition frame mentioned later needs to be generated, although the first video frame in the audio and video corresponding to the dynamic text segment has been generated, it is necessary to wait for the first transition frame to be played before playing the first video frame. Of course, the waiting process is generally not too long, and timely playback can basically be achieved.
[0109] Optionally, in order to achieve a smooth transition of audio and video corresponding to different text segments during playback, the following steps S206-S207 (optional solution one, as shown in Figures 3 and 4) may also be included, or step S208 (optional solution two, as shown in Figure 3), or step S209 (optional solution three, as shown in Figure 4) may be included, which are introduced below respectively.
[0110] Step S206: If the second text segment next to the first text segment in the time domain is a static text segment, the digital human system matches the second audio and video package corresponding to the second text segment from the audio and video library.
[0111] It can be understood that each text segment also corresponds to a timestamp, so the previous text segment and the next text segment of each text segment can be determined in the time domain, and the text type of each text segment is also marked when multiple text segments are split. Therefore, it is easy to determine whether the next second text segment of the first text segment in the time domain is a static text segment or a dynamic text segment.
[0112] The method of matching the second audio and video package corresponding to the second text segment from the audio and video library is the same as the method of matching the first audio and video package corresponding to the first text segment, and will not be repeated here.
[0113] Step S207: the digital human system generates a first transition frame according to the tail x frames of the video portion in the first audio and video package and the head y frames of the video portion in the second audio and video package.
[0114] Specifically, the tail x frame and the head y frame can generate a first transition frame between the tail x frame and the head y frame by interpolation. The first transition frame, including other transition frames mentioned later, can be one frame or multiple frames. Generally speaking, if the content of the x frame is significantly different from that of the y frame, multiple first transition frames can be generated, so that the subsequent playback effect will be better. Optionally, in addition to interpolation, other methods can be used, such as based on a preset generation process, using the tail x frame and the head y frame as input, and obtaining the first transition frame through the generation process; for example, the tail x frame and the head y frame are input into a trained transition frame prediction model to output the first transition frame, wherein the sizes of x and y can be pre-set as needed; of course, there are other methods, which are not listed here one by one.
[0115] In this case, the first transition frame is used to play between the first audio and video package and the second audio and video package.
[0116] As mentioned above, the first transition frame can be played by the digital human system, or the digital human system can send it to the terminal device for playback. The playback method can be video streaming or non-video streaming.
[0117] Step S208: If the second text segment next to the first text segment in the time domain is a dynamic text segment, the digital human system generates a first transition frame based on the tail x frames of the video part in the first audio and video package and the head y frames of the video part in the third audio and video package.
[0118] Specifically, the method of generating the first transition frame may refer to the related description in the previous step S207, which will not be repeated here.
[0119] Among them, the third audio and video package includes an audio and video package corresponding to the second text segment generated according to the second text segment. The method of generating the third audio and video package according to the second text segment (in the case of a dynamic text segment) is the same as the method of generating the first audio and video package according to the first text segment (in the case of a dynamic text segment), and will not be repeated here.
[0120] In this case, the first transition frame is used to play between the first audio and video package and the second audio and video package.
[0121] As mentioned above, the first transition frame can be played by the digital human system, or the digital human system can send it to the terminal device for playback. The playback method can be video streaming or non-video streaming.
[0122] Step S209: If the second text segment next to the first text segment in the time domain is a dynamic text segment, the digital human system generates a reference frame based on the limb and / or facial movement information in the last z frame of the video part in the first audio and video package.
[0123] That is to say, when generating a reference frame, in addition to the movement information of the limbs (such as gestures) and / or face (such as lip shape, expression) in the tail z frame, other information may also be used, such as the digital human image, digital human video background and other information.
[0124] It should be noted that the reference frame is used as the first frame when subsequently generating the video portion of the third audio / video package corresponding to the second text segment. Therefore, the reference frame in step S209 can be generated when the third audio / video package is subsequently generated based on the second text segment, or it can be generated before the third audio / video package is generated based on the second text segment, which is equivalent to a preprocessing operation before generating the third audio / video package.
[0125] It should be noted that steps S206-S209 supplement steps S201-S205. Steps S206-S209 primarily describe how to achieve a smooth transition between the audio and video packages corresponding to the first text segment and the subsequent text segment (i.e., the second text segment), thereby ensuring continuity and smoothness during the playback of the digital human. Therefore, it can be considered an associated operation for generating audio and video packages for the first text segment. In this embodiment of the present application, the processing process for each of the multiple text segments can be similar to that of the first text segment, including steps S201-S205. Except for the last text segment in the time domain, each of the multiple text segments can be similar to the first text segment, including steps S206-S209. Of course, it is also possible to disregard smooth transition and directly play the audio and video packages corresponding to the previous text segment after the audio and video packages corresponding to the next text segment are played. Of course, it is also possible to consider smooth transition but adopt a method other than steps S206-S209.
[0126] In addition, in one implementation, the operations of generating corresponding audio and video packages for different text segments in the above-mentioned multiple text segments can be performed synchronously; in another implementation, the operations of generating corresponding audio and video packages are performed for each text segment in the multiple text segments in sequence according to the time domain order until the operations of generating corresponding audio and video packages are completed for each text segment. In the embodiments of the present application, the specific method can be selected and configured according to the characteristics of the specific application scenario and is not limited here.
[0127] In the method shown in FIG2 , based on the multiple text segments split from the target text, different methods are used to obtain the corresponding audio and video packages for the static and dynamic text segments. For the static text segments, the corresponding audio and video packages are directly searched from the audio and video library, while for the dynamic text, audio and video packages are directly generated. This avoids the problem of high computational overhead caused by regenerating audio and video packages for all audio segments, and also avoids the problem of audio and video packages not meeting the different needs of different scenarios due to searching for audio and video packages from the audio and video library for all audio segments. In other words, the above method can meet the targeted needs of digital human audio and video packages in different scenarios while minimizing computational overhead. Furthermore, after obtaining the audio and video packages corresponding to the text segments, a smooth transition design is used between the audio and video packages corresponding to the current text segment and the audio and video packages corresponding to the next text segment, so that the playback continuity of the previous and next audio and video packages is better.
[0128] The above describes in detail the method of the embodiment of the present application, and the following provides an apparatus of the embodiment of the present application.
[0129] Please refer to Figure 5, which is a structural schematic diagram of a digital human generation device provided in an embodiment of the present application. The device can be the above-mentioned digital human system or a device or module in the digital human system. The device 50 can include a splitting unit 501 and a first generation unit 502, wherein each unit is described in detail as follows.
[0130] A splitting unit 501 is configured to split a target text into a plurality of text segments, wherein the target text includes text in response to a user, and the plurality of text segments include at least one static text segment and at least one dynamic text segment, wherein the dynamic text segment includes text content that changes with a preset factor, and the static text segment includes text content that does not change with a preset factor;
[0131] The first generating unit 502 is configured to generate a first audio and video package corresponding to a first text segment among the multiple text segments when the first text segment belongs to the dynamic text segment, wherein the first audio and video package corresponding to the first text segment is used to be played in the form of a digital human.
[0132] Using this method, based on the multiple text segments split from the target text, only the dynamic text will generate audio and video packages specifically, avoiding the problem of high computational overhead caused by regenerating audio and video packages for all audio segments.
[0133] In a possible implementation, the apparatus further includes:
[0134] A matching unit is used to match a first audio and video package corresponding to a first text segment among the multiple text segments from an audio and video library when the first text segment belongs to the static text segment, wherein the audio and video library includes multiple pre-generated audio and video packages.
[0135] In this method, different methods are used to obtain corresponding audio and video packages for static text segments and dynamic text segments. For static text segments, the corresponding audio and video packages are directly searched from the audio and video library, while for dynamic text, audio and video packages are directly generated. This avoids the problem of high computational overhead caused by regenerating audio and video packages for all audio segments, and also avoids the problem of searching for audio and video packages from the audio and video library for all audio segments, which results in the audio and video packages being unable to meet the different needs of different scenarios. That is, the above method can meet the targeted needs of digital human audio and video packages in different scenarios while reducing computational overhead as much as possible.
[0136] In a possible implementation, the apparatus further includes:
[0137] An output unit is used to output a first audio and video package corresponding to the first text segment.
[0138] In a possible implementation, the apparatus further includes a second generating unit:
[0139] The matching unit is further configured to, when a second text segment that is next to the first text segment in the time domain is a static text segment, match a second audio and video package corresponding to the second text segment from an audio and video library, wherein the audio and video library includes a plurality of pre-generated audio and video packages;
[0140] The second generation unit is used to generate a first transition frame based on the tail x frames of the video part in the first audio and video package and the head y frames of the video part in the second audio and video package, wherein the first transition frame is used to play between the first audio and video package and the second audio and video package.
[0141] In this implementation, after obtaining the audio and video package corresponding to the text segment, a smooth transition design is adopted between the audio and video package corresponding to the current text segment and the audio and video package corresponding to the next text segment, so that the playback continuity of the previous and next audio and video packages is better.
[0142] In a possible implementation, the apparatus further includes:
[0143] The third generation unit is used to generate a first transition frame according to the tail x frames of the video part in the first audio and video package and the head y frames of the video part in the third audio and video package when the next second text segment of the first text segment in the time domain is a dynamic text segment, wherein the third audio and video package includes an audio and video package corresponding to the second text segment generated according to the second text segment, and the first transition frame is used to play between the first audio and video package and the third audio and video package.
[0144] In this implementation, after obtaining the audio and video package corresponding to the text segment, a smooth transition design is adopted between the audio and video package corresponding to the current text segment and the audio and video package corresponding to the next text segment, so that the playback continuity of the previous and next audio and video packages is better.
[0145] In a possible implementation manner, the output unit is further configured to: output the first transition frame.
[0146] In a possible implementation, the apparatus further includes:
[0147] The fourth generation unit is used to generate a reference frame based on the limb and / or facial movement information in the tail z frame of the video part in the first audio and video package when the next second text segment in the time domain of the first text segment is a dynamic text segment. The reference frame is used as the first frame when subsequently generating the video part in the third audio and video package corresponding to the second text segment.
[0148] In this implementation, after obtaining the audio and video package corresponding to the text segment, a smooth transition design is adopted between the audio and video package corresponding to the current text segment and the audio and video package corresponding to the next text segment, so that the playback continuity of the previous and next audio and video packages is better.
[0149] In a possible implementation, the first text segment is any text segment among the multiple text segments except the last text segment.
[0150] In a possible implementation, in terms of generating the first audio and video package corresponding to the first text segment, the first generating unit is specifically configured to:
[0151] synthesizing an audio package corresponding to the first text segment by a speech synthesis module;
[0152] generating a video package based on the first text segment and / or the audio package;
[0153] Timestamps of the audio packet and the video packet are aligned to obtain a first audio and video packet.
[0154] In a possible implementation, the video package includes one or more of the digital human's lip parameters, expression parameters, and action parameters.
[0155] In a possible implementation, the preset factors include one or more of time, application scenario, user type, and user object.
[0156] In a possible implementation, the method further includes: a recognition unit configured to recognize user intention and generate a target text in response to the intention.
[0157] It should be noted that the implementation of each unit may also correspond to the corresponding description of the method embodiment shown in Figure 2, Figure 3 or Figure 4.
[0158] Please refer to Figure 6, which shows a digital human generation device 60 provided in an embodiment of the present application. The device 60 includes a processor 601, a memory 602 and a communication interface 603. The processor 601, the memory 602 and the communication interface 603 are interconnected via a bus.
[0159] The memory 602 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM). The memory 602 is used for storing computer programs and data. The communication interface 603 is used to receive and send data.
[0160] The processor 601 may be one or more central processing units (CPUs). In the case where the processor 601 is a CPU, the CPU may be a single-core CPU or a multi-core CPU.
[0161] The processor 601 in the device 60 is configured to read the computer program code stored in the memory 602 and perform the following operations:
[0162] Splitting the target text to obtain a plurality of text segments, wherein the target text includes text responding to a user, and the plurality of text segments include at least one static text segment and at least one dynamic text segment, wherein the dynamic text segment includes text content that changes with a preset factor, and the static text segment includes text content that does not change with a preset factor;
[0163] If the first text segment among the multiple text segments belongs to the dynamic text segment, a first audio and video package corresponding to the first text segment is generated, wherein the first audio and video package corresponding to the first text segment is used to be played in the form of a digital human.
[0164] Using this method, based on the multiple text segments split from the target text, only the dynamic text will generate audio and video packages specifically, avoiding the problem of high computational overhead caused by regenerating audio and video packages for all audio segments.
[0165] In a possible implementation, the processor is further configured to:
[0166] If the first text segment among the multiple text segments belongs to the static text segment, a first audio and video package corresponding to the first text segment is matched from an audio and video library, wherein the audio and video library includes multiple pre-generated audio and video packages.
[0167] In this method, different methods are used to obtain corresponding audio and video packages for static text segments and dynamic text segments. For static text segments, the corresponding audio and video packages are directly searched from the audio and video library, while for dynamic text, audio and video packages are directly generated. This avoids the problem of high computational overhead caused by regenerating audio and video packages for all audio segments, and also avoids the problem of searching for audio and video packages from the audio and video library for all audio segments, which results in the audio and video packages being unable to meet the different needs of different scenarios. That is, the above method can meet the targeted needs of digital human audio and video packages in different scenarios while reducing computational overhead as much as possible.
[0168] In a possible implementation, the processor is further configured to output the first audio and video package corresponding to the first text segment by calling the communication interface 603 .
[0169] In a possible implementation, the processor is further configured to:
[0170] If a second text segment that is next to the first text segment in the time domain is a static text segment, matching a second audio and video package corresponding to the second text segment from an audio and video library, wherein the audio and video library includes a plurality of pre-generated audio and video packages;
[0171] A first transition frame is generated according to the tail x frames of the video part in the first audio and video package and the head y frames of the video part in the second audio and video package, wherein the first transition frame is used to play between the first audio and video package and the second audio and video package.
[0172] In this implementation, after obtaining the audio and video package corresponding to the text segment, a smooth transition design is adopted between the audio and video package corresponding to the current text segment and the audio and video package corresponding to the next text segment, so that the playback continuity of the previous and next audio and video packages is better.
[0173] In a possible implementation, the processor is further configured to:
[0174] If the next second text segment of the first text segment in the time domain is a dynamic text segment, a first transition frame is generated based on the tail x frames of the video part in the first audio and video package and the head y frames of the video part in the third audio and video package, wherein the third audio and video package includes an audio and video package corresponding to the second text segment generated based on the second text segment, and the first transition frame is used to play between the first audio and video package and the third audio and video package.
[0175] In this implementation, after obtaining the audio and video package corresponding to the text segment, a smooth transition design is adopted between the audio and video package corresponding to the current text segment and the audio and video package corresponding to the next text segment, so that the playback continuity of the previous and next audio and video packages is better.
[0176] In a possible implementation, the processor is further configured to call the communication interface 603 to output the first transition frame.
[0177] In a possible implementation, the processor is further configured to:
[0178] If the next second text segment of the first text segment in the time domain is a dynamic text segment, a reference frame is generated based on the limb and / or facial movement information in the tail z frame of the video part in the first audio and video package, and the reference frame is used as the first frame when generating the video part in the third audio and video package corresponding to the second text segment.
[0179] In this implementation, after obtaining the audio and video package corresponding to the text segment, a smooth transition design is adopted between the audio and video package corresponding to the current text segment and the audio and video package corresponding to the next text segment, so that the playback continuity of the previous and next audio and video packages is better.
[0180] In a possible implementation, the first text segment is any text segment among the multiple text segments except the last text segment.
[0181] In a possible implementation, in terms of generating the first audio and video package corresponding to the first text segment, the processor is specifically configured to:
[0182] synthesizing an audio package corresponding to the first text segment by a speech synthesis module;
[0183] generating a video package based on the first text segment and / or the audio package;
[0184] Timestamps of the audio packet and the video packet are aligned to obtain a first audio and video packet.
[0185] In a possible implementation, the video package includes one or more of the digital human's lip parameters, expression parameters, and action parameters.
[0186] In a possible implementation, the preset factors include one or more of time, application scenario, user type, and user object.
[0187] In a possible implementation, the method further includes: identifying user intention and generating a target text in response to the intention.
[0188] It should be noted that the implementation of each operation may also correspond to the corresponding description of the method embodiment shown in Figure 2, Figure 3 or Figure 4.
[0189] An embodiment of the present application also provides a chip system, which includes at least one processor, a memory and an interface circuit, wherein the memory, the interface circuit and the at least one processor are interconnected through lines, and a computer program is stored in the at least one memory; when the computer program is executed by the processor, the method flow shown in Figure 2, Figure 3 or Figure 4 is implemented.
[0190] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer-readable storage medium is executed on a processor, the method flow shown in FIG. 2 , FIG. 3 , or FIG. 4 is implemented.
[0191] An embodiment of the present application further provides a computer program product, which, when executed on a processor, implements the method flow shown in FIG. 2 , FIG. 3 , or FIG. 4 .
[0192] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by a computer program or computer program-related hardware. The computer program can be stored in a computer-readable storage medium. When executed, the computer program can include the processes in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing computer program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A method for generating a digital human, characterized in that: include: Splitting the target text to obtain a plurality of text segments, wherein the target text includes text responding to the user, and the plurality of text segments include at least one static text segment and at least one dynamic text segment, wherein the dynamic text segment includes text content that changes with a preset factor, and the static text segment includes text content that does not change with a preset factor; If the first text segment among the multiple text segments belongs to the dynamic text segment, a first audio and video package corresponding to the first text segment is generated, wherein the first audio and video package corresponding to the first text segment is used for playing in the form of a digital human.
2. The method according to claim 1, characterized in that The method further comprises: If a first text segment among the multiple text segments belongs to the static text segment, a first audio and video package corresponding to the first text segment is matched from an audio and video library, wherein the audio and video library includes multiple pre-generated audio and video packages.
3. The method according to claim 1 or 2, characterized in that: The method further comprises: Output the first audio and video package corresponding to the first text segment.
4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: If the second text segment next to the first text segment in the time domain is a static text segment, matching a second audio and video package corresponding to the second text segment from an audio and video library, wherein the audio and video library includes a plurality of pre-generated audio and video packages; A first transition frame is generated according to the tail x frames of the video part in the first audio and video package and the head y frames of the video part in the second audio and video package, wherein the first transition frame is used to play between the first audio and video package and the second audio and video package.
5. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: If the next second text segment of the first text segment in the time domain is a dynamic text segment, a first transition frame is generated according to the tail x frames of the video part in the first audio and video package and the head y frames of the video part in the third audio and video package, wherein the third audio and video package includes an audio and video package corresponding to the second text segment generated according to the second text segment, and the first transition frame is used to play between the first audio and video package and the third audio and video package.
6. The method according to claim 4 or 5, characterized in that: Also includes: The first transition frame is output.
7. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: If the next second text segment of the first text segment in the time domain is a dynamic text segment, a reference frame is generated based on the limb and / or facial movement information in the tail z frame of the video part in the first audio and video package, and the reference frame is used as the first frame when generating the video part in the third audio and video package corresponding to the second text segment.
8. The method according to any one of claims 1 to 7, characterized in that: The first text segment is any text segment among the multiple text segments except the last text segment.
9. The method according to any one of claims 1 to 7, characterized in that: The generating a first audio and video package corresponding to the first text segment includes: synthesizing an audio package corresponding to the first text segment by a speech synthesis module; generating a video package based on the first text segment and / or the audio package; Timestamps of the audio packet and the video packet are aligned to obtain a first audio and video packet.
10. The method according to claim 9, characterized in that The video package includes one or more of the lip parameters, expression parameters and action parameters of the digital human.
11. The method according to any one of claims 1 to 9, characterized in that: The preset factors include one or more of time, application scenario, user type, user object, business environment, and business attributes.
12. The method according to any one of claims 1 to 11, characterized in that: Also includes: Identify user intent and generate target text that responds to the intent.
13. A digital human generation device, characterized in that: include: A splitting unit, configured to split the target text to obtain a plurality of text segments, wherein the target text includes text responding to a user, and the plurality of text segments include at least one static text segment and at least one dynamic text segment, wherein the dynamic text segment includes text content that changes with a preset factor, and the static text segment includes text content that does not change with a preset factor; The first generating unit is used to generate a first audio and video package corresponding to a first text segment among the multiple text segments when the first text segment belongs to the dynamic text segment, wherein the first audio and video package corresponding to the first text segment is used to be played in the form of a digital human.
14. The device according to claim 13, characterized in that The device also includes: A matching unit is used to match a first audio and video package corresponding to a first text segment among the multiple text segments from an audio and video library when the first text segment belongs to the static text segment, wherein the audio and video library includes multiple pre-generated audio and video packages.
15. The device according to claim 13 or 14, characterized in that The device also includes: An output unit is used to output a first audio and video package corresponding to the first text segment.
16. The device according to any one of claims 13 to 15, characterized in that: The device also includes a second generating unit: The matching unit is further configured to match a second audio and video package corresponding to the second text segment from an audio and video library when a second text segment next to the first text segment in the time domain is a static text segment, wherein the audio and video library includes a plurality of pre-generated audio and video packages; The second generating unit is used to generate a first transition frame based on the tail x frames of the video part in the first audio and video package and the head y frames of the video part in the second audio and video package, wherein the first transition frame is used to play between the first audio and video package and the second audio and video package.
17. The device according to any one of claims 13 to 15, characterized in that: The device also includes: The third generating unit is used to generate a first transition frame according to the tail x frames of the video part in the first audio and video package and the head y frames of the video part in the third audio and video package when the next second text segment of the first text segment in the time domain is a dynamic text segment, wherein the third audio and video package includes an audio and video package corresponding to the second text segment generated according to the second text segment, and the first transition frame is used to play between the first audio and video package and the third audio and video package.
18. The device according to claim 16 or 17, characterized in that The output unit is also used for: The first transition frame is output.
19. The device according to any one of claims 13 to 15, characterized in that: The device also includes: The fourth generation unit is used to generate a reference frame based on the limb and / or facial movement information in the tail z frame of the video part in the first audio and video package when the next second text segment of the first text segment in the time domain is a dynamic text segment. The reference frame is used as the first frame when subsequently generating the video part in the third audio and video package corresponding to the second text segment.
20. The device according to any one of claims 13 to 19, characterized in that The first text segment is any text segment among the multiple text segments except the last text segment.
21. The device according to any one of claims 13 to 19, characterized in that In terms of generating the first audio and video package corresponding to the first text segment, the first generating unit is specifically used for: synthesizing an audio package corresponding to the first text segment by a speech synthesis module; generating a video package based on the first text segment and / or the audio package; Timestamps of the audio packet and the video packet are aligned to obtain a first audio and video packet.
22. The device according to claim 21, characterized in that The video package includes one or more of the lip parameters, expression parameters and action parameters of the digital human.
23. The device according to any one of claims 13 to 21, characterized in that The preset factors include one or more of time, application scenario, user type, user object, business environment, and business attributes.
24. The device according to any one of claims 13 to 23, characterized in that Also includes: The recognition unit is used to recognize the user's intention and generate a target text in response to the intention.
25. A digital human generation device, characterized in that: The method comprises a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to call the computer program to implement the method according to any one of claims 1 to 12.
26. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is called by a processor, the method according to any one of claims 1 to 12 is implemented.
Citation Information
Patent Citations
Virtual character generation method and device, virtual character display method and device, equipment and medium
CN112102449A
Digital human generation method and device, computer equipment and storage medium
CN114760425A
Business handling method and system based on digital human interaction video, and storage medium
CN116248812A
Video processing method and apparatus, and medium and program product
WO2023045716A1
Cited By
Method for perceiving self appearance, clothes and scene by digital separate body and related product
CN121745149A