A digital human video updating method and device, equipment, medium and product

By comparing the relevance of the dialogue text, a video adjustment plan is determined, and the lower body movement animation sequence is reused or generated. This solves the problems of high cost and poor relevance in updating full-body movement videos of digital humans, and achieves efficient and low-cost video updates.

CN120935407BActive Publication Date: 2026-03-31MOFA (SHANGHAI) INFORMATION TECH CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies require significant hardware resources to modify dialogue text or video scripts each time they generate full-body moving videos of digital humans, resulting in high update costs, poor video relevance before and after updates, and a poor user experience.

Method used

By comparing the current dialogue text with the original dialogue text, a video adjustment plan is determined. The lower body movement animation sequence is reused or regenerated to generate a full-body movement video of the target digital human, ensuring video update efficiency and relevance.

Benefits of technology

It reduces the hardware resource consumption of video updates, improves update efficiency, and ensures the relevance of videos before and after the update and the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120935407B_ABST
    Figure CN120935407B_ABST
Patent Text Reader

Abstract

The application discloses a digital person video updating method and device, equipment, medium and product, and relates to the technical field of video generation. The method comprises the following steps: acquiring current input content comprising current script text; determining the relevance of the original script text and the current script text and a video adjustment scheme based on the text correspondence relationship between the original script text and the current script text; determining the lower body movement animation sequence of the current script text based on the video adjustment scheme; and generating a target digital person full-body movement video based on the lower body movement animation sequence and the current input content. The application determines the video adjustment scheme according to the modification of the script text, then determines the lower body movement animation sequence, and finally determines the target digital person full-body movement video by using the lower body movement animation sequence and the current input content. The qualified lower body movement animation sequence in the initial digital person full-body movement video is reused as much as possible, the video updating efficiency is improved, and the resource consumption of the video updating work is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video generation technology, and in particular to a method, apparatus, device, medium, and product for updating digital human videos. Background Technology

[0002] With the introduction of the concept of "metaverse", the application of digital humans is becoming more and more widespread. For example, digital humans are used to simulate the voice, gestures, expressions, actions, and walking style of product presenters, generating three-dimensional visual (referred to as "3D") full-body moving videos to replace product presenters in explaining products.

[0003] The method described in CN119169157A for generating 3D digital human full-body moving videos involves determining the lower body movement animation of the digital human based on dialogue text and video script, arranging multiple lower body movement animations in chronological order to obtain a complete lower body movement animation sequence, and then generating the 3D digital human full-body moving video based on this complete sequence. If the user needs to change the dialogue text or video script after generating the 3D digital human full-body moving video, the aforementioned video generation process needs to be completely re-executed based on the modified dialogue text and video script to update the 3D digital human full-body moving video. However, each generation of a 3D digital human full-body moving video consumes a significant amount of hardware resources. If the user frequently modifies the dialogue text or video script and continuously generates new 3D digital human full-body moving videos, the consumption of hardware resources will increase exponentially, resulting in high video update costs. Secondly, for the same text information, the generation of the lower body movement animation sequence has a certain degree of randomness, and the difference between the 3D digital human full-body moving videos before and after the update may be quite significant, resulting in poor video relevance and affecting the user's video viewing experience.

[0004] Therefore, designing a low-cost, high-efficiency video update method with minimal changes in qualified content is one of the urgent problems to be solved. Summary of the Invention

[0005] This invention provides a method, apparatus, device, medium, and product for updating digital human videos. The method determines a video adjustment scheme based on the modification of the dialogue text, thereby determining a lower body movement animation sequence. Then, the lower body movement animation sequence and the current input content are used to determine the target digital human full-body movement video. The method reuses qualified lower body movement animation sequences from the initial digital human full-body movement video as much as possible, thereby improving video update efficiency while reducing the resource consumption of video update work.

[0006] According to one aspect of the present invention, a method for updating a digital human video is provided, the method comprising:

[0007] Get the current input content; the current input content includes the current dialogue text;

[0008] Based on the textual correspondence between the current dialogue text and the original dialogue text, the relevance between the original dialogue text and the current dialogue text is determined, as well as the video adjustment scheme for the initial full-body movement video of the digital human; wherein, the original dialogue text is used to generate the initial full-body movement video of the digital human.

[0009] The lower body movement animation sequence of the current dialogue text is determined based on the video adjustment scheme;

[0010] Generate a full-body moving video of the target digital human based on the lower body movement animation sequence and the current input content.

[0011] The digital human video update method of this invention compares the current dialogue text with the original dialogue text to obtain the change information of the dialogue text, and then determines the relevance of the dialogue text and the video adjustment scheme of the initial digital human full-body movement video. Then, it uses the video adjustment scheme to determine the lower body movement animation sequence corresponding to the current dialogue text, and uses the determined lower body movement animation sequence and the current input content to generate the target digital human full-body movement video. Determining the video adjustment scheme based on relevance is to reuse the qualified lower body movement animation sequence in the initial digital human full-body movement video as much as possible, reduce the number of lower body movement animation sequences that need to be regenerated, improve the video update efficiency, and also ensure that the video before and after the update has high relevance and low difference, improve the user's video viewing experience, and also reduce the resource consumption and video update cost of video update work. This solution addresses several issues: Firstly, re-executing the video generation process described in the previous embodiment after each information modification consumes significant hardware resources. Secondly, if users frequently modify dialogue text or video scripts and continuously generate new 3D digital human full-body movement videos, it increases hardware resource consumption, resulting in high video update costs and low efficiency. Thirdly, for the same text information, the generation of lower body movement animation sequences has a certain degree of randomness, leading to potentially significant differences between the 3D digital human full-body movement videos before and after updates, resulting in poor video relevance and a less than ideal user viewing experience.

[0012] According to another aspect of the present invention, a digital human video updating apparatus is provided. The digital human video updating apparatus is used to implement the digital human video updating method in any embodiment of the present invention. The apparatus includes:

[0013] The information determination module is used to obtain the current input content; the current input content includes the current dialogue text.

[0014] The scheme determination module is used to determine the relevance between the original and current dialogue texts based on the textual correspondence between the current and original dialogue texts, as well as the video adjustment scheme for the initial full-body movement video of the digital human; wherein, the original dialogue text is used to generate the initial full-body movement video of the digital human;

[0015] The sequence determination module is used to determine the lower body movement animation sequence of the current dialogue text based on the video adjustment scheme;

[0016] The video update module is used to generate a full-body moving video of the target digital human based on the lower body movement animation sequence and the current input content.

[0017] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0018] At least one processor; and a memory communicatively connected to the at least one processor;

[0019] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to execute the digital human video update method in any embodiment of the present invention.

[0020] According to another aspect of the present invention, a computer-readable storage medium is provided that stores computer instructions for causing a processor to execute and implement a method for updating a digital human video according to any embodiment of the present invention.

[0021] According to another aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements a method for updating digital human video according to any embodiment of the present invention.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in this invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart illustrating a method for updating digital human videos provided by the present invention;

[0025] Figure 2 This is a flowchart illustrating another method for updating digital human videos provided by the present invention;

[0026] Figure 3 This is a schematic diagram of the structure of a digital human video updating device provided by the present invention;

[0027] Figure 4 This is a schematic diagram of the structure of an electronic device provided by the present invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," "initial," "intermediate," "candidate," "alternate," "target," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] Figure 1 This is a flowchart illustrating a method for updating digital human videos provided by the present invention. This embodiment is applicable to updating digital human videos with low cost, high efficiency, and high reusability. The method can be executed by the digital human video updating device provided by the present invention. This device can be implemented in hardware and / or software. In a specific embodiment, the device can be integrated into an electronic device. The following embodiments will illustrate this using the integration of the device into an electronic device as an example. (Refer to...) Figure 1 The method specifically includes the following steps:

[0031] S101. Get the current input content; the current input content includes the current dialogue text.

[0032] The current input content can be understood as the information that the user inputs into the device generating and updating the video (such as a computer) through specific means (e.g., keyboard, mouse, touchscreen, sound capture device, etc.) during the video update process, so that the device can adjust the video that the user is not satisfied with. The current input content is the information input by the user into the device for updating the initial full-body moving digital human video. The initial full-body moving digital human video can be understood as the video obtained after the previous full-body moving digital human video generation / update was completed; that is, the full-body moving digital human video that needs adjustment in this video update, including but not limited to product introduction videos, informational broadcast videos, and knowledge popularization videos that need adjustment. It is worth noting that the digital human will move in specific ways and at specific times in the video. For example, when explaining a specific topic, the digital human moves to the front of the display screen and points to the specific content on the screen, making it easier for the audience to understand the content and attracting the attention of the video viewers. When explaining a specific topic, the digital human moves to the side of the presentation platform, fully displaying the screen to the audience so that they can view the content on the screen, aiming to increase interactivity and improve the overall presentation effect. Under normal circumstances, the initial full-body moving video of the digital human can dynamically present the image of the digital human (the digital human responsible for explaining the product or knowledge) and video materials (static images, dynamic images, videos, documents, slides, etc.) to the user from different perspectives, distances, and movement methods.

[0033] The current input content includes the current dialogue text, which is generally obtained by modifying the original dialogue text according to user requirements. The original dialogue text can be understood as the content of the initial digital human's full-body movement video, and can be used to determine the audio of the initial digital human's full-body movement video. Generally, changes to the dialogue text (e.g., adding, deleting, or modifying text) will affect the timing of the digital human's movement in the video. The movement timing can be understood as the time point on the timeline corresponding to when the digital human starts moving. For example, if the original dialogue text is ABCDEFG and the current dialogue text is AABBCDEFG, assuming that the duration required for the digital human to pronounce each word is the same, the digital human in the initial full-body movement video starts moving when reading the fifth word (i.e., word E). After the dialogue text changes, if other factors are not considered, the digital human should still start moving when reading word E. At this time, the movement timing becomes the time point corresponding to the seventh word on the timeline. The methods for obtaining dialogue text in this invention are diverse. For example, it can be directly input by the user through a client, retrieved from the client's local storage, or downloaded from a server. The form of the dialogue text is also diverse. For example, it can be text, audio (which can be converted into text through speech recognition), or an image (which can be extracted by text recognition), or even a PPT file (from which text can be extracted as dialogue text). The dialogue text can contain only one text segment or multiple text segments. The specific number of text segments and the method of obtaining the dialogue text are related to the video generation logic, and this invention does not limit them.

[0034] To ensure consistent audiovisual effects and maintain visual and auditory consistency in the digital human's appearance, the digital human's voice can be selected based on its image. For example, a female digital human's voice could be a mezzo-soprano or soprano, while a male digital human's voice could be a baritone or tenor. Furthermore, to make the generated video more engaging and vivid, deepening the audience's understanding of the content and leaving a lasting impression, the input content of this invention can also include images. Images can be pictures and / or videos, presented in the generated video as a way to showcase the content. Combined with the digital human's language, tone, and actions, this significantly enhances the video's persuasive effect.

[0035] S102. Based on the textual correspondence between the current dialogue text and the original dialogue text, determine the relevance between the original dialogue text and the current dialogue text, as well as the video adjustment scheme for the initial digital human full-body moving video.

[0036] The original dialogue text is used to generate the initial full-body movement video of the digital human. The relevance between the original and current dialogue text is determined by comparing each character in the current text with a character in the original text, resulting in a similarity score. For example, if 3 out of 100 characters are different, the similarity score is 97%; if 2 out of 10 characters are different, the similarity score is 80%. This similarity score is used to determine the video adjustment scheme, which instructs how to process the lower body movement animation sequence of the digital human in the initial full-body movement video to obtain the corresponding lower body movement animation sequence for the current dialogue text. For example, for parts with minimal variation (high similarity), a lower body movement animation reuse scheme is used, reusing the lower body movement animation sequence from the initial full-body movement video to quickly and cost-effectively obtain a new animation sequence. For parts with significant variation (low similarity), a lower body movement animation generation scheme is used to regenerate the lower body movement animation sequence, ensuring visual quality and avoiding phenomena such as slippage or abnormal movement.

[0037] This invention sets up different video update schemes to specifically handle different types of changes in dialogue text, with the aim of saving video update resources and improving video update efficiency while ensuring the visual effect of the video.

[0038] S103. Determine the lower body movement animation sequence of the current dialogue text based on the video adjustment scheme.

[0039] The video generation process includes: 1) determining the corresponding audio by combining the video script and dialogue text; 2) determining the position of the digital human's movement mode on the audio timeline; 3) determining the lower body movement animation sequence of the digital human by combining the start position, end position, and position of the digital human's movement mode on the audio timeline, thereby obtaining the full-body movement video of the digital human.

[0040] In generating video, this invention first generates corresponding audio based on the video script (including pause markers and preset action markers corresponding to the dialogue text) and the dialogue text. The audio exists on a timeline, namely the audio timeline. Different moments on the audio timeline correspond to different text, animations, materials, and other information. It can be understood that the full-body moving video of the digital human generated by this application mainly involves the digital human explaining within the video frame. Audio is the main component of the full-body moving video of the digital human. Therefore, the audio timeline can serve as the timeline of the full-body moving video of the digital human, and is a unified time scale for various types of information such as audio, video, animation, and video materials, also known as the "timeline".

[0041] The current input also includes the current video script. Based on the current dialogue text and the current video script, the timeline of the updated video can be determined. Using the new timeline as a time scale, and combining the textual correspondence between the current dialogue text and the original dialogue text, the movement animation corresponding to the current dialogue text can be determined. Each movement animation has a different playback start point and movement duration on the timeline. By splicing the movement animations according to their playback start points and movement durations, the lower body movement animation sequence corresponding to the current dialogue text can be obtained.

[0042] On the one hand, when the original dialogue text and the current dialogue text are highly correlated, the video adjustment scheme is a lower body movement animation reuse scheme, and the lower body movement animation sequence of the current dialogue text is determined to be the lower body movement animation sequence of the original dialogue text. On the other hand, when the original dialogue text and the current dialogue text are not highly correlated, the video adjustment scheme is a lower body movement animation generation scheme, and the lower body movement animation sequence is generated according to the animation sequence generation scheme, and the generated lower body movement animation sequence is determined to be the lower body movement animation sequence of the current dialogue text.

[0043] S104. Generate a full-body moving video of the target digital human based on the lower body movement animation sequence and the current input content.

[0044] Generating a full-body moving video of a target digital human based on a lower-body movement animation sequence and the current input content can be understood as using a video generation model to process the lower-body movement animation sequence and the current input content to obtain the full-body moving video of the target digital human. For example: 1) Using a video generation model to process the current input content and the lower-body movement animation sequence to obtain a full-body moving animation sequence of the digital human, and then rendering the full-body moving animation sequence of the digital human to obtain the full-body moving video of the target digital human; 2) Using a video generation model to process the lower-body movement animation sequence and the current input content to obtain an upper-body movement animation sequence of the digital human, splicing and merging the upper-body movement animation sequence and the lower-body movement animation sequence of the digital human to obtain a full-body moving sequence of the target digital human, and then rendering the full-body moving sequence of the target digital human to obtain the full-body moving video of the target digital human.

[0045] The technical solution of the above embodiment compares the current dialogue text with the original dialogue text to obtain the change information of the dialogue text, and then determines the relevance of the dialogue text and the video adjustment scheme of the initial digital human full-body moving video. Then, it uses the video adjustment scheme to determine the lower body moving animation sequence corresponding to the current dialogue text, and uses the determined lower body moving animation sequence and the current input content to generate the target digital human full-body moving video. Determining the video adjustment scheme based on relevance is to reuse the qualified lower body moving animation sequence in the initial digital human full-body moving video as much as possible, reduce the number of lower body moving animation sequences that need to be regenerated, improve the video update efficiency, and also ensure that the video before and after the update has high relevance and low difference, improve the user's video viewing experience, and also reduce the resource consumption and video update cost of video update work. This solution addresses several issues: Firstly, re-executing the video generation process described in the previous embodiment after each information modification consumes significant hardware resources. Secondly, if users frequently modify dialogue text or video scripts and continuously generate new 3D digital human full-body movement videos, it increases hardware resource consumption, resulting in high video update costs and low efficiency. Thirdly, for the same text information, the generation of lower body movement animation sequences has a certain degree of randomness, leading to potentially significant differences between the 3D digital human full-body movement videos before and after updates, resulting in poor video relevance and a less than ideal user viewing experience.

[0046] Figure 2 This is a flowchart illustrating another method for updating digital human videos provided by the present invention. Based on the above embodiments, this embodiment provides a preferred video updating method, specifically, as follows: Figure 2 As shown, the method includes:

[0047] S201. Get the current input content; the current input content includes the current dialogue text.

[0048] Similar to S101, the current input content is the information input by the user into the computer (the device that generates and updates the video) during the video update process, through a specific method (e.g., keyboard, mouse, touchscreen, sound capture device, etc.), so that the computer can adjust the video that the user is not satisfied with. The current input content is the information input by the user into the computer for updating the initial digital human full-body moving video. The initial digital human full-body moving video is the video obtained after the previous digital human full-body moving video generation / update was completed; that is, the digital human full-body moving video that needs adjustment in this video update, including but not limited to product introduction videos, informational broadcast videos, and knowledge popularization videos that need adjustment. The current input content includes the current script text, which is obtained by modifying the original script text according to the user's needs. The original script text is the presentation content of the initial digital human full-body moving video and can be used to determine the presentation audio of the initial digital human full-body moving video.

[0049] S202. According to the pre-set text segmentation rules, the original dialogue text and the current dialogue text are segmented to obtain at least one original dialogue text fragment and at least one current dialogue text fragment.

[0050] To improve video update efficiency, this invention can divide the original and current dialogue texts and process them in parallel according to their types. Pre-defined text division rules serve as the basis for dividing the dialogue text content, including but not limited to division by sentence and division by paragraph (i.e., natural paragraphs). The divided current dialogue text can be divided into at least one current dialogue text segment, and the original dialogue text can be divided into at least one original dialogue text segment. There is a one-to-one correspondence between the current dialogue text segment and the original dialogue text segment. The number of divided dialogue text segments is related to the division criteria and the amount of dialogue text, and this invention does not impose any limitations on this.

[0051] S203. Process at least one current dialogue text segment and at least one current dialogue text segment to obtain at least one text relevance and at least one video segment adjustment scheme.

[0052] Processing at least one current dialogue text fragment and at least one current dialogue text fragment yields at least one text relevance score and at least one video clip adjustment scheme. This can be understood as comparing each current dialogue text fragment with its corresponding original dialogue text fragment to obtain the text relevance score between each current dialogue text fragment and its corresponding original dialogue text fragment, as well as the corresponding video clip adjustment scheme. Specifically, the text relevance score can be used to indicate the type of video clip adjustment scheme. For example, when the text relevance score is high, it is considered that the current dialogue text fragment and its corresponding original dialogue text fragment are relatively similar, and the video clip adjustment scheme is a lower body movement animation reuse scheme, reusing the lower body movement animation sequence corresponding to the original dialogue text fragment as much as possible. When the text relevance score is low, it is considered that the current dialogue text fragment and its corresponding original dialogue text fragment are significantly different, and the reuse method cannot provide a good video effect; the video clip adjustment scheme is a lower body movement animation generation scheme.

[0053] For any set of original and current dialogue text fragments, the original and current dialogue text fragments are processed to obtain text relevance and video clip adjustment schemes. This includes: calculating the text editing distance between the original and current dialogue text fragments and determining whether the text editing distance is less than a preset editing distance threshold; if the text editing distance is less than the preset editing distance threshold, the text relevance is determined to be similar, and the video clip adjustment scheme is a lower body motion animation reuse scheme; if the text editing distance is not less than the preset editing distance threshold, the text relevance is determined to be dissimilar, and the video clip adjustment scheme is a lower body motion animation generation scheme.

[0054] Edit distance refers to the minimum number of single-character editing operations (e.g., insertion, deletion, replacement) required to convert one string into another. Taking text paragraph comparison as an example, text edit distance treats each paragraph as a string and then compares the words of the current text paragraph with those of the original text paragraph, calculating the distance between them. Similarly, taking sentence comparison as an example, each sentence is treated as a string, and the words of the current text paragraph are compared with those of the original text sentence, calculating the distance between them. The edit distance threshold can be understood as a value that measures the degree of change between the current and original dialogue text segments (which can represent the correlation between the two). This is used to determine the video update strategy. For example, for current and original dialogue text segments with small changes (text edit distance less than the preset edit distance threshold), the text correlation between the two is considered similar. The lower body movement animation reuse method is used to determine the lower body movement animation sequence of the original dialogue text, and then the lower body movement animation sequence of the current dialogue text is obtained, saving processing resources and improving the efficiency of animation sequence determination. For current and original dialogue text segments with large changes (text edit distance not less than the preset edit distance threshold), the text correlation between the two is considered very different. The lower body movement animation regeneration method is needed to obtain the lower body movement animation sequence of the current dialogue text to ensure video quality. The goal is to save processing resources as much as possible while ensuring the video update effect.

[0055] S204. In response to the lower body movement animation reuse scheme, process the current dialogue text fragments with similar text relevance to obtain a reused lower body movement animation.

[0056] S204 utilizes a lower body movement animation reuse scheme to process all current dialogue text fragments with similar text relevance and their corresponding original dialogue text fragments, reusing the lower body movement animations corresponding to the original dialogue text fragments as much as possible, thus obtaining the lower body movement animations of that part of the current dialogue text fragments, i.e., reusable lower body movement animations.

[0057] Optionally, the current input also includes the current video script. Processing current dialogue text fragments with similar text relevance yields reusable lower body movement animations. This includes: determining the current video script and video reuse period corresponding to the current dialogue text fragments with similar text relevance; based on the current video script, determining whether the video reuse period satisfies the digital human movement condition; if the video reuse period satisfies the digital human movement condition, then determining the reusable lower body movement animation based on the video reuse period and the initial lower body movement animation sequence corresponding to the initial digital human full-body movement video; if the video reuse period does not satisfy the digital human movement condition, then deleting the lower body movement animation of the video reuse period.

[0058] The current video script corresponding to the current dialogue text fragment can be understood as the script used to describe the current dialogue text fragment among all current video scripts. The video script can be understood as the guiding content for video production; that is, the video script is used to generate video footage. Besides pause markers and preset action markers, the video script also includes information such as display content markers corresponding to the dialogue text, shot type markers, shot transition markers, and camera movement markers. The video reuse period corresponding to the current dialogue text fragment can be understood as the time period occupied by the current dialogue text fragment with similar text relevance on the timeline. Specifically, the timeline includes immovable time periods, movable time periods, and indeterminate time periods. During immovable time periods, the digital human cannot move; during movable time periods, the digital human can move; and during indeterminate time periods, whether the digital human can move depends on the duration of the digital human's movement. The conditions for digital human movement can be understood as the rules governing a digital human's ability to move. For example, if the video reuse period corresponding to the current dialogue text segment is a movable time segment on the timeline, or if the overlap with an undetermined time segment is less than the pre-set static measurement duration, it proves that the digital human can move while explaining that part of the dialogue text segment. The method of movement (i.e., the lower body movement animation sequence) needs to be confirmed based on the time segment corresponding to that part of the dialogue text segment and the initial digital human full-body movement video. Similarly, if the video reuse period corresponding to the current dialogue text segment is a non-movable time segment, or if the overlap with an undetermined time segment exceeds the pre-set static measurement duration, it is considered that the digital human cannot move, and the lower body movement animation needs to be deleted.

[0059] Specifically, the non-movable time periods include: 1) time periods that overlap with preset actions (key actions, silence), full-screen shots, and picture-in-picture shots; 2) time periods that overlap with medium shots, close-ups, extreme close-ups, etc. for a duration exceeding a first preset threshold (e.g., 1s, 2s, etc.); 3) time periods that overlap with panoramic shots for a duration exceeding a second preset threshold (e.g., 5s, 6s, etc.). The specific time thresholds are not limited in this invention.

[0060] Furthermore, based on the video reuse period and the initial lower body movement animation sequence corresponding to the initial digital human full-body movement video, the reused lower body movement animation is determined, including: based on the correspondence between the current dialogue text and the original dialogue text, candidate movement periods corresponding to the video reuse period are selected from the initial digital human full-body movement video; the movement information of the video reuse period is determined according to the start time, duration and mode of the candidate movement period, and the reused lower body movement animation is determined according to the movement information of the video reuse period.

[0061] The candidate video segment is the time period occupied by the original dialogue text segment corresponding to the current dialogue text segment in the initial full-body movement video of the digital human. For example, the initial full-body movement video of the digital human contains one minute and thirty seconds and has a total of three original dialogue text segments. The video segment of original dialogue text segment 1 is the first thirty seconds, the video segment of original dialogue text segment 2 is the middle thirty seconds, and the video segment of original dialogue text segment 3 is the last thirty seconds. The current dialogue text segment corresponds to original dialogue text segment 1. Then, the candidate movement segment is the time period from the start time to the thirtieth second in the initial full-body movement video of the digital human. The reused lower body movement animation is the movement animation obtained by migrating the movement animation in the candidate movement segment to the video reuse segment. If the starting point of the movement animation in the candidate movement segment of the initial full-body movement video of the digital human is the tenth second on the timeline, the movement duration is three seconds, and the movement method is from point A to point B, the movement information of the video reuse segment is used to indicate that the starting movement time of the digital human on the timeline is the tenth second, and the digital human needs to move from point A to point B within three seconds.

[0062] Taking a text paragraph or sentence as an example, reusing lower body movement animation can be achieved by directly using the lower body movement animation sequence of the original text segment corresponding to the current text segment. However, it is worth noting that when reusing lower body movement animation, if the movement duration of the animation covers the non-movable time period, for example, if the candidate movement period of the original text segment 2 is from the 30th to the 60th second, the video reuse period of the current text segment 2 is from the 35th to the 65th second, and the 40th to 45th seconds on the timeline are the non-movable time period, and the starting point of the animation on the timeline in the candidate movement period is the 35th second with a movement duration of 5 seconds, according to the text correspondence between the original text segment 2 and the current text segment 2, the starting point of the reused lower body movement animation on the timeline is the 40th second with a movement duration of 5 seconds, which covers the non-movable time period, this invention will directly delete the lower body movement animation. It is worth noting that after removing the lower body movement animation, it is necessary to determine whether the total number of movements in the overall video (the proportion of movements in the entire video) meets the standard. If it does not meet the standard, then a walking scheme needs to be adaptively added when generating the lower body movement animation to ensure the overall interactivity of the video.

[0063] Taking reusable text segments as an example, reusing lower-body movement animation involves using the correspondence between each character in the original and current dialogue segments. The lower-body movement animation corresponding to each character in the original dialogue segment is used as the lower-body movement animation for its corresponding character in the current dialogue segment. It's worth noting that, on the one hand, for text in the current dialogue segment that lacks lower-body movement animation, the lower-body movement animation corresponding to the characters before and after that character can be used to determine its lower-body movement animation. On the other hand, to ensure that the digital human's movement does not conflict with camera angles and preset actions (characteristic movements, key actions, silences, etc.), it is necessary to filter the digital human's movement periods according to the video script, identifying unsuitable periods and deleting the lower-body movement animation for those periods.

[0064] Unsuitable movement periods include: 1) Movement periods that overlap with preset actions, full-screen shots, and picture-in-picture shots; 2) Movement periods that overlap with medium shots, close-ups, extreme close-ups, etc. for a duration exceeding the first time threshold (e.g., 1s); 3) Movement periods that overlap with panoramic shots for a duration exceeding the second time threshold (e.g., 5s).

[0065] Taking reusable sentence segments as an example, reusing lower body movement animation involves using the lower body movement animation of the original dialogue text segment of the sentence as the lower body movement animation of the current dialogue text segment of the sentence. Reusing lower body movement animation using sentence segments also requires filtering movement periods, deleting lower body animations from unsuitable movement periods to ensure that the digital human's movement does not conflict with the camera angles and preset actions.

[0066] S205. In response to the lower body movement animation generation scheme, process the current dialogue text fragments with different text relevance to obtain a generated lower body movement animation.

[0067] S205 uses a lower body movement animation generation scheme to process all current dialogue text fragments with different text relevance, regenerates multiple lower body movement animations, selects the better lower body movement animation from them, and finally obtains the lower body movement animation of that part of the current dialogue text fragment, that is, a generative lower body movement animation.

[0068] Optionally, the current input content also includes the current video script. Processing the current dialogue text fragments with significantly different text relevance yields a generated lower body movement animation. This includes: determining the current video script and video generation period corresponding to the current dialogue text fragments with significantly different text relevance; determining whether the positions of the lower body movement animations corresponding to the start and end times of the video generation period are the same; if the positions of the lower body movement animations are different, then determining the generated lower body movement animation based on the animation position change information, preset lower body movement animation addition rules, and the duration of the video generation period; if the positions of the lower body movement animations are the same, then determining the generated lower body movement animation based on the current cumulative number of steps corresponding to the current dialogue text, preset walking conditions, and preset lower body movement animation addition rules.

[0069] The video generation period corresponding to the current dialogue text fragment can be understood as the time period on the timeline corresponding to the current dialogue text fragment with significantly different text relevance. The position of the lower body movement animation corresponding to the start time of the video generation period can be understood as the stopping position of a lower body movement animation before the video generation period, and the position of the lower body movement animation corresponding to the end time of the video generation period can be understood as the starting position of a lower body movement animation after the video generation period. The difference between the two positions represents the positional change of the digital human when narrating the current dialogue text fragment, that is, whether the digital human walks when narrating the current dialogue text fragment, and if so, the corresponding walking mode can also be obtained.

[0070] The preset rules for adding lower body movement animations indicate when to add animations, such as the interval between two animations or the frequency of movement. Specifically, if the positions of the lower body movement animations at the start and end of the video generation period are different, it indicates that the digital human has moved. In this case, the lower body movement animation sequence can be added at the appropriate time based on the duration required to explain the current dialogue text segment and the preset rules for adding lower body movement animations. If the positions of the lower body movement animations at the start and end of the video generation period are the same, it means that the digital human may or may not have moved. The lower body movement animation sequence corresponding to the current dialogue text segment needs to be determined based on the current cumulative number of steps, the pre-set walking conditions, and the rules for adding lower body movement animations. For example, if the current cumulative number of steps (total number of steps in the current dialogue text) is greater than the number of steps in the pre-set walking conditions, then it is determined that no lower body movement animation sequence needs to be added, and the lower body movement animation sequence corresponding to the current dialogue text segment is static. If the current cumulative number of steps is less than the number of steps in the pre-set walking conditions, then it is necessary to determine the number of missing steps in the lower body movement animation sequence (the difference between the current cumulative number of steps and the pre-set number of steps). Using the movement animation addition rules, the dialogue text in the current dialogue text segment suitable for adding a lower body movement animation sequence is determined, and the starting point and duration of the movement animation are determined in combination with the time period corresponding to the dialogue text on the timeline, so as to add the lower body movement animation.

[0071] S206. Using the time markers of the motion animation, the reusable lower body motion animation and the generative lower body motion animation are spliced ​​together to obtain the lower body motion animation sequence of the current dialogue text.

[0072] The time stamps in motion animations can be understood as the start or end times of the animations, indicating their sequential order so that they can be stitched together to form a complete lower-body motion animation sequence. The specific stitching method involves sorting the motion animations according to their time stamps and then interleaving them together to obtain the complete lower-body motion animation sequence.

[0073] S207. Input the lower body movement animation sequence and the current input content into the full-body animation sequence generation model for processing to obtain the full-body movement animation sequence.

[0074] The full-body animation sequence generation model can be understood as a pre-trained algorithm that can generate full-body movement animation sequences. By inputting the lower body movement animation sequence and the current input content into the full-body animation sequence generation model, the full-body movement animation sequence can be output.

[0075] It's worth noting that a complete digital human lower body animation sequence includes not only lower body movement animation but also lower body standing animation (static standing). This means the lower body movement animation sequence typically only corresponds to a portion of the timeline. Therefore, the full-body animation sequence generation model first processes the lower body movement animation sequence, filling in the standing animation segments on the timeline that don't correspond to lower body movement animation to obtain a complete lower body animation sequence. Then, based on the lower body animation sequence and the current input content, it generates the full-body movement animation sequence.

[0076] The space between two adjacent lower body movement animations can be filled using several pre-designed lower body standing animations. These different lower body standing animations have slight differences, and each animation has a different duration. When filling the video segment between two lower body movement animations, it's necessary to ensure that the digital human's standing posture is the same as the ending action of the previous lower body movement animation and the starting action of the next. Under this premise, a lower body standing animation is randomly selected and concatenated with the previous one. If this doesn't fill the video segment between the two lower body movement animations, another lower body standing animation is randomly selected and concatenated with the previous one, until the video segment between the two lower body movement animations is filled. If the duration of a randomly selected lower body standing animation exceeds the video segment between the two lower body movement animations, the lower body standing animation is trimmed based on that video segment, achieving the filling of lower body standing animations of arbitrary length.

[0077] Furthermore, for the lower body movement animation sequence composed of reusable lower body movement animations, determining the full-body movement animation sequence also includes: determining a specified upper body movement animation based on the current video script corresponding to the current dialogue text; using the reuse method of the lower body movement animation sequence, selecting reusable upper body movement animations from the upper body movement animation sequence of the initial digital human full-body movement video; using the time markers of the movement animations, splicing the specified upper body movement animations and the reusable upper body movement animations to obtain the upper body movement animation sequence; and generating the full-body movement animation sequence based on the lower body movement animation sequence and the upper body movement animation sequence.

[0078] The current video script may contain specific action restrictions or special vocabulary markers. When the digital human interprets text with these markers, it needs to perform specific actions. For example, it needs to smile when saying "Hello everyone," wave when saying "Goodbye," and point to the screen when saying "Please look at the big screen." Actions directly specified by the current video script that the digital human must produce are designated upper-body movement animations. Reusing lower-body movement animation sequences allows for the selection of reusable upper-body movement animations from the initial full-body movement video. This can be understood as parsing the reusable upper-body movement animations corresponding to the reused lower-body movement animations from the initial full-body movement video based on the correspondence between the original and current dialogue text, combined with the timeline. The splicing method for upper-body movement animation sequences is similar to that for lower-body movement animation sequences and will not be elaborated upon here. Generating a full-body animation sequence based on lower-body and upper-body movement animation sequences can be understood as combining these sequences to obtain the full-body animation sequence. For example, an animation sequence combination model can be used to process the lower-body and upper-body movement animation sequences to output the full-body animation sequence. The method of generating upper-body movement animation sequences by combining "specification" and "reuse" can minimize the cost of generating full-body animation sequences while meeting user needs.

[0079] S208. Generate a full-body motion video of the target digital human based on the full-body motion animation sequence.

[0080] Generating a full-body moving video of a target digital human based on a full-body moving animation sequence can be understood as combining the full-body moving animation sequence of the digital human with the audio information generated from the current input content to obtain a complete full-body moving video of the digital human.

[0081] The main objective of this invention is to reuse lower-body movement animation sequences from previous sequences as much as possible when generating them after the user edits the input content, thereby reducing the workload of video updates and the consumption of hardware resources. For parts that cannot be directly used, the lower-body movement animation sequences are regenerated to fill in the gaps, minimizing the differences between the two sequences generated before and after editing. This ensures minimal change in the display effect of the 3D digital human video before and after the update, maintaining video consistency while improving update efficiency.

[0082] The technical solution of the above embodiment compares the current dialogue text with the original dialogue text to obtain the change information of the dialogue text, and then determines the relevance of the dialogue text and the video adjustment scheme of the initial digital human full-body moving video. Then, it uses the video adjustment scheme to determine the lower body moving animation sequence corresponding to the current dialogue text, and uses the determined lower body moving animation sequence and the current input content to generate the target digital human full-body moving video. Determining the video adjustment scheme based on relevance is to reuse the qualified lower body moving animation sequence in the initial digital human full-body moving video as much as possible, reduce the number of lower body moving animation sequences that need to be regenerated, improve the video update efficiency, and also ensure that the video before and after the update has high relevance and low difference, improve the user's video viewing experience, and also reduce the resource consumption and video update cost of video update work. This solution addresses the issues of consuming significant hardware resources by re-executing the video generation process described in the previous embodiment after each information modification; the increased hardware resource consumption and low video update efficiency caused by users frequently modifying dialogue text or video scripts and continuously generating new 3D digital human full-body movement videos; the randomness in generating lower body movement animation sequences for the same text information, potentially leading to noticeable differences between the updated and unupdated 3D digital human full-body movement videos; and the resulting poor video relevance and user viewing experience.

[0083] Figure 3 This is a schematic diagram of the structure of a digital human video updating device provided by the present invention. Figure 3 As shown, the device includes: an information determination module 301, a scheme determination module 302, a sequence determination module 303, and a video update module 304.

[0084] The information determination module 301 is used to obtain the current input content; the current input content includes the current dialogue text.

[0085] The scheme determination module 302 is used to determine the relevance between the original dialogue text and the current dialogue text, as well as the video adjustment scheme for the initial digital human full-body moving video, based on the text correspondence between the current dialogue text and the original dialogue text; wherein, the original dialogue text is used to generate the initial digital human full-body moving video.

[0086] The sequence determination module 303 is used to determine the lower body movement animation sequence of the current dialogue text based on the video adjustment scheme.

[0087] The video update module 304 is used to generate a full-body movement video of the target digital human based on the lower body movement animation sequence and the current input content.

[0088] Optionally, the scheme determination module 302 is specifically used to: divide the original dialogue text and the current dialogue text according to the pre-set text division rules to obtain at least one original dialogue text fragment and at least one current dialogue text fragment; process at least one current dialogue text fragment and at least one current dialogue text fragment to obtain at least one text relevance and at least one video segment adjustment scheme.

[0089] Optionally, for any set of original dialogue text fragments and current dialogue text fragments, the scheme determination module 302 is specifically used to: calculate the text editing distance between the original dialogue text fragments and the current dialogue text fragments, and determine whether the text editing distance is less than a preset editing distance threshold; if the text editing distance is less than the preset editing distance threshold, the text relevance is determined to be similar, and the video fragment adjustment scheme is a lower body movement animation reuse scheme; if the text editing distance is not less than the preset editing distance threshold, the text relevance is determined to be dissimilar, and the video fragment adjustment scheme is a lower body movement animation generation scheme.

[0090] Optionally, the sequence determination module 303 is specifically used for: responding to the lower body movement animation reuse scheme, processing current dialogue text fragments with similar text relevance to obtain a reused lower body movement animation; responding to the lower body movement animation generation scheme, processing current dialogue text fragments with different text relevance to obtain a generated lower body movement animation; and using the time marker of the movement animation, splicing the reused lower body movement animation and the generated lower body movement animation to obtain the lower body movement animation sequence of the current dialogue text.

[0091] Optionally, the current input content also includes the current video script. The sequence determination module 303 is specifically used to: determine the current video script and video reuse period corresponding to the current dialogue text fragment with similar text relevance; based on the current video script, determine whether the video reuse period meets the digital human movement condition; if the video reuse period meets the digital human movement condition, determine the reused lower body movement animation based on the video reuse period and the initial lower body movement animation sequence corresponding to the initial digital human full-body movement video; if the video reuse period does not meet the digital human movement condition, delete the lower body movement animation of the video reuse period.

[0092] Optionally, the sequence determination module 303 is specifically used to: based on the correspondence between the current dialogue text and the original dialogue text, filter out candidate movement periods corresponding to the video reuse period from the initial digital human full-body movement video; determine the movement information of the video reuse period according to the start time, duration and mode of movement of the candidate movement period, and determine the reused lower body movement animation according to the movement information of the video reuse period.

[0093] Optionally, the current input content also includes the current video script. The sequence determination module 303 is specifically used to: determine the current video script and video generation period corresponding to the current dialogue text fragments with different text relevance; determine whether the positions of the lower body movement animations corresponding to the start and end times of the video generation period are the same; if the positions of the lower body movement animations are different, then determine the generated lower body movement animation based on the change information of the animation position, the preset lower body movement animation addition rules, and the duration of the video generation period; if the positions of the lower body movement animations are the same, then determine the generated lower body movement animation based on the current cumulative number of walks corresponding to the current dialogue text, the preset walking conditions, and the preset lower body movement animation addition rules.

[0094] Optionally, the video update module 304 is specifically used to: input the lower body movement animation sequence and the current input content into the full-body animation sequence generation model for processing to obtain a full-body movement animation sequence; and generate a full-body movement video of the target digital human based on the full-body movement animation sequence.

[0095] Optionally, for a sequence of lower body movement animations composed of reusable lower body movement animations, the video update module 304 is further configured to: determine a specified upper body movement animation based on the current video script corresponding to the current dialogue text; filter reusable upper body movement animations from the upper body movement animation sequence of the initial digital human full-body movement video using the reuse method of the lower body movement animation sequence; splice the specified upper body movement animation and the reusable upper body movement animation using the time markers of the movement animations to obtain an upper body movement animation sequence; and generate a full-body movement animation sequence based on the lower body movement animation sequence and the upper body movement animation sequence.

[0096] The digital human video updating device provided in this embodiment can execute the digital human video updating method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0097] Figure 4 This is a schematic diagram of the structure of an electronic device provided by the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0098] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the read-only memory (ROM) 12 or loaded from storage unit 18 into the random access memory (RAM) 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0099] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0100] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the method for updating digital human video.

[0101] In some embodiments, the method for updating digital human video can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the method for updating digital human video described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the method for updating digital human video by any other suitable means (e.g., by means of firmware).

[0102] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0103] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0104] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0105] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0106] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0107] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0108] In one embodiment, the present invention also includes a computer program product comprising a computer program that, when executed by a processor, implements the method for updating digital human video according to any embodiment of the present invention.

[0109] In the implementation of a computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​as well as conventional procedural programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0110] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0111] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method of updating a digital human video, characterized by, The method comprises: acquiring current input content; the current input content comprises current script text; based on the text correspondence relationship between the current script text and the original script text, determining the relevance of the original script text and the current script text, and based on the relevance, determining a video adjustment scheme of an initial digital human full-body movement video; wherein the original script text is used to generate the initial digital human full-body movement video, if the relevance is high, the video adjustment scheme is a lower body movement animation reuse scheme, and if the relevance is low, the video adjustment scheme is a lower body movement animation generation scheme; based on the video adjustment scheme, determining a lower body movement animation sequence of the current script text; based on the lower body movement animation sequence and the current input content, generating a target digital human full-body movement video.

2. The method of claim 1, wherein, The method comprises: dividing the original script text and the current script text according to a pre-set text division rule to obtain at least one original script text segment and at least one current script text segment; processing the at least one original script text segment and the at least one current script text segment to obtain at least one text relevance, and based on the at least one text relevance, determining at least one video segment adjustment scheme.

3. The method of claim 2, wherein, For any group of original script text segments and current script text segments, processing the original script text segments and the current script text segments to obtain a text relevance, and based on the text relevance, determining a video segment adjustment scheme, comprises: calculating the text edit distance of the original script text segment and the current script text segment, and determining whether the text edit distance is less than a preset edit distance threshold; if the text edit distance is less than the preset edit distance threshold, it is determined that the text relevance is similar, and the video segment adjustment scheme is a lower body movement animation reuse scheme; if the text edit distance is not less than the preset edit distance threshold, it is determined that the text relevance is different, and the video segment adjustment scheme is a lower body movement animation generation scheme.

4. The method of claim 3, wherein, The method comprises: in response to the lower body movement animation reuse scheme, processing the current script text segments with similar text relevance to obtain a reuse type lower body movement animation; in response to the lower body movement animation generation scheme, processing the current script text segments with different text relevance to obtain a generated lower body movement animation; using the time identifier of the movement animation, splicing the reuse type lower body movement animation and the generated lower body movement animation to obtain a lower body movement animation sequence of the current script text.

5. The method of claim 4, wherein, The current input content further comprises a current video script, and the method comprises: determine, based on the current video script, whether the video reuse time period satisfies a digital person movement condition; if the video reuse time period satisfies the digital person movement condition, determine the reuse type lower body movement animation based on the video reuse time period and an initial lower body movement animation sequence corresponding to an initial digital person full body movement video; if the video reuse time period does not satisfy the digital person movement condition, delete the lower body movement animation of the video reuse time period. The determination of the reuse type lower body movement animation based on the video reuse time period and the initial lower body movement animation sequence corresponding to the initial digital person full body movement video comprises:

6. The method of claim 5, wherein, screening a candidate movement time period corresponding to the video reuse time period from the initial digital person full body movement video based on the correspondence between the current dialogue text and the original dialogue text; determining movement information of the video reuse time period according to a start movement time, a movement time length and a movement mode of the candidate movement time period, and determining the reuse type lower body movement animation according to the movement information of the video reuse time period. The current input content further comprises a current video script, and the processing of the current dialogue text segment with the text relevance degree being different to obtain the generated type lower body movement animation comprises:

7. The method of claim 4, wherein, determining a current video script corresponding to the current dialogue text segment with the text relevance degree being different and a video generation time period; determining a current video script corresponding to the current dialogue text segment with the text relevance degree being different and a video generation time period; determining whether the positions of the lower body movement animations corresponding to the start time and the end time of the video generation time period are the same; if the positions of the lower body movement animations are different, determining the generated type lower body movement animation based on change information of the animation positions, a preset lower body movement animation adding rule and a time length of the video generation time period; if the positions of the lower body movement animations are the same, determining the generated type lower body movement animation based on a current cumulative walking number corresponding to the current dialogue text, a preset walking condition and a preset lower body movement animation adding rule.

8. The method of claim 1, wherein, The generation of the target digital person full body movement video based on the lower body movement animation sequence and the current input content comprises: inputting the lower body movement animation sequence and the current input content into a full body animation sequence generation model to obtain a full body movement animation sequence; generating the target digital person full body movement video based on the full body movement animation sequence.

9. The method of claim 8, wherein, For the lower body movement animation sequence composed of the reuse type lower body movement animation, the determination of the full body movement animation sequence further comprises: determining a designated type upper body movement animation based on a current video script corresponding to the current dialogue text; screening a reuse type upper body movement animation from an upper body movement animation sequence of the initial digital person full body movement video by using a reuse mode of the lower body movement animation sequence; splicing the designated type upper body movement animation and the reuse type upper body movement animation by using time identifiers of movement animations to obtain an upper body movement animation sequence; generate the full-body movement animation sequence based on the lower-body movement animation sequence and the upper-body movement animation sequence.

10. An apparatus for updating a digital human video, characterized by comprising: The method for updating the digital human video of any one of claims 1 to 9, the apparatus for updating the digital human video comprising: an information determination module configured to obtain current input content, the current input content including current script text; a scheme determination module configured to determine a relevance of the original script text and the current script text based on a text correspondence relationship between the original script text and the current script text, and determine a video adjustment scheme of an initial digital human full-body movement video based on the relevance, wherein the original script text is used to generate the initial digital human full-body movement video, if the relevance is high, the video adjustment scheme is a lower-body movement animation reuse scheme, and if the relevance is low, the video adjustment scheme is a lower-body movement animation generation scheme; a sequence determination module configured to determine a lower-body movement animation sequence of the current script text based on the video adjustment scheme; a video update module configured to generate a target digital human full-body movement video based on the lower-body movement animation sequence and the current input content.

11. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the method for updating the digital human video of any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing a processor to implement the method for updating the digital human video of any one of claims 1 to 9 when executed by the processor.

13. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method for updating the digital human video of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Virtual human video generation method and device, equipment and storage medium

    CN119169157A

  • Text broadcasting method and device, electronic equipment and storage medium

    CN110941954A

  • Digital human generation method and system

    CN117726730A