Digital human high-quality video generation method and system based on data driving

By analyzing user interaction behavior and historical records, the interaction parameters of digital human videos are intelligently adjusted, solving the problem of lack of creativity and flexibility in digital human video generation, and improving video quality and user satisfaction.

CN120689477AActive Publication Date: 2025-09-23DAOYOUDAO TECH GRP CO LTD

Patent Information

Application Number
CN202511171899.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-09-23
Estimated Expiration
2045-08-21

AI Technical Summary

Technical Problem

The existing technology for generating digital human videos lacks creativity and flexibility, and cannot adapt to the actual characteristics and needs of the audience, affecting the viewing experience and satisfaction.

Method used

By obtaining feature videos of historical digital human videos and target videos, analyzing user interaction behaviors and historical search records, calculating the warning level, and intelligently adjusting interaction parameters to generate high-quality videos.

Benefits of technology

It improves the quality and attractiveness of digital human videos and enhances users' viewing experience and satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689477A_ABST
    Figure CN120689477A_ABST
Patent Text Reader

Abstract

The invention discloses a digital human high-quality video generation method and system based on data driving, and relates to the technical field of data analysis, and the method comprises the steps: obtaining a historical digital human video and a current target video, and extracting a feature video in the digital human video; obtaining a watching user watching the feature video, and obtaining a target user according to an interaction behavior and a watching duration between the watching user and the feature video; analyzing behaviors of the target user when the target user watches the feature video, and extracting audience users corresponding to the target video; and calculating the early warning degree of the target video according to the adjustment record of the parameter value of the interaction parameter in the video changed by the audience user, and performing intelligent prompting. The audience users of the digital human video are obtained through the feature video and the adjustment record, and the actual characteristics of the audience users are intelligently analyzed, so that the video content with high quality and strong attraction is effectively generated, and the watching experience and satisfaction of the user group watching the video are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data analysis, and in particular to a method and system for generating high-quality digital human videos based on data drive. Background Art

[0002] Digital Humans are virtual characters created through artificial intelligence technology. They integrate cutting-edge technologies such as 3D modeling, speech synthesis, and natural language processing. They are widely used in short video creation. As an emerging human-computer interaction medium and digital life form, they possess highly interactive capabilities and intelligent features. They have brought unprecedented changes to the production of spoken-word short videos and are gradually becoming a powerful tool in the hands of many creators. Currently, when creating short videos, video generation relies on preset templates, lacks creativity and flexibility, and generates digital human videos according to the fixed interaction parameters preset in the templates. Although there is a simple manual adjustment of the interaction parameters, it does not combine the characteristics of the audience and historical adjustment records for in-depth analysis, which will result in the digital human video being unable to adapt to the actual characteristics and needs of the audience, affecting the viewing experience and satisfaction of the user group watching the video. Summary of the Invention

[0003] The purpose of the present invention is to provide a method and system for generating high-quality digital human videos based on data drive, so as to solve the problems raised in the prior art.

[0004] To achieve the above object, the present invention provides the following technical solutions: A data-driven method for generating high-quality digital human videos comprises the following steps: Step S100: Acquire historical digital human videos generated using digital human video templates on the video creation platform, as well as a target video currently generated based on the video template; extract feature videos from the digital human video based on the type of video template corresponding to the target video and keywords in the text used to generate the target video; Step S200: Obtain users who watch the feature video on the video playback platform, extract the clips in which the digital human appears and speaks from the feature video, and obtain the target user based on the viewing time and interaction behavior of the user with the feature video; Step S300: Analyze the target user's pause and rewind behaviors when watching the feature video, and combine the titles of the target user's historical searches and playbacks to obtain the target user's keyword set; extract the target video's corresponding audience users from the target user based on the target video's corresponding content and keyword set; Step S400: Based on the adjustment records of the parameter values ​​of the interactive parameters in the video when the audience user watches the feature video, and the pending parameter values ​​of the interactive parameters currently to be applied to the target video, the warning level of the target video is calculated, and intelligent prompts are given to the target video and the pending parameter values ​​according to the warning level.

[0005] Furthermore, step S100 includes: Step S110: Obtain historical digital human videos generated using digital human video templates on the video creation platform, wherein each digital human video is a video that has been published on the video playback platform; obtain the target video V currently generated based on the video template Tem and the text content Tex; Get the type corresponding to each video template, including the style, theme and applicable scene of the video template, and generate a corresponding type set; use the type set of the video template Tem as S Tem , obtain the video template WTem corresponding to a certain digital human video W in history, and use the type set of the video template WTem as S WTem , according to the set S Tem and set S WTem The number of the same type N, the digital human video W is the first reliability X1=1-e -N , e is a natural constant; Step S120: extract all keywords in the text content Tex, remove the same keywords, and obtain a set S Tex , extract all the keywords in the text content WTex corresponding to the digital human video W, remove the same keywords, and get the set S WTex ; Get the set S Tex A keyword KW1 in the set S WTex The similarity between each keyword in and keyword KW1, the maximum similarity among them is used as the target value of keyword KW1, and then the set S is obtained. Tex The target value of each keyword in the , and add up to get the average value, which is the second reliability X2 of the digital human video W as the feature video; In this solution, the feature video is a video that can be used to select suitable and target users for the target video, that is, a video that is relatively similar to the target video. However, since not every video is a feature video, it is necessary to select the feature video. In this step, the feature video is obtained by analyzing the content of the video template: 1. Determine the first reliability of the digital human video W according to the type corresponding to the video template. Since X1=1-e -NAmong them, X1 increases as N increases, and when N > 0, the value of X1 ranges from 0 to 1. Then the value of the first reliability X1 ranges from 0 to 1. 2. Through the similarity of the corresponding keywords in the text content, if there is a greater similarity, then the text content of the digital human video W is more similar to the text content of the target video. Since calculating the similarity between two keywords is a natural language processing technology, the value ranges from 0 to 1, and the specific implementation process will not be elaborated here.

[0006] Pre-set the weights corresponding to the first reliability X1 and the second reliability X2 to obtain the total reliability X0; if X0 > K1 and |X1 - X2| < K2 are satisfied, where K1 is the first numerical threshold and K2 is the second numerical threshold, then the digital human video W is taken as the feature video.

[0007] Further, step S200 includes: Step S210: Obtain a certain viewing user WU who views a certain feature video VID on the video playback platform, and set the initial behavior value of the viewing user WU to 0. If the viewing user WU likes the feature video VID, the behavior value is increased by h1. If the user comments, the behavior value is increased by h2. If the user favorites, the behavior value is increased by h3, where h1 + h2 + h3 = 1, and then the final behavior value H is obtained. Step S220: Extract the segments where the digital human appears on camera and gives an oral broadcast from the feature video VID, and obtain the total duration D0 of the segments. According to the viewing duration D of the viewing user WU for the videos within the segments WU , and the obtained behavior value H, obtain the target degree G of the viewing user WU WU =H×D WU / D0. If the target degree G WU is greater than the preset degree threshold, then the viewing user WU is taken as the target user of the feature video VID, and then all the target users corresponding to each feature video are obtained.

[0008] Further, step S300 includes: Step S310: Establish a keyword set S with an empty element corresponding to a certain target user TU TU ; Obtain the playback progress time when the target user TU executes the pause action while viewing the corresponding feature video FV, extract the text of the statement spoken by the digital human at the playback progress time, and input the keywords in the statement text into the set S TU ; Obtain the initial playback progress time T2 before the target user TU executes the rewind action and the final playback progress time T1 after the rewind action while viewing the video FV. Among them, the time T1 is before T2. Extract all the statement texts spoken by the digital human from time T1 to time T2, and input the keywords in all the statement texts into the set S TU ; Get the title of the video that the target user TU searched and played in the past C days starting from the time when the target user TU watched the feature video FV, and input the keywords in the title into the set S TU Then we get the final keyword set S TU ; In this solution, the keywords in the keyword set are words that users are interested in. When watching a digital human video, if the user encounters content that the digital human says that interests them, they will usually pause to watch the subtitles or rewind to consider. In this case, the keywords of the corresponding sentence text during the pause and rewind can be entered into the keyword set. When users conduct historical searches, they are usually attracted to the titles of the videos they are searching for. This is because such titles are also of interest to users. In this case, the keywords in the titles can also be entered into the keyword set to obtain the final keyword set.

[0009] Step S320: The set S is obtained based on the text content Tex corresponding to the target video. Tex , get the set S Tex A keyword KW2 in the set S TU The similarity between each keyword in and keyword KW2 is calculated, and the maximum similarity is used as the target value of keyword KW2, and then the set S is obtained. Tex The target value of each keyword in is added and the average value is obtained to obtain the correlation degree Y. If the correlation degree Y is greater than the correlation degree threshold, the target user TU is regarded as the audience user of the target video.

[0010] Furthermore, step S400 includes: randomly extracting M adjustment records of the parameter values ​​of the interaction parameters in the video when the audience user watches the feature video, and obtaining the final parameter value after the interaction parameter of each adjustment record is changed, and obtaining the warning level of the target video according to the pending parameter value of the target video , where e is a natural constant, P V is the undetermined parameter value of the target video, P m It is the final parameter value corresponding to the mth adjustment record. If the warning level Z is greater than the warning level threshold, an intelligent prompt is given for the target video and the pending parameter value.

[0011] Interaction parameters include speech speed, speech volume, and video brightness. For the target video, the adjustment records of the target users are an important basis for determining whether the target video's pending parameter values ​​are reasonable. For example, the target video's target users are usually middle-aged and elderly people, and most of them adjust the speech volume to a higher level in order to watch the video more appropriately. In this case, when the target video's pending parameter values ​​deviate significantly from the actual parameter values ​​corresponding to the adjustment records, that is, |P V -P m| is larger, then the warning level should also be larger.

[0012] A data-driven high-quality digital human video generation system, including a feature video extraction module, a target user acquisition module, an audience user extraction module and an interaction parameter analysis module; Feature video extraction module: used to obtain historical digital human videos generated using digital human video templates on the video creation platform, as well as target videos currently generated based on video templates; extract feature videos from digital human videos based on the type of video template corresponding to the target video and the keywords in the text content used to generate the target video; Target user acquisition module: This module is used to acquire users who watch the featured video on the video playback platform. It extracts the clips of the digital human appearing and speaking from the featured video, and then determines the target user based on the viewing time and interaction behavior of the user with the featured video. Audience user extraction module: This module analyzes the target user's pause and rewind behaviors when watching a featured video, and combines the titles of the videos that the target user has searched and played in the past to obtain the target user's keyword set. Based on the target video's corresponding copy and keyword set, the module extracts the audience user corresponding to the target video from the target user. Interaction parameter analysis module: used to calculate the warning level of the target video based on the adjustment records of the parameter values ​​of the interaction parameters in the video when the audience users watch the feature video, as well as the pending parameter values ​​of the interaction parameters currently to be applied to the target video, and provide intelligent prompts for the target video and the pending parameter values ​​based on the warning level.

[0013] Furthermore, the feature video extraction module includes a first reliability calculation unit and a feature video extraction unit; The first reliability calculation unit is used to obtain digital human videos generated in the past using digital human video templates on the video creation platform, and obtain the target video currently generated based on the video template and copy content; obtain the type corresponding to each video template, and generate a corresponding type set to obtain the first reliability of the digital human video as a feature video; Feature video extraction unit: used to extract all keywords in the copy content, establish a set, and obtain the second reliability of the digital human video as a feature video; based on the first reliability and the second reliability, obtain the total reliability, and then obtain the feature video in the digital human video based on the total reliability.

[0014] Furthermore, the audience user extraction module includes a keyword set establishment unit and an audience user extraction unit; Keyword set establishment unit: used to establish a keyword set with an empty element corresponding to the target user, analyze the target user's pause and rewind behavior when watching the feature video, and combine the titles of the target user's historical search and playback videos to obtain the target user's final keyword set; Audience user extraction unit: used to extract the audience users corresponding to the target video from the target users based on the copy content and keyword set corresponding to the target video.

[0015] Compared with the existing technology, the present invention has the following advantages: It provides a data-driven method and system for generating high-quality digital human videos, including: obtaining historical digital human videos and current target videos, extracting feature videos from the digital human videos; obtaining users who watch the feature videos, and obtaining target users based on the interaction behavior and viewing time between the users and the feature videos; analyzing the target users' behavior while watching the feature videos, and extracting the target video's corresponding audience users; and calculating the target video's warning level based on the adjustment records of the target users' changes to the parameter values ​​of the interactive parameters in the videos, and providing intelligent prompts. The present invention obtains the target users of the digital human videos through the feature videos and adjustment records, and intelligently analyzes the actual characteristics of the target users, effectively generating high-quality, attractive video content, and improving the viewing experience and satisfaction of the user groups watching the videos. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A flow chart of a method for generating high-quality digital human videos based on data-driven methods according to the present invention; Figure 2 This is a structural diagram of a data-driven high-quality digital human video generation system of the present invention. DETAILED DESCRIPTION

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0018] Example: Figure 1 As shown, the present invention provides a technical solution for a data-driven method for generating high-quality digital human videos, comprising the following steps: Step S100: Acquire historical digital human videos generated using digital human video templates on the video creation platform, as well as a target video currently generated based on the video template; extract feature videos from the digital human video based on the type of video template corresponding to the target video and keywords in the text used to generate the target video; Step S110: Obtain historical digital human videos generated using digital human video templates on the video creation platform, wherein each digital human video is a video that has been published on the video playback platform; obtain the target video V currently generated based on the video template Tem and the text content Tex; Get the type corresponding to each video template, including the style, theme and applicable scene of the video template, and generate a corresponding type set; use the type set of the video template Tem as S Tem , obtain the video template WTem corresponding to a certain digital human video W in history, and use the type set of the video template WTem as S WTem , according to the set S Tem and set S WTem The number of the same type N, the digital human video W is the first reliability X1=1-e -N , e is a natural constant; Step S120: extract all keywords in the text content Tex, remove the same keywords, and obtain a set S Tex , extract all the keywords in the text content WTex corresponding to the digital human video W, remove the same keywords, and get the set S WTex ; Get the set S Tex A keyword KW1 in the set S WTex The similarity between each keyword in and keyword KW1, the maximum similarity among them is used as the target value of keyword KW1, and then the set S is obtained. Tex The target value of each keyword in the , and add up to get the average value, which is the second reliability X2 of the digital human video W as the feature video; In this solution, the feature video is a video that can be used to select suitable and target users for the target video, that is, a video that is relatively similar to the target video. However, since not every video is a feature video, it is necessary to select the feature video. In this step, the feature video is obtained by analyzing the content of the video template: 1. Determine the first reliability of the digital human video W according to the type corresponding to the video template. Since X1=1-e -NAmong them, X1 increases as N increases, and when N > 0, the value of X1 ranges from 0 to 1. Then the value of the first reliability X1 ranges from 0 to 1; 2. Through the similarity of the corresponding keywords in the text content, if there is a greater similarity, then the text content of the digital human video W is more similar to the text content of the target video. Since obtaining the similarity between two keywords is a natural language processing technology and its value ranges from 0 to 1, the specific implementation process will not be elaborated.

[0019] Pre-set the weights corresponding to the first reliability X1 and the second reliability X2 to obtain the total reliability X0; if X0 > K1 and |X1 - X2| < K2 are satisfied, where K1 is the first numerical threshold and K2 is the second numerical threshold, then the digital human video W is used as the feature video.

[0020] In this embodiment, let the weights corresponding to the first reliability X1 and the second reliability X2 be WX1 and WX2 respectively, and obtain the total reliability X0 = WX1 × X1 + WX2 × X2. Here, setting X0 > K1 and |X1 - X2| < K2 is because the total reliability X0 represents the overall reliability value and is an important factor in measuring the digital human video W as a feature video. When the total reliability X0 is too small, it means that the digital human video W is unreliable and cannot be used as a feature video. And |X1 - X2| represents the absolute value of the difference between the two reliabilities. When one is large and the other is small, it means that the bias of the reliability is unbalanced and cannot be used as a feature video. Since the values of X1 and X2 both range from 0 to 1, in this embodiment, let K1 be 0.6 and K2 be 0.3.

[0021] Step S200: Obtain the viewing users who watch the feature video on the video playback platform, extract the segments where the digital human appears on camera and conducts oral broadcasts from the feature video, and obtain the target users according to the viewing duration of the viewing users and their interaction behaviors with the feature video; Step S210: Obtain a certain viewing user WU who watches a certain feature video VID on the video playback platform, set the initial behavior value of the viewing user WU to 0. If the viewing user WU likes the feature video VID, the behavior value is increased by h1. If the user comments, the behavior value is increased by h2. If the user collects, the behavior value is increased by h3, and h1 + h2 + h3 = 1, and then obtain the final behavior value H; Step S220: Extract the segments where the digital human appears on camera and conducts oral broadcasts from the feature video VID, and obtain the total duration D0 of the segments. According to the viewing duration D of the viewing user WU for the video within the segments WU , and the obtained behavior value H, obtain the target degree G of the viewing user WU WU =H×D WU / D0. If the target degree G WUIf the degree is greater than a preset threshold, the viewing user WU is taken as the target user of the feature video VID, and then all target users corresponding to each feature video are obtained.

[0022] The target user obtained here is determined based on the interactive behavior and viewing time with the feature video. When a user likes, comments, and collects a feature video and watches it for a long time, it means that the content of the feature video is suitable for the user, so the user can be used as the target user of the feature video; however, the user is not necessarily the target video's audience user, so the following analysis is needed to determine the target video's audience user.

[0023] Step S300: Analyze the target user's pause and rewind behaviors when watching the feature video, and combine the titles of the target user's historical searches and playbacks to obtain the target user's keyword set; extract the target video's corresponding audience users from the target user based on the target video's corresponding content and keyword set; Step S310: Create a keyword set S with an empty element corresponding to a target user TU TU ; Get the playback progress time when the target user TU performs the pause action when watching the corresponding feature video FV, extract the sentence text spoken by the digital person during the playback progress time, and input the keywords in the sentence text into the set S TU middle; Get the initial playback progress time T2 before the user TU performs the rewind action when watching the video FV, and the final playback progress time T1 after the rewind action, where time T1 is before T2. Extract all the sentences spoken by the digital person from time T1 to time T2, and input the keywords in all the sentence texts into the set S TU middle; Get the title of the video that the target user TU searched and played in the past C days starting from the time when the target user TU watched the feature video FV, and input the keywords in the title into the set S TU Then we get the final keyword set S TU ; In this solution, the keywords in the keyword set are words that users are interested in. When watching a digital human video, if the user encounters content that the digital human says that interests them, they will usually pause to watch the subtitles or rewind to consider. In this case, the keywords of the corresponding sentence text during the pause and rewind can be entered into the keyword set. When users conduct historical searches, they are usually attracted to the titles of the videos they are searching for. This is because such titles are also of interest to users. In this case, the keywords in the titles can also be entered into the keyword set to obtain the final keyword set.

[0024] Step S320: The set S is obtained based on the text content Tex corresponding to the target video. Tex , get the set S Tex A keyword KW2 in the set S TU The similarity between each keyword in and keyword KW2 is calculated, and the maximum similarity is used as the target value of keyword KW2, and then the set S is obtained. Tex The target value of each keyword in is added and the average value is obtained to obtain the correlation degree Y. If the correlation degree Y is greater than the correlation degree threshold, the target user TU is regarded as the audience user of the target video.

[0025] Step S400: Based on the adjustment records of the parameter values ​​of the interactive parameters in the video when the audience user watches the feature video, and the pending parameter values ​​of the interactive parameters currently to be applied to the target video, the warning level of the target video is calculated, and intelligent prompts are given to the target video and the pending parameter values ​​according to the warning level.

[0026] Step S400 includes: randomly extracting M adjustment records of the parameter values ​​of the interaction parameters in the video when the audience user watches the feature video, and obtaining the final parameter value after the interaction parameter of each adjustment record is changed, and obtaining the warning level of the target video according to the pending parameter value of the target video. , where e is a natural constant, P V is the undetermined parameter value of the target video, P m It is the final parameter value corresponding to the mth adjustment record. If the warning level Z is greater than the warning level threshold, an intelligent prompt is given for the target video and the pending parameter value.

[0027] Interaction parameters include speech speed, speech volume, and video brightness. For the target video, the adjustment records of the target users are an important basis for determining whether the target video's pending parameter values ​​are reasonable. For example, the target video's target users are usually middle-aged and elderly people, and most of them adjust the speech volume to a higher level in order to watch the video more appropriately. In this case, when the target video's pending parameter values ​​deviate significantly from the actual parameter values ​​corresponding to the adjustment records, that is, |P V -P m | is larger, then the warning level should also be larger. Since the warning level Z ranges from 0 to 1, in this solution, the warning level threshold is 0.6. When the warning level Z is greater than 0.6, intelligent prompts are given for the target video and the undetermined parameter values.

[0028] The present invention also provides a data-driven digital human high-quality video generation system, such as Figure 2 Shown, including:

[0029] Feature video extraction module: used to obtain historical digital human videos generated using digital human video templates on the video creation platform, as well as target videos currently generated based on video templates; extract feature videos from digital human videos based on the type of video template corresponding to the target video and the keywords in the text content used to generate the target video; Target user acquisition module: This module is used to acquire users who watch the featured video on the video playback platform. It extracts the clips of the digital human appearing and speaking from the featured video, and then determines the target user based on the viewing time and interaction behavior of the user with the featured video. Audience user extraction module: This module analyzes the target user's pause and rewind behaviors when watching a featured video, and combines the titles of the videos that the target user has searched and played in the past to obtain the target user's keyword set. Based on the target video's corresponding copy and keyword set, the module extracts the audience user corresponding to the target video from the target user. Interaction parameter analysis module: used to calculate the warning level of the target video based on the adjustment records of the parameter values ​​of the interaction parameters in the video when the audience users watch the feature video, as well as the pending parameter values ​​of the interaction parameters currently to be applied to the target video, and provide intelligent prompts for the target video and the pending parameter values ​​based on the warning level.

[0030] The feature video extraction module includes a first reliability calculation unit and a feature video extraction unit; The first reliability calculation unit is used to obtain digital human videos generated in the past using digital human video templates on the video creation platform, and obtain the target video currently generated based on the video template and copy content; obtain the type corresponding to each video template, and generate a corresponding type set to obtain the first reliability of the digital human video as a feature video; Feature video extraction unit: used to extract all keywords in the copy content, establish a set, and obtain the second reliability of the digital human video as a feature video; based on the first reliability and the second reliability, obtain the total reliability, and then obtain the feature video in the digital human video based on the total reliability.

[0031] The audience user extraction module includes a keyword set building unit and an audience user extraction unit; Keyword set establishment unit: used to establish a keyword set with an empty element corresponding to the target user, analyze the target user's pause and rewind behavior when watching the feature video, and combine the titles of the target user's historical search and playback videos to obtain the target user's final keyword set; Audience user extraction unit: used to extract the audience users corresponding to the target video from the target users based on the copy content and keyword set corresponding to the target video.

[0032] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. A data-driven method for generating high-quality digital human videos, characterized in that: It includes the following steps: Step S100: Obtain the digital human videos generated from the digital human video templates on the historical video creation platform, and the target video currently generated according to the video template; Extract the feature videos in the digital human videos according to the type of the video template corresponding to the target video and the keywords in the text content of the generated target video; Step S200: Obtain the viewing users who watch the feature videos on the video playback platform, extract the segments where the digital human appears and conducts oral broadcasts from the feature videos, and obtain the target users according to the viewing duration of the viewing users and their interaction behaviors with the feature videos; Step S300: Analyze the behaviors of the target users to pause and rewind when watching the feature videos, and combine the titles of the videos that the target users have searched for and played historically to obtain the keyword set of the target users; according to the text content corresponding to the target video and the keyword set, extract the audience users corresponding to the target video from the target users; Step S400: Calculate the warning level of the target video according to the adjustment record of the parameter values of the interaction parameters changed in the video by the audience users when watching the feature videos and the pending parameter values of the interaction parameters to be applied to the target video currently, and perform intelligent prompts on the target video and the pending parameter values according to the warning level.

2. The method for generating high-quality digital human videos based on data drive according to claim 1, characterized in that: Step S100 includes: Step S110: Obtain the digital human videos generated from the digital human video templates on the historical video creation platform, where each digital human video is a video that has been published on the video playback platform; obtain the target video V generated according to the video template Tem and the text content Tex; Get the type corresponding to each video template, including the style, theme and applicable scene of the video template, and generate a corresponding type set; use the type set of the video template Tem as S Tem , obtain the video template WTem corresponding to a certain digital human video W in history, and use the type set of the video template WTem as S WTem , according to the set S Tem and set S WTem The number of the same type N, the digital human video W is the first reliability X1=1-e -N , e is a natural constant; Step S120: extract all keywords in the text content Tex, remove the same keywords, and obtain a set S Tex , extract all the keywords in the text content WTex corresponding to the digital human video W, remove the same keywords, and get the set S WTex ; Get the set S Tex A keyword KW1 in the set S WTex The similarity between each keyword in and the keyword KW1 is calculated, and the maximum similarity is used as the target value of the keyword KW1, thereby obtaining the set S Tex The target value of each keyword in the , and add up to get the average value, which is the second reliability X2 of the digital human video W as the feature video; 3. The method for generating high-quality digital human videos based on data drive according to claim 1, characterized in that: Preset the weights corresponding to the first reliability X1 and the second reliability X2 to obtain the total reliability X0; if X0>K1 and |X1-X2|<K2 are satisfied, where K1 is the first numerical threshold and K2 is the second numerical threshold, then use the digital human video W as the feature video. Step S200 includes: Step S220: Extract the segment where the digital human appears and speaks from the feature video VID, and obtain the total duration D0 of the segment. WU , and the obtained behavior value H, the target degree G of the viewing user WU is obtained WU =H×D WU / D0, if the target level G WU If the degree is greater than a preset threshold, the viewing user WU is taken as the target user of the feature video VID, and then all target users corresponding to each feature video are obtained.

4. The method for generating high-quality digital human videos based on data drive according to claim 1, characterized in that: Step S210: Obtain a certain viewing user WU who watches a certain feature video VID on the video playback platform, set the initial behavior value of the viewing user WU to 0, if the viewing user WU likes the feature video VID, increase the behavior value by h1, if comments, increase the behavior value by h2, if favorites, increase the behavior value by h3, and h1+h2+h3=1, and then obtain the final behavior value H; Step S310: Create a keyword set S with an empty element corresponding to a target user TU TU ; Get the target user TU when watching the corresponding feature video FV, perform the pause action of the playback progress time, extract the sentence text spoken by the digital person at the playback progress time, and input the keywords in the sentence text into the set S TU middle; Obtain the initial playback progress time T2 before the user TU performs the rewind action when watching the video FV, and the final playback progress time T1 after the rewind action, where time T1 is before T2, extract all the sentence texts spoken by the digital human from time T1 to time T2, and input the keywords in all the sentence texts into the set S TU middle; Get the title of the video that the target user TU searched and played in the past C days starting from the time when the target user TU watched the feature video FV, and input the keywords in the title into the set S TU Then we get the final keyword set S TU ; Step S320: The set S is obtained based on the text content Tex corresponding to the target video. Tex , get the set S Tex A keyword KW2 in the set S TU The similarity between each keyword in and the keyword KW2 is calculated, and the maximum similarity is used as the target value of the keyword KW2, thereby obtaining the set S Tex The target value of each keyword in is added and the average value is obtained to obtain the correlation degree Y. If the correlation degree Y is greater than the correlation degree threshold, the target user TU is regarded as the audience user of the target video.

5. The method for generating high-quality digital human videos based on data drive according to claim 1, characterized in that: Step S400 includes: randomly extracting M adjustment records of the parameter values ​​of the interaction parameters in the video when the audience user watches the feature video, and obtaining the final parameter value after the interaction parameter of each adjustment record is changed, and obtaining the warning level of the target video according to the pending parameter value of the target video. , where e is a natural constant, P V is the undetermined parameter value of the target video, P m It is the final parameter value corresponding to the mth adjustment record. If the warning level Z is greater than the warning level threshold, an intelligent prompt is given to the target video and the pending parameter value.

6. A digital human high-quality video generation system, used to execute the data-driven digital human high-quality video generation method according to any one of claims 1 to 5, characterized in that: Step S300 includes: The system includes a feature video extraction module, a target user acquisition module, an audience user extraction module, and an interaction parameter analysis module; Feature video extraction module: used to obtain the digital human videos generated from the digital human video templates on the historical video creation platform, and the target video currently generated according to the video template; extract the feature videos in the digital human videos according to the type of the video template corresponding to the target video and the keywords in the text content of the generated target video; Target user acquisition module: used to obtain the viewing users who watch the feature videos on the video playback platform, extract the segments where the digital human appears and conducts oral broadcasts from the feature videos, and obtain the target users according to the viewing duration of the viewing users and their interaction behaviors with the feature videos; Audience user extraction module: used to analyze the target user's pause and rewind behavior when watching the feature video, and combine the titles of the target user's historical search and playback videos to obtain the target user's keyword set; based on the target video's corresponding copy content and the keyword set, extract the audience user corresponding to the target video from the target user; Interaction parameter analysis module: used to calculate the warning level of the target video based on the adjustment records of the parameter values ​​of the interaction parameters in the video when the audience users watch the feature video, as well as the pending parameter values ​​of the interaction parameters currently to be applied to the target video, and provide intelligent prompts for the target video and the pending parameter values ​​according to the warning level.

7. The high-quality digital human video generation system according to claim 6, characterized in that: The feature video extraction module includes a first reliability calculation unit and a feature video extraction unit; The first reliability calculation unit is used to obtain digital human videos generated in the past using digital human video templates on the video creation platform, and obtain the target video currently generated based on the video template and copy content; Obtaining the type corresponding to each video template and generating a corresponding type set to obtain the first reliability of the digital human video as the feature video; Feature video extraction unit: used to extract all keywords in the copy content, establish a set, and obtain the second reliability of the digital human video as a feature video; based on the first reliability and the second reliability, obtain the total reliability, and then obtain the feature video in the digital human video based on the total reliability.

8. The high-quality digital human video generation system according to claim 6, characterized in that: The audience user extraction module includes a keyword set establishment unit and an audience user extraction unit; Keyword set establishment unit: used to establish a keyword set with an empty element corresponding to the target user, analyze the target user's pause and rewind behavior when watching the feature video, and combine the titles of the target user's historical search and playback videos to obtain the target user's final keyword set; Audience user extraction unit: used to extract audience users corresponding to the target video from target users based on the text content corresponding to the target video and the keyword set.

Citation Information

Patent Citations

  • Digital human video processing method, electronic equipment and medium

    CN117376597A

  • Method and system for generating digital human video

    CN119180895A

  • Digital human video generation method and device based on large model, intelligent agent, electronic equipment and storage medium

    CN120302122A

  • Digital life individuation implementation method, device, equipment and medium

    CN120353929A

  • Method and apparatus for generating digital person, method and apparatus for training model, and device and medium

    WO2023240943A1

Cited By

  • Unmanned aerial vehicle intelligent video data analysis system and method based on cloud platform

    CN121438180A

  • Unmanned aerial vehicle intelligent video data analysis system and method based on cloud platform

    CN121438180B