An Interactive Rendering and Testing Method for Virtual Digital Humans

By selecting expression units based on video frame category and the number of synchronized keywords during the rendering process of virtual digital humans, the problem of insufficient expression rendering accuracy of virtual digital humans is solved, and a more natural and adaptive rendering effect is achieved.

CN121414935BActive Publication Date: 2026-05-05GUANGZHOU SIZHI TIMES TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU SIZHI TIMES TECH CO LTD
Filing Date
2025-12-25
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

In existing technologies, virtual digital humans suffer from insufficient semantic analysis accuracy during interactive rendering, and their expression generation is simplistic and lacks dynamic correlation adjustment, resulting in poor expression rendering accuracy.

Method used

By acquiring target video frames and their interactive audio, the system determines whether to optimize rendering based on the video frame category, and uses text matching rendering or motion analysis rendering. It selects facial expression units based on parameters such as the number of synchronized keywords and amplitude deviation, and adjusts them to improve the rendering effect.

Benefits of technology

It improves the continuity and consistency of virtual digital human expressions and movements, ensures the fit and adaptability of expressions with the scene, avoids incorrect or missed optimization, and improves rendering accuracy and naturalness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121414935B_ABST
    Figure CN121414935B_ABST
Patent Text Reader

Abstract

This invention relates to the field of image processing technology, and more particularly to an interactive rendering and testing method for virtual digital humans. The method includes: acquiring a target video frame and the corresponding interactive voice; determining the video frame category based on the existence of a preceding video frame, and determining whether to perform rendering optimization based on data fine-grainedness or inter-frame variation coefficients based on the target video frame's video frame category; in the rendering optimization, determining whether to perform text matching rendering or motion analysis rendering based on the target video frame's video frame category to select an expression unit; and rendering the target video frame based on the expression weights corresponding to the expression units. This invention can improve the expression rendering accuracy of virtual digital humans.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an interactive rendering and testing method for virtual digital humans. Background Technology

[0002] As the application scenarios of virtual digital humans continue to expand, problems such as the disconnect between expressions and context and the distortion of details in the interactive rendering process are gradually exposed. This leads to blurry deviations in micro-movements of the body or facial details, ultimately resulting in stiff visual effects and insufficient realism. Therefore, how to improve the rendering accuracy of virtual digital humans is an urgent problem to be solved by those skilled in the art.

[0003] Chinese Patent Publication No. CN115953521A discloses a remote digital human rendering method, apparatus, and system. The method includes: calculating the inverse document frequency (IVF) of each text based on the size of a preset text set and the length of each text in the text set, and using the IVF to train a neural network model for semantic analysis; generating speech data in response to user input data received from a remote digital human device; performing semantic analysis on the speech data using the neural network model; and rendering the remote digital human based on the results of the semantic analysis to obtain video frames of the remote digital human; synchronizing the speech data and the video frames, and pushing the synchronized speech data and video frames to the remote digital human device. However, the above solution has the following problems: insufficient semantic analysis accuracy, simplistic expression generation, and lack of dynamic correlation adjustment, resulting in poor expression rendering accuracy. Summary of the Invention

[0004] To address this, the present invention provides an interactive rendering and testing method for virtual digital humans, which overcomes the problems of insufficient semantic analysis accuracy, single expression generation and lack of dynamic correlation adjustment in the prior art, resulting in poor expression rendering accuracy.

[0005] To achieve the above objectives, the present invention provides an interactive rendering and testing method for virtual digital humans, comprising:

[0006] Acquire the target video frame and the corresponding interactive voice;

[0007] The video frame category is determined based on the existence of a preceding video frame, and whether to perform rendering optimization is determined based on the video frame category of the target video frame, based on the fine-grained data or the inter-frame variation coefficient. In the rendering optimization, text matching rendering or motion analysis rendering is performed based on the video frame category of the target video frame to determine the selection of expression units.

[0008] In text matching rendering, matching expansion or search expansion is determined based on the number of synchronized keywords to determine the selection of emoji units;

[0009] In motion analysis and rendering, the number of selected facial expression units is determined based on the amplitude deviation, and the number of selected facial expression units is adjusted based on the facial expression transition delay coefficient. The selection of facial expression units is also determined based on the action boundary value, the facial expression unit fit or trajectory coordination.

[0010] Under abnormal test conditions, a weight reference value is determined based on the sum of the weights of the selected facial expression units, and the weight reference value is adjusted according to the number of selected facial expression units.

[0011] Rendering is performed on the target video frame based on the facial expression weights corresponding to the facial expression units.

[0012] Furthermore, for target video frames whose video frame category is Class 1 video frames, whether to perform rendering optimization is determined based on fine-grained data.

[0013] For a target video frame that is a Class II video frame, determine whether to perform rendering optimization based on the inter-frame variation coefficient.

[0014] The target video frame categories include a category of video frames that do not have preceding video frames and a category of video frames that do have preceding video frames.

[0015] Furthermore, text matching and rendering are performed on target video frames that are classified as Class 1 video frames;

[0016] Motion analysis and rendering are performed on target video frames that are classified as Class II video frames.

[0017] Furthermore, for target video frames with a number of synchronized keywords greater than or equal to the preset number of synchronized keywords, matching expansion is performed;

[0018] In the matching expansion, the selected expression unit is determined based on the degree of synchronous correlation, the degree of influence of interval, or the degree of convergence of matching changes.

[0019] The synchronization keywords are determined based on keyword matching degree and frequency of occurrence.

[0020] Furthermore, based on the degree of synchronous correlation, the selection of facial expression units is determined according to the degree of interval influence or the degree of convergence of matching changes, including:

[0021] For target video frames with a synchronization correlation degree greater than or equal to the preset synchronization correlation degree, the selected expression unit is determined based on the interval influence degree.

[0022] For target video frames with a synchronization correlation degree less than the preset synchronization correlation degree, the selected expression unit is determined based on the similarity of the matching changes.

[0023] Furthermore, for target video frames where the number of synchronized keywords is less than the preset number of synchronized keywords, the search is expanded;

[0024] In the search expansion process, the search depth is determined based on the initial search effectiveness of the target keywords, the matching keywords are determined based on the matching coefficient, and the emoji units to be selected are determined based on the matching keywords.

[0025] The matching coefficient is determined based on the proportion of the floating page and the page effectiveness.

[0026] Furthermore, the number of selected facial expression units is determined based on the amplitude deviation.

[0027] The number of selected facial expression units is positively correlated with the degree of amplitude deviation.

[0028] Furthermore, the number of selected facial expression units is reduced based on the facial expression transition delay coefficient;

[0029] The decrease in the number of selected facial expression units is positively correlated with the facial expression transition delay coefficient.

[0030] Furthermore, based on the action boundary value, the selection of facial expression units is determined according to the fit of the facial expression unit or the trajectory coordination, including:

[0031] For target video frames with motion boundary values ​​less than preset motion boundary values, select the appropriate facial expression unit based on the fit of the facial expression unit.

[0032] For target video frames whose action boundary values ​​are greater than or equal to preset action boundary values, the selected facial expression unit is determined based on the trajectory coordination degree.

[0033] Furthermore, under abnormal testing conditions, the weight reference value is adjusted by decreasing it based on the number of selected facial expression units;

[0034] The abnormal test condition is that the abnormal reference value is greater than the preset abnormal reference value, and the decrease in the weight reference value is positively correlated with the number of selected expression units.

[0035] Compared with the prior art, the beneficial effect of the present invention is that, in the technical solution of the present invention, the existence of historical frame information available for comparison is effectively reflected by the video frame category, and then the determination of whether to perform rendering optimization is based on the fine-grained data or the inter-frame change coefficient according to the video frame category of the target video frame, so that the determination of rendering optimization is more in line with the actual application scenario, thereby avoiding incorrect optimization or missed optimization, and thus improving the continuity and consistency of digital human expressions and movements.

[0036] Furthermore, in this invention, text matching rendering or motion analysis rendering is performed by determining the video frame category of the target video frame, which makes the determination of the selected expression unit more in line with the actual application scenario. This is conducive to accurately matching the expression requirements in different scenarios, avoiding the problem of expression and scene disconnect caused by a single rendering method, and thus improving the naturalness and adaptability of the virtual digital human's facial expressions.

[0037] Furthermore, this invention effectively reflects the semantic richness of the target text and historical effective reference text by using the number of synchronized keywords. Based on the number of synchronized keywords, it determines whether to perform matching expansion or search expansion. This helps to accurately determine whether existing historical data can support the effective selection of emoji units. When the number of synchronized keywords is sufficient, matching expansion is performed based on historical data to ensure the fit between emojis and semantics. When the number of synchronized keywords is insufficient, external adaptation data is supplemented through search expansion to avoid gaps in emoji selection due to data gaps, thereby improving the comprehensiveness and adaptability of emoji unit selection.

[0038] Furthermore, in this invention, the action boundary value effectively reflects the relative intensity of the limb movement amplitude in the target video frame relative to the historical effective action. Then, based on the action boundary value, the expression unit is selected according to the expression unit fit degree or trajectory coordination degree. This is conducive to accurately matching the degree of expression adaptation under different movement amplitudes, avoiding the disconnect between expression and limb dynamics, and thus improving the degree of adaptation between expression unit selection and limb movement. Attached Figure Description

[0039] Figure 1 This is a schematic diagram of the interactive rendering and testing method for virtual digital humans according to the present invention;

[0040] Figure 2 This is a flowchart illustrating the present invention's determination of whether to perform rendering optimization based on the video frame category of the target video frame and on fine-grained data or inter-frame variation coefficients.

[0041] Figure 3 This is a flowchart illustrating the present invention for determining text matching rendering or motion analysis rendering based on the video frame category of the target video frame.

[0042] Figure 4 This is a flowchart illustrating the matching expansion or search expansion based on the number of synchronized keywords in this invention. Detailed Implementation

[0043] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0044] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0045] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.

[0046] Please see Figures 1 to 4 As shown, the present invention provides an interactive rendering and testing method for virtual digital humans, comprising:

[0047] Acquire the target video frame and the corresponding interactive voice;

[0048] The video frame category is determined based on the existence of a preceding video frame, and whether to perform rendering optimization is determined based on the video frame category of the target video frame, based on the fine-grained data or the inter-frame variation coefficient. In the rendering optimization, text matching rendering or motion analysis rendering is performed based on the video frame category of the target video frame to determine the selection of expression units.

[0049] In text matching rendering, matching expansion or search expansion is determined based on the number of synchronized keywords to determine the selection of emoji units;

[0050] In motion analysis and rendering, the number of selected facial expression units is determined based on the amplitude deviation, and the number of selected facial expression units is adjusted based on the facial expression transition delay coefficient. The selection of facial expression units is also determined based on the action boundary value, the facial expression unit fit or trajectory coordination.

[0051] Under abnormal test conditions, a weight reference value is determined based on the sum of the weights of the selected facial expression units, and the weight reference value is adjusted according to the number of selected facial expression units.

[0052] Rendering is performed on the target video frame based on the facial expression weights corresponding to the facial expression units.

[0053] The application scenario of this invention is the rendering optimization of virtual digital humans. This invention has several historical records, each of which records at least one historical process of rendering optimization of virtual digital humans, including data granularity, inter-frame variation coefficients, number of synchronization keywords, and amplitude deviation. Each historical record also has a corresponding qualified mark, which records whether the rendering optimization process of virtual digital humans meets the user's needs. The qualified mark can be recorded manually. It is understood that the user can determine whether the rendering optimization process of virtual digital humans meets the requirements based on self-defined indicators. Self-defined indicators can be, but are not limited to, rendering time, which will not be elaborated here. The rendering optimization time is the time required to perform rendering optimization on the target video frame.

[0054] The interactive voice corresponding to the target video frame is the voice corresponding to the video in which the target video frame is located. The interactive voice input by the user is converted into text through ASR technology and recorded as speech-to-text.

[0055] The present invention provides an expression unit matching library, which contains several setting keywords. Each setting keyword corresponds to several expression units. The setting keywords include, but are not limited to, happiness and sadness. When the setting keyword is happiness, the corresponding expression units are smiling and curved eyebrows and eyes. When the setting keyword is sadness, the corresponding expression units are drooping corners of the mouth and raised inner corners of the eyebrows. This is the content that is easy for those skilled in the art to understand, and will not be elaborated further.

[0056] In this invention, the target video frame is the initial video frame that has been rendered by the user. The facial expression units generated in the initial video frame are recorded as initial facial expression units. Facial expression units include, but are not limited to, smiling, eyebrow and eye curves, and downturned corners of the mouth. Each facial expression unit corresponds to several facial points, and each facial point corresponds to an allowable change distance threshold. The allowable change distance threshold for a single facial point is the change distance reference value corresponding to the facial point when the facial point presents the maximum natural effect of the facial expression unit. For a single target video frame and for a single facial expression unit, if all facial points corresponding to the facial expression unit change relative to the base frame, then the facial expression unit is recorded as an initial facial expression unit.

[0057] Specifically, for a target video frame whose video frame category is Class I video frame, whether to perform rendering optimization is determined based on fine-grained data.

[0058] For a target video frame that is a Class II video frame, determine whether to perform rendering optimization based on the inter-frame variation coefficient.

[0059] The target video frame categories include a category of video frames that do not have preceding video frames and a category of video frames that do have preceding video frames.

[0060] Specifically, video frames without preceding video frames are categorized as first-class video frames, while video frames with preceding video frames are categorized as second-class video frames.

[0061] The preceding video frame is a video frame in the video containing the target video frame that is earlier in time and adjacent to the target video frame;

[0062] Obtain the audio-video synchronization timeline of the interactive voice, determine the timestamp of the target video frame (e.g., the target video frame is at the position of "00:01:23.456" in the video, i.e., 1 minute 23 seconds 456 milliseconds), extract a 2-second audio segment before and after the timestamp, extract keywords that can represent the core semantics of the voice through NLP technology, and record them as target keywords. It can be understood that if the target video frame is the earliest or latest video frame of the video corresponding to the interactive voice, then the extracted audio segment corresponding to the target video frame is 1 second.

[0063] The word2vec model is used to convert two keywords into word vectors. The cosine similarity between the two word vectors is recorded as the keyword matching degree between the two keywords. The value range of the keyword matching degree is [0,1].

[0064] The preset keyword matching degree can be determined by the user according to the actual application scenario. The greater the user's demand for improving the accuracy of the expression rendering, the higher the preset keyword matching degree will be. One preset keyword matching degree is provided, with a preset keyword matching degree of 0.8.

[0065] Data granularity = total number of initial facial expression units in the target video frame / average number of facial expression units corresponding to each video frame in the historical record that can meet user needs;

[0066] Understandably, the base frame is a video frame of the digital human's face without any muscle contraction or facial expression, and with the limbs in a natural standing posture. Several facial points are set for the digital human's face. It should be noted that the facial points must cover the five core areas of the face: facial contour, eyebrows, eyes, nose, and lips. Each facial point has a unique and fixed coordinate position in the base frame. The specific positions of the facial points include, but are not limited to, the apex of the left / right eyebrow, the center of the left / right pupil, and the midpoint of the lower lip. The greater the accuracy of the user's determination of the expression test difference, the more facial points are set. One possible value for the number of facial points is 68. A point is randomly selected from each joint of the digital human. These points are limb points, including the head, shoulders, elbows, wrists, hips, knees, and ankles.

[0067] Inter-frame variation coefficient = limb point variation representation value corresponding to the target video frame / facial point variation representation value corresponding to the target video frame; facial point variation representation value corresponding to a single video frame is the sum of variation distance reference values ​​corresponding to each varied facial point in that video frame; limb point variation representation value corresponding to a single video frame is the sum of variation distance reference values ​​corresponding to each varied limb point in that video frame.

[0068] The video frame is a rectangular image. A rectangular coordinate system is established with the lower left corner of the rectangle as the origin, the straight line extending to the right along the bottom edge of the rectangle from the origin as the x-axis, and the straight line extending upward along the left side of the rectangle from the origin as the y-axis. The position of the face point or the position of the limb point are the coordinates of the face point or the limb point in the rectangular coordinate system.

[0069] For a single face point or limb point in the target video frame, if the position of the face point or limb point is different from the corresponding face point or limb point in the base frame, then the face point or limb point is recorded as a changed face point or changed limb point.

[0070] The reference value for the change distance corresponding to a single changed limb point (or changed face point) is the shortest distance between the position of the changed limb point (or changed face point) in the target video frame and the position of the changed limb point (or changed face point) in the preceding video frame corresponding to the target video frame.

[0071] When determining whether to perform rendering optimization based on data granularity, if the data granularity is greater than or equal to the preset data granularity, then no rendering optimization is needed; if the data granularity is less than the preset data granularity, then rendering optimization is performed.

[0072] When determining whether to perform rendering optimization based on the inter-frame variation coefficient, if the inter-frame variation coefficient is greater than or equal to the preset inter-frame variation coefficient, then rendering optimization is performed; if the inter-frame variation coefficient is less than the preset inter-frame variation coefficient, then rendering optimization is not required.

[0073] The values ​​of preset data granularity and preset inter-frame variation coefficient can be determined by the user according to the actual application scenario. The greater the user's demand for improving the accuracy of rendering expression richness, the larger the value of preset data granularity and the smaller the value of preset inter-frame variation coefficient. A method for determining the values ​​of preset data granularity and preset inter-frame variation coefficient is provided. The historical records of rendering optimization are detected, and the average value of the data granularity and the average value of the inter-frame variation coefficient corresponding to the historical records that meet the user's needs are respectively recorded as preset data granularity and preset inter-frame variation coefficient.

[0074] Specifically, text matching and rendering are performed on target video frames whose video frame category is Class 1.

[0075] Motion analysis and rendering are performed on target video frames that are classified as Class II video frames.

[0076] Specifically, for target video frames with a number of synchronized keywords greater than or equal to the preset number of synchronized keywords, matching and expansion are performed;

[0077] In the matching expansion, the selected expression unit is determined based on the degree of synchronous correlation, the degree of influence of interval, or the degree of convergence of matching changes.

[0078] The synchronization keywords are determined based on keyword matching degree and frequency of occurrence.

[0079] Specifically, the speech-to-text corresponding to the target video frame is recorded as the target text, and the speech-to-text corresponding to the historical records that can meet the user's needs is recorded as the reference text. Each keyword in the reference text is extracted using NLP technology, and each keyword in the reference text is recorded as the reference keyword.

[0080] Synchronization keywords are determined based on keyword matching degree and frequency of occurrence, including: reference keywords with a keyword matching degree greater than the target keyword matching degree are recorded as similar keywords; other keywords in the reference text containing similar keywords are recorded as keywords to be supplemented; and keywords to be supplemented that appear more often than the preset frequency of occurrence are recorded as synchronization keywords.

[0081] The number of synchronization keywords corresponding to the target video frame is the total number of different synchronization keywords appearing in the reference text;

[0082] The preset number of synchronized keywords can be determined by the user based on the actual application scenario. The smaller the preset number of synchronized keywords, the greater the user's need for matching expansion. A method for setting the preset number of synchronized keywords is provided, which detects the user's historical matching expansion history and records the average number of synchronized keywords corresponding to the historical history that can meet the user's needs as the preset number of synchronized keywords.

[0083] The number of occurrences of a single keyword to be supplemented is the number of reference texts containing that keyword;

[0084] Specifically, the selection of facial expression units is determined based on the degree of synchronous correlation, the degree of influence of interval, or the degree of convergence of matching changes, including:

[0085] For target video frames with a synchronization correlation degree greater than or equal to the preset synchronization correlation degree, the selected expression unit is determined based on the interval influence degree.

[0086] For target video frames with a synchronization correlation degree less than the preset synchronization correlation degree, the selected expression unit is determined based on the similarity of the matching changes.

[0087] Specifically, the synchronization correlation degree corresponding to the target video frame is the average of the sub-correlation degrees between each synchronization keyword and the target text;

[0088] The sub-association degree between a single synchronous keyword and the target text is the maximum value among the keyword matching degrees of that synchronous keyword and all keywords in the target text.

[0089] The user can determine the preset synchronization correlation value according to the actual application scenario. The greater the user's demand for improving the accuracy of the rendering effect, the smaller the preset synchronization correlation value should be. One preset synchronization correlation value is provided, which is 0.65.

[0090] When selecting facial expression units based on the interval influence, facial expression units that appear in the video frames corresponding to similar keywords in the reference text containing synchronous keywords with an interval influence less than the preset interval influence are selected as facial expression units; the video frames corresponding to a single similar keyword are the video frames in the reference text containing synchronous keywords that use that similar keyword as the target keyword.

[0091] For a single synchronized keyword, the interval influence of that synchronized keyword = the time interval value corresponding to the target text / the average of the time interval values ​​corresponding to all reference texts containing that synchronized keyword;

[0092] The time interval value corresponding to the target text is the duration between the earliest video frame corresponding to the keyword with the highest keyword matching degree with the synchronous keyword in the video where the target video frame is located and the earliest video frame corresponding to the target keyword, in seconds;

[0093] The time interval value corresponding to a single reference text containing the synchronization keyword is the duration between the earliest video frame corresponding to the synchronization keyword in the video corresponding to the reference text and the earliest video frame corresponding to a similar keyword in the reference text.

[0094] When selecting facial expression units based on the similarity of matching changes, the facial expression units that appear in the video frames corresponding to the similar keywords in the reference text where the matching change similarity is greater than the preset matching change similarity are selected as the facial expression units.

[0095] The similarity of collocation changes corresponding to a single synchronous keyword is the average of the similarity of sub-changes corresponding to each reference text in which the synchronous keyword appears;

[0096] The sub-change convergence degree corresponding to a single reference text containing the synchronous keyword = the number of facial expression units appearing in all video frames between the video frame containing the synchronous keyword and the video frame containing the target keyword in the reference text / the total number of facial expression units appearing in all video frames between the video frame containing the synchronous keyword and the video frame containing the target keyword in the reference text.

[0097] Users can determine the values ​​of preset interval influence and preset combination change convergence based on the actual application scenario. The greater the user's demand for improving the accuracy of digital human expression rendering fit, the smaller the value of preset interval influence and the larger the value of preset combination change convergence. One set of preset interval influence and preset combination change convergence values ​​is provided: preset interval influence is 1.2 and preset combination change convergence is 0.7.

[0098] Specifically, for target video frames where the number of synchronized keywords is less than the preset number of synchronized keywords, the search is expanded;

[0099] In the search expansion process, the search depth is determined based on the initial search effectiveness of the target keywords, the matching keywords are determined based on the matching coefficient, and the emoji units to be selected are determined based on the matching keywords.

[0100] The matching coefficient is determined based on the proportion of the floating page and the page effectiveness.

[0101] Specifically, the initial search validity = the number of matching links appearing on the initial search page / the threshold for the number of matching links, where the threshold for the number of matching links is 5;

[0102] For a single link, if the link contains keywords other than the target keyword in the target text, then the link is recorded as a matching link;

[0103] Search depth is the time it takes to crawl the webpage for the initial search page corresponding to the target keyword. Search depth = initial search effectiveness × search depth threshold, and the search depth threshold is 3 seconds.

[0104] The page obtained by searching for the target keywords is recorded as the initial search page, and the initial search page and the pages obtained by web crawling based on the initial search page are recorded as the search pages.

[0105] Other keywords appearing on the search page besides the target keyword are recorded as search keywords;

[0106] For a single search keyword, the search pages that appear for that keyword are recorded as floating pages. The matching coefficient for that search keyword is calculated as follows: (Floating page percentage / Average floating page percentage for all search keywords) × First weight coefficient + (Page validity / Average page validity for all search keywords) × Second weight coefficient. Both the first and second weight coefficients are 0.5.

[0107] Population percentage = Number of pop pages corresponding to the search keyword / Total number of search pages;

[0108] Page effectiveness = Keyword distance reference value / Average keyword distance reference value for each search keyword + Number of effective keywords / Average number of effective keywords for each search keyword;

[0109] The keyword distance reference value corresponding to a single search keyword is the average of the distance reference values ​​corresponding to all search pages where the search keyword appears. The distance reference value corresponding to a single search page where the search keyword appears is the shortest distance from the target keyword to the search keyword on that search page.

[0110] The number of valid keywords corresponding to a single search keyword is the average of the reference values ​​of the number of keywords corresponding to each search page where the search keyword appears. The reference value of the number of keywords corresponding to a single search page where the search keyword appears is the number of keywords in the target text that appears on that search page.

[0111] When determining matching keywords based on the matching coefficient, search keywords with a matching coefficient greater than or equal to the preset matching coefficient are recorded as matching keywords;

[0112] The user can determine the value of the preset matching coefficient according to the actual application scenario. The greater the user's demand for improving the accuracy of the expression rendering, the larger the value of the preset matching coefficient. One preset matching coefficient value is provided, which is 0.75.

[0113] Based on matching keywords, determine the emoji units to be selected, detect the emoji units corresponding to each set keyword to be selected, and set keywords to be selected are set keywords whose keyword matching degree with any matching keyword is greater than the preset keyword matching degree. Emoji units whose occurrence rate is greater than the preset occurrence rate are selected emoji units.

[0114] The occurrence percentage of a single emoji unit = the number of set keywords that appear in that emoji unit / the total number of set keywords. The default occurrence percentage can be determined by the user based on the actual application scenario. The greater the user's demand for improving the accuracy of emoji rendering, the larger the default occurrence percentage will be. One default occurrence percentage is provided, which is 0.72.

[0115] Understandably, the number of synchronized keywords effectively reflects the richness of semantic association between the target text and historical reference text. When the number of synchronized keywords is greater than or equal to the preset number of synchronized keywords, it indicates that the semantic association between the target text and the historical effective text is sufficient, and there is information that can be mined. Therefore, matching expansion is performed to supplement the emoji units that are closely related to semantics, thereby improving the richness and accuracy of the rendering effect. When the number of synchronized keywords is less than the preset number of synchronized keywords, it indicates that the matching degree between the target text and the existing historical data is low, and the existing data cannot provide sufficient support for rendering. Therefore, search expansion is performed to obtain suitable emoji units, ensuring that the rendering effect meets the basic interaction requirements and avoiding rendering deviations or the problem of no matching emojis available due to insufficient data.

[0116] Specifically, the number of facial expression units to be selected is determined based on the degree of amplitude deviation;

[0117] The number of selected facial expression units is positively correlated with the degree of amplitude deviation.

[0118] Specifically, amplitude deviation = limb point change representation value corresponding to the target video frame - limb point change representation value corresponding to the preceding video frame of the target video frame.

[0119] Select the number of facial expression units as n1, where n1 is the smallest integer greater than or equal to n10, and n10 = amplitude deviation / preset amplitude deviation × facial expression unit number threshold. The facial expression unit number threshold is 8.

[0120] The value of the preset amplitude deviation can be determined by the user according to the actual application scenario. The greater the user's demand for improving the accuracy of rendering facial expressions, the smaller the value of the preset amplitude deviation should be. A method for determining the preset amplitude deviation is provided, which is to record the average value of the amplitude deviation corresponding to each type of video frame in the historical record that can meet the user's needs as the preset amplitude deviation.

[0121] Specifically, the number of selected facial expression units is reduced based on the facial expression transition delay coefficient;

[0122] The decrease in the number of selected facial expression units is positively correlated with the facial expression transition delay coefficient.

[0123] Specifically, users can obtain the required duration of each generated video frame during the historical rendering process. The expression transition delay coefficient = the average of the required duration of each video frame generated earlier than the target video frame / the average number of expression units corresponding to each video frame generated earlier than the target video frame - the average of the required duration of each video frame generated in the historical record that can meet the user's needs / the average number of expression units corresponding to each video frame in the historical record that can meet the user's needs.

[0124] The reduction value for the number of facial expression units is selected as n2, where n2 is the smallest integer greater than or equal to n20. n20 = facial expression transition delay coefficient / average of the facial expression transition delay coefficients of each type of video frame corresponding to the historical records that can meet user needs × adjustment threshold, where the adjustment threshold is 3.

[0125] The final number of selected facial expression units is n, where n = n1 - n2;

[0126] Specifically, the selection of facial expression units is determined based on the action boundary value, the fit of the facial expression unit, or the trajectory coordination, including:

[0127] For target video frames with motion boundary values ​​less than preset motion boundary values, select the appropriate facial expression unit based on the fit of the facial expression unit.

[0128] For target video frames whose action boundary values ​​are greater than or equal to preset action boundary values, the selected facial expression unit is determined based on the trajectory coordination degree.

[0129] Specifically, the action boundary value = the limb point change representation value corresponding to the target video frame / the maximum value of the limb point change representation values ​​corresponding to each video frame in the historical record that can meet the user's needs;

[0130] The user can determine the value of the preset action boundary value according to the actual application scenario. The larger the value of the preset action boundary value, the greater the user's need to determine the expression unit based on the fit of the expression unit. One preset action boundary value is provided, which is 0.6.

[0131] The video frames in the history that meet the user’s needs and contain all the initial expression units corresponding to the target video frame are recorded as the first historical video frames. The other expression units in the first historical video frames, excluding the initial expression units, are recorded as the first historical expression units.

[0132] The fit of a single historical emoji unit is equal to the number of historical video frames in which that historical emoji unit appears, divided by the total number of historical video frames.

[0133] The second type of video frame in the historical records that meets the user’s needs and whose body point coordination with the target video frame is greater than the preset body point coordination is recorded as the second historical video frame. The other facial expression units in the second historical video frame, excluding the initial facial expression unit, are recorded as the second historical facial expression unit.

[0134] The trajectory coordination degree corresponding to a single second historical expression unit = 1 / (standard deviation of the facial point change representation value corresponding to each second video frame in which the second historical expression unit appears + 1).

[0135] The coordination degree of limb points corresponding to any two video frames = the number of identical changing limb points in the two video frames / the number of changing limb points corresponding to the target video frame × the number weight coefficient + (1 - the absolute value of the difference between the limb point change representation values ​​corresponding to the two video frames / the limb point change representation value corresponding to the target video frame) × the difference weight coefficient, where the number weight coefficient is 0.6 and the difference weight coefficient is 0.4.

[0136] The user can determine the preset limb point coordination degree based on the actual application scenario. The greater the user's demand for improving the accuracy of facial expression rendering fit, the smaller the preset limb point coordination degree will be. One preset limb point coordination degree is provided, with a preset limb point coordination degree of 0.65.

[0137] When selecting facial expression units based on their fit, the first historical facial expression units are selected in descending order of their fit, until the number of selected first historical facial expression units reaches n.

[0138] When determining the selection of facial expression units based on trajectory coherence, the second historical facial expression units are selected in descending order of trajectory coherence until the number of selected second historical facial expression units reaches n.

[0139] It is understandable that the action boundary value effectively reflects the relative intensity of the limb movement amplitude in the target video frame relative to the historical effective action. When the action boundary value is less than the preset action boundary value, it means that the target limb movement amplitude is small. At this time, the dynamic characteristics of the action itself have a weaker constraint on the expression. The user's perception focus is more inclined to the matching degree between the expression and the semantics. Therefore, it is necessary to prioritize the expression unit that has the highest frequency of matching with the initial expression in the historical data. That is, the expression unit is selected based on the fit of the expression unit.

[0140] When the action boundary value is greater than or equal to the preset action boundary value, it indicates that the target limb movement is large. At this time, the dynamic trajectory of the movement has a significant impact on the naturalness of the overall interaction. The expression needs to form a stable coordination with the movement trajectory. Therefore, the expression unit with the highest stability matching the movement trajectory should be selected first. That is, the expression unit is selected based on the trajectory coordination degree.

[0141] Specifically, under abnormal testing conditions, the weight reference value is reduced based on the number of selected facial expression units;

[0142] The abnormal test condition is that the abnormal reference value is greater than the preset abnormal reference value, and the decrease in the weight reference value is positively correlated with the number of selected expression units.

[0143] Specifically, if the target video frame is a Class I video frame, then the abnormal reference value = CPU utilization rate / CPU utilization rate threshold; the CPU utilization rate threshold is 50%, and the CPU utilization rate is the CPU utilization rate before the rendering of the target video frame begins, collected through the PDH interface; if the target video frame is a Class II video frame, then the abnormal reference value = the average of the generation time required for each video frame whose generation time is earlier than the target video frame / the average of the generation time required for each video frame in the historical records that can meet the user's needs;

[0144] The decrease in the weight reference value = number of selected emoji units / preset number of selected emoji units × weight threshold, where the weight threshold is 0.2;

[0145] Users can determine the preset abnormal reference value and the preset number of selected facial expression units according to the actual application scenario. The greater the user's demand for improving the accuracy of rendering smoothness, the smaller the preset abnormal reference value and the preset number of selected facial expression units should be. A method for determining the preset abnormal reference value and the preset number of selected facial expression units is provided. The historical records of adjusting the weight reference value by reducing it according to the number of selected facial expression units are detected. The average value of the abnormal reference value and the average value of the number of selected facial expression units corresponding to the historical records that meet the user's needs are respectively recorded as the preset abnormal reference value and the preset number of selected facial expression units.

[0146] The methods for setting the expression weight for a single selected expression unit in a target video frame include:

[0147] If the target video frame does not have a preceding video frame, or the selected expression unit does not appear in the preceding video frame, then the expression weight takes the initial base value of 0.2.

[0148] If the selected expression unit has appeared in the preceding video frame, then the semantic continuity σ is calculated, where σ = the keyword matching degree between the target keyword corresponding to the target video frame and the target keyword corresponding to the preceding video frame.

[0149] If σ≤0.6, it is considered a sudden change in emotion, and the weight of the expression is reset to 0.2;

[0150] If σ > 0.6, then adaptive fine-tuning is performed, slightly increasing when σ > 0.8 and slightly decreasing when σ < 0.8, with the change range always ≤ 0.1.

[0151] The sub-weight of a single facial point in a single expression unit = the reference value of the change distance of the facial point corresponding to the expression unit in the target video frame / the allowable change distance threshold of the facial point corresponding to the expression unit; it should be noted that the sub-weights of each facial point corresponding to a single expression unit are equal, and the expression weight of a single expression unit = the sub-weight of any facial point in the expression unit; the maximum value of the expression weight is 1;

[0152] The decrease in the expression weight corresponding to a single selected expression unit = the decrease in the weight reference value / the number of selected expression units; the expression weight corresponding to the initial expression unit in the target video frame remains unchanged;

[0153] The target video frame is rendered based on the facial weights corresponding to each facial expression unit. It should be noted that if two or more facial expression units need to drive the same facial point, only the facial expression unit with the highest weight is used to render that facial point, and the driving effect of the other facial expression units at that point is completely suppressed, thereby avoiding motion amplitude conflicts.

[0154] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.

Claims

1. A method for interactive rendering and testing of virtual digital humans, characterized in that, include: Acquire the target video frame and the corresponding interactive voice; The video frame category is determined based on the existence of a preceding video frame, and whether to perform rendering optimization is determined based on the video frame category of the target video frame, based on the fine-grained data or the inter-frame variation coefficient. In the rendering optimization, text matching rendering or motion analysis rendering is performed based on the video frame category of the target video frame to determine the selection of expression units. In text matching rendering, matching expansion or search expansion is determined based on the number of synchronized keywords to determine the selection of emoji units; In motion analysis and rendering, the number of selected facial expression units is determined based on the amplitude deviation, and the number of selected facial expression units is adjusted based on the facial expression transition delay coefficient. The selection of facial expression units is also determined based on the action boundary value, the facial expression unit fit or trajectory coordination. Under abnormal test conditions, a weight reference value is determined based on the sum of the weights of the selected facial expression units, and the weight reference value is adjusted according to the number of selected facial expression units. Rendering is performed on the target video frame based on the facial expression weights corresponding to the facial expression units; Amplitude deviation = limb point change representation value corresponding to the target video frame - limb point change representation value corresponding to the preceding video frame of the target video frame; the limb point change representation value corresponding to a single video frame is the sum of the change distance reference values ​​corresponding to each changed limb point in that single video frame. Expression transition delay coefficient = average of the generation time required for each video frame whose generation time is earlier than the target video frame / average of the number of expression units corresponding to each video frame whose generation time is earlier than the target video frame - average of the generation time required for each video frame in the history that can meet the user's needs / average of the number of expression units corresponding to each video frame in the history that can meet the user's needs. The fit of a single historical emoji unit is equal to the number of historical video frames in which that historical emoji unit appears, divided by the total number of historical video frames. The trajectory coordination degree corresponding to a single second historical expression unit = 1 / (standard deviation of the facial point change representation value corresponding to each second video frame in which the second historical expression unit appears + 1); the facial point change representation value corresponding to a single video frame is the sum of the change distance reference values ​​corresponding to each changed facial point in that single video frame.

2. The interactive rendering and testing method for virtual digital humans according to claim 1, characterized in that, For target video frames that belong to the same video frame category, determine whether to perform rendering optimization based on fine-grained data. For a target video frame that is a Class II video frame, determine whether to perform rendering optimization based on the inter-frame variation coefficient. The video frame categories include a category of video frames that do not have preceding video frames and a category of video frames that do have preceding video frames.

3. The interactive rendering and testing method for virtual digital humans according to claim 2, characterized in that, For target video frames whose video frame category is Class 1, perform text matching and rendering; Motion analysis and rendering are performed on target video frames that are classified as Class II video frames.

4. The interactive rendering and testing method for virtual digital humans according to claim 1, characterized in that, Matching and expanding is performed on target video frames where the number of synchronized keywords is greater than or equal to the preset number of synchronized keywords; In the matching expansion, the selected expression unit is determined based on the degree of synchronous correlation, the degree of influence of interval, or the degree of convergence of matching changes. The synchronization keywords are determined based on keyword matching degree and frequency of occurrence; The interval influence of a single synchronized keyword = the time interval value of the target text / the average time interval value of each reference text containing the synchronized keyword; The convergence of collocation changes corresponding to a single synchronized keyword is the average of the convergence of sub-changes corresponding to each reference text that contains the synchronized keyword; the convergence of sub-changes corresponding to a single reference text that contains the synchronized keyword = the number of facial expression units that appear in all video frames between the video frame corresponding to the synchronized keyword and the video frame corresponding to the target keyword in the reference text / the total number of facial expression units that appear in all video frames between the video frame corresponding to the synchronized keyword and the video frame corresponding to the target keyword in the reference text.

5. The interactive rendering and testing method for virtual digital humans according to claim 4, characterized in that, The selection of facial expression units is determined based on the degree of synchronous correlation, the degree of influence of interval, or the degree of convergence of matching changes, including: For target video frames with a synchronization correlation degree greater than or equal to the preset synchronization correlation degree, the selected expression unit is determined based on the interval influence degree. For target video frames with a synchronization correlation degree less than the preset synchronization correlation degree, the selected expression unit is determined based on the similarity of the matching changes.

6. The interactive rendering and testing method for virtual digital humans according to claim 1, characterized in that, For target video frames where the number of synchronized keywords is less than the preset number of synchronized keywords, the search is expanded; In the search expansion process, the search depth is determined based on the initial search effectiveness of the target keywords, the matching keywords are determined based on the matching coefficient, and the emoji units to be selected are determined based on the matching keywords. The matching coefficient is determined based on the proportion of the floating page and the page effectiveness.

7. The interactive rendering and testing method for virtual digital humans according to claim 1, characterized in that, The number of facial expression units to be selected is determined based on the amplitude deviation. The number of selected facial expression units is positively correlated with the degree of amplitude deviation.

8. The interactive rendering and testing method for virtual digital humans according to claim 7, characterized in that, The number of selected facial expression units is reduced based on the facial expression transition delay coefficient. The decrease in the number of selected facial expression units is positively correlated with the facial expression transition delay coefficient.

9. The interactive rendering and testing method for virtual digital humans according to claim 8, characterized in that, The selection of facial expression units is determined based on the action boundary value, the fit of the facial expression unit, or the trajectory coordination. This includes: For target video frames with motion boundary values ​​less than preset motion boundary values, select the appropriate facial expression unit based on the fit of the facial expression unit. For target video frames whose action boundary values ​​are greater than or equal to preset action boundary values, the selected facial expression unit is determined based on the trajectory coordination degree.

10. The interactive rendering and testing method for virtual digital humans according to claim 1, characterized in that, Under abnormal testing conditions, the weight reference value is adjusted to decrease based on the number of selected facial expression units; The abnormal test condition is that the abnormal reference value is greater than the preset abnormal reference value, and the decrease in the weight reference value is positively correlated with the number of selected expression units.

Citation Information

Patent Citations

  • Digital human video generation method based on multi-modal large model

    CN120472059A

  • Digital human generation method based on multi-modal large model

    CN120543710A