Digital Human Quality Assessment Method and Device
By analyzing the human characteristics, key points of movement, and audio-visual synchronization of digital humans, the problem of incomplete quality assessment of digital humans in existing technologies has been solved, and a comprehensive and accurate assessment of the quality of digital humans has been achieved.
Patent Information
- Application Number
- CN202511151652.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-08-18
AI Technical Summary
Existing technologies do not provide a comprehensive assessment of the quality of digital humans, lacking evaluation of the digital human's character, movements, and audio-visual synchronization.
By analyzing the characteristics of the person in the reference data and the person in the digital human video, the difference values of the key points of the human body parts in the motion-driven video, the difference values of the pixel motion vectors in the digital human video, and the audio-video synchronization of the digital human audio, the scores of person consistency, motion consistency, motion coherence, audio reverberation and audio-video synchronization are calculated, and these scores are weighted and summed to obtain a comprehensive digital human quality score.
It enables a more comprehensive assessment of digital human quality, accurately evaluating the consistency and coherence of digital human video movements, as well as audio-visual synchronization, thus providing a more accurate assessment of digital human quality.
Smart Images

Figure CN120726537B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a method and apparatus for assessing the quality of digital humans. Background Technology
[0002] As digital human technology is widely used in many fields such as virtual customer service, intelligent assistants and virtual anchors, the quality requirements for generated digital humans are increasing. As a result, some tools have emerged to assist humans in automatically evaluating digital humans.
[0003] In existing related technologies, the aforementioned automatic evaluation tools mainly assess the quality of the video and audio corresponding to the digital human, but lack assessment of the digital human's character, movements, and audio-visual synchronization, resulting in an incomplete assessment of the digital human's quality. Summary of the Invention
[0004] This invention provides a method and apparatus for assessing the quality of digital humans, in order to solve the problem that the assessment of digital human quality in the prior art is not comprehensive enough.
[0005] This invention provides a method for assessing the quality of a digital human, comprising:
[0006] By analyzing the characteristics of the individuals in the reference data and the digital human videos, the consistency score of the individuals in the digital human videos is obtained.
[0007] Analyze the differences in motion key points of corresponding human body parts in motion-driven videos and digital human videos, and determine the motion consistency score of the digital human video based on the differences in motion key points of corresponding human body parts.
[0008] The motion vector difference value of each pixel in the digital human video between each two adjacent frames is analyzed, and the motion coherence score of the digital human video is determined by the difference value of each motion vector.
[0009] The audio of the digital human is analyzed to obtain the audio reverberation score and audio quality score.
[0010] The synchronization between the speaking actions in the digital human video and the speech in the digital human audio is analyzed to obtain the audio-video synchronization score of the digital human.
[0011] The digital human quality score is obtained by weighting and summing the character consistency score, action consistency score, action coherence score, audio reverberation score, audio quality score, and audio-video synchronization score.
[0012] The digital human video is generated based on the person's reference data and motion-driven video.
[0013] According to the present invention, a method for assessing the quality of a digital human is provided, wherein the reference data for the human figure is a human figure image;
[0014] Analyzing the characteristics of individuals in the reference data and digital human videos, a consistency score for the individuals in the digital human videos is obtained, including:
[0015] Extract the first human feature from the image of the person;
[0016] Extract the character feature sequence of the person in each frame of the digital human video;
[0017] Calculate the similarity between each second character feature in the character feature sequence and the first character feature to obtain a similarity sequence;
[0018] The mean similarity of each similarity in the similarity sequence is calculated, and the mean similarity is used as the consistency score of the person.
[0019] According to the present invention, a method for assessing the quality of a digital human is provided, wherein the reference data for the human is a video of the human.
[0020] Analyzing the characteristics of individuals in the reference data and digital human videos, a consistency score for the individuals in the digital human videos is obtained, including:
[0021] Based on the duration of the digital human video, the person video and the digital human video are time-aligned;
[0022] Feature extraction is performed on each frame of the character video to obtain the first character feature sequence;
[0023] Feature extraction is performed on the characters in each frame of the digital human video to obtain a second character feature sequence;
[0024] Calculate the similarity of the character features in the corresponding time frames of the first character feature sequence and the second character feature sequence to obtain a similarity sequence;
[0025] The mean similarity of each similarity in the similarity sequence is calculated, and the mean similarity is used as the consistency score of the person.
[0026] According to a digital human quality assessment method provided by the present invention, the method analyzes the differences in motion key points of corresponding human body parts in motion-driven videos and digital human videos, and determines the motion consistency score of the digital human video based on the differences in motion key points of corresponding human body parts, including:
[0027] Extract the coordinates of multiple first motion key points from the motion-driven video. ,in, Indicates the first action in the motion-driven video iHuman body parts in the frame j The x-coordinate of the key point Indicates the first action in the motion-driven video i Human body parts in the frame j The ordinate of the key points;
[0028] Extract multiple second action key point coordinates from the digital human video. ,in, Indicates the first in the digital human video i Human body parts in the frame j The x-coordinate of the key point Indicates the first in the digital human video i Human body parts in the frame j The ordinate of the key points;
[0029] Based on the coordinates of the key points of the first action Coordinates of the second key point of the corresponding human body part The difference is used to determine the consistency score of the action.
[0030] According to the present invention, a digital human quality assessment method is provided, based on the coordinates of a first action key point. Coordinates of the second key point of the corresponding human body part The difference is used to determine the consistency score of the action, including:
[0031] For frames at corresponding times in the motion-driven video and the digital human video, the sum of squared differences between the coordinates of the first motion keypoint and the coordinates of the second motion keypoint corresponding to each human body part is calculated using the following formula:
[0032] ;
[0033] The action consistency score is obtained by summing the squared differences of the corresponding frames, taking the square root of the sum, and then summing the squared differences of the sums of the squared differences of the frames. :
[0034] ;
[0035] in, M This indicates the number of motion keypoints corresponding to each frame. N Indicates the number of frames.
[0036] According to a digital human quality assessment method provided by the present invention, the method analyzes the difference value of the motion vector of each pixel in a digital human video between each two adjacent frames, and determines the motion coherence score of the digital human video based on the difference value of each motion vector, including:
[0037] Calculate the optical flow field between two adjacent frames in the digital human video, wherein the optical flow field includes the motion direction and / or motion velocity of each pixel between adjacent frames;
[0038] Calculate the angle difference and / or velocity difference of each pixel in the direction of motion between adjacent frames;
[0039] The motion continuity score is determined based on the angle difference and / or speed difference.
[0040] According to a digital human quality assessment method provided by the present invention, the synchronization between speaking actions in the digital human video and speech in the digital human audio is analyzed to obtain an audio-video synchronization score for the digital human, including:
[0041] Determine the first start time and the first end time of each speaking action segment in the digital human video. The speaking action segment is divided into N consecutive closed-mouth frames. The first start time is the time of the first open-mouth frame of the character in the speaking action segment, and the first end time is the time of the first closed-mouth frame in the N consecutive closed-mouth frames after the speaking action segment.
[0042] Determine the second start time and the second end time of each speech segment in the digital human audio, wherein the speech segment is divided into segments based on the duration of N frames after the speech stops;
[0043] For the corresponding speech action segment and speech segment, calculate the difference between the first start time and the second start time, and calculate the difference between the first end time and the second end time;
[0044] The mean of all differences is determined as the audio-video synchronization score.
[0045] The present invention also provides a digital human quality assessment device, comprising the following units:
[0046] The character consistency analysis unit is used to analyze the character reference data and the characteristics of each character in the digital human video to obtain the character consistency score of the digital human video.
[0047] The motion consistency analysis unit is used to analyze the difference values of the motion key points of each character in the motion-driven video and the digital human video in the corresponding human body parts, and to determine the motion consistency score of the digital human video based on the difference values of the motion key points of each corresponding human body part.
[0048] The motion coherence analysis unit is used to analyze the difference value of the motion vector of each pixel in the digital human video between each two adjacent frames, and to determine the motion coherence score of the digital human video based on the difference value of each motion vector.
[0049] The audio analysis unit is used to analyze the digital human's audio and obtain the audio reverberation score and audio quality score of the digital human's audio.
[0050] The audio-visual synchronization analysis unit is used to analyze the synchronization between the speaking actions in the digital human video and the speech in the digital human audio, and to obtain the audio-visual synchronization score of the digital human.
[0051] The quality score evaluation unit is used to weight and sum the character consistency score, action consistency score, action coherence score, audio reverberation score, audio quality score, and audio-video synchronization score to obtain the digital human quality score.
[0052] The digital human video is generated based on the person's reference data and motion-driven video.
[0053] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the program to implement the digital human quality assessment method as described above.
[0054] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the digital human quality assessment method as described above.
[0055] The digital human quality assessment method and apparatus provided by this invention analyzes aspects such as character consistency, motion consistency, motion coherence, audio reverberation, audio quality, and audio-video synchronization, obtaining corresponding scores. These scores are then weighted and summed to obtain a comprehensive digital human quality score, thus achieving a more comprehensive assessment of digital human quality compared to existing methods. Furthermore, by comparing the differences in motion key points of corresponding body parts in motion-driven videos and digital human videos, the motion consistency score of the digital human video is determined based on these differences. Additionally, by comparing the differences in motion vectors of each pixel in the digital human video between adjacent frames, the motion coherence score of the digital human video is determined based on these differences, thereby achieving a more accurate assessment of the digital human video quality. Attached Figure Description
[0056] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0057] Figure 1 This is a flowchart illustrating the digital human quality assessment method provided by the present invention.
[0058] Figure 2This is a schematic diagram illustrating the principle of the digital human quality assessment method provided by this invention.
[0059] Figure 3 This is a schematic diagram of the structure of the digital human quality assessment device provided by the present invention.
[0060] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0061] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0062] The digital human quality assessment method of this invention, such as... Figure 1 and Figure 2 As shown, the procedure includes steps S110 to S160.
[0063] Step S110: Analyze the features of the people in the reference data and the digital human video to obtain the person consistency score of the digital human video. The reference data is one of the inputs for generating the digital human, which can be a person image or a person video. When generating the digital human, the person in the digital human video must be consistent with the person in the reference data, mainly in terms of facial consistency. If they are inconsistent, or if the appearance of the people is significantly different, the quality of the generated digital human will be poor.
[0064] In this step, the features of people in the reference data and the features of people in the generated digital human video can be compared. For example, feature similarity can be calculated, and the statistical value of the similarity (e.g., mean or variance) can be used as the person consistency score of the digital human video. .
[0065] Step S120: Analyze the differences in motion keypoints of corresponding body parts between the characters in the motion-driven video and the digital human video, and determine the motion consistency score of the digital human video based on the differences in motion keypoints of corresponding body parts. The motion-driven video is another input for generating the digital human; that is, the digital human video is generated based on character reference data and the motion-driven video. The aforementioned character reference data provides a reference for the characters in the digital human video, ensuring that the characters in the digital human video are consistent with the characters in the character reference data. The motion-driven video drives the actions of the characters in the digital human video, providing motion references so that the characters in the digital human video perform the actions of the characters in the motion-driven video.
[0066] In this step, a motion consistency score for the digital human video is obtained by comparing the consistency of the movements of the characters in the motion-driven video and the digital human video. For example, the differences in motion keypoints of corresponding body parts for each character in the two videos can be compared, and the statistical values of these differences (e.g., mean or variance) are used as the motion consistency score for the digital human video. .
[0067] Step S130: Analyze the difference value of the motion vector of each pixel in the digital human video between each adjacent frame, and determine the motion coherence score of the digital human video based on the difference value of each motion vector. The more coherent the movements of the person in the digital human video, the higher the quality of the digital human video.
[0068] For example, the motion coherence score of the digital human video can be obtained by comparing the difference values of the motion vectors of each pixel in the video between each two adjacent frames, and using the statistical values (e.g., mean or variance) of each difference value. .
[0069] Step S140: Analyze the digital human audio to obtain the audio reverberation score of the digital human audio. and audio quality score The audio format of digital humans is usually common audio formats such as WAV and MP3. In this step, audio reverberation analysis and audio quality analysis can be performed on the digital human audio to obtain an audio reverberation score. and audio quality score .
[0070] Step S150: Analyze the synchronization between the speaking actions in the digital human video and the speech in the digital human audio to obtain the audio-video synchronization score of the digital human. Audio-video synchronization reflects whether the speaking actions of the character in the digital human video match the corresponding speech. That is, there should be no situation where the character has opened his mouth to speak, but the corresponding sound lags behind the speaking action by a large amount of time, or the sound has been broadcast, but the character's speaking action lags behind the corresponding sound by a large amount of time.
[0071] For example, in this step, the audio-visual synchronization score of the digital human video can be obtained by comparing the difference between the start time and the end time of each speech action segment and the corresponding speech segment, and using the statistical value (e.g., mean or variance) of each difference value. .
[0072] Step S160: The digital human quality score is obtained by weighted summing of the character consistency score, action consistency score, action coherence score, audio reverberation score, audio quality score, and audio-video synchronization score. The weights for each score are set according to the actual situation, ensuring that the sum of the absolute values of all weights is 1. Therefore, the final digital human quality score is obtained. for:
[0073] .
[0074] in, , , , , and These are the weights of the character consistency score, action consistency score, action continuity score, audio reverberation score, audio quality score, and audio-visual synchronization score.
[0075] It should be noted that the scores for motion consistency, motion coherence, and audio-visual synchronization are all based on the statistical values of their respective differences. Since smaller differences indicate better quality, and to align with the principle that higher scores equate to higher quality, the weights of these scores are negative. For example, the weight of the character consistency score... The weight of the action consistency score is 0.2. The weight of the motion continuity score is -0.2. The weight of the audio reverb score is -0.15. The weight of the audio quality score is 0.1. The weight of the audio-video synchronization score is 0.2. The final digital human quality score is -0.15, calculated using the weights mentioned above for each score. for:
[0076] .
[0077] Digital Human Quality Score A higher value indicates better quality. This means that before calculating the digital human quality score, each score undergoes dimensionality elimination and normalization, normalizing it to the range of [0,1], [0,10], or [0,100] to ensure a consistent digital human quality score. It also falls within the range of [0,1], [0,10], or [0,100]. Furthermore, a quality pass / fail score threshold can be obtained through manual evaluation. A score greater than or equal to this threshold indicates that the generated digital human is of acceptable quality.
[0078] The digital human quality assessment method in this embodiment analyzes character consistency, motion consistency, motion coherence, audio reverberation, audio quality, and audio-video synchronization, obtaining corresponding scores. These scores are then weighted and summed to obtain a comprehensive digital human quality score, thus achieving a more comprehensive assessment of digital human quality compared to existing methods. Furthermore, it compares the differences in motion key points of corresponding body parts in motion-driven videos and digital human videos, using these differences to determine the motion consistency score of the digital human video. Additionally, it compares the differences in motion vectors of each pixel in the digital human video between adjacent frames, using these differences to determine the motion coherence score of the digital human video, thereby achieving a more accurate assessment of the digital human's video quality.
[0079] In some embodiments, the SRMR (Speech to Noise Ratio in the Modulation Domain) algorithm can be used to perform audio reverberation analysis on digital human audio. SRMR is a reverberation detection and quantification method that is mainly based on the modulation energy characteristics of speech signals. It evaluates the reverberation level by analyzing the modulation energy distribution of speech signals in different frequency bands, obtaining the ratio of speech to reverberation modulation energy, thereby quantifying the reverberation level of the speech signal. This can then be used to detect whether the speech signal has undergone déverification processing. Specifically, the digital human audio is input into SRMR to obtain the speech to reverberation modulation energy ratio of the digital human audio, and this ratio is used as the audio reverberation score. .
[0080] Audio quality analysis is primarily achieved using UTMOSV2, a sound quality assessment model for evaluating the Mean Opinion Score (MOS). It integrates spectral and speech features during training, using a pre-trained EfficientNetV2 as the image feature extractor and a pre-trained wav2vec 2.0 as the speech feature extractor. The features and data domain encodings from both are fused and input into a fully connected layer to predict the MOS. This effectively improves the accuracy and relevance of speech naturalness prediction and enhances the model's ability to evaluate high-quality speech. Specifically, the digital human's audio is input into UTMOSV2, and UTMOSV2 outputs the MOS, which is used as the audio quality score. .
[0081] It should be noted that before performing audio reverberation analysis and audio quality analysis, it is preferable to normalize the digital human audio and adjust the audio amplitude to the standard range of -6dB to -12dB to eliminate differences in audio amplitude and ensure the accuracy of subsequent analysis.
[0082] In some embodiments, the person reference data is a person image, that is, a static person image. Based on this, step S110 specifically includes:
[0083] Step 1: Extract the first person feature from the person image. Since person consistency mainly judges whether the person in the digital human video and the person in the person image are the same person, the first person feature includes facial features. For example, a face detection algorithm can be used to extract the first person feature, such as using the ArcFace model to extract facial features from the person image as the first person feature. ArcFace introduces additive angular margin loss, which makes the interval between different categories of facial features in the feature space larger. In this way, ArcFace significantly improves the discrimination ability of face recognition, making the model more robust to complex situations and enhancing the model's recognition accuracy.
[0084] Step 2: Extract the character feature sequence from each frame of the digital human video. In this step, the ArcFace model can also be used to extract facial features from each frame as secondary character features, forming a character feature sequence from each frame.
[0085] Step 3: Calculate the similarity between each second character feature in the character feature sequence and the first character feature to obtain a similarity sequence. Specifically, the cosine similarity algorithm can be used to calculate the similarity between each second character feature and the first character feature to obtain a similarity sequence.
[0086] Step 4: Calculate the mean similarity for each similarity in the similarity sequence, and use the mean similarity as the consistency score for the person. .
[0087] In this embodiment, the similarity between the features of the person in each frame of the digital human video and the features of the person in the human image is calculated, and the average similarity is used as the consistency score of the person. The larger the average similarity, the better the consistency of the person, thereby achieving an accurate assessment of the consistency of the person.
[0088] In some embodiments, the person reference data is a person video, that is, the person reference data is a dynamic person video. Based on this, step S110 specifically includes:
[0089] Step 1: Based on the duration of the digital human video, align the person video and the digital human video in terms of time. Alignment means matching the duration of the person video with the duration of the digital human video. If the person video is longer than the digital human video, simply extract a video segment of the same length. If the person video is shorter than the digital human video, repeatedly splice multiple person videos until they are the same length as the digital human video.
[0090] Step 2: Extract features from each frame of the person video to obtain a first person feature sequence. Since person consistency mainly judges whether the person in the digital human video and the person in the real person video are the same person, the person features in the first person feature sequence include facial features. For example, a face detection algorithm can be used to extract the first person features. For instance, the ArcFace model can be used to extract the facial features from each frame of the person video as the first person features, and the first person features of each frame form the first person feature sequence.
[0091] Step 3: Extract features from the people in each frame of the digital human video to obtain a second person feature sequence. Alternatively, the ArcFace model can be used to extract facial features from each frame as the second person features, forming a person feature sequence from the second person features of each frame.
[0092] Step 4: Calculate the similarity of character features in corresponding time frames of the first and second character feature sequences to obtain a similarity sequence. The cosine similarity algorithm can be used to calculate the similarity of character features in each corresponding frame.
[0093] Step 5: Calculate the mean similarity for each similarity in the similarity sequence, and use the mean similarity as the consistency score for the person. .
[0094] In this embodiment, the similarity between the human features in each frame of the digital human video and the human features in the corresponding frame of the human video is calculated, and the average similarity is used as the consistency score of the human. The larger the average similarity, the better the consistency of the human, thereby achieving an accurate assessment of the consistency of the human.
[0095] In some embodiments, step S120 specifically includes:
[0096] Step 1: Extract the coordinates of multiple first motion key points from the motion-driven video. ,in, Indicates the first action in the motion-driven video i Human body parts in the frame j The x-coordinate of the key point Indicates the first action in the motion-driven videoi Human body parts in the frame j The vertical coordinates of key points. For example, for the action of speaking, in each frame, the body parts include at least the upper lip, lower lip, left cheek, and right cheek.
[0097] For example, motion detection algorithms can be used to extract the coordinates of multiple first motion key points from motion-driven videos. For example, the Unipose model can be used. The Unipose model is a general-purpose keypoint detection model that can detect keypoints of any object. In this step, the motion-driven video is input into the Unipose model, which will output the coordinates of multiple first motion keypoints for each frame of the motion-driven video. .
[0098] Step 2: Extract the coordinates of multiple second action key points from the digital human video. ,in, Indicates the first in the digital human video i Human body parts in the frame j The x-coordinate of the key point Indicates the first in the digital human video i Human body parts in the frame j The ordinates of key points. In this step, the Unipose model can also be used to extract the coordinates of multiple secondary action key points in each frame of the digital human video. .
[0099] Step 3: Based on the coordinates of the first action key points Coordinates of the second key point of the corresponding human body part The difference is used to determine the consistency score of the action. The smaller the difference, the higher the consistency of motion between the digital human video and the motion-driven video, and the better the quality.
[0100] In this embodiment, by comparing the differences in the coordinates of the key points of the corresponding human body parts in each frame of the motion-driven video and the digital human video, the motion consistency score of the digital human video is determined by the difference, thereby achieving an accurate assessment of motion consistency.
[0101] Furthermore, based on the coordinates of the key points of the first action Coordinates of the second key point of the corresponding human body part The difference is used to determine the consistency score of the action, including:
[0102] For frames at corresponding times in the motion-driven video and the digital human video, the sum of squared differences between the coordinates of the first motion keypoint and the coordinates of the second motion keypoint corresponding to each human body part is calculated using the following formula:
[0103] .
[0104] The action consistency score is obtained by summing the squared differences of the corresponding frames, taking the square root of the sum, and then summing the squared differences of the sums of the squared differences of the frames. :
[0105] .
[0106] in, M This indicates the number of motion keypoints corresponding to each frame. N Indicates the number of frames.
[0107] In this embodiment, the statistical value of the difference in the coordinates of the key points of the action is obtained by using the above formula. The calculation is simple and has low complexity. The calculated action consistency score can accurately reflect the action consistency between the digital human video and the action-driven video.
[0108] In some embodiments, step S130 specifically includes:
[0109] Step 1: Calculate the optical flow field between two adjacent frames in the digital human video. The optical flow field includes the motion direction and / or motion velocity of each pixel between adjacent frames. The motion vector of each pixel between two adjacent frames is characterized by the motion direction and / or motion velocity of each pixel between adjacent frames in the optical flow field.
[0110] For example, in this step, the StreamFlow model can be used to calculate the optical flow field between two adjacent frames in the digital human video. StreamFlow is a multi-frame optical flow estimation method that excels at using efficient spatiotemporal relationship mining techniques to identify the optical flow between multiple video frames, thereby obtaining a highly accurate optical flow field. Specifically, the digital human video is input into the StreamFlow model, which outputs the optical flow field between adjacent frames in the digital human video.
[0111] Step 2: Calculate the angle difference and / or velocity difference of each pixel in the direction of motion between adjacent frames. The smaller the angle and the smaller the velocity difference, the higher the motion continuity.
[0112] Step 3: Determine the motion continuity score based on the angle difference and / or velocity difference. For example, the average of all angle differences across all pixels can be taken as the first score, and the average of all velocity differences across all pixels can be taken as the second score, i.e., the motion continuity score. It contains two components: the first score Second score You can choose one of the sub-items as the motion continuity score, or you can use both sub-items as the motion continuity score. Using both sub-items as the motion continuity score allows for a more accurate assessment of motion continuity. When both sub-items are used as the motion continuity score, then the above calculation... The formula can be written as:
[0113] .
[0114] in, .For example:
[0115] .
[0116] In this embodiment, by calculating the optical flow field between two adjacent frames in the digital human video, the angle difference and / or velocity difference of the movement direction of pixels in the optical flow field between each adjacent frame can characterize the continuity of the action. Therefore, determining the action continuity score based on the angle difference and / or velocity difference can accurately reflect the action continuity in the digital human video, thereby achieving an accurate evaluation of the action continuity in the digital human video.
[0117] In some embodiments, step S150 specifically includes:
[0118] Step 1: Determine the first start and end times of each speaking action segment in the digital human video. Each speaking action segment is divided into N consecutive closed-mouth frames. The first start time is the moment of the first open-mouth frame in the speaking action segment, and the first end time is the moment of the first closed-mouth frame among the N consecutive closed-mouth frames following the speaking action segment. During speech, there are pauses of a certain duration between sentences. Corresponding to speaking actions, this is represented by the duration of mouth closure, i.e., the number of closed-mouth frames. Therefore, in this step, the speaking action segment is divided into N consecutive closed-mouth frames (e.g., 3-5, determined by the specific speaking speed).
[0119] Step Two: Determine the second start time and second end time of each speech segment in the digital human audio. The speech segment is divided into segments based on the duration of N frames following the cessation of speech. During human speech, there are pauses of a certain duration between sentences. In speech, this corresponds to the duration of the speech pause, which is the duration of N (no speech) frames following the cessation of speech. Therefore, in this step, the speech segment is divided into segments based on the duration of N (no speech) frames following the cessation of speech.
[0120] Step 3: For the corresponding speech action segment and audio segment, calculate the difference between the first start time and the second start time, and calculate the difference between the first end time and the second end time. After the first two steps, starting from the moment the video began playing, the video was divided into speech action segments in chronological order, and the audio was divided into corresponding audio segments. Furthermore, the start and end times of each corresponding speech action segment and audio segment have been determined. Therefore, in this step, for each corresponding speech action segment and audio segment, the difference between the first start time and the second start time, as well as the difference between the first end time and the second end time, can be calculated. The smaller the difference, the better the audio-video synchronization.
[0121] Step 4: Determine the mean of all differences as the audio-visual synchronization score.
[0122] In this embodiment, by determining the start and end times of the speaking action segment in the digital human video and the corresponding speech segment in the digital human audio, the difference between the corresponding start times and the difference between the corresponding end times is calculated. The smaller the difference, the better the audio-video synchronization. Therefore, using the mean of the difference as the audio-video synchronization score can achieve a more accurate evaluation of audio-video synchronization.
[0123] The digital human quality assessment device provided by the present invention is described below. The digital human quality assessment device described below can be referred to in correspondence with the digital human quality assessment method described above.
[0124] The digital human quality assessment device of this invention, as described in the embodiments, is as follows: Figure 3 As shown, it includes:
[0125] The character consistency analysis unit 310 is used to analyze the character reference data and the characteristics of each character in the digital human video to obtain the character consistency score of the digital human video.
[0126] The motion consistency analysis unit 320 is used to analyze the difference values of the motion key points of each character in the motion-driven video and the digital human video in the corresponding human body parts, and to determine the motion consistency score of the digital human video based on the difference values of the motion key points of each corresponding human body part.
[0127] The motion coherence analysis unit 330 is used to analyze the difference value of the motion vector of each pixel in the digital human video between each two adjacent frames, and to determine the motion coherence score of the digital human video based on the difference value of each motion vector.
[0128] The audio analysis unit 340 is used to analyze the digital human's audio and obtain the audio reverberation score and audio quality score of the digital human's audio.
[0129] The audio-visual synchronization analysis unit 350 is used to analyze the synchronization between the speaking actions in the digital human video and the speech in the digital human audio, and to obtain the audio-visual synchronization score of the digital human.
[0130] The quality score evaluation unit 360 is used to weight and sum the character consistency score, action consistency score, action continuity score, audio reverberation score, audio quality score, and audio-video synchronization score to obtain the digital human quality score.
[0131] The digital human video is generated based on the person's reference data and motion-driven video.
[0132] In some embodiments, the person reference data is a person image, and the person consistency analysis unit 310 includes:
[0133] The first person feature extraction unit is used to extract the first person feature of the person in the person image.
[0134] The character feature sequence extraction unit is used to extract the character feature sequence of the characters in each frame of the digital human video.
[0135] The similarity calculation unit is used to calculate the similarity between each second character feature in the character feature sequence and the first character feature, thereby obtaining a similarity sequence.
[0136] The similarity mean calculation unit is used to calculate the average similarity of each similarity in the similarity sequence, and to use the average similarity as the consistency score of the person.
[0137] In some embodiments, the person reference data is a person video; the person consistency analysis unit 310 includes:
[0138] The time alignment unit is used to align the person video and the digital human video in time based on the duration of the digital human video.
[0139] The first character feature extraction unit is used to extract features from each frame of the character video to obtain the first character feature sequence.
[0140] The second character feature extraction unit is used to extract features from the characters in each frame of the digital human video to obtain a second character feature sequence.
[0141] The similarity calculation unit is used to calculate the similarity of the character features of the corresponding time frames in the first character feature sequence and the second character feature sequence to obtain a similarity sequence.
[0142] The similarity mean calculation unit is used to calculate the average similarity of each similarity in the similarity sequence, and to use the average similarity as the consistency score of the person.
[0143] In some embodiments, the action consistency analysis unit 320 includes:
[0144] The first key point extraction unit is used to extract the coordinates of multiple first action key points from the action-driven video. ,in, Indicates the first action in the motion-driven video i Human body parts in the frame j The x-coordinate of the key point Indicates the first action in the motion-driven video i Human body parts in the frame j The ordinate of the key point.
[0145] The second key point extraction unit is used to extract multiple second action key point coordinates from the digital human video. ,in, Indicates the first in the digital human video i Human body parts in the frame j The x-coordinate of the key point Indicates the first in the digital human video i Human body parts in the frame j The ordinate of the key point.
[0146] Action consistency score determination unit, used to determine the score based on the coordinates of the first action key points. Coordinates of the second key point of the corresponding human body part The difference is used to determine the consistency score of the action.
[0147] In some embodiments, the action consistency score determination unit is specifically used for:
[0148] For frames at corresponding times in the motion-driven video and the digital human video, the sum of squared differences between the coordinates of the first motion keypoint and the coordinates of the second motion keypoint corresponding to each human body part is calculated using the following formula:
[0149] .
[0150] The action consistency score is obtained by summing the squared differences of the corresponding frames, taking the square root of the sum, and then summing the squared differences of the sums of the squared differences of the frames. :
[0151] .
[0152] in, M This indicates the number of motion keypoints corresponding to each frame. N Indicates the number of frames.
[0153] In some embodiments, the motion coherence analysis unit 330 includes:
[0154] The optical flow field calculation unit is used to calculate the optical flow field between two adjacent frames in the digital human video. The optical flow field includes the motion direction and / or motion speed of each pixel between adjacent frames.
[0155] The motion difference calculation unit is used to calculate the angle difference of the motion direction and / or the velocity difference of the motion speed of each pixel between adjacent frames.
[0156] The motion continuity score determination unit is used to determine the motion continuity score based on the included angle difference and / or speed difference.
[0157] In some embodiments, the audio / video synchronization analysis unit 350 includes:
[0158] The first moment determination unit is used to determine the first start moment and the first end moment of each speaking action segment in the digital human video. The speaking action segment is divided into N consecutive closed-mouth frames. The first start moment is the moment of the first open-mouth frame of the character in the speaking action segment, and the first end moment is the moment of the first closed-mouth frame in the N consecutive closed-mouth frames after the speaking action segment.
[0159] The second time determination unit is used to determine the second start time and the second end time of each speech segment in the digital human audio, wherein the speech segment is divided into segments based on the duration of N frames after the speech stops.
[0160] The time difference calculation unit is used to calculate the difference between the first start time and the second start time for the corresponding speech action segment and speech segment, and to calculate the difference between the first end time and the second end time.
[0161] A synchronization score determination unit is used to determine the mean of all differences as the audio-video synchronization score.
[0162] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a digital human quality assessment method, which includes:
[0163] By analyzing the characteristics of the individuals in the reference data and the digital human videos, a consistency score for the individuals in the digital human videos is obtained.
[0164] The differences in motion key points of corresponding human body parts in motion-driven videos and digital human videos are analyzed, and the motion consistency score of the digital human video is determined based on the differences in motion key points of corresponding human body parts.
[0165] The motion vector difference value of each pixel in the digital human video between each adjacent two frames is analyzed, and the motion coherence score of the digital human video is determined by the difference value of each motion vector.
[0166] The audio of the digital human is analyzed to obtain the audio reverberation score and audio quality score.
[0167] The synchronization between the speaking actions in the digital human video and the speech in the digital human audio is analyzed to obtain the audio-video synchronization score of the digital human.
[0168] The digital human quality score is obtained by weighting and summing the scores for character consistency, action consistency, action coherence, audio reverberation, audio quality, and audio-video synchronization.
[0169] The digital human video is generated based on the person's reference data and motion-driven video.
[0170] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0171] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is able to execute the digital human quality assessment method provided by the above methods, the method comprising:
[0172] By analyzing the characteristics of the individuals in the reference data and the digital human videos, a consistency score for the individuals in the digital human videos is obtained.
[0173] The differences in motion key points of corresponding human body parts in motion-driven videos and digital human videos are analyzed, and the motion consistency score of the digital human video is determined based on the differences in motion key points of corresponding human body parts.
[0174] The motion vector difference value of each pixel in the digital human video between each adjacent two frames is analyzed, and the motion coherence score of the digital human video is determined by the difference value of each motion vector.
[0175] The audio of the digital human is analyzed to obtain the audio reverberation score and audio quality score.
[0176] The synchronization between the speaking actions in the digital human video and the speech in the digital human audio is analyzed to obtain the audio-video synchronization score of the digital human.
[0177] The digital human quality score is obtained by weighting and summing the scores for character consistency, action consistency, action coherence, audio reverberation, audio quality, and audio-video synchronization.
[0178] The digital human video is generated based on the person's reference data and motion-driven video.
[0179] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the digital human quality assessment method provided by the methods described above, the method comprising:
[0180] By analyzing the characteristics of the individuals in the reference data and the digital human videos, a consistency score for the individuals in the digital human videos is obtained.
[0181] The differences in motion key points of corresponding human body parts in motion-driven videos and digital human videos are analyzed, and the motion consistency score of the digital human video is determined based on the differences in motion key points of corresponding human body parts.
[0182] The motion vector difference value of each pixel in the digital human video between each adjacent two frames is analyzed, and the motion coherence score of the digital human video is determined by the difference value of each motion vector.
[0183] The audio of the digital human is analyzed to obtain the audio reverberation score and audio quality score.
[0184] The synchronization between the speaking actions in the digital human video and the speech in the digital human audio is analyzed to obtain the audio-video synchronization score of the digital human.
[0185] The digital human quality score is obtained by weighting and summing the scores for character consistency, action consistency, action coherence, audio reverberation, audio quality, and audio-video synchronization.
[0186] The digital human video is generated based on the person's reference data and motion-driven video.
[0187] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0188] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0189] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for assessing the quality of a digital human, characterized in that, include: By analyzing the characteristics of the individuals in the reference data and the digital human videos, the consistency score of the individuals in the digital human videos is obtained. Analyze the differences in motion key points of corresponding human body parts in motion-driven videos and digital human videos, and determine the motion consistency score of the digital human video based on the differences in motion key points of corresponding human body parts. The motion vector difference value of each pixel in the digital human video between each two adjacent frames is analyzed, and the motion coherence score of the digital human video is determined by the difference value of each motion vector. The audio of the digital human is analyzed to obtain the audio reverberation score and audio quality score. The synchronization between the speaking actions in the digital human video and the speech in the digital human audio is analyzed to obtain the audio-video synchronization score of the digital human. The digital human quality score is obtained by weighting and summing the character consistency score, action consistency score, action coherence score, audio reverberation score, audio quality score, and audio-video synchronization score. The digital human video is generated based on the person reference data and motion-driven video. This includes analyzing the differences in motion key points of corresponding body parts in motion-driven videos and digital human videos, and determining the motion consistency score of the digital human video based on the differences in motion key points of corresponding body parts, including: Extract the coordinates of multiple first motion key points from the motion-driven video. ,in, Indicates the first action in the motion-driven video i Human body parts in the frame j The x-coordinate of the key point Indicates the first action in the motion-driven video i Human body parts in the frame j The ordinate of the key points; Extract multiple second action key point coordinates from the digital human video. ,in, Indicates the first in the digital human video i Human body parts in the frame j The x-coordinate of the key point Indicates the first in the digital human video i Human body parts in the frame j The ordinate of the key points; Based on the coordinates of the key points of the first action Coordinates of the second key point of the corresponding human body part The difference is used to determine the consistency score of the action; Among them, based on the coordinates of the key points of the first action Coordinates of the second key point of the corresponding human body part The difference is used to determine the consistency score of the action, including: For frames at corresponding times in the motion-driven video and the digital human video, the sum of squared differences between the coordinates of the first motion keypoint and the coordinates of the second motion keypoint corresponding to each human body part is calculated using the following formula: ; The action consistency score is obtained by summing the squared differences of the corresponding frames, taking the square root of the sum, and then summing the squared differences of the sums of the squared differences of the frames. : ; in, M This indicates the number of motion keypoints corresponding to each frame. N Indicates the number of frames.
2. The digital human quality assessment method according to claim 1, characterized in that, The reference data for the person is a person image; Analyzing the characteristics of individuals in the reference data and digital human videos, a consistency score for the individuals in the digital human videos is obtained, including: Extract the first human feature from the image of the person; Extract the character feature sequence of the person in each frame of the digital human video; Calculate the similarity between each second character feature in the character feature sequence and the first character feature to obtain a similarity sequence; The mean similarity of each similarity in the similarity sequence is calculated, and the mean similarity is used as the consistency score of the person.
3. The digital human quality assessment method according to claim 1, characterized in that, The reference data for the individuals is video footage of them. Analyzing the characteristics of individuals in the reference data and digital human videos, a consistency score for the individuals in the digital human videos is obtained, including: Based on the duration of the digital human video, the person video and the digital human video are time-aligned; Feature extraction is performed on each frame of the character video to obtain the first character feature sequence; Feature extraction is performed on the characters in each frame of the digital human video to obtain a second character feature sequence; Calculate the similarity of the character features in the corresponding time frames of the first character feature sequence and the second character feature sequence to obtain a similarity sequence; The mean similarity of each similarity in the similarity sequence is calculated, and the mean similarity is used as the consistency score of the person.
4. The digital human quality assessment method according to claim 1, characterized in that, Analyze the difference values of motion vectors for each pixel in a digital human video between each pair of adjacent frames, and determine the motion coherence score of the digital human video based on the difference values of each motion vector, including: Calculate the optical flow field between two adjacent frames in the digital human video, wherein the optical flow field includes the motion direction and / or motion velocity of each pixel between adjacent frames; Calculate the angle difference and / or velocity difference of each pixel in the direction of motion between adjacent frames; The motion continuity score is determined based on the angle difference and / or speed difference.
5. The digital human quality assessment method according to claim 1, characterized in that, Analyzing the synchronization between the speaking actions in the digital human's video and the speech in the digital human's audio, an audio-video synchronization score for the digital human is obtained, including: Determine the first start time and the first end time of each speaking action segment in the digital human video. The speaking action segment is divided into N consecutive closed-mouth frames. The first start time is the time of the first open-mouth frame of the character in the speaking action segment, and the first end time is the time of the first closed-mouth frame in the N consecutive closed-mouth frames after the speaking action segment. Determine the second start time and the second end time of each speech segment in the digital human audio, wherein the speech segment is divided into segments based on the duration of N frames after the speech stops; For the corresponding speech action segment and speech segment, calculate the difference between the first start time and the second start time, and calculate the difference between the first end time and the second end time; The mean of all differences is determined as the audio-video synchronization score.
6. A digital human quality assessment device, characterized in that, include: The character consistency analysis unit is used to analyze the character reference data and the characteristics of each character in the digital human video to obtain the character consistency score of the digital human video. The motion consistency analysis unit is used to analyze the difference values of the motion key points of each character in the motion-driven video and the digital human video in the corresponding human body parts, and to determine the motion consistency score of the digital human video based on the difference values of the motion key points of each corresponding human body part. The motion coherence analysis unit is used to analyze the difference value of the motion vector of each pixel in the digital human video between each two adjacent frames, and to determine the motion coherence score of the digital human video based on the difference value of each motion vector. The audio analysis unit is used to analyze the digital human's audio and obtain the audio reverberation score and audio quality score of the digital human's audio. The audio-visual synchronization analysis unit is used to analyze the synchronization between the speaking actions in the digital human video and the speech in the digital human audio, and to obtain the audio-visual synchronization score of the digital human. The quality score evaluation unit is used to weight and sum the character consistency score, action consistency score, action coherence score, audio reverberation score, audio quality score, and audio-video synchronization score to obtain the digital human quality score. The digital human video is generated based on the person reference data and motion-driven video. The action consistency analysis unit includes: The first key point extraction unit is used to extract the coordinates of multiple first action key points from the action-driven video. ,in, Indicates the first action in the motion-driven video i Human body parts in the frame j The x-coordinate of the key point Indicates the first action in the motion-driven video i Human body parts in the frame j The ordinate of the key points; The second key point extraction unit is used to extract multiple second action key point coordinates from the digital human video. ,in, Indicates the first in the digital human video i Human body parts in the frame j The x-coordinate of the key point Indicates the first in the digital human video i Human body parts in the frame j The ordinate of the key points; Action consistency score determination unit, used to determine the score based on the coordinates of the first action key points. Coordinates of the second key point of the corresponding human body part The difference is used to determine the consistency score of the action; The action consistency score determination unit is specifically used for: For frames at corresponding times in the motion-driven video and the digital human video, the sum of squared differences between the coordinates of the first motion keypoint and the coordinates of the second motion keypoint corresponding to each human body part is calculated using the following formula: ; The action consistency score is obtained by summing the squared differences of the corresponding frames, taking the square root of the sum, and then summing the squared differences of the sums of the squared differences of the frames. : ; in, M This indicates the number of motion keypoints corresponding to each frame. N Indicates the number of frames.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the digital human quality assessment method as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the digital human quality assessment method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Video quality evaluation method and device, equipment and storage medium
CN118885821A
Digital human binding evaluation method
CN119068158A