Digital human quality assessment method and device
By analyzing digital human videos and audios from multiple aspects, and calculating scores for character consistency, action consistency, action continuity, and audio-video synchronization, the problem of incomplete digital human quality assessment in existing technologies is solved, achieving a more comprehensive quality assessment.
Patent Information
- Application Number
- CN202511151652.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-18
AI Technical Summary
The existing technology for evaluating the quality of digital humans is not comprehensive enough, and lacks evaluation of the digital human's character itself, character movements, and audio and video synchronization.
By analyzing the characteristics of the character reference data and the digital human video, the character consistency score is calculated; by analyzing the differences in the key points of the action in the action-driven video and the digital human video, the action consistency score is calculated; by analyzing the differences in the motion vectors of the digital human video, the action coherence score is calculated; by analyzing the audio quality and synchronization of the digital human audio, the audio and video synchronization score is calculated; finally, these scores are weighted and summed to obtain the digital human quality score.
It achieves a more comprehensive evaluation of the quality of digital humans, can accurately evaluate the consistency and coherence of the movements of digital human videos, as well as the synchronization of audio and video, and provides a more comprehensive quality assessment method.
Smart Images

Figure CN120726537A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a digital human quality assessment method and device. Background Art
[0002] As digital human technology is widely used in many fields such as virtual customer service, intelligent assistants and virtual anchors, the quality requirements for the generated digital humans are increasing. Therefore, some tools have emerged to assist humans in automatically evaluating digital humans.
[0003] In existing related technologies, the above-mentioned automatic evaluation tools mainly evaluate the quality of digital humans in terms of the quality of the corresponding video and audio of the digital humans. They lack the evaluation of the digital humans themselves, their movements, and the synchronization of audio and video, resulting in an incomplete evaluation of the quality of digital humans. Summary of the Invention
[0004] The present invention provides a digital human quality assessment method and device, which are used to solve the problem that digital human quality assessment in the prior art is not comprehensive enough.
[0005] The present invention provides a digital human quality assessment method, comprising: Analyzing the character reference data and the respective character features in the digital human video to obtain a character consistency score for the digital human video; Analyzing the difference values of the action key points of the corresponding human body parts in the action-driven video and the digital human video, and determining the action consistency score of the digital human video based on the difference values of the action key points of the corresponding human body parts; Analyzing the difference value of the motion vector of each pixel between two adjacent frames in the digital human video, and determining the motion coherence score of the digital human video based on the difference value of each motion vector; Analyze the digital human audio to obtain the audio reverberation score and audio quality score of the digital human audio; Analyzing the synchronization between the speaking action in the digital human video and the voice in the digital human audio to obtain an audio and video synchronization score of the digital human; A weighted summation of the character consistency score, action consistency score, action continuity score, audio reverberation score, audio quality score, and audio-video synchronization score is performed to obtain a digital human quality score; The digital human video is generated based on the character reference data and the action-driven video.
[0006] According to a digital human quality assessment method provided by the present invention, the human reference data is a human image; Analyzing the character reference data and the respective character features in the digital human video to obtain the character consistency score of the digital human video includes: Extracting a first character feature of a person in the character image; Extracting a character feature sequence of a character in each frame of the digital human video; Calculating the similarity between each second character feature in the character feature sequence and the first character feature to obtain a similarity sequence; A similarity mean is calculated for each similarity in the similarity sequence, and the similarity mean is used as the character consistency score.
[0007] According to a digital human quality assessment method provided by the present invention, the human reference data is a human video; Analyzing the character reference data and the respective character features in the digital human video to obtain the character consistency score of the digital human video includes: Based on the duration of the digital human video, time-align the character video and the digital human video; Extracting features of the characters in each frame of the character video to obtain a first character feature sequence; Extracting features of the character in each frame of the digital human video to obtain a second character feature sequence; Calculating similarities between the character features of frames at corresponding times in the first character feature sequence and the second character feature sequence to obtain a similarity sequence; A similarity mean is calculated for each similarity in the similarity sequence, and the similarity mean is used as the character consistency score.
[0008] According to the present invention, a digital human quality assessment method is provided, which analyzes the difference values of the action key points of the corresponding human body parts in the action-driven video and the digital human video, and determines the action consistency score of the digital human video based on the difference values of the action key points of the corresponding human body parts, including: Extracting multiple first action key point coordinates from the action-driven video ,in, Indicates the action driving the video i Body parts in the frame j The horizontal coordinates of the key points, Indicates the action driving the video i Body parts in the frame j The vertical coordinate of the key point; Extracting multiple second action key point coordinates from the digital human video ,in, Indicates the first i Body parts in the frame j The horizontal coordinates of the key points, Indicates the first iBody parts in the frame j The vertical coordinate of the key point; Based on the coordinates of the first action key point The coordinates of the second key point of the corresponding human body part The difference between and is used to determine the action consistency score.
[0009] According to a digital human quality assessment method provided by the present invention, based on the coordinates of the first action key point The coordinates of the second key point of the corresponding human body part The difference between , and the action consistency score is determined, including: For the frames of the action-driven video and the digital human video at corresponding times, the sum of squares of the differences between the coordinates of the first action key point and the coordinates of the second action key point corresponding to each human body part is calculated according to the following formula: ; The action consistency score is obtained by taking the square root of the sum of the squared difference values corresponding to each frame. : ; in, M Indicates the number of action key points corresponding to each frame, N Indicates the frame number.
[0010] According to the present invention, a digital human quality assessment method is provided, which analyzes the difference value of the motion vector of each pixel between two adjacent frames in the digital human video, and determines the motion coherence score of the digital human video based on the difference value of each motion vector, including: Calculating an optical flow field between two adjacent frames in the digital human video, wherein the optical flow field includes a motion direction and / or a motion speed of each pixel between adjacent frames; Calculating the angle difference and / or speed difference of the motion direction of each pixel between adjacent frames; The action continuity score is determined based on the angle difference and / or speed difference.
[0011] According to a digital human quality assessment method provided by the present invention, the synchronization between the speaking action in the digital human video and the voice in the digital human audio is analyzed to obtain the digital human's audio and video synchronization score, including: Determine a first start time and a first end time of each speaking action segment in the digital human video, wherein the speaking action segment is segmented by N consecutive mouth-closed frames, the first start time being the time of the first mouth-opening frame of the character in the speaking action segment, and the first end time being the time of the first mouth-closed frame among the N consecutive mouth-closed frames following the speaking action segment; Determining a second start time and a second end time of each speech segment in the digital human audio, wherein the speech segment is segmented by a duration of N frames after the speech stops; For the corresponding speaking action segment and speech segment, calculate the difference between the first start time and the second start time, and calculate the difference between the first end time and the second end time; The average of all the differences is determined as the audio-video synchronization score.
[0012] The present invention also provides a digital human quality assessment device, comprising the following units: A character consistency analysis unit, configured to analyze the character reference data and the respective character features in the digital human video to obtain a character consistency score for the digital human video; An action consistency analysis unit is used to analyze the difference values of the action key points of the corresponding human body parts of the respective characters in the action-driven video and the digital human video, and determine the action consistency score of the digital human video based on the difference values of the action key points of the corresponding human body parts; a motion continuity analysis unit for analyzing the difference in motion vectors of each pixel in the digital human video between two adjacent frames, and determining a motion continuity score of the digital human video based on the difference in motion vectors; An audio analysis unit, used to analyze the digital human audio to obtain an audio reverberation score and an audio quality score of the digital human audio; an audio and video synchronization analysis unit, configured to analyze the synchronization between the speaking action in the digital human video and the voice in the digital human audio, and obtain an audio and video synchronization score for the digital human; A quality score evaluation unit, configured to perform a weighted summation of the character consistency score, action consistency score, action continuity score, audio reverberation score, audio quality score, and audio-video synchronization score to obtain a digital human quality score; The digital human video is generated based on the character reference data and the action-driven video.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the program, the digital human quality assessment method as described above is implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described digital human quality assessment methods.
[0015] The digital human quality assessment method and device provided by the present invention analyzes character consistency, motion consistency, motion continuity, audio reverberation, audio quality, and audio-video synchronization, obtains corresponding scores, and then performs a weighted summation of these scores to produce a comprehensive digital human quality score. This provides a more comprehensive assessment of digital human quality than existing assessment methods. Furthermore, the method compares the difference values of each character's motion key points at corresponding body parts in the action-driven video and the digital human video, determining the motion consistency score of the digital human video based on the difference values of each corresponding body part's motion key points. Furthermore, the method compares the difference values of each pixel's motion vector between two adjacent frames in the digital human video, determining the motion continuity score of the digital human video based on the difference values of each motion vector, thereby achieving a more accurate assessment of the digital human's video quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 It is a flow chart of the digital human quality assessment method provided by the present invention.
[0018] Figure 2 It is a schematic diagram of the principle of the digital human quality assessment method provided by the present invention.
[0019] Figure 3 It is a structural diagram of the digital human quality assessment device provided by the present invention.
[0020] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0021] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0022] The digital human quality assessment method of the embodiment of the present invention is as follows: Figure 1 and Figure 2 As shown, the process includes the following steps S110 to S160.
[0023] Step S110: Analyze the character features in the character reference data and the digital human video to obtain a character consistency score for the digital human video. The character reference data is one of the inputs for generating a digital human and can be either a character image or a character video. When generating a digital human, the character in the digital human video must be consistent with the character in the character reference data, primarily with regard to the face. If they are inconsistent, or if the character's appearance differs significantly, the generated digital human will be of poor quality.
[0024] In this step, the features of the character in the character reference data and the features of the character in the generated digital human video can be compared, for example, by calculating the feature similarity, and using the statistical value of the similarity (for example, mean or variance) as the character consistency score of the digital human video. .
[0025] Step S120: Analyze the difference values between the key action points of the corresponding human body parts in the action-driven video and the digital human video, and determine the action consistency score of the digital human video based on the difference values of the key action points of the corresponding human body parts. The action-driven video is another input for generating the digital human; that is, the digital human video is generated based on the character reference data and the action-driven video. The aforementioned character reference data provides a reference for the character in the digital human video, ensuring that the character in the digital human video is consistent with the character in the character reference data. The action-driven video is used to drive the character's movements in the digital human video, providing a reference for the character in the digital human video so that the character in the digital human video performs the movements of the character in the action-driven video.
[0026] In this step, the action consistency score of the digital human video is obtained by comparing the actions of the characters in the action-driven video and the digital human video. For example, the difference values of the key points of the corresponding body parts of the characters in the two videos can be compared, and the statistical value of each difference value (for example, mean or variance) is used as the action consistency score of the digital human video. .
[0027] Step S130: Analyze the difference in motion vectors between two adjacent frames of each pixel in the digital human video, and determine the motion coherence score of the digital human video based on the difference in motion vectors. The more coherent the character's motion in the digital human video, the higher the quality of the digital human video.
[0028] For example, the motion vector difference between two adjacent frames of each pixel in the digital human video can be compared, and the statistical value of each difference value (for example, mean or variance) can be used as the motion coherence score of the digital human video. .
[0029] Step S140: Analyze the digital human audio to obtain the audio reverberation score of the digital human audio and audio quality score The format of the digital human audio is usually WAV, MP3 and other common audio formats. In this step, the audio reverberation analysis and audio quality analysis can be performed on the digital human audio to obtain the audio reverberation score. and audio quality score .
[0030] Step S150: Analyze the synchronization between the speaking action in the digital human video and the audio in the digital human audio to obtain the digital human's audio-video synchronization score. Audio-video synchronization reflects whether the digital human's speaking action matches the corresponding audio in the video. This means that the character should not open their mouth to speak but the corresponding audio lags behind the speaking action by a significant amount, or the audio has already been broadcast but the character's speaking action lags behind the corresponding audio by a significant amount.
[0031] For example, in this step, the difference between the start time and the end time of each speech action segment and the corresponding voice segment can be compared, and the statistical value (for example, mean or variance) of each difference value can be used as the audio-video synchronization score of the digital human video. .
[0032] Step S160: Weighted sum of the character consistency score, action consistency score, action continuity score, audio reverberation score, audio quality score and audio-video synchronization score to obtain the digital human quality score. The weights corresponding to each score are set according to the actual situation, ensuring that the absolute value of each weight sums to 1. Thus, the final digital human quality score is obtained. for: .
[0033] in, 、 、 、 、 and They are the weights of character consistency score, action consistency score, action continuity score, audio reverberation score, audio quality score and audio-video synchronization score.
[0034] It should be noted that the scores for action consistency, action continuity and audio-video synchronization are all based on the statistical values of the corresponding difference values. Since the smaller the difference value, the better the quality, in order to conform to the idea that the higher the score, the better the quality, the weights of the action consistency score, action continuity score and audio-video synchronization score are negative. For example: the weight of the character consistency score is 0.2, the weight of the action consistency score The weight of the action continuity score is -0.2 -0.15, the weight of the audio reverberation score The weight of the audio quality score is 0.1 The weight of the audio and video synchronization score is 0.2 When the above weights are applied to each score, the final digital human quality score is for: .
[0035] Digital Human Quality Rating It is understood that before calculating the digital human quality score, each score is dimensionally eliminated and normalized to the range of [0,1], [0,10] or [0,100], so that the digital human quality score It is also in the range of [0,1], [0,10] or [0,100]. And a qualified score threshold can be obtained through manual evaluation. When the score is greater than or equal to this threshold, it means that the quality of the generated digital human is qualified.
[0036] The digital human quality assessment method of this embodiment analyzes character consistency, motion consistency, motion coherence, audio reverberation, audio quality, and audio-video synchronization, obtaining corresponding scores. These scores are then weighted and summed to produce a comprehensive digital human quality score. This method achieves a more comprehensive assessment of digital human quality than existing assessment methods. Furthermore, the method compares the difference values of each character's motion key points at corresponding body parts in the action-driven video and the digital human video, determining the motion consistency score of the digital human video based on these differences. Furthermore, the method compares the difference values of the motion vectors of each pixel between two adjacent frames in the digital human video, determining the motion coherence score of the digital human video based on these differences, thereby achieving a more accurate assessment of the digital human's video quality.
[0037] In some embodiments, the SRMR (Speech to Noise Ratio in the Modulation Domain) algorithm can be used to perform audio reverberation analysis on digital human audio. SRMR is a reverberation detection and quantification method that is mainly based on the modulation energy characteristics of the speech signal. It evaluates the reverberation degree by analyzing the modulation energy distribution of the speech signal in different frequency bands, and obtains the speech and reverberation modulation energy ratio, thereby quantifying the reverberation of the speech signal. This can then be used to detect whether the speech signal has been dereverberated. Specifically, the digital human audio is input into the SRMR to obtain the speech and reverberation modulation energy ratio of the digital human audio, and the speech and reverberation modulation energy ratio is used as the audio reverberation score. .
[0038] Audio quality analysis is mainly achieved through UTMOSV2, a sound quality assessment model used to evaluate the Mean Opinion Score (MOS). It integrates spectral features and speech features for training, uses pre-trained EfficientNetV2 as an image feature extractor, and pre-trained wav2vec 2.0 as a speech feature extractor, and fuses the features and data domain encoding of the two and inputs them into the fully connected layer to predict MOS. This effectively improves the accuracy and relevance of speech naturalness prediction and enhances the model's ability to evaluate high-quality speech. Specifically, digital human audio is input into UTMOSV2, UTMOSV2 outputs MOS, and MOS is used as the audio quality score. .
[0039] It should be noted that before performing audio reverberation analysis and audio quality analysis, it is preferred to normalize the digital human audio and adjust the audio amplitude to the standard range of -6dB to -12dB to eliminate audio amplitude differences and ensure the accuracy of subsequent analysis.
[0040] In some embodiments, the character reference data is a character image, that is, the character reference data is a static character image. Based on this, step S110 specifically includes: Step 1: Extract the first character feature of the character in the character image. Since character consistency is mainly to judge whether the character in the digital human video and the character in the character image are the same person, the first character feature includes facial features. Exemplarily, a face detection algorithm can be used to extract the first character feature, for example: using the ArcFace model to extract facial features in the character image as the first character feature. ArcFace introduces additive angular margin loss to make the intervals between facial features of different categories larger in the feature space. In this way, ArcFace significantly improves the discriminative ability of face recognition, makes the model more robust in the face of complex situations, and enhances the recognition accuracy of the model.
[0041] Step 2: Extracting a character feature sequence of the character in each frame of the digital human video. In this step, an ArcFace model can also be used to extract facial features in each frame as second character features, and the second character features of each frame form a character feature sequence.
[0042] Step 3: Calculate the similarity between each second character feature in the character feature sequence and the first character feature to obtain a similarity sequence. Specifically, a cosine similarity algorithm can be used to calculate the similarity between each second character feature and the first character feature to obtain a similarity sequence.
[0043] Step 4: Calculate the mean similarity of each similarity in the similarity sequence, and use the mean similarity as the character consistency score .
[0044] In this embodiment, the similarity between the character features of the character in each frame of the digital human video and the character features in the character image is calculated, and the mean similarity is used as the character consistency score. The larger the mean similarity, the better the character consistency, thereby achieving accurate evaluation of character consistency.
[0045] In some embodiments, the character reference data is a character video, that is, the character reference data is a dynamic character video. Based on this, step S110 specifically includes: Step 1: Using the duration of the digital human video as a benchmark, time-align the character video and the digital human video. Alignment means aligning the character video's duration with the digital human video's duration. If the character video is longer than the digital human video, a video clip of the same duration is captured. If the character video is shorter than the digital human video, multiple character videos are repeatedly spliced together until they are the same length as the digital human video.
[0046] Step 2: Extract features from each frame of the character video to obtain a first character feature sequence. Because character consistency primarily determines whether the character in the digital human video and the character video are the same person, the character features in the first character feature sequence include facial features. Exemplarily, a face detection algorithm can be used to extract the first character features. For example, an ArcFace model can be used to extract facial features from each frame of the character video as the first character features. The first character features of each frame form a first character feature sequence.
[0047] Step 3: Extract features of the characters in each frame of the digital human video to obtain a second character feature sequence. In this step, the ArcFace model can also be used to extract facial features in each frame as the second character features, and the second character features of each frame form a character feature sequence.
[0048] Step 4: Calculate the similarity of the character features of the frames at corresponding times in the first character feature sequence and the second character feature sequence to obtain a similarity sequence. The similarity of the character features of each corresponding frame can be calculated using a cosine similarity algorithm.
[0049] Step 5: Calculate the mean similarity of each similarity in the similarity sequence, and use the mean similarity as the character consistency score .
[0050] In this embodiment, the similarity between the character features of the character in each frame of the digital human video and the character features in the corresponding frame of the character video is calculated, and the mean similarity is used as the character consistency score. The larger the mean similarity, the better the character consistency, thereby achieving accurate evaluation of character consistency.
[0051] In some embodiments, step S120 specifically includes: Step 1: Extract multiple first action key point coordinates from the action-driven video ,in, Indicates the action driving the video i Body parts in the frame j The horizontal coordinates of the key points, Indicates the action driving the video i Body parts in the frame j For example, for a speaking action, in each frame, the body parts include at least the upper lip, lower lip, left cheek, and right cheek.
[0052] For example, a motion detection algorithm can be used to extract the coordinates of multiple first motion key points from the motion-driven video. For example, the Unipose model is used. The Unipose model is a universal key point detection model that can detect the key points of any object. In this step, the action-driven video is input into the Unipose model, which will output the coordinates of multiple first action key points in each frame of the action-driven video. .
[0053] Step 2: Extract multiple second action key point coordinates from the digital human video ,in, Indicates the first i Body parts in the frame j The horizontal coordinates of the key points, Indicates the first i Body parts in the frame j In this step, the Unipose model can also be used to extract the coordinates of multiple second action key points in each frame of the digital human video. .
[0054] Step 3: Based on the coordinates of the first action key point The coordinates of the second key point of the corresponding human body part The difference between ,The smaller the difference, the higher the motion consistency between the digital human video and the action-driven video, and the better the quality.
[0055] In this embodiment, by comparing the difference in the coordinates of the action key points of the corresponding human body parts in each frame of the action-driven video and the digital human video, the action consistency score of the digital human video is determined based on the difference, thereby achieving accurate evaluation of the action consistency.
[0056] Furthermore, based on the coordinates of the first action key point The coordinates of the second key point of the corresponding human body part The difference between , and the action consistency score is determined, including: For the frames of the action-driven video and the digital human video at corresponding times, the sum of squares of the differences between the coordinates of the first action key point and the coordinates of the second action key point corresponding to each human body part is calculated according to the following formula: .
[0057] The action consistency score is obtained by taking the square root of the sum of the squared difference values corresponding to each frame. : .
[0058] in, M Indicates the number of action key points corresponding to each frame, N Indicates the frame number.
[0059] In this embodiment, the above formula is used to obtain the statistical value of the difference value of the coordinates of the key points of the action. The calculation is simple and the complexity is low. The calculated action consistency score can accurately reflect the action consistency between the digital human video and the action driving video.
[0060] In some embodiments, step S130 specifically includes: Step 1: Calculate the optical flow field between two adjacent frames in the digital human video, where the optical flow field includes the motion direction and / or motion speed of each pixel between adjacent frames, wherein the motion vector of each pixel between two adjacent frames is represented by the motion direction and / or motion speed of each pixel in the optical flow field between adjacent frames.
[0061] For example, in this step, the StreamFlow model can be used to calculate the optical flow field between two adjacent frames in the digital human video. StreamFlow is a multi-frame optical flow estimation method that excels at identifying optical flow between multiple video frames using efficient spatiotemporal relationship mining techniques, thereby obtaining a highly accurate optical flow field. Specifically, the digital human video is input into the StreamFlow model, which then outputs the optical flow fields of adjacent frames in the digital human video.
[0062] Step 2: Calculate the angle difference and / or speed difference of each pixel's motion direction between adjacent frames. The smaller the angle and the smaller the speed difference, the higher the motion continuity.
[0063] Step 3: Based on the angle difference and / or speed difference, determine the motion continuity score. For example, the average of all angle differences of all pixels can be taken as the first score, and the average of all speed differences of all pixels can be taken as the second score, i.e., the motion continuity score Contains two components: the first score and second score , you can choose one of them as the action continuity score, or you can choose both as the action continuity score. When both are used as the action continuity score, a more accurate action continuity assessment can be achieved. In the case where both are used as the action continuity score, the above calculation The formula can be written as: .
[0064] in, .For example: .
[0065] In this embodiment, by calculating the optical flow field between two adjacent frames in the digital human video, the angle difference in the motion direction and / or the speed difference of the pixels in the optical flow field between each adjacent frame can be used to represent the coherence of the movement. Therefore, determining the movement coherence score based on the angle difference and / or speed difference can accurately reflect the movement coherence in the digital human video, thereby achieving an accurate assessment of the movement coherence in the digital human video.
[0066] In some embodiments, step S150 specifically includes: Step 1: Determine the first start and end times of each speaking action segment in the digital human video. The speaking action segments are segmented by N consecutive closed-mouth frames. The first start time is the moment of the character's first open-mouth frame in the speaking action segment, and the first end time is the moment of the first closed-mouth frame among the N consecutive closed-mouth frames following the speaking action segment. When a person speaks, there are pauses of a certain length between sentences. This is reflected in the speaking action as the duration of the closed mouth, i.e., the number of closed-mouth frames. Therefore, in this step, the speaking action segments are segmented by N consecutive closed-mouth frames (for example, 3-5, determined by the specific speaking speed).
[0067] Step 2: Determine the second start and end times of each speech segment in the digital human audio. Speech segments are segmented by the duration of N frames after speech stops. When a person speaks, there are pauses of a certain length between sentences. This is represented by the duration of the pause, i.e., the duration of N (no speech) frames after speech stops. Therefore, in this step, speech segments are segmented by the duration of N (no speech) frames after speech stops.
[0068] Step 3: For the corresponding speech action segments and voice segments, calculate the difference between the first start time and the second start time, and calculate the difference between the first end time and the second end time. After the first two steps, starting from the moment the video begins playing, the video is divided into speech action segments in chronological order, and the audio is divided into corresponding voice segments. The start and end times of the corresponding speech action segments and voice segments have been determined. Therefore, in this step, for each corresponding speech action segment and voice segment, the difference between the first start time and the second start time, as well as the difference between the first end time and the second end time, can be calculated. The smaller the difference, the better the audio and video synchronization.
[0069] Step 4: Determine the mean of all differences as the audio and video synchronization score.
[0070] In this embodiment, by determining the start and end times of the speaking action segment in the digital human video and the corresponding voice segment in the digital human audio, the difference between the corresponding start and end times is calculated. The smaller the difference, the better the audio and video synchronization. Therefore, using the average of the differences as the audio and video synchronization score can achieve a more accurate assessment of audio and video synchronization.
[0071] The digital human quality assessment device provided by the present invention is described below. The digital human quality assessment device described below and the digital human quality assessment method described above can be referenced to each other.
[0072] The digital human quality assessment device of the embodiment of the present invention is as follows Figure 3 Shown, including: The character consistency analysis unit 310 is used to analyze the character reference data and the respective character features in the digital human video to obtain a character consistency score of the digital human video.
[0073] The action consistency analysis unit 320 is used to analyze the difference values of the action key points of the corresponding human body parts in the action driving video and the digital human video, and determine the action consistency score of the digital human video based on the difference values of the action key points of the corresponding human body parts.
[0074] The motion continuity analysis unit 330 is used to analyze the difference value of the motion vector of each pixel in the digital human video between two adjacent frames, and determine the motion continuity score of the digital human video based on the difference value of each motion vector.
[0075] The audio analysis unit 340 is used to analyze the digital human audio to obtain an audio reverberation score and an audio quality score of the digital human audio.
[0076] The audio and video synchronization analysis unit 350 is used to analyze the synchronization between the speaking action in the digital human video and the voice in the digital human audio to obtain the audio and video synchronization score of the digital human.
[0077] The quality score evaluation unit 360 is used to perform weighted summation of the character consistency score, action consistency score, action continuity score, audio reverberation score, audio quality score and audio-video synchronization score to obtain a digital human quality score.
[0078] The digital human video is generated based on the character reference data and the action-driven video.
[0079] In some embodiments, the character reference data is a character image, and the character consistency analysis unit 310 includes: The first character feature extraction unit is used to extract a first character feature of a character in the character image.
[0080] The character feature sequence extraction unit is used to extract the character feature sequence of the character in each frame of the digital human video.
[0081] The similarity calculation unit is used to calculate the similarity between each second character feature in the character feature sequence and the first character feature to obtain a similarity sequence.
[0082] The similarity mean calculation unit is used to calculate the similarity mean of each similarity in the similarity sequence, and use the similarity mean as the character consistency score.
[0083] In some embodiments, the character reference data is a character video; the character consistency analysis unit 310 includes: The time alignment unit is used to time-align the character video and the digital human video based on the duration of the digital human video.
[0084] The first character feature extraction unit is used to extract features of the character in each frame of the character video to obtain a first character feature sequence.
[0085] The second character feature extraction unit is used to extract features of the character in each frame of the digital human video to obtain a second character feature sequence.
[0086] The similarity calculation unit is used to calculate the similarity of the character features of the frames at corresponding times in the first character feature sequence and the second character feature sequence to obtain a similarity sequence.
[0087] The similarity mean calculation unit is used to calculate the similarity mean of each similarity in the similarity sequence, and use the similarity mean as the character consistency score.
[0088] In some embodiments, the action consistency analysis unit 320 includes: A first key point extraction unit is used to extract multiple first action key point coordinates from the action-driven video ,in, Indicates the action driving the video i Body parts in the frame j The horizontal coordinates of the key points, Indicates the action driving the video i Body parts in the frame j The vertical coordinate of the key point.
[0089] The second key point extraction unit is used to extract multiple second action key point coordinates from the digital human video ,in, Indicates the first i Body parts in the frame j The horizontal coordinates of the key points, Indicates the first i Body parts in the frame j The vertical coordinate of the key point.
[0090] An action consistency score determination unit is configured to determine an action consistency score based on the coordinates of the first action key point. The coordinates of the second key point of the corresponding human body part The difference between and is used to determine the action consistency score.
[0091] In some embodiments, the action consistency score determination unit is specifically configured to: For the frames of the action-driven video and the digital human video at corresponding times, the sum of squares of the differences between the coordinates of the first action key point and the coordinates of the second action key point corresponding to each human body part is calculated according to the following formula: .
[0092] The action consistency score is obtained by taking the square root of the sum of the squared difference values corresponding to each frame. : .
[0093] in, M Indicates the number of action key points corresponding to each frame, N Indicates the frame number.
[0094] In some embodiments, the action continuity analysis unit 330 includes: The optical flow field calculation unit is used to calculate the optical flow field between two adjacent frames in the digital human video, where the optical flow field includes the motion direction and / or motion speed of each pixel between adjacent frames.
[0095] The motion difference calculation unit is used to calculate the angle difference and / or speed difference of the motion direction of each pixel between adjacent frames.
[0096] The action continuity score determining unit is configured to determine the action continuity score based on the angle difference and / or speed difference.
[0097] In some embodiments, the audio and video synchronization analysis unit 350 includes: The first moment determination unit is used to determine the first start moment and the first end moment of each speaking action segment in the digital human video, where the speaking action segment is segmented into N consecutive closed-mouth frames. The first start moment is the moment of the character's first open-mouth frame in the speaking action segment, and the first end moment is the moment of the first closed-mouth frame among the N consecutive closed-mouth frames after the speaking action segment.
[0098] The second time determination unit is used to determine the second start time and the second end time of each voice segment in the digital human audio, and the voice segment is segmented according to the duration of N frames after the voice stops.
[0099] The time difference calculation unit is used to calculate the difference between the first start time and the second start time, and the difference between the first end time and the second end time for the corresponding speaking action segment and voice segment.
[0100] The synchronization score determining unit is configured to determine an average of all differences as the audio and video synchronization score.
[0101] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the digital human quality assessment method, which includes: The character features of the character reference data and the digital human video are analyzed to obtain a character consistency score of the digital human video.
[0102] The difference values of the action key points of the corresponding human body parts of the respective characters in the action-driven video and the digital human video are analyzed, and the action consistency score of the digital human video is determined based on the difference values of the action key points of the corresponding human body parts.
[0103] The difference value of the motion vector of each pixel between two adjacent frames in the digital human video is analyzed, and the motion continuity score of the digital human video is determined based on the difference value of each motion vector.
[0104] The digital human audio is analyzed to obtain an audio reverberation score and an audio quality score of the digital human audio.
[0105] The synchronization between the speaking action in the digital human video and the voice in the digital human audio is analyzed to obtain the audio and video synchronization score of the digital human.
[0106] The character consistency score, action consistency score, action continuity score, audio reverberation score, audio quality score and audio-video synchronization score are weighted and summed to obtain the digital human quality score.
[0107] The digital human video is generated based on the character reference data and the action-driven video.
[0108] Furthermore, the logic instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0109] On the other hand, the present invention further provides a computer program product, comprising a computer program, which may be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the digital human quality assessment method provided by the above methods, which includes: The character features of the character reference data and the digital human video are analyzed to obtain a character consistency score of the digital human video.
[0110] The difference values of the action key points of the corresponding human body parts of the respective characters in the action-driven video and the digital human video are analyzed, and the action consistency score of the digital human video is determined based on the difference values of the action key points of the corresponding human body parts.
[0111] The difference value of the motion vector of each pixel between two adjacent frames in the digital human video is analyzed, and the motion continuity score of the digital human video is determined based on the difference value of each motion vector.
[0112] The digital human audio is analyzed to obtain an audio reverberation score and an audio quality score of the digital human audio.
[0113] The synchronization between the speaking action in the digital human video and the voice in the digital human audio is analyzed to obtain the audio and video synchronization score of the digital human.
[0114] The character consistency score, action consistency score, action continuity score, audio reverberation score, audio quality score and audio-video synchronization score are weighted and summed to obtain the digital human quality score.
[0115] The digital human video is generated based on the character reference data and the action-driven video.
[0116] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the digital human quality assessment method provided by the above methods, the method comprising: The character features of the character reference data and the digital human video are analyzed to obtain a character consistency score of the digital human video.
[0117] The difference values of the action key points of the corresponding human body parts of the respective characters in the action-driven video and the digital human video are analyzed, and the action consistency score of the digital human video is determined based on the difference values of the action key points of the corresponding human body parts.
[0118] The difference value of the motion vector of each pixel between two adjacent frames in the digital human video is analyzed, and the motion continuity score of the digital human video is determined based on the difference value of each motion vector.
[0119] The digital human audio is analyzed to obtain an audio reverberation score and an audio quality score of the digital human audio.
[0120] The synchronization between the speaking action in the digital human video and the voice in the digital human audio is analyzed to obtain the audio and video synchronization score of the digital human.
[0121] The character consistency score, action consistency score, action continuity score, audio reverberation score, audio quality score and audio-video synchronization score are weighted and summed to obtain the digital human quality score.
[0122] The digital human video is generated based on the character reference data and the action-driven video.
[0123] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0124] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A digital human quality assessment method, characterized in that: include: Analyzing the character reference data and the respective character features in the digital human video to obtain a character consistency score for the digital human video; Analyzing the difference values of the action key points of the corresponding human body parts in the action-driven video and the digital human video, and determining the action consistency score of the digital human video based on the difference values of the action key points of the corresponding human body parts; Analyzing the difference value of the motion vector of each pixel between two adjacent frames in the digital human video, and determining the motion coherence score of the digital human video based on the difference value of each motion vector; Analyze the digital human audio to obtain the audio reverberation score and audio quality score of the digital human audio; Analyzing the synchronization between the speaking action in the digital human video and the voice in the digital human audio to obtain an audio and video synchronization score of the digital human; A weighted summation of the character consistency score, action consistency score, action continuity score, audio reverberation score, audio quality score, and audio-video synchronization score is performed to obtain a digital human quality score; The digital human video is generated based on the character reference data and the action-driven video.
2. The digital human quality assessment method according to claim 1, characterized in that: The character reference data is a character image; Analyzing the character reference data and the respective character features in the digital human video to obtain the character consistency score of the digital human video includes: Extracting a first character feature of a person in the character image; Extracting a character feature sequence of a character in each frame of the digital human video; Calculating the similarity between each second character feature in the character feature sequence and the first character feature to obtain a similarity sequence; A similarity mean is calculated for each similarity in the similarity sequence, and the similarity mean is used as the character consistency score.
3. The digital human quality assessment method according to claim 1, characterized in that: The character reference data is a character video; Analyzing the character reference data and the respective character features in the digital human video to obtain the character consistency score of the digital human video includes: Based on the duration of the digital human video, time-align the character video and the digital human video; Extracting features of the characters in each frame of the character video to obtain a first character feature sequence; Extracting features of the character in each frame of the digital human video to obtain a second character feature sequence; Calculating similarities between the character features of frames at corresponding times in the first character feature sequence and the second character feature sequence to obtain a similarity sequence; A similarity mean is calculated for each similarity in the similarity sequence, and the similarity mean is used as the character consistency score.
4. The digital human quality assessment method according to claim 1, characterized in that: Analyzing the difference values of the action key points of the corresponding human body parts of the respective characters in the action-driven video and the digital human video, and determining the action consistency score of the digital human video based on the difference values of the action key points of the corresponding human body parts, including: Extracting multiple first action key point coordinates from the action-driven video ,in, Indicates the action driving the video i Body parts in the frame j The horizontal coordinates of the key points, Indicates the action driving the video i Body parts in the frame j The vertical coordinate of the key point; Extracting multiple second action key point coordinates from the digital human video ,in, Indicates the first i Body parts in the frame j The horizontal coordinates of the key points, Indicates the first i Body parts in the frame j The vertical coordinate of the key point; Based on the coordinates of the first action key point The coordinates of the second key point of the corresponding human body part The difference between and is used to determine the action consistency score.
5. The digital human quality assessment method according to claim 4, characterized in that: Based on the coordinates of the first action key point The coordinates of the second key point of the corresponding human body part The difference between , and the action consistency score is determined, including: For the frames of the action-driven video and the digital human video at corresponding times, the sum of squares of the differences between the coordinates of the first action key point and the coordinates of the second action key point corresponding to each human body part is calculated according to the following formula: ; The action consistency score is obtained by taking the square root of the sum of the squared difference values corresponding to each frame. : ; in, M Indicates the number of action key points corresponding to each frame, N Indicates the frame number.
6. The digital human quality assessment method according to claim 1, characterized in that: Analyzing the difference value of the motion vector of each pixel between two adjacent frames in the digital human video, and determining the motion coherence score of the digital human video based on the difference value of each motion vector, including: Calculating an optical flow field between two adjacent frames in the digital human video, wherein the optical flow field includes a motion direction and / or a motion speed of each pixel between adjacent frames; Calculating the angle difference and / or speed difference of the motion direction of each pixel between adjacent frames; The action continuity score is determined based on the angle difference and / or speed difference.
7. The digital human quality assessment method according to claim 1, characterized in that: Analyzing the synchronization between the speaking action in the digital human video and the voice in the digital human audio to obtain the digital human audio and video synchronization score includes: Determine a first start time and a first end time of each speaking action segment in the digital human video, wherein the speaking action segment is segmented by N consecutive mouth-closed frames, the first start time being the time of the first mouth-opening frame of the character in the speaking action segment, and the first end time being the time of the first mouth-closed frame among the N consecutive mouth-closed frames following the speaking action segment; Determining a second start time and a second end time of each speech segment in the digital human audio, wherein the speech segment is segmented by a duration of N frames after the speech stops; For the corresponding speaking action segment and voice segment, calculate the difference between the first start time and the second start time, and calculate the difference between the first end time and the second end time; The average of all the differences is determined as the audio-video synchronization score.
8. A digital human quality assessment device, characterized in that: include: A character consistency analysis unit, configured to analyze the character reference data and the respective character features in the digital human video to obtain a character consistency score for the digital human video; An action consistency analysis unit is used to analyze the difference values of the action key points of the corresponding human body parts of the respective characters in the action-driven video and the digital human video, and determine the action consistency score of the digital human video based on the difference values of the action key points of the corresponding human body parts; a motion continuity analysis unit for analyzing the difference in motion vectors of each pixel in the digital human video between two adjacent frames, and determining a motion continuity score of the digital human video based on the difference in motion vectors; An audio analysis unit, configured to analyze the digital human audio to obtain an audio reverberation score and an audio quality score of the digital human audio; an audio and video synchronization analysis unit, configured to analyze the synchronization between the speaking action in the digital human video and the voice in the digital human audio, and obtain an audio and video synchronization score for the digital human; A quality score evaluation unit, configured to perform a weighted summation of the character consistency score, action consistency score, action continuity score, audio reverberation score, audio quality score, and audio-video synchronization score to obtain a digital human quality score; The digital human video is generated based on the character reference data and the action-driven video.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the digital human quality assessment method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the digital human quality assessment method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Video quality evaluation method and device, equipment and storage medium
CN118885821A
Digital human binding evaluation method
CN119068158A
Virtual human action driving method and device based on language analysis and storage medium
CN120147488A
Voice interaction method and system of mobile digital human
CN120279896A
4D digital human quality evaluation method and system based on multi-feature fusion
CN120431453A
Cited By
Live person live broadcast identification method and related device
CN120954111A