Quality detection method and device of digital human video, electronic equipment and storage medium

By generating a video set and using multiple detection dimensions and objects to perform quality detection of digital human videos, the problems of inefficiency and subjective deviation in the prior art are solved, efficient and accurate digital human video quality detection is achieved, and the quality of target videos is ensured to the optimal quality of the target videos.

CN120238647APending Publication Date: 2025-07-01BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510495193.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the prior art, the quality detection of digital human videos is inefficient and subjective deviations, making it difficult to fully cover various quality problems.

Method used

By generating a video set, the digital human video is used to perform quality detection of the digital human video using multiple target detection dimensions and detection objects, the quality detection results are obtained, and the target digital human video is determined based on the results.

Benefits of technology

It realizes comprehensive and accurate quality inspection of digital human videos, ensures the optimal quality of the target videos, and improves the detection efficiency and objectivity of the detection results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238647A_ABST
    Figure CN120238647A_ABST
Patent Text Reader

Abstract

The invention provides a digital human video quality detection method and device, electronic equipment and a storage medium, and relates to the technical field of data processing, in particular to the field of computer vision, deep learning and large models. According to the specific implementation scheme, video synthesis materials are obtained, a video set is generated based on the video synthesis materials, and the video set comprises at least two digital human videos with the same video content; a detection task is obtained, the detection task comprises one or more target detection dimensions, and the target detection dimensions comprise at least one detection object; according to the detection task, performing quality detection on the digital human video in the video set to obtain a quality detection result of the digital human video; and determining a target digital human video from the video set according to a quality detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technologies, specifically to the fields of computer vision, deep learning, and large models, and particularly to a method, apparatus, electronic device, and storage medium for quality detection of digital human videos. Background Art

[0002] There are various effect problems in the videos synthesized by digital humans. For example, in a certain frame of the video, the digital human suddenly disappears, the screen is black / white / flowered, the video is stuck, the clarity of the digital human is insufficient resulting in a blurred picture, the digital human's audio and video are out of sync, the synthesized content is missing words or characters, the subtitles and sound of the synthesized video are out of sync, etc. Currently, these problems mainly rely on manual inspection or user feedback, which not only has low efficiency, but also has drawbacks such as subjective deviation and incomplete coverage. Moreover, as the video scale expands, the limitations become more and more prominent. Summary of the Invention

[0003] The present disclosure provides a method, apparatus, electronic device, and storage medium for quality detection of digital human videos.

[0004] According to one aspect of the present disclosure, there is provided a method for quality detection of digital human videos, including:

[0005] Obtaining video synthesis materials and generating a video set based on the video synthesis materials, where the video set includes at least two digital human videos with the same video content;

[0006] Obtaining a detection task, where the detection task includes one or more target detection dimensions, and the target detection dimension includes at least one detection object;

[0007] According to the detection task, performing quality detection on the digital human videos in the video set to obtain the quality detection results of the digital human videos;

[0008] Determining a target digital human video from the video set according to the quality detection results.

[0009] According to another aspect of the present disclosure, there is provided a quality detection apparatus for digital human videos, including:

[0010] A first acquisition module for obtaining video synthesis materials and generating a video set based on the video synthesis materials, where the video set includes at least two digital human videos with the same video content;

[0011] A second acquisition module for obtaining a detection task, where the detection task includes one or more target detection dimensions, and the target detection dimension includes at least one detection object;

[0012] A quality detection module, configured to perform quality detection on the digital human videos in the video set according to the detection task, and obtain quality detection results of the digital human videos;

[0013] A screening module, configured to determine target digital human videos from the video set according to the quality detection results.

[0014] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in the embodiments of the first aspect.

[0018] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the embodiments of the first aspect.

[0019] According to a fifth aspect of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described in the embodiments of the first aspect are implemented.

[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0022] Figure 1 is a schematic diagram of a method for quality detection of digital human videos provided by an embodiment of the present disclosure;

[0023] Figure 2 is a schematic diagram of another method for quality detection of digital human videos provided by an embodiment of the present disclosure;

[0024] Figure 3 is a schematic diagram of another method for quality detection of digital human videos provided by an embodiment of the present disclosure;

[0025] Figure 4 is a schematic diagram of another method for quality detection of digital human videos provided by an embodiment of the present disclosure;

[0026] Figure 5 It is an exemplary logic diagram of a method for detecting the quality of a digital human video provided by an embodiment of the present disclosure;

[0027] Figure 6 It is a structural diagram of a device for detecting the quality of a digital human video provided by an embodiment of the present disclosure;

[0028] Figure 7 It shows a schematic block diagram of an electronic device for implementing an embodiment of the present disclosure. Detailed implementation manners

[0029] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0030] Data processing is the collection, storage, retrieval, processing, transformation, and transmission of data. The basic purpose is to extract and derive valuable and meaningful data for certain specific people from a large amount of data that may be chaotic and difficult to understand.

[0031] Computer vision is an important branch of artificial intelligence. Its core goal is to endow machines with human-like visual perception capabilities so that they can extract information from multi-dimensional visual data such as images and videos and make decisions.

[0032] Deep learning is a machine learning method based on deep neural network models. Its core feature is the ability to automatically extract multi-level feature representations from data and is widely used in fields such as computer vision and natural language processing.

[0033] A large model refers to a machine learning model with a large number of parameters and a complex computational structure. It is usually constructed by a deep neural network and has billions or even hundreds of billions of parameters. The purpose is to improve the expressive ability and prediction performance of the model and be able to handle more complex tasks and data.

[0034] Figure 1 It is a schematic diagram of a method for detecting the quality of a digital human video provided by an embodiment of the present disclosure. As Figure 1 shown, the method includes the following steps:

[0035] S101, Obtain video synthesis materials and generate a video set based on the video synthesis materials.

[0036] Optionally, the video synthesis materials may include materials required for generating a video, such as background audio, digital human image, digital human narration text, subtitle format, and special effects, and generate a digital human video according to the video synthesis materials.

[0037] In this embodiment, at least two digital human videos are generated by using different video generation means based on the same video synthesis materials to obtain a video set; it can be understood that each digital human video in the video set is generated based on the same video synthesis materials, so the video content such as the digital human image and digital human narration text between the digital human videos in the video set is the same; however, different digital human videos are obtained based on different video generation means, so there are differences in the generation quality between the digital human videos.

[0038] S102, Obtain a detection task.

[0039] It can be understood that the detection task is a task for quality detection of digital human videos; the detection task may include one or more target detection dimensions, and each target detection dimension includes at least one detection object.

[0040] Optionally, the target detection dimensions may include a picture dimension, an audio dimension, and a subtitle dimension, etc. The detection objects of the picture dimension may include picture clarity, picture smoothness, and picture display, etc. The detection objects of the audio dimension may include whether the audio and video are synchronized, whether the audio is complete, and the detection objects of the subtitle dimension may be whether the subtitle is synchronized with the audio.

[0041] S103, According to the detection task, perform quality detection on the digital human videos in the video set to obtain the quality detection results of the digital human videos.

[0042] Optionally, quality detection can be performed on each digital human video in the video set according to the detection task. In other embodiments, quality detection can also be performed only on some of the digital human videos in the video set. Quality detection of digital human videos means detecting each digital human video according to the target detection dimensions in the detection task. For example, based on the detection objects corresponding to each target detection dimension, quality detection is performed on each digital human video to obtain the detection results of the digital human video under the detection object. In this embodiment, quality detection is performed on the digital human video under each detection object, and the sub-detection results of each digital human video under each target detection dimension are obtained according to the detection results under all detection objects. According to the sub-detection results of the digital human video under each target detection dimension, the quality detection result of the digital human video is obtained. The quality detection result at least includes quality problems of the digital human video under different detection objects, such as low picture clarity, picture display lag, or subtitle-audio out-of-sync, etc.

[0043] S104, determine the target digital human video from the video set according to the quality detection result.

[0044] It can be understood that by performing quality detection on each digital human video in the video set based on the detection task, the quality detection result corresponding to each digital human video can be obtained; optionally, the quality detection results of all digital human videos in the video set can be compared, and the digital human video corresponding to the best quality detection result can be determined as the target digital human video. The digital human video corresponding to the best quality detection result is a digital human video without quality problems or a digital human video with the fewest quality problems to ensure the optimal quality of the target digital human video.

[0045] In this embodiment, multiple digital human videos with the same video content are generated through different video generation methods, and quality detection is performed on the digital human videos according to the detection task. Through multiple target detection dimensions and the detection objects under each target detection dimension, comprehensive quality detection is performed on the digital human videos to obtain a more comprehensive and accurate quality detection result of the detected digital human videos. Then, according to the quality detection result, the target digital human video with the best quality detection result is determined from multiple digital human videos in the video set to ensure the optimal quality of the target digital human video and improve the display effect of the target digital human video.

[0046] Figure 2 It is a schematic diagram of another method for quality detection of digital human videos provided by an embodiment of the present disclosure. As Figure 2 shown, the method includes the following steps:

[0047] S201, obtain video synthesis materials and generate a video set based on the video synthesis materials.

[0048] In the embodiments of the present disclosure, the implementation method of step S201 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made herein and will not be elaborated further.

[0049] S202, obtain a detection task.

[0050] In the embodiments of the present disclosure, the implementation method of step S202 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made herein and will not be elaborated further.

[0051] S203, determine the execution mode of the target detection dimension according to the detection task.

[0052] Optionally, the execution mode of the target detection dimension may include parallel execution and serial execution. If the target detection dimensions in the detection task are independent of each other and have no correlation, the target detection dimensions can be executed in parallel, that is, multiple target detection dimensions perform quality detection simultaneously to improve the quality detection efficiency. If there are target detection dimensions with sequential correlations in the detection task, serial execution is used for quality detection of the target detection dimensions to ensure the accuracy of the quality detection results. Or when there are many target detection dimensions but the detection ability is limited, to avoid exceeding the processing load, serial execution is used to perform quality detection on the target detection dimensions in sequence to ensure the smooth progress of the quality detection.

[0053] S204, perform quality detection on the digital human video according to the execution mode of the target detection dimension and the detection object, and obtain the quality detection result of the digital human video.

[0054] In some implementations, the target detection dimensions can be detected and traversed according to the execution mode; it can be understood that if the execution mode is parallel execution, the detection is performed on multiple target detection dimensions simultaneously; if the execution model is serial execution, the target detection dimensions are detected one by one.

[0055] For the currently traversed target detection dimension, based on the detection process of the detection object under the target detection dimension, perform quality detection on the digital human video. Optionally, based on the detection process of each detection object under the target detection dimension, perform quality detection on the digital human video to obtain a more comprehensive quality detection result of the digital human video. The detection process refers to the analysis process of the digital human video under each detection object. For example, when detecting the detection object of black screen in the picture quality dimension, each frame of the digital human video to be detected can be extracted first, and then the to-be-detected picture is grayscale processed to obtain a gray image, and finally, based on the color features in the gray image, it is determined whether there is a black screen problem. Perform quality detection on the digital human video based on the detection processes of each detection object to obtain the quality detection result of the digital human video.

[0056] In some implementations, during the quality inspection of digital humans, different detection thread groups can also be called according to the number of digital human videos in the video set. For example, if quality inspection is required for each digital human video in the video set, the same number of detection thread groups as the number of digital human videos needs to be called, and each detection thread group performs quality inspection on the digital human videos. A detection thread group can include one or more detection threads, and a detection thread can be used to perform traversal detection of the target detection dimension in the detection task. For example, when the target detection dimension is to perform detection in parallel, multiple detection threads in the detection thread group simultaneously perform quality inspection on different target detection dimensions. When the target detection dimension is to perform detection serially, a single detection thread in the detection thread group performs sequential quality inspection on the target detection dimension.

[0057] S205. Determine the target digital human video from the video set according to the quality inspection result.

[0058] In the embodiments of the present disclosure, the implementation method of step S205 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made here and will not be elaborated further.

[0059] In this embodiment, during the detection of digital human videos, the execution mode of the target detection dimension is determined according to the detection task. When there is an association relationship between the number of target detection dimensions, to ensure the effectiveness of the quality inspection result or to ensure that the processing load is not exceeded even when the number of target detection dimensions is large, the target detection dimensions are serially executed; when the target detection dimensions do not exceed the processing load and are independent of each other, to improve the quality inspection efficiency, the target detection dimensions are executed in parallel. The quality inspection strategy is flexibly adjusted, and the target detection dimensions are detected and traversed based on the execution mode. For the traversed target detection dimensions, the quality inspection of digital human videos is performed according to the detection process of the detection object. This embodiment can also perform quality inspection by calling threads to obtain the quality inspection result, and determine the target digital human video according to the quality inspection result, which can be applicable to both complex scenarios with multiple target detection dimensions and simple scenarios with fewer target detection dimensions, improving the flexibility of the quality inspection process while ensuring the accuracy of the quality inspection result.

[0060] Figure 3 is a schematic diagram of another method for quality inspection of digital human videos provided by the embodiments of the present disclosure. As Figure 3 shown, the method includes the following steps:

[0061] S301. Obtain video synthesis materials and generate a video set based on the video synthesis materials.

[0062] In the embodiments of the present disclosure, the implementation method of step S301 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made herein and will not be elaborated further.

[0063] S302. Obtain a detection task.

[0064] In the embodiments of the present disclosure, the implementation method of step S302 can be implemented in any one of the embodiments of the present disclosure, and no limitation is made herein and will not be elaborated further.

[0065] S303. Determine the number of digital human videos in the video set.

[0066] S304. Invoke a detection thread group that is consistent with the number of videos.

[0067] The detection thread group includes at least one detection thread. Different detection thread groups correspond to different digital human videos to perform quality detection on different digital human videos simultaneously by the detection thread group, thereby improving the detection efficiency.

[0068] Optionally, the detection threads in each detection thread group can be the same or different. In this embodiment, the number of detection threads in the detection thread group can be determined according to the target detection dimension and the execution mode of the target detection dimension, where the execution mode of the target detection dimension can include serial execution and parallel execution.

[0069] Exemplarily, if the execution mode of the target detection dimension is serial execution, the number of detection threads in the detection thread group can be 1, that is, 1 detection thread sequentially executes all the target detection dimensions in sequence.

[0070] Exemplarily, if the execution mode of the target detection dimension is serial execution, the number of detection threads in the detection thread group can also be determined according to the number of detection objects of the target detection dimension. For example, the maximum number of detection objects corresponding to the target detection dimension is used as the number of detection threads in the detection thread group. Suppose there are target detection dimension A, target detection dimension B, and target detection dimension C, where target detection dimension A includes 3 detection objects, target detection dimension B includes 5 detection objects, and target detection dimension C includes 4 detection objects. Then the number of detection threads in the current detection thread group can be 5.

[0071] Exemplarily, if the execution mode of the target detection dimension is parallel execution, the number of detection threads can be determined according to the number of target detection dimensions. For example, if there are 3 target detection dimensions, the detection thread group includes 3 detection threads, and each detection thread performs quality detection under one target detection dimension.

[0072] Exemplary illustration, if the execution mode of the target detection dimension is parallel execution, the number of detection threads in the detection thread group can also be determined according to the number of detection objects of the target detection dimension. For example, the number of all detection objects corresponding to all target detection dimensions is used as the number of detection threads, that is, each detection thread performs the quality detection of one detection object.

[0073] S305, send the corresponding digital human video to the detection thread group, and the detection thread performs quality detection on the corresponding digital human video according to the detection task to obtain the quality detection result.

[0074] In some implementations, if the execution mode of the target detection dimension is serial execution and there is only one detection thread in the detection thread group, then this detection thread sequentially performs quality detection on each target detection dimension to determine the detection result of the digital human video under each target detection dimension, thereby obtaining the quality detection result of the digital human.

[0075] In some implementations, if the execution mode of the target detection dimension is serial execution and the number of detection threads in the detection thread group is the same as the maximum number of detection objects corresponding to the target detection dimension. For example, there are target detection dimension A, target detection dimension B, and target detection dimension C, where target detection dimension A includes 3 detection objects, target detection dimension B includes 5 detection objects, and target detection dimension C includes 4 detection objects. Then the number of detection threads in the current detection thread group is 5. During actual quality detection, any 3 of the 5 detection threads perform quality detection on the 3 detection objects of target detection dimension A. After the detection of target detection dimension A is completed, the 5 detection threads perform quality detection on the 5 detection objects of target detection dimension B. After the detection of target detection dimension B is completed, any 4 of the 5 detection threads perform quality detection on the 4 detection objects of target detection dimension C, so as to achieve the quality detection of target detection dimension A, target detection dimension B, and target detection dimension C in serial order and obtain the quality detection result of the digital human video.

[0076] In some implementations, if the execution mode of the target detection dimension is parallel execution and the number of detection threads in the detection thread group is the same as the number of target detection dimensions, then quality detection of the target detection dimension is performed based on each detection thread, that is, N detection threads simultaneously perform quality detection on N target detection dimensions to obtain the quality detection result of the digital human video, where N is the number of target detection dimensions.

[0077] In some implementations, if the execution mode of the target detection dimension is parallel execution, the detection thread group includes the same number of detection threads as the number of all detection objects corresponding to all target detection dimensions. For example, if the number of all detection objects corresponding to all target detection dimensions is M, then the quality inspection of M detection objects is simultaneously performed according to M detection threads, and the quality inspection result of the digital human is obtained.

[0078] S306. Determine the target digital human video from the video set according to the quality inspection result.

[0079] In the embodiments of the present disclosure, the implementation method of step S306 can be implemented in any one of the embodiments of the present disclosure, which is not limited herein and will not be elaborated further.

[0080] In this embodiment, the detection thread group can be called according to the number of digital human videos in the video set, and the quality inspection of different digital human videos is simultaneously performed by each detection thread group to improve the quality inspection efficiency. Also, according to the number of detection objects in the target detection dimension and the execution mode of the target detection dimension, the number of detection threads in each detection thread group can be determined, and the quality inspection of the detection objects is performed by different numbers of detection threads. The quality inspection process is more flexible, ensuring the quality inspection effect while improving the timeliness and efficiency of the quality inspection, determining the target digital human video with a more efficient and accurate quality inspection result, and improving the display effect of the target digital human video.

[0081] Figure 4 It is a schematic diagram of another method for quality inspection of digital human videos provided by the embodiments of the present disclosure. As Figure 4 shown, the method includes the following steps:

[0082] S401. Obtain video synthesis materials and generate a video set based on the video synthesis materials.

[0083] In the embodiments of the present disclosure, the implementation method of step S401 can be implemented in any one of the embodiments of the present disclosure, which is not limited herein and will not be elaborated further.

[0084] S402. Obtain a detection task.

[0085] In the embodiments of the present disclosure, the implementation method of step S402 can be implemented in any one of the embodiments of the present disclosure, which is not limited herein and will not be elaborated further.

[0086] S403. Extract picture frames from the digital human video.

[0087] Optionally, different frame picture data can be extracted from the digital human video according to the key frame extraction algorithm to obtain multiple picture frames.

[0088] S404. Process the frame of the image according to the detection process of the detection object to determine the quality evaluation information of the frame of the image.

[0089] In some implementations, the grayscale image of the frame of the image can be determined according to the detection process of the detection object. It can be understood that determining the grayscale image of the frame of the image is actually converting the frame of the image into a grayscale image, and the grayscale image only includes single-channel brightness information, which can improve the processing speed of the image.

[0090] Optionally, the mean and variance of the three RGB channels corresponding to the grayscale image can be obtained as the first quality evaluation information of the grayscale image; that is, the R mean, G mean, B mean, R variance, G variance, and B variance corresponding to the red R channel, green G channel, and blue B channel in the grayscale image are obtained as the first quality evaluation information of the grayscale image.

[0091] Alternatively, edge detection can be performed on the grayscale image to obtain the edge variance of the grayscale image as the first quality evaluation information of the grayscale image; optionally, the edge variance of the grayscale image can be obtained by using a Laplace operator, and the edge variance is used as the first quality evaluation information of the grayscale image.

[0092] Alternatively, face detection can be performed on the grayscale image to obtain a face detection list of the grayscale image as the first quality evaluation information of the grayscale image; optionally, the CascadeClassifier component in the Open Source Computer Vision Library (OpenCV) can be used to load the Haar cascade classifier for face feature detection; specifically, the detectMultiscale function of the Haar cascade classifier for detecting multiple scales can be used to detect the faces in the grayscale image to obtain a list containing the bounding boxes of all detected faces, that is, the face detection list, as the first quality evaluation information of the grayscale image.

[0093] Alternatively, the gradient information of the grayscale image can also be obtained, and the clarity score of the grayscale image can be determined according to the gradient information as the first quality evaluation information of the grayscale image; optionally, algorithms such as the Sobel operator or Prewitt operator can be used to obtain the gradient information of the grayscale image. The gradient information is used to reflect the intensity and direction of the pixel brightness change. The average gradient of the grayscale image can be determined according to the gradient information of the grayscale image, and the average gradient is used as the clarity score to obtain the first quality evaluation information. The larger the average gradient, the higher the clarity score, and the clearer the grayscale image.

[0094] That is to say, in this embodiment, the first quality evaluation information may be the mean and variance of the RGB three channels corresponding to the grayscale image, the edge variance of the grayscale image, the face detection list of the grayscale image, or the clarity score of the grayscale image, so as to more comprehensively detect the picture quality.

[0095] In some implementations, the frame information of the current picture frame and the previous picture frame may also be obtained according to the detection process; in this embodiment, the frame information may be the relevant attribute information of the picture frame, such as the frame number, timestamp of the picture frame, and grayscale information or pixel matrix in the picture frame, etc.

[0096] Optionally, the pixel difference between the current picture frame and the previous picture frame may be determined according to the frame information as the second quality evaluation information of the current picture frame; the pixel difference between each pixel point of the picture frame and the previous picture frame may be calculated respectively, so as to obtain the pixel difference between the current picture frame and the previous picture frame as the second quality evaluation information of the current picture frame.

[0097] Or, according to the frame information, the frame number difference between the current picture frame and the previous picture frame may be determined as the second quality evaluation information of the current picture frame; that is, calculate the difference between the frame number of the current picture frame and the frame number of the previous picture frame to obtain the frame number difference as the second quality evaluation information of the current picture frame.

[0098] That is to say, in this embodiment, the second quality evaluation information may be the pixel difference between the current picture frame and the previous picture frame or the frame number difference between the current picture frame and the previous picture frame, and the accuracy and video smoothness are analyzed according to the frame information of the picture frame.

[0099] In some implementations, the audio-visual synchronization of the picture frame may also be detected to obtain the first confidence level of the picture frame; that is, to detect whether the audio and the picture in the picture frame are synchronized. For example, the synchronization network syncnet proposed in the prior art is used to detect the audio-visual synchronization to obtain the confidence level LSE-C indicating whether the audio and the picture are synchronized, that is, the first confidence level.

[0100] Furthermore, the audio and lip synchronization of the picture frame may be detected to obtain the second confidence level of the picture frame; optionally, the synchronization network syncnet proposed in the prior art may be used to detect the audio and lip synchronization to obtain the matching degree LSE-D of the audio and the lip, that is, the second confidence level; in this embodiment, the audio-visual synchronization detection method may be implemented based on code and packaged in the form of an sdk for the detection service to call. After inputting the picture frame to be detected, the output result includes the first confidence level and the second confidence level.

[0101] Further, based on the first confidence level and the second confidence level, third quality assessment information of the picture frame is obtained; in this embodiment, the third quality assessment information includes the first confidence level and the second confidence level, covering information on whether the audio and the picture are synchronized and whether the audio and the lip shape are synchronized, and performing audio-visual synchronization detection more accurately.

[0102] In some implementations, initial subtitles can also be extracted from the picture frame, and interfering text can be removed from the initial subtitles to obtain subtitle text; optionally, a video subtitle extractor (video-subtitle-extractor) can be used to extract the subtitles in the picture frame to obtain the initial subtitles, and in combination with a video watermark subtitle remover (video-subtitle-remover), interfering text such as watermarks and logo text in the initial subtitles can be removed to obtain complete subtitle text.

[0103] Optionally, the subtitle text can include one or more first strings, and time information for each first string, such as the start time and end time of the first string.

[0104] Further, human voice audio is extracted from the digital human video, and speech recognition is performed on the human voice audio to obtain audio text; optionally, a dynamic image expert group (Fast Forward Moving Picture Experts Group, FFmpeg) can be used to extract an audio file from the digital human video, and further, the demucs model is used to separate the human voice and background music in the audio file to obtain the human voice audio. Further, according to a pre-trained large model, text content is extracted from the human voice video, and a second string segmented by characters is output to obtain the audio text. The audio text can also include time information for each second string, such as the start time and end time corresponding to each second string.

[0105] Further, synchronization detection is performed on the subtitle text and the audio text to obtain fourth quality assessment information of the picture frame. Optionally, the first string in the subtitle text can be matched with the second string in the audio text. For example, character matching is performed between the second string and the first string, that is, the same characters are matched, and time information matching is performed on the mutually matched characters, so as to obtain the fourth quality assessment information. In this embodiment, the fourth quality assessment information at least includes the character matching result and the character time information matching result between the first string and the second string. By double determination of characters and time, the accuracy of subtitle and audio synchronization detection is improved.

[0106] S405, determine the preset quality assessment conditions for the detection object.

[0107] It can be understood that the quality assessment conditions are conditions for judging the specific quality results of the quality assessment information. In this embodiment, the preset quality assessment conditions can be one or more conditions. For example, it can include the first quality assessment condition corresponding to the first quality assessment information, the second quality assessment condition corresponding to the second quality assessment information, the third quality assessment condition corresponding to the third quality assessment information, and the fourth quality assessment condition corresponding to the fourth quality assessment information.

[0108] In some implementations, the first quality assessment condition can include a black screen judgment condition, a white screen judgment condition, a color distortion judgment condition, a digital human face disappearance judgment condition, and a low definition judgment condition; the second quality assessment condition can include a freezing judgment condition and a frame skipping judgment condition; the third quality assessment condition can be an audio-video out-of-sync judgment condition; the fourth quality assessment video can be a subtitle-audio out-of-sync judgment condition.

[0109] S406. Determine the sub-quality detection result of the picture frame according to the quality assessment information and the quality assessment conditions.

[0110] Optionally, the first sub-quality detection result can be determined according to the first quality assessment information and the first quality assessment condition; the second sub-quality detection result can be determined according to the second quality assessment information and the second quality assessment condition; correspondingly, the third sub-quality detection result can also be determined according to the third quality assessment information and the third quality assessment condition, and the fourth sub-quality detection result can be determined according to the fourth quality assessment information and the fourth quality assessment condition, where the sub-quality detection result of the picture frame is used to obtain the quality detection result of the digital human video.

[0111] For the first quality assessment information:

[0112] If the first quality assessment information is the mean and variance of the RGB three channels corresponding to the gray image, it can be judged whether the means of the RGB three channels are all less than the set black threshold, and whether the variances are all lower than the preset variance threshold. If the means of the RGB three channels are all less than the set black threshold, and the variances are all lower than the preset variance threshold, it is determined that the current first quality assessment information meets the black screen judgment condition in the first quality assessment condition, and it is determined that the first sub-quality detection result is that the picture is black.

[0113] In this embodiment, it can also be judged according to the means and variances of the RGB three channels whether the means of the RGB three channels are all greater than or equal to the set white threshold, and the variances are all lower than the preset variance threshold. If the means of the RGB three channels are all greater than or equal to the set white threshold, and the variances are all lower than the preset variance threshold, it is determined that the current first quality assessment information meets the white screen judgment condition in the first quality assessment condition, and it is determined that the first sub-quality detection result is that the picture is white.

[0114] It can be understood that if the means of the three RGB channels are not all less than the set black threshold, or if the variances are not all lower than the preset variance threshold, the first sub-quality detection result is that there is no black screen in the picture; correspondingly, if the means of the three RGB channels are not all greater than or equal to the set white threshold, or if there are variances that are all lower than the preset variance threshold, the first sub-quality detection result is that there is no white screen in the picture.

[0115] If the first quality assessment information is the edge variance of a grayscale image, determine whether the current edge variance is less than the minimum edge variance threshold or greater than the maximum edge variance threshold. In response to the edge variance being less than the minimum edge variance threshold or the edge variance being greater than the maximum edge variance threshold, it is determined that the current first quality assessment information meets the screen freeze judgment condition in the first quality assessment conditions, and it is determined that the first sub-quality detection result is a screen freeze.

[0116] It can be understood that if the edge variance is between the minimum edge variance threshold and the maximum edge variance threshold, it is determined that the first sub-quality detection result is that there is no screen freeze in the picture.

[0117] If the first quality assessment information is a face detection list, determine whether the length of the face detection list is zero. In response to the length of the face detection list being zero, it is determined that the face has disappeared, that is, the first quality assessment information meets the digital face disappearance judgment condition in the first quality assessment conditions, and the first sub-quality detection result is digital face disappearance; in some implementations, the first sub-quality detection result can also record the picture frame or frame number when the digital face disappears, as well as the timestamp of the picture frame, for easy problem tracing.

[0118] It can be understood that in response to the length of the face detection list not being zero, the first sub-quality detection result is that there is no digital face message problem.

[0119] If the first quality assessment information is the clarity score of a grayscale image, determine whether the current clarity score is lower than the preset clarity threshold. In response to the clarity score being lower than the preset clarity threshold, it indicates that the current clarity does not meet the requirements, and it is determined that the first quality assessment information of the current picture frame meets the low clarity judgment condition in the first quality assessment conditions, and the first sub-quality detection result is low clarity; in some implementations, the first sub-quality detection result can also record the picture frame or frame number of the current low clarity, as well as the timestamp of the picture frame, for easy tracing and optimization of unclear pictures.

[0120] It can be understood that if the clarity score of the current picture frame is greater than or equal to the preset clarity threshold, it is determined that the first sub-quality detection result is that the clarity of the picture frame is normal.

[0121] Regarding the second quality assessment information:

[0122] If the second quality assessment information is the pixel difference between the current video frame and the previous video frame, determine whether the pixel difference is less than the set difference threshold. If the pixel difference is less than the set difference threshold, it indicates that the similarity between the current video frame and the previous video frame is extremely high. Further determine whether the pixel differences of three consecutive frames are all lower than the set difference threshold. If the pixel differences of three consecutive frames are all lower than the set difference threshold, then determine that the second quality assessment information meets the frame freeze judgment condition in the second quality assessment criteria, and determine that the second sub-quality detection result is frame freeze. The second sub-quality detection result may also include the timestamp of the current video frame with frame freeze.

[0123] It can be understood that if the pixel differences of all video frames of the digital human video are greater than or equal to the set difference threshold, it is determined that there is no frame freeze in the digital human video, and the second sub-quality detection result is no frame freeze problem.

[0124] If the second quality assessment information is the frame number difference between the current video frame and the previous video frame, determine whether the frame number difference is 1. If the frame sequence number of the current frame is not the frame sequence number of the previous video frame + 1, that is, the frame number difference is not 1, then determine that the second quality assessment information meets the frame skipping judgment condition in the second quality assessment criteria, and determine that the second sub-quality detection result is frame skipping; in some implementations, the second sub-quality detection result may also record the timestamp of the current video frame with frame skipping for anomaly location.

[0125] It can be understood that if the frame sequence number of the current frame is the frame sequence number of the previous video frame + 1, that is, the frame number difference is 1, then it is determined that there is no frame skipping problem in the digital human video, and the second sub-quality detection result is no frame skipping problem; if the frame sequence numbers of all video frames in the digital human video increase, that is, the frame number difference between every two video frames is 1, then it is determined that there is no frame skipping in the digital human video.

[0126] Regarding the third quality assessment information: Based on the first confidence level and the second confidence level in the third quality assessment information, determine whether the first confidence level and the second confidence level are respectively within the corresponding threshold ranges. In response to either the first confidence level or the second confidence level not being within the corresponding threshold range, determine that the third quality assessment information meets the audio-video out-of-sync judgment condition in the third quality assessment criteria, and the third sub-quality detection result is audio-video out-of-sync.

[0127] It can be understood that if both the first confidence level and the second confidence level in the third quality assessment information are within the corresponding threshold ranges, then determine that the third sub-quality detection result is audio-video in-sync.

[0128] Regarding the fourth quality assessment information: Based on the character matching result and the matching result of character time information between the first string and the second string in the fourth quality assessment information, determine whether all characters in the first string and the second string can match each other, and the time information of the mutually matching characters is the same. If there are characters in the first string and the second string that do not match each other, or the characters match each other but the corresponding time information of the characters is different, then determine that the fourth quality assessment information meets the subtitle-audio out-of-sync judgment condition in the fourth quality assessment condition, and determine that the fourth sub-quality detection result is subtitle-audio out-of-sync; In some implementations, the fourth sub-quality detection result may also include the characters and character time information that did not match successfully. In this embodiment, not matching successfully includes two cases: character mismatch and different character time information.

[0129] It can be understood that if all characters in the first string and the second string can match each other, and the time information corresponding to the mutually matching characters is also exactly the same, then determine that the fourth sub-quality detection result is subtitle-audio in-sync.

[0130] Based on the comparison and judgment between each quality assessment information and the quality assessment condition, determine the first sub-quality detection result, the second sub-quality detection result, the third sub-quality detection result, and the fourth sub-quality detection result of the picture frame.

[0131] In some implementations, the quality detection result of the digital human video includes at least the first sub-quality detection result, the second sub-quality detection result, the third sub-quality detection result, and the fourth sub-quality detection result of the picture frame.

[0132] In some implementations, the video synthesis material may also include the digital human's broadcast text. Extract the human voice audio from the digital human video and perform speech recognition on the human voice audio to obtain the audio text. It can be understood that the digital human's broadcast text is the original text that the digital human needs to broadcast, and the audio text is the text that the digital human actually broadcasts in the digital human video.

[0133] Preprocess the audio text to obtain the first text, and preprocess the broadcast text to obtain the second text; The preprocessing in this embodiment may be to remove all Chinese and English punctuation marks in the text, leaving only the pure text. That is, remove the Chinese and English punctuation marks from the audio text and the broadcast text respectively to obtain the pure text first text and second text.

[0134] Further, content integrity detection is performed based on the first text and the second text to obtain the content detection result of the digital human video. That is, taking the second text of the broadcast text as the standard, all text contents in the first text are detected to determine whether the content of the first text is complete, and a content detection result with higher accuracy is obtained. The content detection result at least includes an indication result of content completeness or incompleteness, and may also include the text similarity between the first text and the second text when the content is incomplete, as well as the number of different characters and the positions of different characters between the first text and the second text.

[0135] In this embodiment, the content detection result is used to obtain the quality detection result of the digital human video. That is, the quality detection result of the digital human video at least further includes the content detection result; therefore, the quality detection result of the digital human video in this embodiment at least includes the first sub-quality detection result, the second sub-quality detection result, the third sub-quality detection result, the fourth sub-quality detection result of the picture frame and the content detection result.

[0136] It should be noted that in this embodiment, when obtaining the quality detection result, quality detection can be performed according to different execution modes of the target detection dimension. For example, each target detection dimension can be sequentially executed in a serial execution mode, or each target detection dimension can be simultaneously executed in a parallel execution mode. Specifically, any one of the embodiments of the present disclosure can be used to implement this, and no limitation is made here and no further description is given.

[0137] It should be noted that in this embodiment, when obtaining the quality detection result, quality detection can also be performed according to the detection threads in the called detection thread group. Specifically, any one of the embodiments of the present disclosure can be used to implement this, and no limitation is made here and no further description is given.

[0138] S407, determine the target digital human video from the video set according to the quality detection result.

[0139] The first sub-quality detection result, the second sub-quality detection result, the third sub-quality detection result, the fourth sub-quality detection result and the content detection result are summarized to obtain the quality detection result.

[0140] Exemplarily, assume that the first sub-quality detection result indicates that there is a black screen in the picture, the second sub-quality detection result indicates that there is no freezing or frame skipping, the third sub-quality detection result indicates that the audio and video are synchronized, the fourth sub-quality detection result indicates that the subtitles and audio are synchronized, and the content detection result indicates that the content is complete. Then the final quality detection result is that there is a black screen in the picture, and it includes the picture frame or frame number where the black screen in the picture is located, so as to quickly locate and process the picture frame with quality problems.

[0141] Optionally, a quality comparison graph corresponding to the video set can be generated according to the quality detection results of the digital human videos; that is, a quality comparison graph is generated from the quality detection results of each digital human video in the video set, so as to facilitate the comparison of the quality detection results of all digital human videos and more intuitively determine the digital human video with the optimal quality as the target digital human video.

[0142] Optionally, an optimization prompt can also be generated according to the quality detection results of the target digital human video; the large model optimizes the target digital human video according to the optimization prompt to generate the final digital human video; that is, after determining the target digital human video with the optimal quality, it can also be optimized according to the quality detection results of the digital human video. For example, if the clarity score of a certain frame in the target digital human video with the optimal quality is low, an optimization prompt can be generated to optimize the clarity of the frame with a low clarity score in the target digital human video through the large model, so as to obtain a better-quality final digital human video and improve the display effect of the digital human video.

[0143] In this embodiment, the corresponding RGB three-channel mean and variance, the edge variance of the grayscale image, the face detection list of the grayscale image, or the clarity score of the grayscale image of the frame are determined as the first quality evaluation information, the pixel difference between the current frame and the previous frame or the frame number difference between the current frame and the previous frame is used as the second quality evaluation information, the first confidence level of audio-visual synchronization and the second confidence level of audio and lip shape are used as the third quality evaluation information, and the character and time information matching between the first string of the subtitle text and the second string of the audio text is used as the fourth quality evaluation information. The quality detection results are determined by comparing the quality evaluation information with the corresponding quality evaluation conditions respectively. It can detect problems such as black screen, white screen, colorful screen, disappearance of digital human face, low clarity of the screen, screen freezing, frame skipping, audio-video out of sync, and subtitle-audio out of sync. It can also perform content integrity detection based on the second text of the broadcast text and the first text of the audio text to determine whether there is a quality problem of incomplete content, providing a more rich and accurate quality detection process. The quality detection results of the digital human video obtained almost cover all possible quality problems, thus ensuring that the effect of the target digital human video selected based on the quality detection results is better; furthermore, the target digital human video can be optimized based on the large model to improve the display effect of the digital human video, making the target digital human video more in line with user needs.

[0144] Figure 5 It is an exemplary logic diagram of a method for detecting the quality of a digital human video provided by an embodiment of the present disclosure. As Figure 5As shown in the figure, a video set is generated based on video synthesis materials. The video set includes digital human videos with the same video content such as Video 1, Video 2, and Video 3. The picture quality of the digital human videos is detected, that is, black screen, white screen, mosaic, and disappearance of the digital human face are detected to determine whether there are picture quality problems. If there are, the specific quality problems, the picture frames and timestamps with quality problems are recorded, and video smoothness analysis is performed; if there are no picture quality problems, video smoothness analysis is performed, that is, stuttering detection and frame skipping detection. If there are smoothness problems, the specific stuttering or frame skipping problems, the picture frames and timestamps with problems are recorded, and multimodal detection is performed; if there are no smoothness problems, multimodal synchronization detection is directly performed. Multimodal synchronization detection includes audio-video synchronization detection and subtitle-audio synchronization detection. If there are asynchronous problems, the specific asynchronous problems, the picture frames and timestamps with problems are recorded, and content integrity verification is performed; if there are no asynchronous problems, content integrity verification is directly performed, that is, integrity detection of the digital human's broadcast text. If there are problems with content integrity, the timestamps of missing or omitted characters are recorded and video clarity evaluation is performed; if the content integrity is okay, video clarity evaluation is directly performed, that is, it is detected whether there are blur or jagged problems in the clarity score. If there are blur or jagged problems, the clarity scores of the picture frames with problems are recorded, and the final aggregated quality detection results are output. If there are no blur or jagged problems, the final aggregated quality detection results are directly output to facilitate quickly locating the problem. It can be understood that in the specific implementation process, picture quality problems, video smoothness analysis, multimodal synchronization detection, content integrity, and clarity evaluation can all be executed simultaneously or in combination, which will not be elaborated here.

[0145] Figure 6 It is a schematic structural diagram of a quality detection device for digital human videos provided by an embodiment of the present disclosure. As Figure 6 shown, the quality detection device 600 for digital human videos includes:

[0146] The first acquisition module 601 is used to acquire video synthesis materials and generate a video set based on the video synthesis materials. The video set includes at least two digital human videos with the same video content;

[0147] The second acquisition module 602 is used to acquire a detection task. The detection task includes one or more target detection dimensions, and the target detection dimension includes at least one detection object;

[0148] The quality detection module 603 is used to perform quality detection on the digital human videos in the video set according to the detection task to obtain the quality detection results of the digital human videos;

[0149] A screening module 604, configured to determine a target digital human video from a video collection according to the quality detection result.

[0150] In some implementations, the quality detection module 603 includes:

[0151] Determine the execution mode of the target detection dimension according to the detection task;

[0152] Perform quality detection on the digital human video according to the execution mode of the target detection dimension and the detection object, and obtain the quality detection result of the digital human video.

[0153] In some implementations, the quality detection module 603 includes:

[0154] Determine the number of videos of digital humans in the video collection;

[0155] Invoke a detection thread group consistent with the number of videos, where the detection thread group includes at least one detection thread, and different detection thread groups correspond to detecting different digital human videos;

[0156] Send the corresponding digital human video to the detection thread group, and the detection thread performs quality detection on the corresponding digital human video according to the detection task.

[0157] In some implementations, the quality detection module 603 includes:

[0158] Determine the number of detection threads in the detection thread group according to the target detection dimension and the execution mode of the target detection dimension.

[0159] In some implementations, the quality detection module 603 includes:

[0160] Perform detection traversal on the target detection dimension according to the execution mode;

[0161] For the currently traversed target detection dimension, perform quality detection on the digital human video based on the detection process of the detection object under the target detection dimension, and obtain the quality detection result of the digital human video.

[0162] In some implementations, the quality detection module 603 includes:

[0163] Extract picture frames from the digital human video;

[0164] Process the picture frames according to the detection process of the detection object to determine the quality evaluation information of the picture frames;

[0165] Determine the preset quality evaluation conditions of the detection object;

[0166] Determine the sub-quality detection result of the video frame according to the quality assessment information and quality assessment conditions, where the sub-quality detection result of the video frame is used to obtain the quality detection result of the digital human video.

[0167] In some implementations, the quality detection module 603 includes:

[0168] Determine the grayscale image of the video frame according to the detection process of the detection object;

[0169] Obtain the mean and variance of the three RGB channels corresponding to the grayscale image as the first quality assessment information of the grayscale image; or,

[0170] Perform edge detection on the grayscale image to obtain the edge variance of the grayscale image as the first quality assessment information of the grayscale image; or,

[0171] Perform face detection on the grayscale image to obtain the face detection list of the grayscale image as the first quality assessment information of the grayscale image; or,

[0172] Obtain the gradient information of the grayscale image and determine the clarity score of the grayscale image according to the gradient information as the first quality assessment information of the grayscale image.

[0173] In some implementations, the quality detection module 603 includes:

[0174] Obtain the frame information of the current video frame and the previous video frame according to the detection process;

[0175] Determine the pixel difference between the current video frame and the previous video frame according to the frame information as the second quality assessment information of the current video frame; or,

[0176] Determine the frame number difference between the current video frame and the previous video frame according to the frame information as the second quality assessment information of the current video frame.

[0177] In some implementations, the quality detection module 603 includes:

[0178] Perform audio-visual synchronization detection on the video frame to obtain the first confidence level of the video frame;

[0179] Perform audio-lip synchronization detection on the video frame to obtain the second confidence level of the video frame;

[0180] Obtain the third quality assessment information of the video frame according to the first confidence level and the second confidence level.

[0181] In some implementations, the quality detection module 603 includes:

[0182] Extract the initial caption from the video frame and remove the interfering text from the initial caption to obtain the caption text;

[0183] Extract the human voice audio from the digital human video, and perform speech recognition on the human voice audio to obtain the audio text;

[0184] Perform synchronous detection on the subtitle text and the audio text to obtain the fourth quality assessment information of the video frames.

[0185] In some implementations, the quality detection module 603 includes:

[0186] Extract the human voice audio from the digital human video, and perform speech recognition on the human voice audio to obtain the audio text;

[0187] Preprocess the audio text to obtain the first text, and preprocess the broadcast text to obtain the second text;

[0188] Perform content integrity detection based on the first text and the second text to obtain the content detection result of the digital human video, and the content detection result is used to obtain the quality detection result of the digital human video.

[0189] In some implementations, the device 600 further includes:

[0190] Generate a quality comparison chart corresponding to the video set according to the quality detection result of the digital human video.

[0191] In some implementations, the device 600 further includes:

[0192] Generate an optimization prompt word according to the quality detection result of the target digital human video;

[0193] Optimize the target digital human video through the large model according to the optimization prompt word to generate the final digital human video.

[0194] In this embodiment, multiple digital human videos with the same video content are generated through different video generation methods, and the quality of each digital human video is detected according to the detection task. Through multiple target detection dimensions and multiple detection objects included in each target detection dimension, a comprehensive quality detection of each digital human video is performed to obtain a comprehensive and accurate quality detection result of the digital human video. Thus, according to the quality detection result, the target digital human video is determined from the video set to ensure that the quality of the target digital human video is optimal and improve the display effect of the target digital human video.

[0195] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0196] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0197] Figure 7 A schematic block diagram of an electronic device for implementing an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0198] As Figure 7 shown, the device 700 includes a computing unit 701 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0199] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as, for example, a keyboard, a mouse, etc.; an output unit 707, such as, for example, various types of displays, speakers, etc.; a storage unit 708, such as, for example, a magnetic disk, an optical disk, etc.; and a communication unit 709, such as, for example, a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0200] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the method for detecting the quality of a digital human video. For example, in some embodiments, the method for detecting the quality of a digital human video can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the method for detecting the quality of a digital human video described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the method for detecting the quality of a digital human video by any other suitable means (e.g., by means of firmware).

[0201] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0202] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0203] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0204] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0205] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0206] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0207] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0208] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A quality detection method for a digital human video, wherein: The method comprises: Acquire video synthesis material, and generate a video set based on the video synthesis material, wherein the video set includes at least two digital human videos with the same video content; Acquire a detection task, where the detection task includes one or more target detection dimensions, and the target detection dimension includes at least one detection object; According to the detection task, quality detection is performed on the digital human videos in the video set to obtain quality detection results of the digital human videos; According to the quality detection result, a target digital human video is determined from the video set.

2. The method according to claim 1, wherein: The step of performing quality detection on the digital human videos in the video set according to the detection task to obtain quality detection results of the digital human videos includes: Determining, according to the detection task, an execution mode of the target detection dimension; According to the execution mode of the target detection dimension and the detection object, the quality detection of the digital human video is performed to obtain the quality detection result of the digital human video.

3. The method according to claim 1, wherein: The step of performing quality detection on the digital human videos in the video set according to the detection task to obtain quality detection results of the digital human videos includes: Determine the number of digital human videos in the video collection; Calling a detection thread group that is consistent with the number of the videos, wherein the detection thread group includes at least one detection thread, and different detection thread groups correspond to detecting different digital human videos; The corresponding digital human video is sent to the detection thread group, and the detection thread performs quality detection on the corresponding digital human video according to the detection task.

4. The method according to claim 3, wherein: The calling of the detection thread group having the same number as the video comprises: The number of detection threads in the detection thread group is determined according to the target detection dimension and the execution mode of the target detection dimension.

5. The method according to claim 2, wherein: The step of performing quality detection on the digital human video according to the execution mode of the target detection dimension and the detection object to obtain the quality detection result of the digital human video includes: Performing detection traversal on the target detection dimension according to the execution mode; For the target detection dimension currently traversed, based on the detection process of the detection object under the target detection dimension, the quality detection of the digital human video is performed to obtain the quality detection result of the digital human video.

6. The method according to claim 5, wherein: The detection process based on the detection object under the target detection dimension, performing quality detection on the digital human video, includes: Extracting picture frames from the digital human video; Processing the picture frame according to the detection process of the detection object to determine quality assessment information of the picture frame; Determining a preset quality assessment condition of the detection object; The sub-quality detection result of the picture frame is determined according to the quality evaluation information and the quality evaluation condition, wherein the sub-quality detection result of the picture frame is used to obtain the quality detection result of the digital human video.

7. The method according to claim 5, wherein: The step of processing the picture frame according to the detection process of the detection object to determine the quality assessment information of the picture frame includes: Determining a gray image of the picture frame according to a detection process of the detection object; Obtaining the mean and variance of the RGB three channels corresponding to the gray image as first quality assessment information of the gray image; or, Performing edge detection on the gray image to obtain edge variance of the gray image as first quality assessment information of the gray image; or, Performing face detection on the gray image to obtain a face detection list of the gray image as first quality assessment information of the gray image; or, Gradient information of the gray image is acquired, and a clarity score of the gray image is determined according to the gradient information as first quality assessment information of the gray image.

8. The method according to claim 5, wherein: The step of processing the picture frame according to the detection process of the detection object to determine the quality assessment information of the picture frame includes: According to the detection process, frame information of the current picture frame and the previous picture frame is obtained; Determine, according to the frame information, a pixel difference between the current picture frame and the previous picture frame as second quality assessment information of the current picture frame; or According to the frame information, a frame number difference between the current picture frame and the previous picture frame is determined as second quality evaluation information of the current picture frame.

9. The method according to claim 5, wherein: The step of processing the picture frame according to the detection process of the detection object to determine the quality assessment information of the picture frame includes: Performing audio-video synchronization detection on the picture frame to obtain a first confidence level of the picture frame; Performing audio and lip synchronization detection on the picture frame to obtain a second confidence level of the picture frame; Third quality assessment information of the picture frame is obtained according to the first confidence level and the second confidence level.

10. The method according to claim 5, wherein: The step of processing the picture frame according to the detection process of the detection object to determine the quality assessment information of the picture frame includes: Extracting initial subtitles from the picture frame, and removing interfering text from the initial subtitles to obtain subtitle text; Extracting human voice audio from the digital human video, and performing speech recognition on the human voice audio to obtain an audio text; The subtitle text and the audio text are synchronously detected to obtain fourth quality assessment information of the picture frame.

11. The method according to claim 3, wherein: The video synthesis material includes the broadcast text of the digital human, and the detection process based on the detection object under the target detection dimension performs quality detection on the digital human video, including: Extracting human voice audio from the digital human video, and performing speech recognition on the human voice audio to obtain an audio text; Preprocessing the audio text to obtain a first text, and preprocessing the broadcast text to obtain a second text; A content integrity check is performed according to the first text and the second text to obtain a content check result of the digital human video, and the content check result is used to obtain a quality check result of the digital human video.

12. The method according to claim 1, wherein: After obtaining the quality detection result of the digital human video, the method further includes: A quality comparison graph corresponding to the video set is generated according to the quality detection result of the digital human video.

13. The method according to claim 1, wherein: After determining the target digital human video from the video set according to the quality detection result, the method further includes: generating optimized prompt words according to the quality detection result of the target digital human video; The target digital human video is optimized according to the optimization prompt words through the large model to generate a final digital human video.

14. A quality detection device for a digital human video, comprising: A first acquisition module, used for acquiring video synthesis materials, and generating a video set based on the video synthesis materials, wherein the video set includes at least two digital human videos with the same video content; A second acquisition module, used to acquire a detection task, wherein the detection task includes one or more target detection dimensions, and the target detection dimension includes at least one detection object; A quality detection module, used to perform quality detection on the digital human videos in the video set according to the detection task, and obtain quality detection results of the digital human videos; The screening module is used to determine the target digital human video from the video set according to the quality detection result.

15. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 13.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-13.

17. A computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 13.

Citation Information

Cited By

  • Communication video quality real-time automatic detection method and system based on AI technology

    CN121056683A

  • Three-dimensional digital human detection method and device based on large model, electronic equipment, medium, program product and three-dimensional digital human

    CN121169888A

  • Methods, devices, electronic equipment, media, and software products for 3D digital human detection based on large models, and 3D digital humans.

    CN121169888B