Video quality detection method and device, electronic equipment and storage medium

By segmenting and multi-dimensionally detecting videos, the problem of low efficiency and poor accuracy in video quality detection in existing technologies is solved, achieving efficient and accurate video quality assessment and automatic screening.

CN122454473APending Publication Date: 2026-07-24BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2026-03-17
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies for video quality detection based on artificial intelligence rely on manual review, which is inefficient and makes it difficult to obtain the generated materials, resulting in inaccurate detection results and an inability to assess video quality in a timely and effective manner.

Method used

By segmenting the target video, extracting image sequences, and performing quality assessments based on dimensions such as image clarity, temporal coherence, behavioral reasonableness, and semantic reasonableness, the quality assessment results of the video segments are generated, avoiding reliance on generated materials as a reference.

Benefits of technology

It improves the accuracy and efficiency of video quality detection, enables precise assessment and automatic screening of video quality, and ensures the consistency and reliability of assessment standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122454473A_ABST
    Figure CN122454473A_ABST
Patent Text Reader

Abstract

The present disclosure provides a quality detection method and device of a video, an electronic device and a storage medium, relates to the technical field of computers, in particular to the technical fields of computer vision, deep learning, large models, artificial intelligence and the like, and can be applied to scenarios such as digital people and content generation based on artificial intelligence. The specific implementation scheme is: performing segmentation processing on a target video to obtain a plurality of video segments; extracting a plurality of images in each video segment to obtain an image sequence corresponding to each video segment; detecting each image sequence according to a plurality of preset detection dimensions to obtain first detection results of each video segment under each preset detection dimension, wherein the preset detection dimensions include at least one of an image clarity dimension, a time sequence coherence dimension, a behavior rationality dimension and a semantic rationality dimension; and for each video segment, generating a quality evaluation result of the target video according to the first detection results under each preset detection dimension.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, particularly to the fields of computer vision, deep learning, large models, and artificial intelligence, and can be applied to scenarios such as digital humans and AI-based content generation. Specifically, it relates to a video quality detection method, device, electronic device, and storage medium. Background Technology

[0002] With the rapid development of artificial intelligence (AI) technology, AI-generated video technology has been widely applied. For example, digital human generation models can produce video content on a large scale and are widely used in fields such as virtual anchors and customer service. However, AI-generated videos often suffer from various quality issues, affecting their usability and user viewing experience. Therefore, quality control of these videos is crucial. Summary of the Invention

[0003] This disclosure provides a method, apparatus, electronic device, and storage medium for video quality detection.

[0004] According to one aspect of this disclosure, a video quality detection method is provided, comprising: The target video is segmented to obtain multiple video clips; Extract several images from each video segment to obtain the image sequence corresponding to each video segment; Each image sequence is detected according to multiple preset detection dimensions to obtain the first detection result of each video segment under each preset detection dimension. The preset detection dimensions include at least one of image clarity dimension, temporal coherence dimension, behavioral rationality dimension and semantic rationality dimension. For each video segment, a quality assessment result for the target video is generated based on the first detection result under each preset detection dimension.

[0005] According to another aspect of this disclosure, a video quality inspection apparatus is provided, comprising: The image acquisition module is used to segment the target video to obtain multiple video segments; The image acquisition module is also used to extract several images from each video segment to obtain the image sequence corresponding to each video segment; The detection module is used to detect each image sequence according to multiple preset detection dimensions to obtain the first detection result of each video segment under each preset detection dimension. The preset detection dimensions include at least one of image clarity dimension, temporal coherence dimension, behavioral rationality dimension and semantic rationality dimension. The processing module is used to generate a quality assessment result for the target video based on the first detection result under each preset detection dimension for each video segment.

[0006] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the video quality detection method described above.

[0007] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to perform the video quality detection method described above.

[0008] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the video quality detection method described above.

[0009] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0010] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is a schematic flowchart of a video quality detection method provided in an embodiment of this disclosure; Figure 2 This is a schematic diagram of a video cropping and frame extraction method provided in this disclosure. Figure 3 This is a schematic diagram of a timing coherence curve provided in an embodiment of this disclosure; Figure 4 This is a schematic diagram of another timing coherence curve provided in an embodiment of this disclosure; Figure 5 This is a schematic diagram of a behavior rationality detection method provided in an embodiment of this disclosure; Figure 6 This is a schematic diagram of a video quality detection process provided in an embodiment of this disclosure; Figure 7 This is a schematic diagram of a video quality detection device provided in an embodiment of this disclosure; Figure 8 This is a block diagram of an electronic device according to an embodiment of the present disclosure. Detailed Implementation

[0011] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0012] The terms such as "first," "second," and "third" used herein are used only to distinguish one entity (or operation) from another, and are not intended to require or imply a specific order or relationship between these entities (or operations). Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0013] The collection, storage, use, processing, transmission, provision, and disclosure of any type of information, such as user personal information, in this technical solution comply with relevant laws and regulations and do not violate public order and good morals.

[0014] Before describing the technical solutions provided by the embodiments of this disclosure, in order to facilitate the understanding of the embodiments of this disclosure, this disclosure first specifically explains the problems existing in the related technologies.

[0015] With the rapid development of artificial intelligence (AI) technology, AI-generated video technology has been widely applied. For example, digital human generation models can produce video content on a large scale and are widely used in fields such as virtual anchors and customer service. However, AI-generated videos often suffer from various quality issues, affecting their usability and user viewing experience.

[0016] Technical personnel have found that current video quality assessments primarily rely on manual review, especially since AI-powered video production on a large scale still requires significant manpower for individual review, resulting in low efficiency. While some automated video quality detection solutions exist, many rely on the source material used in video generation. For example, they assess video quality by comparing the similarity of the actions of people and animals in the source material to the synthesized video. However, for those reviewing video quality, obtaining the source material is difficult, hindering effective and timely review. Other video detection solutions use only portions of the video as criteria for quality assessment. Since these portions may also contain low-quality elements, the reliability of the results is poor, making accurate quality judgments difficult.

[0017] Therefore, there is an urgent need for a video quality inspection method to improve the accuracy and efficiency of the review results.

[0018] In view of this, this disclosure provides a video quality inspection method, apparatus, electronic device, and storage medium to improve the accuracy of video quality inspection and the efficiency of review.

[0019] The video quality detection method disclosed herein, such as Figure 1 As shown, the method may include steps 101 and 104.

[0020] Step 101: Segment the target video to obtain multiple video segments; Step 102: Extract several images from each video segment to obtain an image sequence corresponding to each video segment; Step 103: Detect each image sequence according to multiple preset detection dimensions to obtain the first detection result of each video segment under each preset detection dimension, wherein the preset detection dimensions include at least one of image clarity dimension, temporal coherence dimension, behavioral rationality dimension and semantic rationality dimension. Step 104: For each video segment, generate a quality assessment result for the target video based on the first detection result under each preset detection dimension.

[0021] Specifically, steps 101 and 102 involve a target video that requires quality inspection, specifically a video file or video stream, particularly videos generated by artificial intelligence models, such as digital human videos or virtual anchor videos. The video content includes, but is not limited to, dynamic objects such as people and animals, as well as static objects such as everyday items, plants, and decorative items.

[0022] The target video comprises multiple video segments, meaning each video segment originates from the target video. Multiple video segments can be obtained from the target video by following preset segmentation rules. Optionally, the preset segmentation rules can be based on fixed duration segmentation, scene switching segmentation, or coded keyframe interval segmentation.

[0023] A multi-frame image is a set of several frames extracted from each video segment. The frame extraction strategy can be equal-interval frame extraction, random frame extraction, or adaptive frame extraction based on the variation of frame content. In this disclosure, the set of multi-frame images extracted from each video segment and arranged in chronological order constitutes the image sequence, thus obtaining the image sequence corresponding to each video segment.

[0024] Optionally, segmenting the target video refers to dividing the target video into multiple video segments along the timeline. These multiple video segments, arranged chronologically, sequentially cover the playback interval of the target video. Optionally, if the total duration of the target video cannot be divided evenly by a preset duration, the remaining duration at the end can be used to form a tail segment with a duration shorter than the preset duration, or the tail segment can be merged into the preceding video segment to ensure the integrity of the video segments.

[0025] By segmenting the target video, the quality detection granularity can be further refined to the level of video segments, which can reduce false judgments caused by detecting the entire segment and facilitate the accurate location of the time interval where quality problems occurred.

[0026] When extracting a preset number of images from each video segment, the process can be performed using an equal-interval frame extraction method, which means selecting the corresponding frames at fixed time intervals within the time range of each video segment; or it can be performed using a fixed-interval frame extraction method, which means selecting the corresponding frames at fixed frame number intervals, thereby obtaining the image sequence corresponding to each video segment.

[0027] The image sequence corresponding to each video segment serves as the common input for multiple preset detection dimensions, enabling subsequent detections of image clarity, temporal coherence, behavioral rationality, and semantic rationality dimensions to be performed based on the same object, ensuring the comparability of detection results across dimensions and facilitating statistical analysis.

[0028] By constructing image sequences, subsequent detection can not only be judged based on the quality of a single frame image, but also comprehensively evaluated based on inter-frame changes, which helps to accurately detect problems such as blur, out-of-focus, frame skipping, and semantic anomalies.

[0029] Next, in step 103, different preset detection dimensions correspond to different video quality detection methods. Among them, the preset detection dimensions include at least one of the following: image clarity dimension, temporal coherence dimension, behavioral rationality dimension, and semantic rationality dimension.

[0030] For example, in the image sharpness dimension, the degree of detail discernibility in each frame of an image sequence can be detected, especially for effective detection of image quality issues such as blurriness and defocus. For example, each frame can be detected based on sharpness evaluation operators, which include, but are not limited to, the Laplacian operator, gradient energy operator, wavelet multi-scale sharpness analysis, etc.

[0031] In the temporal coherence dimension, the continuity and consistency of image sequences along the time axis can be detected, thereby effectively detecting frame skipping issues in videos. Optionally, the detection result in this dimension can be determined by adjacent frame similarity, frame skipping detection, or inter-frame motion consistency.

[0032] In terms of behavioral rationality, the reasonableness of the movement, interaction, or state changes of a target object in an image sequence can be detected. For example, the rationality of the limbs in a digital human video can be detected. In some optional embodiments, the focus can be on detecting abnormalities in the hands, such as whether there are distortions, defects, twists, or other abnormalities in the hands. Similarly, the rationality of the movement speed and turning changes of vehicles and animals can also be detected, specifically whether there are sudden accelerations, instantaneous displacements, abrupt changes in movement direction, or other abnormalities that do not conform to physics or logic.

[0033] In terms of semantic rationality, semantic understanding and logical judgment can be performed on image content to detect whether the semantic content of the image sequence is consistent, whether there are unreasonable events or text anomalies, such as objects appearing or disappearing out of thin air, sudden changes in the attributes of the same object, and abnormal text content in the image.

[0034] In this disclosure, step 104 involves performing detection processes corresponding to each preset detection dimension on the same image sequence, thereby quickly obtaining the first detection result of the image sequence under each preset detection dimension, which serves as the result of the video segment's video quality in the corresponding dimension. Optionally, the first detection result can be a quality level label, or information used to distinguish video quality, such as a quality assessment score.

[0035] Based on this, the detection results for each preset detection dimension are summarized to obtain the corresponding detection result for the video segment. When summarizing the first detection results for different dimensions, the higher the proportion of information indicating excellent video quality, the higher the quality of the corresponding video segment.

[0036] Finally, the detection results for each video segment are further aggregated to determine the quality assessment result of the target video.

[0037] When the detection results for a video segment are represented by a quality level label, the quality detection results for the video segment can be further summarized to determine the quality of each video segment, analyze the proportion of video segments of different quality, and then output qualified, unqualified, or multi-level quality levels. When the detection results for a video segment are represented by a score, the quality level corresponding to the video quality can be determined by taking the average value, taking the minimum value, or statistically analyzing the proportion of video segments below the threshold.

[0038] In this disclosure, based on the quality assessment results, videos can be classified into quality levels, and low-quality videos can be promptly intercepted and alerted in subsequent video distribution, recommendation or storage stages, as well as high-quality videos can be automatically selected and low-quality videos can be filtered, thereby controlling video quality in real time.

[0039] In this disclosure, by segmenting the target video, the granularity of video quality detection is further refined to the segment level, facilitating precise location of the time point where quality issues occur, thereby improving the accuracy of video quality detection. Furthermore, throughout the entire detection process, there is no need to rely on the source material used to generate the video or on parts of the video as references. The video quality is evaluated from multiple dimensions, including image quality, frame-by-frame coherence, character plausibility, and physical logic plausibility, and the quality evaluation results are output. Since the detection results from different dimensions can complement each other, the occurrence of missed detections and false detections can be effectively reduced, thus improving the reliability of the detection results. Simultaneously, the quality detection scheme provided in this disclosure can be applied to any video, enabling efficient review of a large number of videos. In addition, quality evaluation based on objective algorithms and models ensures the consistency of evaluation standards, helping video producers to form a referable video quality evaluation standard.

[0040] In some embodiments of this disclosure, the duration of each video segment is a preset duration. The target video is segmented to obtain multiple video segments, including: performing sliding window segmentation on the target video according to the preset duration and the preset sliding step size to obtain multiple video segments of preset duration, wherein the preset sliding step size is less than the preset duration.

[0041] Specifically, each video segment has a preset duration, which allows for a uniform detection granularity across all segments, facilitating subsequent comparisons of detection scores across different segments. Optionally, the preset duration can be 2 seconds, 3 seconds, or 5 seconds, etc.

[0042] By successively moving a window along the timeline of the target video and capturing video content within the window's coverage area, multiple video clips are obtained. By setting the preset sliding step size to be less than the preset duration, adjacent video clips can have overlapping time intervals, such as... Figure 2As shown, ensuring that the continuity detection of the video does not miss any anomalies at the boundaries can improve the detection probability of short-term, sporadic quality problems.

[0043] In some embodiments of this disclosure, each image sequence is detected according to multiple preset detection dimensions to obtain the first detection result of each video segment under each preset detection dimension, including: when the preset detection dimension includes the image clarity dimension, detecting the clarity of each frame of the image sequence to obtain the clarity score of each frame of the image; and determining the first detection result of the video segment under the image clarity dimension based on the number of images in the image sequence whose clarity scores are within a preset range.

[0044] For example, calculating image sharpness can involve applying edge enhancement or gradient operators to each frame of the image to obtain response information, and then statistically calculating the response information to obtain a sharpness score. For instance, the variance, energy, and / or mean of the response map can be calculated to obtain the sharpness score. It should be noted that the embodiments of this disclosure do not limit the specific implementation of calculating image sharpness, as long as it can output a sharpness score that characterizes the degree of sharpness. A higher sharpness score indicates a sharper image frame, while a lower sharpness score indicates a blurrier image frame or that is out of focus.

[0045] The effective range of the sharpness score can be defined based on the preset sharpness range. Optionally, the preset range can be in the form of a single threshold, such as the sharpness score being greater than or equal to a first threshold, or in the form of a double threshold, such as the sharpness score being between the first threshold and the second threshold, to characterize that the sharpness of the image is acceptable.

[0046] After determining the sharpness score for each frame, initial quality level labels can be used to characterize the video's quality in that dimension. Initial quality level labels can be categorized as acceptable, unacceptable, or have multiple quality levels.

[0047] Optionally, a score can also be used to reflect the quality of the video in this dimension. Specifically, the number of frames in the image sequence that meet the preset range constraints can be counted to obtain a quantity value. This quantity value is then mapped to the first detection result of the video segment in the image clarity dimension. For example, the quantity value can be directly used as the first detection result, or a score obtained by normalizing the quantity value can be used as the first detection result, or segmented first detection results can be determined based on the interval where the quantity value lies. This allows the first detection result to reflect the overall level of the proportion of clear frames in the image clarity dimension of the video segment.

[0048] Optionally, when using a score, the range of the first detection result of the video clip in the image clarity dimension can be limited to between 0 and 1, so as to facilitate the subsequent fusion of the first detection results of each dimension and avoid the inconsistency of the units causing one dimension to naturally suppress or amplify other dimensions.

[0049] In this disclosure, the image sharpness is detected, and the number of images with sharpness scores within a preset range is counted. Based on this, the first detection result of the video segment in the image sharpness dimension is determined, so that the first detection result can directly reflect the coverage of clear and usable images in the video segment, thereby having a higher detection capability for problems such as short-term blur, phased defocus, and intermittent sharpness fluctuations.

[0050] In this disclosure, when the preset detection dimension includes the temporal coherence dimension, the system detects whether there are frame skips in the image sequence and generates detection results.

[0051] If the detection result indicates the presence of frame skipping, the first detection result of the video segment in the temporal coherence dimension is determined as the first quality label; if the detection result indicates the absence of frame skipping, the first detection result of the video segment in the temporal coherence dimension is determined as the second quality label, and the second quality label is greater than the first quality label.

[0052] Specifically, frame skipping refers to the phenomenon where content abruptly changes between two adjacent frames, which does not conform to the normal playback interval, thus disrupting the continuity of the image.

[0053] In an optional implementation, frame skipping detection in an image sequence can be based on frame content abrupt change detection. Specifically, this involves calculating the Structural Similarity Index (SSIM) value between adjacent frames to determine the degree of structural similarity between the two images. Structural information refers to the image information perceived by the human eye, such as brightness, contrast, and consistency of texture / structure. A normal video's SSIM curve should change smoothly; a sudden drop indicates frame skipping. Figure 3 As shown, the left figure is the curve without frame skipping, and the right figure is the curve with frame skipping.

[0054] In another alternative implementation, the detection of frame skipping in an image sequence can be based on motion vector analysis. Specifically, an optical flow algorithm is used to calculate the motion vectors between adjacent frames. The motion amplitude in a normal video should change continuously; a sudden increase indicates frame skipping. Combined with... Figure 4 As shown, the left figure is the curve without frame skipping, and the right figure is the curve with frame skipping.

[0055] Continue to combine Figure 4As shown, there is a pulse between two frames where frame skipping occurs. A threshold is set based on the pulse size. Taking motion vector analysis as an example, the threshold can be set to 1.5 to quickly determine whether frame skipping exists in the image sequence.

[0056] In this disclosure, only one detection method can be used to detect temporal continuity, or two or more detection methods can be used to detect temporal continuity, and the detection results of different detection methods can complement each other.

[0057] When the detection results indicate the presence of frame skipping, quality level labels can be used to characterize the video's quality in that dimension. For example, a first quality label is used to indicate lower video quality, reflecting the negative impact of frame skipping on the viewing experience, while a second quality label is used to indicate the video's quality level, reflecting the good temporal continuity of the video segment.

[0058] Alternatively, the score can be used directly as a quality label. For example, when the detection result indicates that there is frame skipping, the first detection result is set to a lower first score; when the detection result indicates that there is no frame skipping, the first detection result of the temporal coherence dimension is set to a higher second score. The first and second scores can be preset constants, such as the first score being 0 or 0.2 and the second score being 1 or 0.8. There are no specific restrictions here.

[0059] In this disclosure, automatic identification of temporal continuity defects in video segments is achieved, enabling the rapid detection of frame skipping, a factor that significantly affects the viewing experience.

[0060] In this disclosure, when the preset detection dimension includes the behavior rationality dimension, each image sequence is input into the pose estimation model to identify the limb behavior of the behavior object in each frame image through the pose estimation model, and the region image corresponding to the limb behavior is input. The region image corresponding to each frame is input into the limb behavior anomaly detection model, so that the limb behavior anomaly detection model outputs the first detection result under the behavior rationality dimension.

[0061] Specifically, the behavioral object refers to the target object in the image sequence that needs to be tested for behavioral rationality, such as people or animals.

[0062] Limb behavior refers to the posture or action state characterized by the relative positional relationship of the limb key points of the action object in a single frame image, such as standing, walking, raising hands, clenching fists, bending over, etc., but this embodiment does not limit the specific action category.

[0063] Pose estimation models can perform pose recognition on behavioral objects in a single frame image. Their output includes at least the limb key points of the behavioral object, the confidence of the key points, and the skeleton structure information composed of the key points.

[0064] Specifically, the pose estimation model automatically determines the region containing limb keypoints in the image and crops the corresponding region image from the image. Based on the region image, the behavior object can be locally focused, which helps to reduce background interference and reduce input redundancy of subsequent anomaly detection models, making subsequent anomaly judgments of limb behavior more focused on the pixel region related to the behavior.

[0065] The abnormal body behavior detection model can determine whether there is an abnormality in body behavior. Its input region image is output as the detection score under the dimension of behavior rationality, so as to characterize the degree of rationality of the video segment corresponding to the image sequence under the dimension of behavior rationality.

[0066] Optionally, the first detection result output by the model can be configured as a numerical result, with a value range of [0, 1]. The higher the detection score, the more the limb behavior conforms to the preset behavior pattern, and the lower the detection score, the higher the degree of abnormality of the limb behavior.

[0067] The first detection result output by the model can also be configured as a label. For example, the first detection result can be a label indicating whether it is qualified, unqualified, or a multi-level quality level.

[0068] By performing region image extraction and anomaly detection model inference on each frame of the image sequence, the frame-level detection result corresponding to each frame can be obtained. Furthermore, the detection results of multiple frames can be fused to obtain the first detection result of the image sequence under the behavioral rationality dimension.

[0069] As a concrete example, in detecting abnormalities in the hands, a human pose estimation model can be used to detect key points in each frame of the image and locate the skeletal key points of the person. For example... Figure 5 As shown, the hand region image is cropped from the original image based on the detected coordinates of key hand points. Next, the cropped hand region image is input into a specially trained hand anomaly detection model to determine if there are any abnormalities such as distortion, incompleteness, or twisting in the hand.

[0070] In this disclosure, by performing pose recognition and anomaly detection in separate steps, it is beneficial to improve the accuracy of recognition in situations such as out-of-focus, occlusion, and complex backgrounds, and enhance the ability to detect behavioral anomalies such as missing limbs, pose jumps, and sudden movements, thereby improving the reliability of video quality detection results.

[0071] In this disclosure, with the preset detection dimensions including the semantic rationality dimension, the analysis prompt information of each image sequence and the item display logic is input into the visual semantic analysis model. The visual semantic analysis model analyzes the rationality of the item display logic in each frame of the image and outputs the first detection result under the semantic rationality dimension.

[0072] Specifically, the object display logic refers to the logic of the rationality of the display of objects in an image. For example, whether the appearance, disappearance, change of position, change of quantity, and change of form of objects are reasonable and consistent with the scene shown in the image. Among them, objects can be commodities, props, utensils, signs, or other visual objects, and there are no restrictions here.

[0073] Visual semantic analysis models are models capable of semantic understanding and logical judgment of video content, such as multimodal visual language models. Multimodal quality inspection models can be built using pre-trained visual language models as their base model; for example, the base model could be Qwen2.5-VL-7B.

[0074] In an optional embodiment, taking the quality detection of digital human videos as an example, to adapt the multimodal quality inspection model to the digital human video quality inspection task, supervised fine-tuning (SFT) can be performed on the digital human video quality inspection dataset. The digital human video quality inspection dataset can include a large number of digital human video segments as digital human video samples, and manually annotated quality assessment captions corresponding to each digital human video sample. The quality assessment captions are used to characterize the quality assessment conclusions and / or quality problem descriptions of the corresponding digital human videos, serving as supervisory signals for the supervised fine-tuning, thereby enabling the multimodal quality inspection model to output quality assessment results for digital human videos.

[0075] By using image sequences as visual input and analysis prompts as textual input, the visual semantic analysis model can perform semantic understanding and logical judgment on each frame of the image sequence under the guidance of the analysis prompts.

[0076] For example, the analysis prompts include: the category information of the item to be analyzed, the judgment points corresponding to the item display logic, and the definition information of the abnormal type, and outputs a score to represent the semantic rationality of the image sequence. The score is a value in the range of [0, 1]. The higher the detection score, the more reasonable the item display logic is, and the lower the detection score, the higher the degree of semantic abnormality.

[0077] Optionally, the analysis prompts may also guide or instruct the model to represent the semantic rationality of the image sequence in the form of quality level labels.

[0078] In this disclosure, by inputting the image sequence and the analysis prompts of the item display logic into the visual semantic analysis model, the visual semantic analysis model can combine inter-frame semantic association and business logic constraints to identify higher-level semantic anomalies, such as items appearing or disappearing out of thin air, items' positions or quantities changing abruptly without basis, items' shapes or key attributes changing abnormally, thereby improving the ability to detect anomalies in complex content and thus achieving automatic detection of whether the item display is reasonable.

[0079] Furthermore, when the first detection result is in the form of a score, the detection score under the semantic rationality dimension can be normalized to a unified range, such as [0, 1], so that the semantic rationality detection result can be uniformly integrated with the detection scores of the image clarity dimension, temporal coherence dimension, and behavioral rationality dimension to form a comparable second quality score.

[0080] Based on this, by performing the detection process corresponding to each preset detection dimension on the same image sequence, the detection score of the image sequence under each preset detection dimension can be quickly obtained, which serves as a numerical result of the video quality of the video segment under the corresponding dimension.

[0081] In some embodiments of this disclosure, the first detection result includes a first quality score of the video segment; for each video segment, a quality assessment result of the target video is generated based on the first detection result under each preset detection dimension, including: for each video segment, calculating the first quality score under each preset detection dimension according to preset calculation rules to obtain a second quality score for each video segment; and determining the quality assessment result of the target video based on the second quality score of each video segment.

[0082] Specifically, a second quality score for each video segment is obtained by fusing the first quality scores under each preset detection dimension. Finally, the second quality score of each video segment is further statistically analyzed to determine the quality assessment result of the target video.

[0083] Statistical analysis of video clips, such as averaging, minimizing, or statistically analyzing the proportion of video clips below a threshold, outputs a quality rating of "qualified," "unqualified," or multiple quality levels. This yields a quality assessment result for the target video within the overall context.

[0084] Based on the overall quality assessment of the target video, low-quality videos are promptly intercepted and alerted during subsequent video distribution, recommendation, or storage processes. High-quality videos are automatically selected while low-quality videos are filtered out, thereby controlling video quality in real time.

[0085] For each video segment, a second quality score is determined based on the detection scores under each preset detection dimension. Specifically, this may include: when there are multiple preset detection dimensions, for each video segment, the first quality score under each preset detection dimension is weighted and calculated according to the weight corresponding to each preset detection dimension to obtain the second quality score of each video segment.

[0086] Specifically, for the same video segment, the detection scores of the video segment under each preset detection dimension are calculated along with their corresponding weights to obtain a second quality score. The second quality score can be used as a comprehensive score of the overall quality of the video segment.

[0087] For example, taking the quality assessment results as a score, when the detection dimensions include image clarity, temporal coherence, behavioral plausibility, and semantic plausibility, the video quality detection process is as follows: Figure 6 As shown.

[0088] In some embodiments, the weighted calculation can employ a weighted summation technique, where the detection scores for each preset detection dimension are multiplied by their corresponding weights and then summed to obtain the second quality score. This allows the second quality score to directly reflect the comprehensive result of the detection scores for each dimension under weight constraints. Optionally, the weights corresponding to each preset detection dimension can be configured according to actual application requirements.

[0089] In this disclosure, a second quality score for each video segment is obtained by weighted calculation based on the first quality score under each preset detection dimension, thereby achieving multi-dimensional comprehensive quantification of video segment quality and enabling different types of quality defects to be considered simultaneously under a unified scoring framework.

[0090] In some embodiments of this disclosure, the quality assessment result includes a quality score of the target video; determining the quality assessment result of the target video based on the second quality score of each video segment includes: determining the lowest second quality score among the second quality scores corresponding to multiple video segments as the quality score of the target video.

[0091] Specifically, after determining the second quality score for each video segment, the scores between the various video segments can be compared, and the second quality score with the smallest value can be selected as the quality score of the target video.

[0092] Based on this, during the video quality detection process, as long as any video segment in the target video has a low second quality score, the impact of the low-quality segment can be reflected in the quality score of the target video. This allows the quality assessment results to accurately reflect whether there is low-quality content in the target video, thereby reducing the occurrence of low-quality content being masked by high-quality content due to average statistics, and improving the ability to identify quality defects that are short-term but have a serious impact on the user experience.

[0093] Optionally, the video segment corresponding to the lowest second quality score can also be output as a low-quality localization result to indicate the time interval in the target video where the quality problem is most prominent.

[0094] In some embodiments of this disclosure, the quality assessment results also include quality analysis results. When the preset detection dimension includes a semantic rationality dimension, the visual semantic analysis model is also used to output the scoring reasons for each video segment. The method also includes: obtaining the scoring reasons corresponding to each video segment whose second quality score is lower than a preset score, and obtaining the quality analysis results.

[0095] Specifically, when the visual semantic analysis model detects the semantic rationality dimension of a video segment, it can also output explanatory content to explain why the video segment is judged as low or abnormal in the semantic rationality dimension.

[0096] The scoring reasons may include information such as a description of the anomaly type, the range of frame numbers in which the anomaly occurred, and a description of the key evidence corresponding to the anomaly.

[0097] As a concrete example, the detection results output by the model can include two parts: a score and a reasoning for the score, for example: Detection score: 0.11.

[0098] Reasons for scoring: 1. Objects appearing and disappearing out of thin air: The transparent spherical pendant on the side of the bag exists in frames 1-7, suddenly disappears in frames 8-10, and at the same time, a yellow pendant appears out of thin air in frames 9-10...

[0099] 2. Object consistency defect: A small red label suddenly appears in the middle of the green part of the bag between frames 3 and 5, and the label disappears in frame 8, without a reasonable explanation...

[0100] 3. Abnormal distortion of objects: In frames 8-10, the outline of the bag shows an unnatural blurry texture effect... This does not conform to the laws of real physics.

[0101] In this disclosure, after calculating the second quality score for each video segment, video segments with second quality scores lower than a preset score are selected, and the scoring reasons are extracted from the output of the visual semantic analysis model corresponding to these video segments. The extracted scoring reasons are then summarized to form the quality assessment result of the target video.

[0102] In this disclosure, by summarizing the reasons for the low scores of video segments and providing explanatory information corresponding to the reasons for the low scores in the quality assessment results, the main quality problems of the target video can be automatically summarized and located to the specific time interval of the video segment. This can effectively improve the efficiency of video quality detection and also improve the interpretability of quality assessment.

[0103] Corresponding to the video quality detection method provided in this disclosure, this disclosure also provides a video quality detection device, including an image acquisition module 701, a detection module 702, and a processing module 703.

[0104] The image acquisition module 701 is used to segment the target video to obtain multiple video segments; The image acquisition module 701 is also used to extract several images from each video segment to obtain an image sequence corresponding to each video segment; The detection module 702 is used to detect each image sequence according to multiple preset detection dimensions to obtain the first detection result of each video segment under each preset detection dimension, wherein the preset detection dimensions include at least one of image clarity dimension, temporal coherence dimension, behavioral rationality dimension and semantic rationality dimension. The processing module 703 is used to generate a quality assessment result of the target video for each video segment based on the first detection result under each preset detection dimension.

[0105] In embodiments of this disclosure, the image acquisition module 701 is further configured to perform sliding window segmentation on the target video according to a preset duration and a preset sliding step size to obtain multiple video segments of preset duration, wherein the preset sliding step size is less than the preset duration.

[0106] In embodiments of this disclosure, the detection module 702 is further configured to detect the sharpness of each frame of the image sequence when the preset detection dimension includes the image sharpness dimension, and obtain a sharpness score for each frame of the image. The detection module 702 is further configured to determine the first detection result of the video segment in the image clarity dimension based on the number of images in the image sequence whose clarity scores are within a preset range.

[0107] In embodiments of this disclosure, the detection module 702 is further configured to detect whether there are frame skips in the image sequence when the preset detection dimension includes a temporal coherence dimension, and generate a first detection result; The detection module 702 is further configured to determine the first detection result of the video segment in the temporal coherence dimension as the first quality label when the detection result indicates that there is frame skipping; The detection module 702 is further configured to determine, when the detection result indicates that there is no frame skipping, the first detection result of the video segment in the temporal coherence dimension as the second quality label, wherein the video quality indicated by the second quality label is higher than the video quality indicated by the first quality label.

[0108] In the embodiments of this disclosure, the detection module 702 is further configured to, when the preset detection dimension includes the behavior rationality dimension, input each image sequence into the pose estimation model, so as to identify the limb behavior of the behavior object in each frame image through the pose estimation model, and input the region image corresponding to the limb behavior; The detection module 702 is also used to input the region image corresponding to each frame image into the limb behavior abnormality detection model, so as to output the first detection result under the behavior rationality dimension through the limb behavior abnormality detection model.

[0109] In embodiments of this disclosure, the detection module 702 is further configured to, when the preset detection dimension includes a semantic rationality dimension, input the analysis prompt information of each image sequence and the item display logic into a visual semantic analysis model, analyze the rationality of the item display logic in each frame of the image through the visual semantic analysis model, and output the first detection result under the semantic rationality dimension.

[0110] In embodiments of this disclosure, the first detection result includes a first quality score of the video segment; The processing module 703 is further configured to calculate the first quality score under each preset detection dimension according to preset calculation rules for each video segment, thereby obtaining the second quality score of each video segment; and determine the quality assessment result of the target video based on the second quality score of each video segment.

[0111] In the embodiments of this disclosure, the processing module 703 is further configured to, when there are multiple preset detection dimensions, calculate the first quality score under each preset detection dimension according to the weight corresponding to each preset detection dimension for each video segment, so as to obtain the second quality score of each video segment.

[0112] In embodiments of this disclosure, the quality assessment result includes the quality score of the target video; The processing module 703 is further configured to determine the lowest second quality score among the second quality scores corresponding to the plurality of video segments as the quality score of the target video.

[0113] In embodiments of this disclosure, the quality assessment results also include quality analysis results. When the preset detection dimension includes a semantic rationality dimension, the visual semantic analysis model is also used to output the scoring reason for each video segment. The processing module 703 is also used to obtain the scoring reason corresponding to the video segment whose second quality score is lower than the preset score, and to obtain the quality analysis result.

[0114] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0115] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0116] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0117] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0118] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as a video quality detection method. For example, in some embodiments, the video quality detection method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the video quality detection method described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the video quality detection method by any other suitable means (e.g., by means of firmware).

[0119] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0120] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0121] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0122] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0123] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0124] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0125] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0126] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A video quality detection method, comprising: The target video is segmented to obtain multiple video clips; Extract several images from each video segment to obtain an image sequence corresponding to each video segment; Each image sequence is detected according to multiple preset detection dimensions to obtain the first detection result of each video segment under each preset detection dimension. The preset detection dimensions include at least one of image clarity dimension, temporal coherence dimension, behavioral rationality dimension and semantic rationality dimension. For each video segment, a quality assessment result of the target video is generated based on the first detection result under each preset detection dimension.

2. The method according to claim 1, wherein, The duration of each video segment is a preset duration. The segmentation process of the target video yields multiple video segments, including: The target video is segmented by sliding window according to a preset duration and a preset sliding step size to obtain multiple video segments of preset duration, wherein the preset sliding step size is less than the preset duration.

3. The method according to claim 1, wherein, The step of detecting each image sequence according to multiple preset detection dimensions to obtain the first detection result of each video segment under each preset detection dimension includes: When the preset detection dimension includes the image sharpness dimension, the sharpness of each frame in the image sequence is detected to obtain a sharpness score for each frame. The first detection result of the video segment in the image clarity dimension is determined based on the number of images in the image sequence whose sharpness scores are within a preset range.

4. The method according to claim 1, wherein, The step of detecting each image sequence according to multiple preset detection dimensions to obtain the first detection result of each video segment under each preset detection dimension includes: When the preset detection dimension includes the temporal coherence dimension, the system detects whether there are frame skips in the image sequence and generates a first detection result. If the detection result indicates the presence of frame skipping, the first detection result of the video segment in the temporal coherence dimension is determined as the first quality label; If the detection result indicates that there are no frame skips, the first detection result of the video segment in the temporal coherence dimension is determined as the second quality label, and the video quality indicated by the second quality label is higher than the video quality indicated by the first quality label.

5. The method according to claim 1, wherein, The step of detecting each image sequence according to multiple preset detection dimensions to obtain the first detection result of each video segment under each preset detection dimension includes: When the preset detection dimension includes the behavior rationality dimension, each image sequence is input into the pose estimation model to identify the limb behavior of the object in each frame image through the pose estimation model, and the region image corresponding to the limb behavior is input. The region image corresponding to each frame is input into the limb behavior anomaly detection model, so that the limb behavior anomaly detection model outputs the first detection result under the behavior rationality dimension.

6. The method according to claim 1, wherein, The step of detecting each image sequence according to multiple preset detection dimensions to obtain the first detection result of each video segment under each preset detection dimension includes: When the preset detection dimension includes the semantic rationality dimension, the analysis prompt information of each image sequence and the item display logic is input into the visual semantic analysis model. The visual semantic analysis model analyzes the rationality of the item display logic in each frame of the image and outputs the first detection result under the semantic rationality dimension.

7. The method according to claim 6, wherein, The first detection result includes the first quality score of the video segment; For each of the video segments, a quality assessment result for the target video is generated based on the first detection result under each preset detection dimension, including: For each video segment, the first quality score under each preset detection dimension is calculated according to the preset calculation rules to obtain the second quality score of each video segment; The quality assessment result of the target video is determined based on the second quality score of each video segment.

8. The method according to claim 7, wherein, For each video segment, a first quality score is calculated according to a preset calculation rule for each preset detection dimension to obtain a second quality score for each video segment, including: When there are multiple preset detection dimensions, for each video segment, the first quality score under each preset detection dimension is weighted and calculated according to the weight corresponding to each preset detection dimension to obtain the second quality score of each video segment.

9. The method according to any one of claims 7 or 8, wherein, The quality assessment results include the quality score of the target video; The process of determining the quality assessment result of the target video based on the second quality score of each video segment includes: Among the second quality scores corresponding to the multiple video segments, the lowest second quality score is determined as the quality score of the target video.

10. The method according to claim 7, wherein, The quality assessment results also include quality analysis results. When the preset detection dimensions include the semantic rationality dimension, the visual semantic analysis model is also used to output the scoring reasons for each video segment. The method further includes: Obtain the reasons for the scores of video segments whose second quality scores are lower than the preset scores, and obtain the quality analysis results.

11. A video quality inspection device, comprising: The image acquisition module is used to segment the target video to obtain multiple video segments; The image acquisition module is also used to extract several images from each video segment to obtain an image sequence corresponding to each video segment; The detection module is used to detect each image sequence according to multiple preset detection dimensions to obtain the first detection result of each video segment under each preset detection dimension, wherein the preset detection dimensions include at least one of image clarity dimension, temporal coherence dimension, behavioral rationality dimension and semantic rationality dimension. The processing module is used to generate a quality assessment result of the target video for each video segment based on the first detection result under each preset detection dimension.

12. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-10.

13. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-10.

14. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-10.