Video quality detection method, device, equipment and computer readable storage medium

The target object features in the video frame are extracted through image recognition technology, the broken object video frame is determined and its quality is judged, which solves the problem of inaccurate low-quality broken face detection in the prior art, and improves the recall and accuracy of video quality detection.

CN115116103BActive Publication Date: 2025-06-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110286029.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-17
Publication Date
2025-06-06
Estimated Expiration
2041-03-17

AI Technical Summary

Technical Problem

There is a lack of effective low-quality incomplete face detection methods in the prior art, resulting in inaccurate video quality detection and low recall rate.

Method used

Through image recognition technology, the target detection frame, target key points and target attributes of the target object in the video frame to be detected are extracted, the broken object video frame is determined, and the target quality is judged based on the target attributes, and the number of video frames is calculated for detection.

Benefits of technology

It improves the recall rate of videos for target quality incomplete objects, ensures the accuracy of video quality detection, reduces misjudgment, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115116103B_ABST
    Figure CN115116103B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a video quality detection method, device, equipment and computer-readable storage medium, which relate to the field of artificial intelligence technology. The method includes: extracting video frames of a video to be detected to obtain at least two video frames to be detected; performing image recognition on each video frame to be detected to obtain a target detection frame, target key points and target attributes of the target object in each video frame to be detected; determining the incomplete object video frame in at least two video frames to be detected based on the target detection frame and the target key points; determining the number of target quality incomplete object video frames and target quality incomplete object video frames in the incomplete object video frame based on the target attributes of each incomplete object video frame; determining the detection result of the video to be detected based on the number of video frames. Through the present application, it is possible to accurately detect whether the video to be detected is a target quality incomplete object video, thereby improving the recall rate of the target quality incomplete object video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and relate to but are not limited to a video quality detection method, apparatus, device, and computer-readable storage medium. Background Art

[0002] Low-quality incomplete faces are a common low-quality problem in images and videos. They are caused by malicious cropping of images and videos. Malicious cropping means that the producers or editors of images and videos maliciously crop a certain proportion of the original content material in order to circumvent copyright deduplication or other reasons, resulting in obvious incompleteness of the content of the picture. When low-quality incomplete faces appear in the picture, it will seriously affect the user's browsing experience, so this feature needs to be identified and blocked.

[0003] A defective face is only a necessary feature of a low-quality defective face, but not a sufficient feature. Some images with defective faces are not "low-quality" defective faces. To meet the "low-quality" condition, it is also necessary to ensure that the defective face itself is indeed the main subject of the image.

[0004] In the related art, among the image and video quality detection schemes, only target detection, general face detection and face key point detection are disclosed, and there are no related solutions to solve the problems of low-quality images and videos with incomplete faces. Therefore, the related art does not disclose any detection method for low-quality images and videos with incomplete faces. Summary of the invention

[0005] The embodiments of the present application provide a video quality detection method, device, equipment and computer-readable storage medium, which relate to the field of artificial intelligence technology. By determining the number of video frames of target quality defective object video frames based on the target detection frame, target key points and target attributes of the target object in the video frame to be detected obtained by image recognition, it is possible to accurately detect whether the video to be detected is a target quality defective object video, thereby improving the recall rate of the target quality defective object video.

[0006] The technical solution of the embodiment of the present application is implemented as follows:

[0007] The present application provides a video quality detection method, the method comprising:

[0008] Extract video frames from the video to be detected to obtain at least two video frames to be detected;

[0009] Performing image recognition on each of the video frames to be detected to obtain a target detection frame, target key points, and target attributes of the target object in each of the video frames to be detected;

[0010] Determine, according to the target detection frame and the target key point, a defective object video frame in at least two frames of video frames to be detected, wherein the defective object video frame is a video frame in which the target object is a defective object;

[0011] According to the target attribute of each of the defective object video frames, determining target quality defective object video frames and the number of video frames of the target quality defective object video frames in the defective object video frames;

[0012] According to the number of video frames, a detection result of the video to be detected is determined.

[0013] The present application provides a video quality detection device, the device comprising:

[0014] A video frame extraction module, used to extract video frames from the video to be detected, to obtain at least two video frames to be detected;

[0015] An image recognition module is used to perform image recognition on each of the video frames to be detected, and obtain a target detection frame, target key points and target attributes of the target object in each of the video frames to be detected;

[0016] A first determination module is used to determine a defective object video frame in at least two frames of video frames to be detected according to the target detection frame and the target key point, wherein the defective object video frame is a video frame in which the target object is a defective object;

[0017] A second determining module is used to determine, according to the target attribute of each of the defective object video frames, target quality defective object video frames and the number of video frames of the target quality defective object video frames in the defective object video frames;

[0018] The third determination module is used to determine the detection result of the video to be detected according to the number of video frames.

[0019] An embodiment of the present application provides a computer program product or a computer program, wherein the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; wherein a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor is used to execute the computer instructions to implement the above-mentioned video quality detection method.

[0020] An embodiment of the present application provides a video quality detection device, including: a memory, used to store executable instructions; a processor, used to implement the above-mentioned video quality detection method when executing the executable instructions stored in the memory.

[0021] An embodiment of the present application provides a computer-readable storage medium storing executable instructions for causing a processor to execute the executable instructions to implement the above-mentioned video quality detection method.

[0022] The embodiments of the present application have the following beneficial effects: image recognition is performed on each video frame to be detected extracted from the video to be detected, and the target detection frame, target key points and target attributes of the target object in each video frame to be detected are obtained; based on the target detection frame and target key points of the target object in each video frame to be detected, the incomplete object video frame is determined, and in the incomplete object video frame, the target quality incomplete object video frame is determined according to the target attributes of the target object, and then the video quality detection is performed on the video to be detected according to the number of video frames of the target quality incomplete object video frame. In this way, since the target detection frame, target key points and target attributes of the target object in the video frame to be detected obtained by image recognition can accurately determine the number of video frames of the target quality incomplete object video frame, it is possible to accurately detect whether the video to be detected is a target quality incomplete object video, thereby improving the recall rate of the target quality incomplete object video. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a schematic diagram of some low-quality incomplete human faces provided in the embodiments of the present application;

[0024] Figure 2 This is a schematic diagram of a physically defective face rather than a low-quality defective face provided by an embodiment of the present application;

[0025] Figure 3 This is a schematic diagram of another physically defective human face rather than a low-quality defective human face provided in an embodiment of the present application;

[0026] Figure 4 This is an optional architecture diagram of a video quality detection system provided in an embodiment of the present application;

[0027] Figure 5 is a structural diagram of a video quality detection device provided in an embodiment of the present application;

[0028] Figure 6 This is an optional flow chart of the video quality detection method provided in the embodiment of the present application;

[0029] Figure 7 This is an optional flow chart of the video quality detection method provided in the embodiment of the present application;

[0030] Figure 8 It is a structural diagram of a video quality detection model provided in an embodiment of the present application;

[0031] Fig. 9It is a flowchart of a method for training a video quality detection model provided in an embodiment of the present application;

[0032] Fig.10 This is a low-quality incomplete face detection flow chart provided in an embodiment of the present application;

[0033] Fig.11 It is a part of the video frames extracted from the same video using the video frame extraction method provided in the embodiment of the present application;

[0034] Fig.12 It is a schematic diagram of the IoU provided in an embodiment of the present application;

[0035] Fig.13 This is a schematic diagram of the design principle of the incomplete sample bbox definition provided in the embodiment of the present application;

[0036] Fig.14 This is a schematic diagram of a human face with a large proportion of defect provided by an embodiment of the present application;

[0037] Fig.15 This is a schematic diagram of blackening a portion of a face area provided in an embodiment of the present application. DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.

[0039] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments, but it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as those commonly understood by those skilled in the art of the technical field of the embodiments of the present application. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0040] Before describing the video quality detection method of the embodiment of the present application, the video detection method in the related art is first described:

[0041] Low-quality incomplete faces are caused by malicious cropping of the picture. Malicious cropping may cause phenomena such as incomplete subtitles, incomplete icons, and incomplete faces. Because faces appear in most of the content in pictures and videos, low-quality incomplete faces are a very obvious feature.

[0042] Low-quality incomplete faces can be simply divided into upper incomplete, lower incomplete, and left and right incomplete. According to the current definition, upper incomplete means that the eyebrows of the main face are above the edge of the screen, lower incomplete means that the lower lip of the main face is below the lower edge, and left and right incomplete means that the facial area above one eye is outside the left and right edges of the screen. When low-quality incomplete faces appear on the screen, it will seriously affect the user's browsing experience, so this feature needs to be identified and intercepted. Figure 1 is a schematic diagram of some low-quality incomplete human faces provided in the embodiments of the present application, such as Figure 1 As shown in Figures a to d in , there is a human face in each picture, and the facial features of the human face are partially incomplete.

[0043] It is worth noting that an incomplete face is only a necessary feature of a low-quality incomplete face, but not a sufficient feature. Some images with incomplete faces are not "low-quality" incomplete faces. To meet the "low-quality" condition, it is also necessary to ensure that the incomplete face itself is indeed the subject of the image. The relationship between the subject of the image and the face may be varied, such as there may be a face in the image but the subject is other objects rather than the face, there are multiple faces in the image and they are all the subject, there are multiple faces in the image but the subject is only part of the face, etc. Figure 2 Schematic diagram of a physically defective face rather than a low-quality defective face provided by an embodiment of the present application. Figure 2 As shown in Figure a in the figure, there is a face in the picture. Although the face 201 in the picture is a mutilated face, the picture is about selling free-range chicken, that is, the main body of the picture is "free-range chicken with Thai seafood sauce", so Figure a is not a low-quality mutilated face; Figure 2 As shown in Figure b, although the face 202 in the picture is a defective face, the picture is about eating broadcast, that is, the main body of the picture is the food 203 in front of the face 202, so Figure a is not a low-quality defective face.

[0044] Figure 3 is another schematic diagram of a physically defective face rather than a low-quality defective face provided by an embodiment of the present application, such as Figure 3 As shown in Figure a in FIG, although there is a defective face 301 in the picture, the defective face 301 is not the main body of the picture, and the main body of the picture is the complete face 302. Therefore, Figure a is not a low-quality defective face; Figure 3 As shown in Figure b, although there is an incomplete face 303 in the picture, the incomplete face 303 is not the main body of the picture. The main body of the picture is the complete face 304. Therefore, Figure b is not a low-quality incomplete face.

[0045] Currently, there is no direct solution to the problem of low-quality face defects. Public papers in related fields include object detection, face detection, and face key point detection. Object detection is intended to detect and identify the borders of general objects, while face detection is used to detect and identify the borders of faces in the picture. Face key point detection is used to predict the detailed areas of the face, such as eyebrows, eyes, mouth, nose, cheeks, etc. Depending on the coarse and fine granularity, 5 to thousands of key points can be predicted.

[0046] As a core branch of target detection, face detection has made great progress. With deep learning solutions and diverse human samples, it has achieved good detection results in difficult areas such as extreme lighting, large angles, complex makeup, and occlusion. At the same time, face key point detection can also rely on the detected faces to better learn key points. However, the current public models and solutions still have great limitations for the detection of low-quality incomplete faces, mainly in the following three aspects.

[0047] 1) Due to demand constraints, conventional face detection models only detect faces that are mostly within the frame, and do not detect faces that are more outside the frame. Therefore, the recall effect for incomplete faces is poor, and incomplete faces are missed on a large scale. At the same time, the face key point detection model does not perform well for key point detection outside the frame, and there will be obvious key point deviations.

[0048] 2) For the detected incomplete faces, the current solution lacks a good solution to judge whether they are "low quality", that is, it is difficult to accurately judge whether the face is the main expression of the picture. If there is no restriction on this, directly intercepting all the content with incomplete faces will cause very serious accidental injuries.

[0049] 3) Many public articles use public datasets such as WIDER Face, AFLW, and WFLW for training and verification. These datasets are inconsistent with the data distribution of real production environments, so they cannot better fit the data distribution of actual production environments, which will lead to overfitting.

[0050] Based on the above problems existing in the related art, the embodiment of the present application provides a video quality detection method, which modifies the sampling module of the usual face detection. On the one hand, it modifies the positive and negative sample definitions of the conventional face detection, and on the other hand, it deliberately crops the faces in the sample at the time of sampling, so that the model can see as many incomplete faces as possible, improve the recall ability of incomplete faces, and adopts simple data enhancement to solve the problem of poor prediction of super-frame faces by face key points. In addition, the method performs an additional round of screening on the detected incomplete faces, and uses a combination of strategy and model to comprehensively judge whether a incomplete face really belongs to the main content of the picture. At the same time, based on the real online environment, a large number of real full-field video frames and video cover images are taken, and a large number of data sets consistent with the online distribution are constructed at low cost through prompt rapid manual annotation. This data set is used for training and verification, which can better fit the data distribution of the actual production environment and avoid overfitting.

[0051] The video quality detection method provided by the embodiment of the present application first extracts video frames from the video to be detected to obtain at least two video frames to be detected; performs image recognition on each video frame to be detected to obtain the target detection frame, target key points and target attributes of the target object in each video frame to be detected; then, according to the target detection frame and the target key points, in the at least two video frames to be detected, the incomplete object video frame is determined, and the incomplete object video frame is a video frame in which the target object is an incomplete object; according to the target attribute of each incomplete object video frame, the number of video frames of the target quality incomplete object video frame and the target quality incomplete object video frame is determined in the incomplete object video frame; finally, according to the number of video frames, the detection result of the video to be detected is determined. In this way, since the target detection frame, target key points and target attributes of the target object in the video frame to be detected obtained by image recognition can accurately determine the number of video frames of the target quality incomplete object video frame, it is possible to accurately detect whether the video to be detected is a target quality incomplete object video, thereby improving the recall rate of the target quality incomplete object video.

[0052] The following describes an exemplary application of the video quality detection device of the embodiment of the present application. In one implementation, the video quality detection device provided by the embodiment of the present application can be implemented as a laptop, tablet computer, desktop computer, mobile device (e.g., mobile phone, portable music player, personal digital assistant, dedicated messaging device, portable gaming device), intelligent robot, vehicle-mounted computer, wearable electronic device, smart home, VR / AR device, and any other terminal with image display or video playback function; in another implementation, the video quality detection device provided by the embodiment of the present application can also be implemented as a server. Below, an exemplary application of the video quality detection device when implemented as a server will be described.

[0053] See also Figure 4 , Figure 4 It is an optional architecture diagram of the video quality detection system 10 provided in the embodiment of the present application. In order to realize accurate video quality detection of the video to be detected, the video quality detection system 10 provided in the embodiment of the present application includes a terminal 100, a network 200 and a server 300, and a video playback application is running on the terminal 100, and the video playback application can play the video to be detected. In the embodiment of the present application, the user can form or input the video to be detected on the client of the video playback application on the terminal, and the terminal generates a video quality detection request, and the video quality detection request includes the video to be detected. The terminal sends the video quality detection request to the server 300 through the network 200 to request the server 300 to perform video quality detection on the video to be detected, and when the detection result is that the video quality is qualified, the terminal is allowed to play the video to be detected.

[0054] In the embodiment of the present application, the server 300 obtains the video to be detected, extracts the video frames of the video to be detected, and obtains at least two frames of video frames to be detected; then, image recognition is performed on each video frame to be detected to obtain the target detection frame, target key points and target attributes of the target object in each video frame to be detected; then, based on the target detection frame and the target key points, the incomplete object video frame is determined in at least two frames of video frames to be detected; based on the target attributes of each incomplete object video frame, the number of video frames of incomplete object video frames with target quality and incomplete object video frames with target quality is determined in the incomplete object video frame; finally, based on the number of video frames, the detection result of the video to be detected is determined. After obtaining the detection result, the server 300 sends the detection result to the terminal 100, and the terminal 100 can display the detection result on the current interface 100-1. In some embodiments, when the detection result indicates that the video quality is qualified, that is, the video to be detected is not a target quality incomplete object video, the terminal is allowed to play the video to be detected; when the detection result indicates that the video quality is unqualified, that is, the video to be detected is a target quality incomplete object video, the terminal is prohibited from playing the video to be detected.

[0055] The video quality detection method provided in the embodiment of the present application can also be implemented based on a cloud platform and through cloud technology. For example, the server 300 can be a cloud server, and the cloud server performs video quality detection on the video to be detected to obtain the detection result. Alternatively, a cloud storage can also be provided, and the video to be detected and the corresponding detection result can be stored in the cloud storage, so that when the user plays the video to be detected again later, the detection result of the video to be detected can be directly obtained from the cloud storage, without the need for the server to perform re-detection, thereby reducing the data processing capacity of the server.

[0056] It should be noted here that cloud technology refers to a hosting technology that unifies hardware, software, network and other resources in a wide area network or local area network to achieve data computing, storage, processing and sharing. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool that can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The backend services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark, and all need to be transmitted to the backend system for logical processing. Data of different levels will be processed separately. All kinds of industry data require strong system backing support, which can only be achieved through cloud computing.

[0057] In some embodiments, the video quality detection method provided in the embodiments of the present application also involves the field of artificial intelligence technology, and the detection result corresponding to the video to be detected can be determined by artificial intelligence technology, that is, the video to be detected can be extracted by artificial intelligence technology, the image recognition can be performed on the video to be detected by artificial intelligence technology, the incomplete object video frame can be determined by artificial intelligence technology, the target quality incomplete object video frame can be determined by artificial intelligence technology, etc. In some embodiments, a video quality detection model can also be trained by artificial intelligence technology, and the video quality detection method of the embodiment of the present application can be implemented by the video quality detection model, that is, the detection result of the video to be detected is automatically generated by the video quality detection model.

[0058] In the embodiment of the present application, at least it can be realized by machine learning technology and computer vision technology in artificial intelligence technology. Among them, machine learning (ML, Machine Learning) is a multi-field cross-disciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines, specializing in how computers simulate or realize human learning behavior, in order to acquire new knowledge or skills, reorganize the existing knowledge structure to continuously improve its performance. Machine learning is the core of artificial intelligence, and is the fundamental way to make computers intelligent, and its application is spread across all fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning. Computer vision technology (CV, Computer Vision) is a science that studies how to make machines "see", and further, it refers to machine vision such as using cameras and computers to replace human eyes to identify and measure targets, and further do graphic processing to make computer processing more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies, trying to establish an artificial intelligence system that can obtain information from images or multidimensional data. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and map construction, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.

[0059] Figure 5 is a schematic diagram of the structure of a video quality detection device provided in an embodiment of the present application, Figure 5 The video quality detection device shown includes: at least one processor 310, a memory 350, at least one network interface 320 and a user interface 330. The various components in the video quality detection device are coupled together through a bus system 340. It can be understood that the bus system 340 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 340 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 340 is not used in the embodiment of the present invention. Figure 5 Various buses are labeled as bus system 340 .

[0060] The processor 310 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0061] The user interface 330 includes one or more output devices 331 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 330 also includes one or more input devices 332, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0062] The memory 350 may be removable, non-removable or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drive, optical disk drive, etc. The memory 350 may optionally include one or more storage devices physically away from the processor 310. The memory 350 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 350 described in the embodiment of the present application is intended to include any suitable type of memory. In some embodiments, the memory 350 can store data to support various operations, and examples of these data include programs, modules, and data structures or subsets or supersets thereof, as exemplarily described below.

[0063] Operating system 351, including system programs for processing various basic system services and performing hardware-related tasks, such as framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0064] A network communication module 352, for reaching other computing devices via one or more (wired or wireless) network interfaces 320, exemplary network interfaces 320 include: Bluetooth, Wireless Compatibility Authentication (WiFi), and Universal Serial Bus (USB);

[0065] The input processing module 353 is used to detect one or more user inputs or interactions from one of the one or more input devices 332 and translate the detected inputs or interactions.

[0066] In some embodiments, the device provided in the embodiments of the present application can be implemented in software. Figure 5A video quality detection device 354 stored in the memory 350 is shown. The video quality detection device 354 may be a video quality detection device in a video quality detection device, which may be software in the form of a program or a plug-in, and includes the following software modules: a video frame extraction module 3541, an image recognition module 3542, a first determination module 3543, a second determination module 3544, and a third determination module 3545. These modules are logical, and therefore may be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.

[0067] In other embodiments, the device provided in the embodiments of the present application can be implemented in hardware. As an example, the device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the video quality detection method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components.

[0068] The video quality detection method provided in the embodiment of the present application will be described below in combination with the exemplary application and implementation of the video quality detection device provided in the embodiment of the present application, wherein the video quality detection device can be any terminal with image display or video playback function, or it can also be a server, that is, the video quality detection method of the embodiment of the present application can be executed by the terminal, can also be executed by the server, or can also be executed by the terminal interacting with the server.

[0069] See also Figure 6 , Figure 6 This is an optional flow chart of the video quality detection method provided in the embodiment of the present application. Figure 6 The steps shown are explained, and it should be noted that Figure 6 The video quality detection method in the embodiment is a video quality detection method implemented by using a server as an execution subject.

[0070] Step S601: extract video frames from the video to be detected to obtain at least two video frames to be detected.

[0071] Here, the video to be detected may be any type of video, for example, it may be a video shot by a user through a terminal, or a video downloaded by a user, or a video produced by a user through video production software.

[0072] Extracting video frames from a video to be detected refers to extracting a certain number of video frames to be detected from the video to be detected according to a specific extraction method, wherein the extraction method can be equal-interval extraction, that is, extracting the video frames to be detected from the video to be detected according to a certain time interval, or can be unequal-interval extraction, that is, randomly extracting a certain number of video frames to be detected from the video to be detected. It should be noted that in the embodiment of the present application, two consecutive video frames will not be extracted from the video to be detected.

[0073] Step S602 , performing image recognition on each to-be-detected video frame to obtain a target detection frame, target key points and target attributes of the target object in each to-be-detected video frame.

[0074] Here, any image recognition method can be used to perform image recognition on the video frame to be detected to identify the target object in the video frame to be detected, as well as the target detection frame corresponding to the target, at least one target key point corresponding to the target object, and the target attribute of the target object. For example, an image recognition network can be used to perform image recognition on the video frame to be detected to determine whether there is a target object in the video frame to be detected. When there is a target object, the target detection frame, target key points, and target attributes of the target object are identified. Among them, the target detection frame is the detection frame corresponding to the area where the target object is located in the video frame to be detected; the target key point is the key point in the target object, and the characteristic information of the target object can be represented by the key point; the target attribute includes but is not limited to the type of the target object, the information represented by the target object in the video frame to be detected, the clarity of the target object, the relationship between the target object and other objects in the video frame to be detected, the type of the video to be detected, or the content to be expressed by the video to be detected as a whole, etc.

[0075] In some embodiments, the target object may be a face. Accordingly, step S602 may be to perform face recognition on each video frame to be detected to obtain a face video frame having a face; then, perform image recognition on each face video frame to obtain a face detection frame, face key points and face attributes in each face video frame. The face detection frame refers to a detection frame corresponding to the area where the face is located; the face key points may be key points corresponding to the eyes, nose, mouth, etc. in the face; and the face attributes may be attributes of the face itself, for example, the face definition, and the relationship between the face and other faces or objects in the video frame to be detected.

[0076] Step S603, determining a defective object video frame from at least two to-be-detected video frames according to the target detection frame and the target key points, where the defective object video frame is a video frame in which the target object is a defective object.

[0077] Here, it is possible to determine whether the target object in each to-be-detected video frame is an incomplete object based on the coordinates of the target detection frame and the relationship between the target key point and the target detection frame. If the target object is an incomplete object, the to-be-detected video frame is determined to be an incomplete object video frame. In the embodiment of the present application, the incomplete object refers to an object in which part of the image of the target object is not displayed in the to-be-detected video frame.

[0078] In an embodiment of the present application, the target detection frame obtained through image recognition is a predicted detection frame that includes a complete target object. That is to say, for an incomplete object or a non-incomplete object of the same object, the predicted target detection frame is of the same size, except that for an incomplete object, the target detection frame also includes a portion of the area that exceeds the content displayed by the video frame to be detected, and for a non-incomplete object, the target detection frame is an area that completely surrounds the target object in the video to be detected.

[0079] In an embodiment of the present application, since the target detection frame of the incomplete object also includes a portion of the area that exceeds the area displayed by the video frame to be detected, the coordinates of this portion of the area are negative coordinates, and the coordinates corresponding to the portion of the area located in the video frame to be detected are positive coordinates. Therefore, it is possible to determine whether the target object is an incomplete object based on the coordinates of the target detection frame.

[0080] In some embodiments, when the coordinates of the target detection frame have negative coordinates and the predicted target key points exceed the border of the video frame to be detected, that is, are located outside the border of the video frame to be detected, then the target object can be determined as an incomplete object based on the coordinates of the target detection frame, wherein the presence of negative coordinates in the coordinates of the target detection frame indicates that the target object is not fully displayed in the video frame to be detected, but does not mean that the target object is an incomplete object, because the definition of an incomplete object is that certain target key points in the target object are not displayed in the video frame to be detected, and therefore it is necessary to continue to judge based on the target key points, that is, to judge whether the target key points exceed the border of the video frame to be detected.

[0081] Step S604: determining target quality defective object video frames and the number of target quality defective object video frames in the defective object video frames according to the target attribute of each defective object video frame.

[0082] Here, based on the target attributes of each incomplete object video frame, it is determined whether the target object is the subject of expression in the incomplete object video frame. If so, the incomplete object video frame is determined to be a target quality incomplete object video frame. For example, the target quality incomplete object video frame can be a low quality incomplete object video frame.

[0083] Step S605: determining the detection result of the video to be detected according to the number of video frames.

[0084] In some embodiments, step S605 can be implemented by the following steps: determining the video length of the video to be detected, and determining a threshold value for the number of video frames for evaluating the video to be detected based on the video length; when the number of video frames is greater than the threshold value for the number of video frames, it indicates that the number of low-quality and incomplete object video frames is large, and therefore, the detection result is determined to be that the video to be detected is a target quality and incomplete object video; when the number of video frames is less than or equal to the threshold value for the number of video frames, it indicates that the number of low-quality and incomplete object video frames is small, and therefore, the detection result is determined to be that the video to be detected is a non-target quality and incomplete object video, that is, a normal video.

[0085] The video quality detection method provided in the embodiment of the present application performs image recognition on each video frame to be detected extracted from the video to be detected, and obtains the target detection frame, target key points and target attributes of the target object in each video frame to be detected; determines the incomplete object video frame based on the target detection frame and target key points of the target object in each video frame to be detected, and determines the target quality incomplete object video frame based on the target attributes of the target object in the incomplete object video frame, and then performs video quality detection on the video to be detected based on the number of video frames of the target quality incomplete object video frame. In this way, since the target detection frame, target key points and target attributes of the target object in the video frame to be detected obtained by image recognition can accurately determine the number of video frames of the target quality incomplete object video frame, it is possible to accurately detect whether the video to be detected is a target quality incomplete object video, thereby improving the recall rate of the target quality incomplete object video.

[0086] based on Figure 6 , Figure 7 This is an optional flow chart of the video quality detection method provided in an embodiment of the present application. In some embodiments, there are multiple target key points of the target object in each video frame to be detected. The method for determining the incomplete object video frame in step S603 can be implemented by the following steps:

[0087] Step S701, for any video frame to be detected, determine whether the target detection frame of the video frame to be detected has negative coordinates. If the determination result is yes, execute step S702; if the determination result is no, determine that the video frame to be detected is a normal video frame.

[0088] Step S702, determine whether at least one target key point of the video frame to be detected exceeds the border of the video frame to be detected. If the determination result is yes, execute step S703; if the determination result is no, determine that the video frame to be detected is a normal video frame.

[0089] Step S703, determining whether the number of target key points that exceed the border in the video frame to be detected is less than a preset number threshold. If the judgment result is yes, the video frame to be detected is determined to be an incomplete object video frame; if the judgment result is no, the video frame to be detected is determined to be an invalid video frame. Here, an invalid video frame refers to a video frame in which the target object is displayed less in the video frame to be detected and can be ignored. For example, an invalid video frame may be a video frame in which most of the face of a person is outside the video frame, and only a small part of the face, such as only the chin of a person, is displayed. In this case, the video frame can be considered to be an invalid video frame that does not include a face.

[0090] Please continue to refer to Figure 7 In some embodiments, the method for determining the defective object video frame in step S604 can be implemented by the following steps:

[0091] Step S704: for any defective object video frame, when it is determined according to the target attribute that the area where the target object is located is a clear area, the defective object video frame is determined as a target quality defective object video frame.

[0092] Here, because in general, the blurred and fuzzy objects in the image are usually not the main body of the picture compared to the clear objects, the incomplete objects in the clear area can be determined as the target quality incomplete object video frame.

[0093] Step S705: for any defective object video frame, when it is determined according to the target attribute that the video content represented by the defective object video frame is related to the target object, the defective object video frame is determined as a target quality defective object video frame.

[0094] Here, since different screen expressions correspond to different subjects, for example, the images of eating broadcasts are mostly intended to show food and the eating process, so the eyes of people in the picture are not necessary subjects; for example, some kitchen cutting pictures are usually intended to show knife skills and ingredients, so it is acceptable for the facial features of the characters to be cropped. In this case, when the subject corresponding to the screen expression is different from the type of the target object, that is, the target object is not the subject corresponding to the screen expression of the incomplete object video frame, the video content represented by the incomplete object video frame is irrelevant to the target object; on the contrary, when the subject corresponding to the screen expression object is of the same type as the target object, that is, the target object is the subject corresponding to the screen expression of the incomplete object video frame, the video content represented by the incomplete object video frame is relevant to the target object.

[0095] Step S706, extract features from each defective object video frame to obtain an image feature vector of the defective object video frame and an object feature vector of the target object. Step S707, concatenate the image feature vector and the object feature vector to obtain a concatenation matrix. Step S708, perform attention calculation on the concatenation matrix based on the self-attention mechanism to obtain self-attention features. Step S709, determine whether the target object is the main content in the defective object video frame based on the self-attention features. Step S710, when the target object is the main content in the defective object video frame, determine the defective object video frame as a target quality defective object video frame.

[0096] Here, a self-attention mechanism is used to learn and determine whether a defective object video frame is a target quality defective object video frame.

[0097] Step S711: Count the number of the determined target quality defective object video frames as the number of video frames.

[0098] In some embodiments, a pre-trained video quality detection model may be used to detect the video to be detected. Figure 8 is a schematic diagram of the structure of the video quality detection model provided in the embodiment of the present application, such as Figure 8 As shown, the video quality detection model 800 includes a target detection network 801, a defective object recognition network 802 and a video detection network 803; wherein the target detection network 801 is used to perform image recognition on the video frames to be detected in the video to be detected, and obtain the target detection frame, target key points and target attributes of the target object in each video frame to be detected; the defective object recognition network 802 is used to determine whether the video frame to be detected is a target quality defective object video frame based on the target detection frame, target key points and target attributes; the video detection network 803 is used to obtain the detection result corresponding to the video to be detected according to the number of target quality defective object video frames in the video to be detected.

[0099] The following will be combined Figure 8 The structure of the video quality detection model shown illustrates the training method of the video quality detection model provided in the embodiment of the present application. Fig. 9 is a flow chart of a method for training a video quality detection model provided in an embodiment of the present application, such as Fig. 9 As shown, the method comprises the following steps:

[0100] Step S901, constructing a sample data set, wherein the sample data set includes at least two sample video frames, and the sample video frames are obtained by performing video frame extraction and online generation of incomplete samples on the sample video.

[0101] In some embodiments, constructing the sample data set in step S901 may be achieved by the following steps:

[0102] Step S9011, extract video frames from the sample video to obtain at least two sampled video frames.

[0103] Step S9011, determining the incomplete object sampling video frames in at least two sampling video frames, and deleting the incomplete object sampling video frames to obtain an updated video frame set.

[0104] Here, the defective object sampling video frame can also be determined by the following steps:

[0105] Step S91a, performing target detection on each sampled video frame to obtain a target detection result.

[0106] Step S91b: according to the target detection result, the video frames without the target object in at least two sampled video frames are deleted to obtain a sampled video frame set after the video frames are deleted.

[0107] Step S91c, performing key point prediction on the sampled video frames in the sampled video frame set to obtain a detection frame and key points of each sampled video frame.

[0108] Step S91d, obtaining the labeled detection frame, labeled key points and labeled attributes obtained after manual labeling based on the detection frame and key points for each sampled video frame.

[0109] Step S91e, determining the incomplete object sampling video frame in the sampling video frame set according to the marked detection frame, the marked key points and the marked attributes.

[0110] Step S9011, cropping the video frames in the updated video frame set according to a preset cropping ratio to generate incomplete samples. Here, the preset cropping ratio refers to the ratio of the number of video frames cropped in the video frame set, and the cropping ratio can be determined according to the requirements of the incomplete samples in the sample data set.

[0111] In the embodiment of the present application, the video frame can be cropped according to a preset cropping size to crop out some key points of the target object in the video frame, that is, the target object is cropped into a defective object. After generating the defective sample, the cropping size corresponding to each defective sample can also be saved, and the cropping size is used as the annotation data of the defective sample, and is mapped and stored together with the defective sample in the sample data set.

[0112] Step S9011, forming a sample data set according to the video frames and the incomplete samples in the updated video frame set. Here, the sample data set is formed by the video frames in the updated video frame set and the incomplete samples obtained by cutting.

[0113] Step S902: Input each sample video frame into a target detection network for image recognition to obtain a sample target detection frame, sample target key points, and sample target attributes of a sample target object in each sample video frame.

[0114] Step S903, through the incomplete object recognition network, according to the sample object detection frame, the sample object key points and the sample object attributes, it is determined whether the sample video frame is a target quality incomplete object video frame.

[0115] Step S904, obtaining a sample detection result corresponding to the sample video according to the number of target quality defective object video frames in the sample video through a video detection network.

[0116] Step S905, input the sample detection result into a preset loss model to obtain a loss result.

[0117] Here, the preset loss model includes a loss function, and the loss function is used to calculate the similarity between the sample detection result and the real result manually pre-labeled, and the loss result is obtained according to the similarity.

[0118] When the similarity between the sample detection result and the true result is greater than the similarity threshold, it indicates that the current video quality detection model can accurately predict the type of video to be detected, and accurately judge and identify the target quality incomplete object video. Therefore, the current video quality detection model can be stopped from being trained; when the similarity between the sample detection result and the true result is less than or equal to the similarity threshold, it indicates that the current video quality detection model cannot accurately predict the type of video to be detected, and cannot accurately judge and identify the target quality incomplete object video. Therefore, it is necessary to continue to train the current video quality detection model so that the video quality detection model can be more inclined to predict the accurate type of video to be detected.

[0119] In some embodiments, a training constraint condition may be pre-set, and the training constraint condition may be any one of the following: a training duration threshold, a training number threshold, and a training result similarity threshold. When the training constraint condition is a training duration threshold, the timing starts when the video quality detection model is trained, and when the training duration reaches the training duration threshold, the training of the video quality detection model is stopped; when the training constraint condition is a training number threshold, a counter may be pre-set, and the counter is cleared before the video quality detection model is trained, and after the training starts and each time a complete training process is completed, the counter is incremented by one, and when the current count of the counter reaches the training number threshold, the training of the video quality detection model is stopped; when the training constraint condition is a training result similarity threshold, when the similarity between the sample detection result obtained in the current training process and the real result reaches the training result similarity threshold, the training of the video quality detection model is stopped.

[0120] Step S906, according to the loss result, the parameters in the target detection network, the incomplete object recognition network and the video detection network are modified to obtain a trained video quality detection model.

[0121] The training method of the video quality detection model obtained by training in the embodiment of the present application can accurately identify and judge the type of the video to be detected, and accurately determine whether the video to be detected is a low-quality and incomplete object video, thereby improving the recall rate of low-quality and incomplete object videos through the trained video quality detection model.

[0122] In some embodiments, based on Figure 8 The structure of the video quality detection model shown in FIG. 5 , step S602 can also be implemented by the following steps:

[0123] Step S11, determine at least two predefined bounding boxes with different scale parameters. Step S12, extract features of the video frame to be detected through the feature extraction layer in the target detection network to obtain the video frame characteristic vector of the video frame to be detected. Step S13, extract features of the image corresponding to each predefined bounding box in the video frame to be detected through the feature extraction layer, and generate a bounding box characteristic vector corresponding to each predefined bounding box. Step S14, predict the target object in the video frame to be detected based on the video frame characteristic vector and the bounding box characteristic vector, and obtain the target detection box and target key points. Step S15, perform content recognition on the video frame to be detected to obtain target attributes.

[0124] In some embodiments, when training the target detection network, the method may further include the following steps:

[0125] Step S21 , performing partial area occlusion processing on the target object in the sample video frame to obtain a processed sample video frame.

[0126] Step S22, obtaining a predefined bounding box corresponding to the target object.

[0127] Step S23 , cropping the processed sample video frame according to the predefined border to obtain the occluded target object.

[0128] It should be noted that step S22 and step S23 can also be executed before step S21, that is, the target object can be cropped first and then the partial area occlusion processing can be performed, that is, the target object is first cropped to obtain the cropped target object, and then the cropped target object is partially occluded to obtain the occluded target object.

[0129] Step S24, performing key point recognition on the occluded target object to obtain at least one target key point of the target object.

[0130] In some embodiments, when generating a video to be detected, the video generator will add a few seconds of opening or ending credits to the main body of the video content to promote its own column. This part of the content is irrelevant to the main body of the video content. Therefore, the added opening and ending credits can be identified and cropped. The embodiment of the present application provides a method for identifying and cropping opening and ending credits. The method can be performed while extracting video frames of the video to be detected. That is, the above step S601 can be implemented by the following steps:

[0131] Step S31, performing title and ending detection on the video to be detected, and determining the title segment and ending segment in the video to be detected.

[0132] Here, any one of the methods for detecting the opening and ending credits may be adopted to implement the detection, wherein the method for detecting the opening and ending credits may be to determine whether the video content expressed by the video frames in the video to be detected is the same based on the video content expressed by the video frames in the video to be detected, and to determine a video clip located at the beginning of the video to be detected and having a different video content from that expressed in the middle part of the video to be detected as the opening clip, and to determine a video clip located at the end of the video to be detected and having a different video content from that expressed in the middle part of the video to be detected as the ending clip.

[0133] In an embodiment of the present application, when determining the video content expressed by a video frame, artificial intelligence technology can be used to identify the content in the video frame, thereby determining the video content expressed by the video frame.

[0134] Step S32, cropping the opening segment and the ending segment to obtain a cropped video.

[0135] Step S33, determining the number of video frames to be detected based on the length of the cropped video.

[0136] Step S34, according to the number of the extracted video frames to be detected, video frames are extracted from the cropped video by using an evenly spaced frame extraction method to obtain the number of video frames to be detected.

[0137] In the embodiment of the present application, since the opening and ending segments that are irrelevant to the main content of the video are cropped, it is possible to ensure accurate identification and judgment of the main content of the video, thereby accurately determining the type of video to be detected, and analyzing whether the video to be detected is a low-quality and incomplete object video; at the same time, since the opening and ending segments are cropped, the length of the entire video is reduced, the number of extracted video frames to be detected is reduced, thereby reducing the amount of data processing, or, without reducing the number of extracted video frames to be detected, since the extracted video frames to be detected are all from the main part of the video that can express the main content of the video, the detection accuracy of low-quality and incomplete object videos can be further improved, thereby greatly improving the recall rate of low-quality and incomplete object videos.

[0138] The following is an explanation of an exemplary application of the embodiments of the present application in a practical application scenario.

[0139] The embodiment of the present application provides a video quality detection method, which is a method for detecting low-quality incomplete faces in a picture, and can be used as one of the low-quality features of pictures and videos for content interception and filtering. First, the training samples are quickly annotated by using a public general face model to assist in manual annotation; then, the sample definition and sampling strategy design of the general face detection task are modified to achieve efficient detection of incomplete faces by the model, and key point detection of incomplete faces is achieved by data enhancement; then, for video frame pictures, the predicted face detection frame (i.e., the above-mentioned target detection frame), key points (i.e., the above-mentioned target key points), and face feature vectors (corresponding to the above-mentioned target attributes) are predicted by a self-attention model to determine whether the incomplete face is the "main content face" in the picture, and further determine whether a single picture has a low-quality problem of incomplete faces; finally, for videos, the post-strategy design is performed on multiple video frame pictures obtained by prediction to obtain the final video-level incomplete face attributes to determine whether a video has a problem of low-quality incomplete faces.

[0140] The video quality detection method provided in the embodiment of the present application can be applied to the standardization process of Penguin video. Each Penguin video and video cover will be detected by the low-quality incomplete face detection algorithm to determine whether the video has low-quality incomplete faces, and output the three fields of incomplete face area, incomplete face key points, and incomplete face prediction confidence to the video feature field of Penguin. The feature field will be passed along with all the information of the entire standardization process to the downstream business party (for example, Tencent Video, Kandian, QQ Space, Weishi, etc.) for use by the business party. The method of use includes using it as a recommendation sorting feature for the recommendation system, or as a machine-based quality mark for manual re-labeling. On the other hand, the low-quality incomplete face detection algorithm provided in the embodiment of the present application can also be put online to form a video pipeline plug-in, which can provide a unified input format, output format and request method. Any organization that has a demand for the algorithm in the video pipeline process can apply to call the video pipeline plug-in to make quality judgments on videos and pictures.

[0141] The video quality detection method provided in the embodiment of the present application is described below.

[0142] The low-quality incomplete face detection flow chart of a single image (i.e. a frame of video to be detected) is as follows: Fig.10 As shown in the figure, the dashed boxes are data nodes, the thin solid boxes are model or strategy nodes, and the thick solid boxes are improvements to the used model. Fig.10 For the input image 1001, the face detection model Ret inaFace 1002 is first used to optimize the incomplete face, and face detection is performed to obtain a face detection frame 1003 and an image-level face feature 1004. Here, when the face detection model RetinaFace 1002 is used to perform face detection, the incomplete face bounding box (bbox, bounding_box) definition, the incomplete face hyperframe retention and the incomplete face online generation are performed. After obtaining the face detection frame 1003, the face key point prediction model HRNet 1005 is used to perform face key point detection to obtain at least one key point 1006. Here, when performing face key point detection, cropping data enhancement processing is performed. After obtaining the image-level face feature 1004, the self-attention subject judgment model 1007 is used to determine whether the face is the subject of the image 1001. If it is to express a subject, after obtaining the key point 1006, a post-strategy 1008 is used to determine whether the image 1001 is a low-quality incomplete face image.

[0143] Below, we will explain the various components of the video quality detection model step by step.

[0144] 1) Construct a dataset.

[0145] Since the entire video (i.e. the video to be tested) is usually of huge capacity, and the content between adjacent frames in the video is highly similar and redundant, it is inappropriate to use all the frames of the entire video as training data. It is necessary to sparsely extract a certain number of key frames from the video to cover the main pictures and main scenes of the video. The detailed strategy for extracting video frames is as follows:

[0146] The videos in common video software will have a large number of self-media contributions, and some self-media will add a few seconds of opening or ending credits to the main body of the video to promote their own columns. This part of the content is irrelevant to the main body of the video content. The embodiment of the present application adopts a head and ending credit detection algorithm to detect and remove the head and ending parts, and extract video frames in the main part of the video.

[0147] Since we hope that the video frames extracted from the video have a certain degree of representativeness and avoid high similarity, after removing the opening and ending credits, we also use the method of evenly extracting frames at equal intervals to extract video frames. For example, for videos with a main body of more than 60 seconds, 5 frames are evenly extracted (with a minimum interval of 12 seconds); for videos with a main body of more than 30 seconds but not more than 60 seconds, 3 frames are evenly extracted (with a minimum interval of 15 seconds); for videos with a main body of less than 30 seconds, only 1 frame is randomly extracted. Fig.11 As shown in Figures a to d in the figure, some video frames are extracted from the same video using the video frame extraction method provided in the embodiment of the present application.

[0148] For the data after frame extraction, the public face detection model RetinaFace, which is more accurate for predicting regular faces, can be used to predict whether there is a face in the image and the specific area of ​​the face. Because there are a lot of background areas in the detection task, their sampling is enough to serve as negative samples of the model, so all samples without faces can be deleted to reduce the redundancy of training data. After removing the data without faces, the public regular face key point prediction model HRNet is used to predict the key points of all faces, so that a set of detection boxes and key points are obtained for the faces in the samples.

[0149] The following is the start of manual labeling, which requires labeling of three dimensions. For each face, its detection box, key points, and whether it is the subject of the picture expression need to be labeled. If labeling is performed without prompts, the labeling speed is only 200 images / day. Since the embodiment of the present application provides the above-mentioned preset results for key point prediction of all faces, the annotator can make adjustments based on the preset results, such as deleting erroneous detection results, adding missed detection results, and fine-tuning the coordinates of the detection box and key points. If the preset results are accurate enough, no adjustment is required. In this way, the manual labeling speed can be increased to 2000 images / day.

[0150] After manual labeling is completed, the images with incomplete faces in the sample are filtered out by coordinates, and only the images with complete faces are retained. Here, the goal is to detect incomplete faces, but this part of the data is filtered out in the sample because it is impossible to determine where the actual face detection frame boundary is for faces that have been cropped to incompleteness, which will interfere with the task. Therefore, this application will use online automatic generation to input incomplete face samples, which will be explained in the model sampling section below.

[0151] 2) Feature extraction network.

[0152] For face detection tasks, the overall architecture of RetinaFace can be used, and for face key point detection tasks, the overall architecture of HRNet can be used. RetinaFace is an effective conventional face detection task, which consists of a forward network (usually 5 times downsampling), a feature pyramid network (FPN), and a single-stage headless face detection (SSH) multi-scale module. According to the configuration, predefined bounding boxes (anchors) of different stride scales can be selected. Since this task mainly detects the face of the subject in the picture, only three stride scales of 8, 16, and 32 can be used, focusing on the face of the size in the picture. According to the model structure of the feature extraction network, the feature extraction network can generate a vector feature for the image, recorded as img_feature (that is, the video frame feature vector mentioned above). The feature extraction network can also generate a vector feature for each anchor, which is combined into a list such as [anchor1_feature, anchor2_feature, ...] (that is, the border feature vector mentioned above). However, most anchor areas are background negative samples and are not responsible for predicting the border of the face. Therefore, we only need to care about the anchor features that ultimately predict the face, which can be recorded as [face1_feature, face2_feature, ...].

[0153] 3) Sampling and matching design for incomplete face detection.

[0154] For the face detection model, the model learning method is not to directly learn the four manually marked detection frame coordinates through a unique image feature. Because there may be multiple objects and multiple detection frames in a picture, it is impossible to learn multiple different objects through a unified feature. Therefore, for the face detection model, generating multiple local features and how to determine which learning target the local features are responsible for is called sampling matching, which is one of the core designs of the face detection model. In the design of the face detection model, in order to detect incomplete faces, this application does not modify the network structure of the face detection model, but implements this function by modifying the design of sampling matching. For clarity, the following will first describe the conventional face detection process from input to sampling matching, and then describe how to generate incomplete samples online based on complete faces, appropriately match and learn incomplete faces in the embodiments of the present application.

[0155] In conventional face detection, the model side will generate local features of different sizes. These features are called anchors. These features are used as virtual anchor boxes to match the real detection box bounding_box (bbox). The matching of anchor and bbox may have two results: if the anchor does not match the bbox, then the anchor is responsible for predicting that this is a negative sample; if the anchor matches the bbox, then the anchor is responsible for predicting that this is a positive sample, and is also responsible for predicting its own coordinate offset relative to the bbox. How to judge whether the anchor is responsible for a certain bbox? It will be judged based on the intersection over union (IoU) indicator. IoU refers to the ratio of the intersection area of ​​two areas to the union area. For example Fig.12 As shown, it is a schematic diagram of the IoU provided by the embodiment of the present application. Given two prediction ranges BB1 and BB2, the intersection of BB1 and BB2 is represented as BB1∩BB2, which is defined as the overlapping area 121 of BB1 and BB2; BB1∪BB2 is defined as the unified area 122 of BB1 and BB2. The intersection over union (IoU) is Fig.12 In the expression, Fig.12 The ratio between the overlapping area 121 and the uniform area 122 of the medium dark area.

[0156] Usually, if the IoU between the anchor and the bbox is greater than 0.5, the anchor is responsible for this bbox. If the IoU between the anchor and all bboxes is less than 0.3, it is not responsible for any bbox. If it is between the two, the anchor will be ignored and not learned. In addition, because the model does not need to detect faces beyond the border, when the bbox exceeds a certain proportion of the border (such as 30%, 50%), the bbox will be ignored.

[0157] Design of bbox definition for incomplete samples: According to the matching rules, first determine how to design the bbox for the incomplete face. If a incomplete face appears in the picture, should the bbox be defined close to the edge where the face is cropped, or should it be defined beyond the edge as a rectangle with negative coordinates? The latter is chosen in this embodiment of the application, because the former will cause a conflict in matching definitions. Fig.13 is a schematic diagram of the design principle of the incomplete sample bbox definition provided in the embodiment of the present application, such as Fig.13 As shown, the solid-line frame 1301 in Figure a is a bbox defined by the border. Assuming that there is a complete and identical face inside the picture (solid-line frame 1302 in Figure b), the two dotted-line frames 1303 and 1304 cover similar content in the anchors. However, since the area of ​​the bbox (i.e., the solid-line frame 1301) that is attached to the border is significantly smaller than the complete face bbox (i.e., the solid-line frame 1302), the two anchors with similar content (i.e., the two dotted-line frames 1303 and 1304) are one positive sample and the other negative sample, causing learning ambiguity. Therefore, the embodiment of the present application adopts the latter definition, that is, the bbox is defined as the actual face size. Even if the face has exceeded the border, there will be a "supposed" and "imagined" face boundary.

[0158] For conventional face detection models, it is not necessary to detect faces beyond the border, and the bbox will be ignored when it exceeds the border by a certain proportion. However, for the goal of incomplete face detection, the embodiment of the present application requires faces that are beyond the border, so the bbox that is beyond the border cannot be ignored. On the other hand, if in extreme cases, a face is 100% not inside the picture, it is impossible to detect the face anyway. In other words, even for the detection of incomplete faces, it is necessary to retain a part of obvious facial features inside the picture, and it is necessary to set an upper limit on the proportion that exceeds the edge of the picture. However, the facial structure and facial features of different people are not consistent, and it is impossible to use a fixed indicator to judge whether the face "still has obvious features". For this reason, the embodiment of the present application uses pre-annotated key point information. When the face cropped above retains lips in the picture, the face cropped below retains eyes in the picture, and the face cropped on the left and right retains left and right eyes and corners of the mouth in the picture, the sample cannot be ignored, otherwise it is considered that the face has lost reasonable features that can be detected. Fig.14 The face shown cannot be detected by the face detection model because the face is too incomplete and has only a few features.

[0159] The following will explain why all the image samples containing incomplete faces are discarded in the above embodiment, because the bbox definition of the incomplete face must find the face boundary that is beyond the frame. However, for the incomplete face, it is impossible to know where its real boundary is. Therefore, all the labeled complete faces are used to automatically generate incomplete faces online during the training process, which meets both the input of incomplete faces and the input of complete face bbox. The specific operation steps are as follows: the first step is to randomly select one of the faces in the picture with faces; the second step is to control whether the face needs to be deliberately cropped this time through a crop rate (crop_rate), because it is hoped that the model will not only detect the incomplete face, but also retain the ability to recall the normal face, so the crop_rate will be used to balance the input ratio of the two; the third step is that if cropping is not required, the sample generation process is redeemed, if cropping is required, a cropping direction is selected, and the face is cropped to exceed the frame, but still meet the requirements of the incomplete face proposed in the above embodiment; the fourth step is to adjust the coordinates of the cropped bbox according to the direction and degree of cropping, and complete the generation of the incomplete face sample. The above steps are all performed online during the training process, because even for the same picture, sufficient and diverse training samples can be generated through random parameters.

[0160] Through the above steps, the face detection model can receive incomplete face samples that were originally unable to learn, and at the same time, there is no conflict with the learning objectives of the old regular face samples. The face detection model can learn the ability to detect incomplete faces.

[0161] 4) Face key point detection model. The face key point detection model can use HRNet to learn face key points. HRNet accepts an image and a face bbox, cuts the face out of the image according to the bbox, and inputs the face key point detection model to obtain face key points. The usual face key point detection model does not have a good detection capability for incomplete faces. In this case, data enhancement can be used to randomly black out a part of the face area, such as Fig.15 As shown in FIG. 1 , the eye part 1501 of the face in FIG. a is blackened to obtain FIG. b, and the eye part 1502 of the face in FIG. c is blackened to obtain FIG. d, and then the facial key points 1503 are detected. After blackening some areas, the learning objective remains unchanged, simulating the situation when the face is incomplete. It can be seen from the training that the facial key point detection model has a good generalization ability for limited blackening.

[0162] 5) Determine whether the incomplete face is the main subject of the picture.

[0163] After performing incomplete face detection through the above-mentioned face detection module and face key point detection module, the prediction results of the two models can determine whether a predicted face is an incomplete face. However, according to the definition and actual situation, only the incomplete face that belongs to the main body of the picture is a low-quality incomplete face. The following will use model design to determine whether the incomplete face is the facial attribute of the main body of the picture.

[0164] It can be observed that whether a mutilated face is the main subject of the picture can be reflected mainly through two features. On the one hand, it is the attributes of the face itself, and on the other hand, it is the relationship with the overall picture. The attributes of the face itself, that is, the attributes of the face area. Generally, blurred and fuzzy faces are usually not the main body of the picture compared to clear faces; the relationship with the overall picture mainly involves the meaning of the picture. For example, the pictures of eating broadcasts are mostly intended to show food and the eating process, so parts such as people’s eyes are not necessary subjects; for example, some kitchen cutting pictures are usually intended to show knife skills and ingredients, so it is acceptable for the facial features of the characters to be cropped. The above detailed business judgments have been annotated by the review and annotation team through the dimension of "whether it is the main body of the picture" when annotating.

[0165] Since the conclusion of whether the incomplete face is the main subject of the picture is related to its own attributes, the overall meaning of the picture, and possibly the attributes of other faces in the picture, simply reclassifying the vector of the face itself will not have a good effect. In order to solve this problem, the embodiment of the present application adopts a learning method of self-attention mechanism.

[0166] According to the above description, in addition to outputting the predicted detection frame, the face detection model can also output the image feature img_feature and the predicted facial features [face1_feature, face2_feature, ...]. Here, these two features are concatenated into a matrix. Assuming that there are K faces detected in the image and the feature dimension is F, the dimension of the feature matrix M is (K+1)×F. Assuming that p, q, and v are all learnable linear layers, the calculation method of the self-attention feature A is A=softmax(M(p)·M′(q)·M(v)). Among them, "·" represents the inner product; M′ represents the transpose of the M matrix. Under this calculation method, the dimension of A is still (K+1)×F. Under the self-attention feature, the image features and each face feature are associated and fused in pairs. After softmax normalization, they are used as weights to re-weight the face features. In this way, the learning process not only retains the characteristics of the face itself, but also retains the relationship between the face and the image, and the face and other faces.

[0167] After obtaining the self-attention feature A, add a F×2 linear layer at the end to learn whether a face is the main content of an image or not in these two dimensions. It should be noted that the self-attention feature A is of (K+1)×F dimensions. In addition to the K individual face features, there is also a global feature of the image. In learning, the output of the global feature of the image is not supervised and is ignored.

[0168] 6) Image prediction and video post-processing strategies.

[0169] After the face detection model and face key point detection model are trained, input a picture and you can get the following information: detection frames of all faces in the picture, confidence of detection frames, face key points, and probability of being the main content. It is very simple to determine whether a face is incomplete. As long as the predicted detection frame has negative coordinates and the core area of ​​the face key points exceeds the frame (such as eyebrows exceeding the frame, etc.), plus the condition that the probability of the main content is greater than 0.5, you can get an incomplete face that is "the main content", that is, a low-quality incomplete face.

[0170] After predicting the image, further post-strategy operations can be performed on the video. When predicting, a one-frame-per-second frame extraction method can be used. The frequency of one frame per second is not too sparse, and it can also cover most of the video content. There may be temporary transitions in the video. In a short moment, the subject's face moves very briefly beyond the border, becoming a low-quality incomplete face in the sense of a picture. The one-frame-per-second frame extraction method will also reduce the probability of extracting such frames, but it cannot be completely avoided. Therefore, after experimental adjustment, the embodiment of the present application adopts the form of cumulative video frames for video recognition. For videos of less than one minute, a cumulative of 3 frames of low-quality incomplete faces are required; for videos of more than one minute, a cumulative of 5 frames of low-quality incomplete faces are required. Videos that meet the cumulative low-quality incomplete face frames are low-quality incomplete face videos.

[0171] The video quality detection method provided in the embodiment of the present application can greatly reduce the annotation cost and obtain sufficient data for training by predicting the face frame and face key points in the picture and using them as a preset annotation method; by optimizing the face detection model and the face key point detection model, the recall of the face detection model for incomplete faces is improved, so that the conventional face detection model can better predict the area containing incomplete faces, and the conventional face key point model can accurately predict the key points of incomplete faces. The following Table 1 is a comparison of data indicators of different models. As shown in Table 1, the recall of incomplete faces of the optimized face detection model provided in the embodiment of the present application can reach 92%, and the normalized mean square error (NME) of the optimized face key point model is reduced to 7.84. It should be noted that the smaller the NME error, the better.

[0172] Table 1 Comparison of data indicators of different models

[0173] Model branching Data indicators General face detection model Incomplete face recall: 54% Optimized face detection model Incomplete face recall: 92% Through the facial key point model Incomplete face NME: 11.71 Optimized facial key point model Incomplete face NME: 7.84

[0174] In the embodiment of the present application, the subject judgment model (i.e., the video quality detection model) can be used to screen out the faces that are incomplete but not low-quality incomplete. This can meet the standards of actual business and human perception and reduce the false recall of the model. Since this video quality detection model is used for direct interception of low-quality videos, the accuracy rate will be prioritized in its use. The following Table 2 shows the corresponding indicators of each model combination under the premise of ensuring the accuracy rate as much as possible.

[0175] Table 2 Corresponding indicators under each model combination

[0176]

[0177] As shown in Table 2, the solution of the embodiment of the present application was tested on offline samples. A large-scale video cover test of 50,000 video applications was taken, and the recall sample precision was 91%, and the recall rate was 73%; a large-scale video content test of 5,000 video applications was taken, and the recall sample precision was 93%, and the recall rate was 86%, which is obviously higher than the conventional general method.

[0178] The solution of the embodiment of the present application provides low-quality and incomplete face detection functions for images and videos for the online standardization process of video applications from scratch. Under the current video call volume, approximately 400 additional cover images with such low-quality problems and 1,000 additional videos with such low-quality problems can be detected per day compared to not using this function.

[0179] It should be noted that, in some embodiments, the face detection model may not be limited to the RetinaFace network, for example, it may also be replaced by a face detection model such as ASFD. The face key point model is not limited to the HRNet network, for example, it may also be replaced by a face key point detection model such as LAB.

[0180] The following is a description of an exemplary structure of the video quality detection device 354 provided in the embodiment of the present application implemented as a software module. In some embodiments, for example, Figure 5 As shown, the video quality detection device 354 includes:

[0181] The video frame extraction module 3541 is used to extract video frames from the video to be detected to obtain at least two frames of video frames to be detected; the image recognition module 3542 is used to perform image recognition on each of the video frames to be detected to obtain the target detection frame, target key points and target attributes of the target object in each of the video frames to be detected; the first determination module 3543 is used to determine the incomplete object video frame in at least two frames of video frames to be detected based on the target detection frame and the target key points, and the incomplete object video frame is a video frame in which the target object is a incomplete object; the second determination module 3544 is used to determine the target quality incomplete object video frame and the video frame number of the target quality incomplete object video frame in the incomplete object video frame according to the target attributes of each of the incomplete object video frames; the third determination module 3545 is used to determine the detection result of the video to be detected based on the number of video frames.

[0182] In some embodiments, there are multiple target key points of the target object in each of the video frames to be detected; the first determination module is also used for: for any video frame to be detected, when it is determined that the target detection frame of any video frame to be detected has negative coordinates, and at least one of the target key points of any video frame to be detected exceeds the border of the video frame to be detected, and the number of the target key points exceeding the border is less than a preset number threshold.

[0183] In some embodiments, the second determination module is further used for: for any incomplete object video frame, when it is determined according to the target attribute that the area where the target object is located is a clear area, determining any incomplete object video frame as the target quality incomplete object video frame; or, for any incomplete object video frame, when it is determined according to the target attribute that the video content represented by any incomplete object video frame is related to the target object, determining any incomplete object video frame as the target quality incomplete object video frame; and counting the number of incomplete object video frames of the target quality determined as the number of video frames.

[0184] In some embodiments, the second determination module is further used to: perform feature extraction on each of the incomplete object video frames to obtain the image feature vector of the incomplete object video frame and the object feature vector of the target object; splice the image feature vector and the object feature vector to obtain a splicing matrix; perform attention calculation on the splicing matrix based on the self-attention mechanism to obtain self-attention features; determine whether the target object is the main content of the incomplete object video frame according to the self-attention features; when the target object is the main content of the incomplete object video frame, determine the incomplete object video frame as the target quality incomplete object video frame; and count the number of target quality incomplete object video frames determined as the number of video frames.

[0185] In some embodiments, the device further comprises: a processing module for detecting the video to be detected using a video quality detection model; the video quality detection model comprises a target detection network, a defective object recognition network and a video detection network; wherein the video quality detection model is trained by the following steps: constructing a sample data set, the sample data set comprising at least two sample video frames, the sample video frames being obtained by extracting video frames from the sample video and generating defective samples online; inputting each of the sample video frames into the target detection network for image recognition, and obtaining a sample target detection frame, a sample target detection frame of the sample target object in each of the sample video frames, and a sample target detection frame of the sample target object in the sample video frames. The target key points and sample target attributes of the sample target are obtained; through the incomplete object recognition network, according to the sample target detection box, the sample target key points and the sample target attributes, determining whether the sample video frame is a target quality incomplete object video frame; through the video detection network, according to the number of target quality incomplete object video frames in the sample video, obtaining the sample detection result corresponding to the sample video; inputting the sample detection result into a preset loss model to obtain a loss result; according to the loss result, correcting the parameters in the target detection network, the incomplete object recognition network and the video detection network to obtain a trained video quality detection model.

[0186] In some embodiments, the video quality detection model is trained through the following steps: extracting video frames from the sample video to obtain at least two sampled video frames; determining the incomplete object sampled video frames in the at least two sampled video frames, and deleting the incomplete object sampled video frames to obtain an updated video frame set; performing image cropping on the video frames in the updated video frame set according to a preset cropping rate to generate incomplete samples; and forming the sample data set based on the video frames in the updated video frame set and the incomplete samples.

[0187] In some embodiments, the video quality detection model is trained through the following steps: performing target detection on each of the sampled video frames to obtain a target detection result; based on the target detection result, deleting the video frames that do not have the target object in the at least two sampled video frames to obtain a set of sampled video frames after the video frames are deleted; performing key point prediction on the sampled video frames in the set of sampled video frames to obtain a detection frame and key points of each sampled video frame; obtaining a labeled detection frame, labeled key points and labeled attributes obtained after manual labeling of each sampled video frame based on the detection frame and the key points; and determining the incomplete object sampled video frame in the set of sampled video frames based on the labeled detection frame, the labeled key points and the labeled attributes.

[0188] In some embodiments, the video frame extraction module is also used to: perform header and tail detection on the video to be detected, and determine the header segment and tail segment in the video to be detected; crop the header segment and the tail segment to obtain a cropped video; determine the number of extracted video frames to be detected based on the length of the cropped video; and extract video frames from the cropped video according to the number of extracted video frames to be detected, using an equidistant and uniform frame extraction method, to obtain the number of video frames to be detected.

[0189] In some embodiments, the image recognition module is also used to: determine at least two predefined bounding boxes with different scale parameters; perform feature extraction on the video frame to be detected through the feature extraction layer in the target detection network to obtain a video frame characteristic vector of the video frame to be detected; perform feature extraction on the image corresponding to each of the predefined bounding boxes in the video frame to be detected through the feature extraction layer to generate a bounding box characteristic vector corresponding to each of the predefined bounding boxes; predict the target object in the video frame to be detected based on the video frame characteristic vector and the bounding box characteristic vector to obtain the target detection box and the target key points; perform content recognition on the video frame to be detected to obtain the target attributes.

[0190] In some embodiments, when training the target detection network, the device also includes: an occlusion processing module, which is used to perform partial area occlusion processing on the target object in the sample video frame to obtain a processed sample video frame; a predefined border acquisition module, which is used to obtain the predefined border corresponding to the target object; a cropping module, which is used to crop the processed sample video frame according to the predefined border to obtain the occluded target object; a key point recognition module, which is used to perform key point recognition on the occluded target object to obtain at least one target key point of the target object.

[0191] In some embodiments, the target object is a human face, and the image recognition module is further used to: perform face recognition on each of the video frames to be detected to obtain a face video frame with a face; perform image recognition on each of the face video frames to obtain a face detection frame, face key points and face attributes in each of the face video frames.

[0192] In some embodiments, the third determination module is also used to: determine the video length of the video to be detected; determine a threshold value of the number of video frames for evaluating the video to be detected based on the video length; when the number of video frames is greater than the threshold value of the number of video frames, determine that the detection result is that the video to be detected is a target quality incomplete object video.

[0193] It should be noted that the description of the device of the embodiment of the present application is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment, so it is not repeated. For technical details not disclosed in the embodiment of the device, please refer to the description of the method embodiment of the present application for understanding.

[0194] The embodiment of the present application provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned method of the embodiment of the present application.

[0195] The present application embodiment provides a storage medium storing executable instructions, wherein the executable instructions are stored. When the executable instructions are executed by a processor, the processor will execute the method provided by the present application embodiment, for example, Figure 6 The method shown.

[0196] In some embodiments, the storage medium can be a computer-readable storage medium, for example, a ferroelectric random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disk, or a compact disk read-only memory (CD-ROM), etc.; it can also be various devices including one or any combination of the above memories.

[0197] In some embodiments, executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.

[0198] As an example, executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file storing other programs or data, such as one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions). As an example, executable instructions may be deployed to be executed on one computing device, or on multiple computing devices located at one location, or on multiple computing devices distributed at multiple locations and interconnected by a communication network.

[0199] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. A video quality detection method, It is characterized in that The method comprises: Extract video frames from the video to be detected to obtain at least two video frames to be detected; Performing image recognition on each of the video frames to be detected to obtain a target detection frame, target key points and target attributes of the target object in each of the video frames to be detected; the target attributes include the clarity of the target object or the relevance of the target object to the content to be expressed by the entire video to be detected; Determine, according to the target detection frame and the target key point, a defective object video frame in at least two frames of video frames to be detected, wherein the defective object video frame is a video frame in which the target object is a defective object; According to the target attribute of each of the defective object video frames, determining target quality defective object video frames and the number of video frames of the target quality defective object video frames in the defective object video frames; According to the number of video frames, a detection result of the video to be detected is determined.

2. The method according to claim 1, It is characterized in that There are multiple target key points of the target object in each of the to-be-detected video frames; The step of determining a video frame of a defective object in at least two video frames to be detected according to the target detection frame and the target key point comprises: For any video frame to be detected, when it is determined that the target detection frame of any video frame to be detected has negative coordinates, and at least one target key point of any video frame to be detected exceeds the border of the video frame to be detected, and the number of the target key points exceeding the border is less than the preset number threshold, Determine that any of the to-be-detected video frames is the incomplete object video frame.

3. The method according to claim 1, It is characterized in that The step of determining target quality incomplete object video frames and the number of target quality incomplete object video frames in the incomplete object video frames according to the target attribute of each incomplete object video frame comprises: For any defective object video frame, when it is determined according to the target attribute that the area where the target object is located is a clear area, the any defective object video frame is determined as the target quality defective object video frame; or For any defective object video frame, when it is determined according to the target attribute that the video content represented by the any defective object video frame is related to the target object, the any defective object video frame is determined as the target quality defective object video frame; The number of the target quality defective object video frames determined by statistics is the number of video frames.

4. The method according to claim 1, It is characterized in that The step of determining the target quality incomplete object video frame and the number of the target quality incomplete object video frame from the incomplete object video frame comprises: Performing feature extraction on each of the incomplete object video frames to obtain an image feature vector of the incomplete object video frame and an object feature vector of the target object; splicing the image feature vector and the object feature vector to obtain a splicing matrix; Based on the self-attention mechanism, attention calculation is performed on the splicing matrix to obtain a self-attention feature; Determining whether the target object is the main content in the incomplete object video frame according to the self-attention feature; When the target object is the main content in the defective object video frame, determining the defective object video frame as the target quality defective object video frame; The number of the target quality defective object video frames determined by statistics is the number of video frames.

5. The method according to claim 1, It is characterized in that The method further comprises: using a video quality detection model to detect the video to be detected; the video quality detection model comprises a target detection network, a defective object recognition network and a video detection network; The video quality detection model is trained by the following steps: Constructing a sample data set, wherein the sample data set includes at least two sample video frames, wherein the sample video frames are obtained by performing video frame extraction and online generation of incomplete samples on the sample video; Inputting each of the sample video frames into the target detection network for image recognition, and obtaining a sample target detection frame, a sample target key point, and a sample target attribute of a sample target object in each of the sample video frames; Determining, by the incomplete object recognition network, whether the sample video frame is a target quality incomplete object video frame according to the sample object detection frame, the sample object key points and the sample object attributes; Obtaining, by the video detection network, a sample detection result corresponding to the sample video according to the number of target quality defective object video frames in the sample video; Inputting the sample detection results into a preset loss model to obtain a loss result; According to the loss result, the parameters in the target detection network, the incomplete object recognition network and the video detection network are modified to obtain a trained video quality detection model.

6. The method according to claim 5, It is characterized in that The constructing of the sample data set comprises: Extracting video frames from the sample video to obtain at least two sampled video frames; Determine a defective object sampled video frame among the at least two sampled video frames, and delete the defective object sampled video frame to obtain an updated video frame set; Performing image cropping on the video frames in the updated video frame set according to a preset cropping ratio to generate incomplete samples; The sample data set is formed according to the video frames in the updated video frame set and the incomplete samples.

7. The method according to claim 6, It is characterized in that The step of determining the defective object sampled video frame in the at least two sampled video frames comprises: Performing target detection on each of the sampled video frames to obtain a target detection result; According to the target detection result, the video frames without the target object in the at least two sampled video frames are deleted to obtain a sampled video frame set after the video frames are deleted; Performing key point prediction on the sampled video frames in the sampled video frame set to obtain a detection frame and key points of each sampled video frame; Obtaining a labeled detection frame, a labeled key point, and a labeled attribute obtained after manually labeling each sampled video frame based on the detection frame and the key point; The incomplete object sampling video frame is determined in the sampling video frame set according to the labeled detection frame, the labeled key points and the labeled attributes.

8. The method according to claim 1, It is characterized in that The step of extracting video frames from the video to be detected to obtain at least two video frames to be detected includes: Performing title and ending detection on the video to be detected to determine the title segment and the ending segment in the video to be detected; Cropping the opening segment and the ending segment to obtain a cropped video; Determining the number of video frames to be detected according to the duration of the cropped video; According to the number of the extracted video frames to be detected, video frames are extracted from the cropped video in an evenly spaced frame extraction manner to obtain the number of video frames to be detected.

9. The method according to claim 5, It is characterized in that The performing image recognition on each of the video frames to be detected to obtain a target detection frame, target key points and target attributes of the target object in each of the video frames to be detected includes: determining at least two predefined bounding boxes having different scale parameters; Performing feature extraction on the to-be-detected video frame through a feature extraction layer in the target detection network to obtain a video frame characteristic vector of the to-be-detected video frame; Performing feature extraction on the image corresponding to each of the predefined borders in the video frame to be detected through the feature extraction layer, and generating a border feature vector corresponding to each of the predefined borders; Predicting the target object in the to-be-detected video frame according to the video frame characteristic vector and the frame characteristic vector to obtain the target detection frame and the target key points; Content recognition is performed on the video frame to be detected to obtain the target attribute.

10. The method according to claim 9, It is characterized in that When training the target detection network, the method further includes: Performing partial area occlusion processing on the target object in the sample video frame to obtain a processed sample video frame; Obtaining the predefined border corresponding to the target object; According to the predefined border, cropping the processed sample video frame to obtain the occluded target object; Key point recognition is performed on the blocked target object to obtain at least one target key point of the target object.

11. The method according to any one of claims 1 to 10, It is characterized in that The target object is a human face, and image recognition is performed on each of the video frames to be detected to obtain a target detection frame, target key points and target attributes of the target object in each of the video frames to be detected, including: Performing face recognition on each of the to-be-detected video frames to obtain a face video frame having a face; The image recognition is performed on each of the face video frames to obtain a face detection frame, face key points and face attributes in each of the face video frames.

12. The method according to any one of claims 1 to 10, It is characterized in that Determining the detection result of the video to be detected according to the number of video frames includes: Determine the video length of the video to be detected; Determining a video frame quantity threshold for evaluating the video to be detected according to the video duration; When the number of video frames is greater than the video frame number threshold, it is determined that the detection result is that the video to be detected is a target quality incomplete object video.

13. A video quality detection device, It is characterized in that The device comprises: A video frame extraction module, used to extract video frames from the video to be detected, to obtain at least two video frames to be detected; An image recognition module is used to perform image recognition on each of the video frames to be detected, and obtain a target detection frame, target key points and target attributes of the target object in each of the video frames to be detected; the target attributes include the clarity of the target object or the relevance of the target object to the content to be expressed by the entire video to be detected; A first determination module is used to determine a defective object video frame in at least two frames of video frames to be detected according to the target detection frame and the target key point, wherein the defective object video frame is a video frame in which the target object is a defective object; A second determining module is used to determine, according to the target attribute of each of the defective object video frames, target quality defective object video frames and the number of video frames of the target quality defective object video frames in the defective object video frames; The third determination module is used to determine the detection result of the video to be detected according to the number of video frames.

14. A video quality detection device, It is characterized in that include: A memory for storing executable instructions; The processor is used to implement the video quality detection method according to any one of claims 1 to 12 when executing the executable instructions stored in the memory.

15. A computer-readable storage medium, It is characterized in that Executable instructions are stored, which are used to cause a processor to execute the executable instructions to implement the video quality detection method according to any one of claims 1 to 12.

16. A computer program product, It is characterized in that The computer program product includes executable instructions stored in a computer-readable storage medium; Wherein, when the processor of the video quality detection device reads the executable instructions from the computer-readable storage medium and executes the executable instructions, the video quality detection method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Object key point positioning method and device, image processing method and device and storage medium

    CN109684920A

  • Identification photo shooting method, apparatus and device, and storage medium

    CN110602379A