Video uploading verification method, device and equipment and storage medium thereof

By performing frame processing and face detection on uploaded videos, generating verification picture sets and inputting pre-trained models for detection, the accuracy problem caused by single detection features in the prior art is solved, and more accurate video authenticity recognition is achieved.

CN120126224APending Publication Date: 2025-06-10PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510192318.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing fake video detection technology has the problem that the detection characteristics are relatively single, resulting in inaccurate detection results.

Method used

By obtaining the video to be uploaded, performing frame processing and face detection, generating verification picture sets, and inputting them into the pre-trained face detection model, obtaining detection classification results, determining the authenticity of the video, and determining whether to perform upload operations.

Benefits of technology

By combining face detection technology, the authenticity detection model of video is trained, and the changes in face features and feature changes are made full use of the authenticity of the video, the accuracy of the detection results is improved, and the forged video is identified before uploading the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126224A_ABST
    Figure CN120126224A_ABST
Patent Text Reader

Abstract

The embodiment of the invention belongs to the technical field of image verification, is applied to a video authenticity identification scene, and relates to a video uploading verification method and device, equipment and a storage medium thereof. Framing processing is carried out; face detection is carried out on the pictures after framing processing; only the pictures containing the human faces are reserved, and a verification picture set is generated; acquiring a detection classification result output by the face detection model for the verification picture set; based on a detection classification result, determining a verification category of the to-be-uploaded video; and according to the verification category, judging whether to execute an uploading operation. A video authenticity detection model is trained through a face detection technology, face features are fully utilized, and the authenticity of a video to be uploaded by a user is predicted according to feature changes of the face features in the whole video. When the method is applied to a face recognition scene under a financial service, forged video recognition can be automatically and intelligently carried out, the workload of manual recognition is reduced, and the related service efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image verification technology and is applied to the scenario of identifying the authenticity of videos. In particular, it relates to a video upload verification method, device, equipment, and its storage medium. Background Art

[0002] The progress of smartphone cameras and the popularity of the Internet have made the production and dissemination of digital videos unprecedentedly convenient. At the same time, the rapid development of deep learning technology, especially the emergence of generative adversarial networks (GANs), has enabled people to generate highly realistic fake videos.

[0003] Especially in the field of financial-related businesses, such as bank loan services or claim video request services, highly realistic fake videos often pose greater recognition challenges to face recognition in related businesses, meaning that enterprises will bear higher business risks. To address this challenge, deepfake detection technology has become crucial. Past methods for detecting deepfake videos often suffered from problems such as single detection features, strong dependence on datasets, insufficient model robustness, and lack of comprehensiveness. For example, many methods relied only on a single feature for detection, such as only focusing on simple features like blinking, opening the mouth, and turning the head. Therefore, existing fake video detection technologies still have the problem of relatively single detection features, resulting in inaccurate detection results. Summary of the Invention

[0004] The purpose of the embodiments of this application is to propose a video upload verification method, device, equipment, and its storage medium to solve the problem that existing fake video detection technologies still have relatively single detection features, resulting in inaccurate detection results.

[0005] To solve the above technical problems, the embodiments of this application provide a video upload verification method, which adopts the following technical solutions:

[0006] A video upload verification method includes the following steps:

[0007] Obtain the video to be uploaded;

[0008] Perform frame splitting on the video to be uploaded to obtain the pictures after frame splitting;

[0009] Perform face detection on the pictures after frame splitting respectively to obtain face detection results;

[0010] According to the face detection results, only retain the pictures containing faces to generate a verification picture set;

[0011] Input the verification picture set into a pre-trained face detection model to obtain the detection classification results output by the face detection model for the verification picture set;

[0012] Based on the detection and classification results, determine the verification category of the video to be uploaded;

[0013] According to the verification category of the video to be uploaded, determine whether to perform the upload operation.

[0014] Further, the step of performing frame splitting on the video to be uploaded to obtain the pictures after frame splitting specifically includes:

[0015] Perform frame splitting on the video to be uploaded based on a preset frame splitting time interval;

[0016] According to the sequence of frame splitting and a preset time point marking component, sequentially mark the frame sequence value and the frame splitting time point for the pictures obtained by frame splitting;

[0017] Obtain all the pictures that simultaneously have a frame sequence value and a frame splitting time point marking as the pictures after frame splitting.

[0018] Further, the step of performing face detection on the pictures after frame splitting to obtain face detection results specifically includes:

[0019] Adopt object detection technology to perform target face detection on the pictures after frame splitting respectively, and distinguish and mark the pictures containing the target face, wherein the object detection technology includes image object detection technology based on the DeepLab series;

[0020] The step of only retaining the pictures containing faces according to the face detection results to generate a verification picture set specifically includes:

[0021] According to the distinction marking, screen out all the pictures containing the target face;

[0022] Add all the pictures containing the target face to a preset set in the order of the frame sequence value or the frame splitting time point marking to generate the verification picture set.

[0023] Further, the face detection model includes a feature extraction network based on the ResNet convolutional neural network and a verification classification network based on the LSTM recurrent neural network. Before performing the step of inputting the verification picture set into a pre-trained face detection model to obtain the detection and classification results output by the face detection model for the verification picture set, the method further includes:

[0024] Obtain a labeled video data set, wherein the video data set contains positive and negative video samples with a preset first proportional relationship, the positive video samples correspond to real videos, and the negative video samples correspond to forged videos;

[0025] Perform frame-by-frame processing on all video samples in the video dataset to obtain pictures after frame-by-frame processing. Perform face detection on the pictures after frame-by-frame processing to obtain face detection results. According to the face detection results, only retain the pictures containing faces, and generate a training picture set corresponding to each video sample;

[0026] Determine the actual authenticity of each training picture set according to the annotation results of each video sample;

[0027] Input the training picture sets corresponding to all the video samples into the face detection model to be trained;

[0028] Extract the image features of different pictures in all the training picture sets through the feature extraction network based on the ResNet convolutional neural network;

[0029] Extract the temporal features of different pictures in all the training picture sets through the verification classification network based on the LSTM recurrent neural network;

[0030] According to the temporal features and the image features, obtain the image feature change results corresponding to each training picture set;

[0031] Based on the actual authenticity of each training picture set, perform binary classification on the image feature change results corresponding to all the training picture sets to obtain the image feature change results of all the training picture sets that are true and the image feature change results of all the training picture sets that are false;

[0032] According to the image feature change results of all the training picture sets that are true, fit the first image feature change function, and

[0033] According to the image feature change results of all the training picture sets that are false, fit the second image feature change function;

[0034] Deploy the first image feature change function and the second image feature change function as authenticity verification functions to the classification verification nodes of the verification classification network to obtain the pre-trained face detection model.

[0035] Further, before performing the step of deploying the first image feature change function and the second image feature change function as authenticity verification functions to the classification verification nodes of the verification classification network to obtain the pre-trained face detection model, the method further includes:

[0036] Adopt a random sampling method to randomly obtain a video sample with a preset second proportional relationship from the annotated video dataset as a test video set;

[0037] Perform frame splitting on all video samples in the test video set respectively to obtain the pictures after frame splitting. Perform face detection on the pictures after frame splitting respectively to obtain the face detection results. According to the face detection results, only retain the pictures containing faces, and generate a test picture set corresponding to each video sample respectively;

[0038] According to the annotation results of each video sample, determine the actual authenticity of each test picture set as the true result;

[0039] Input the test picture sets corresponding to all the video samples respectively into the face detection model, and obtain the authenticity category output by the face detection model for each test picture set as the test result;

[0040] If the test result is consistent with the true result, the pre-training of the face detection model is completed. Otherwise, adjust the processing parameters of the face detection model, and re-train and test until the test result is consistent with the true result.

[0041] Further, the step of inputting the verification picture set into the pre-trained face detection model to obtain the detection classification result output by the face detection model for the verification picture set specifically includes:

[0042] Extract the image features of different pictures in the verification picture set through the feature extraction network based on the ResNet convolutional neural network;

[0043] Extract the temporal features of different pictures in the verification picture set through the verification classification network based on the LSTM recurrent neural network;

[0044] Obtain the image feature change result corresponding to the verification picture set according to the temporal features and the image features;

[0045] Identify the image feature change function that the image feature change result conforms to through comparison;

[0046] If the image feature change result conforms to the first image feature change function, the detection classification result is a real video;

[0047] If the image feature change result conforms to the second image feature change function, the detection classification result is a forged video.

[0048] Further, the step of determining the verification category of the video to be uploaded based on the detection classification result specifically includes:

[0049] If the detection classification result is a real video, the video to be uploaded is a real video;

[0050] If the detection and classification result is a forged video, then the video to be uploaded is a forged video;

[0051] The step of determining whether to perform an upload operation according to the verification category of the video to be uploaded specifically includes:

[0052] If the video to be uploaded is a genuine video, perform an upload operation;

[0053] If the video to be uploaded is a forged video, perform an upload blocking operation and send a feedback message of not uploading to the target upload end or the video provider end.

[0054] To solve the above technical problems, the embodiment of the present application also provides a video upload verification device, which adopts the following technical solutions:

[0055] A video upload verification device includes:

[0056] A video to be uploaded acquisition module, configured to acquire a video to be uploaded;

[0057] A video frame splitting processing module, configured to perform frame splitting processing on the video to be uploaded to obtain pictures after frame splitting processing;

[0058] A picture face detection module, configured to perform face detection on the pictures after frame splitting processing respectively to obtain face detection results;

[0059] A verification picture set generation module, configured to only retain the pictures containing faces according to the face detection results and generate a verification picture set;

[0060] A model detection and classification module, configured to input the verification picture set into a pre-trained face detection model and obtain the detection and classification results output by the face detection model for the verification picture set;

[0061] A verification category determination module, configured to determine the verification category of the video to be uploaded based on the detection and classification results;

[0062] An upload operation determination module, configured to determine whether to perform an upload operation according to the verification category of the video to be uploaded.

[0063] To solve the above technical problems, the embodiment of the present application also provides a computer device, which adopts the following technical solutions:

[0064] A computer device includes a memory and a processor, and computer-readable instructions are stored in the memory. When the processor executes the computer-readable instructions, the steps of the above-mentioned video upload verification method are implemented.

[0065] To solve the above technical problems, an embodiment of the present application further provides a computer-readable storage medium, adopting the following technical solutions:

[0066] A computer-readable storage medium, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the steps of the video upload verification method as described above are implemented.

[0067] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects:

[0068] In the video upload verification method of the embodiment of the present application, by obtaining the video to be uploaded; performing frame splitting processing on the video to be uploaded to obtain pictures after frame splitting processing; respectively performing face detection on the pictures after frame splitting processing to obtain face detection results; according to the face detection results, only retaining the pictures containing faces to generate a verification picture set; inputting the verification picture set into a pre-trained face detection model to obtain the detection classification results output by the face detection model for the verification picture set; based on the detection classification results, determining the verification category of the video to be uploaded; and according to the verification category of the video to be uploaded, determining whether to perform an upload operation. By combining face detection technology to train a video authenticity detection model, the face features are fully utilized, and according to the face features and the feature changes of the face features in the entire video, the authenticity of the video to be uploaded by the user is predicted, so that forged videos can be identified before the video is uploaded. Applying the method to the face recognition scenario in the financial business can not only automatically identify forged videos, making it more intelligent, but also reduce the workload of manual identification and improve the efficiency of related services. The method can also be applied to the online medical service scenario. By performing face recognition detection in the online medical service scenario and combining the actual medical service results, when errors occur in medical services, such as incorrect drug distribution or incorrect benefit discounts, corrective measures can be taken in a timely manner according to the face recognition scenario to ensure the accuracy of the online medical service scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] To more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0070] Figure 1 is an exemplary system architecture diagram to which the present application can be applied;

[0071] Figure 2 is a flowchart of an embodiment of the video upload verification method according to the present application;

[0072] Figure 3 Yes Figure 2 It is a flowchart of a specific embodiment of step 202 shown;

[0073] Figure 4 It is a flowchart of a specific embodiment of pre-training a face detection model in the video upload verification method described in this application;

[0074] Figure 5 It is a flowchart of a specific embodiment of testing a face detection model in the video upload verification method described in this application;

[0075] Figure 6 Yes Figure 2 It is a flowchart of a specific embodiment of step 205 shown;

[0076] Figure 7 It is a schematic structural diagram of an embodiment of a video upload verification device according to this application;

[0077] Figure 8 It is a schematic structural diagram of an embodiment of a computer device according to this application. Specific Embodiments

[0078] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.

[0079] Referring to "embodiments" herein means that specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0080] To enable those skilled in the technical field to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0081] As Figure 1As shown, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. The terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0082] Users can use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.

[0083] The terminal device 101 may be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, the tablet computer 1012, or the mobile phone 1013, the terminal device 101 may also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop portable computer, a desktop computer, etc.

[0084] The server 103 may be a server providing various services, such as a background server supporting the pages displayed on the terminal device 101.

[0085] It should be noted that the video upload verification method provided by the embodiments of the present application is generally executed by the server. Correspondingly, the video upload verification device is generally set in the server.

[0086] It should be understood that Figure 1 the numbers of the terminal devices, networks, and servers in

[0087] Continue to refer to Figure 2 , which shows a flowchart of an embodiment of the video upload verification method according to the present application. The video upload verification method includes the following steps:

[0088] Step 201, obtain the video to be uploaded.

[0089] In this embodiment, the video to be uploaded, for example, in the field of financial services, in the scenario of online bank loans, it is a face recognition video recorded by a user for loan purposes; it also includes the field of medical services, in the scenario of online medical services, it is a face recognition video recorded for business identity recognition of users.

[0090] It should be understood that the video to be uploaded is mainly a video in which a user performs identity recognition based on the face when handling relevant business;

[0091] Of course, with the rapid development of the Internet short video industry, the video to be uploaded can also be a video work created by the user himself, and at least the user himself is photographed in the video to be uploaded.

[0092] By obtaining the video to be uploaded to perform authenticity detection before the video to be uploaded is uploaded, for the field of financial services, the scenario of online bank loans or the scenario of online medical services, through authenticity detection, forged videos can be discovered in time to avoid losses to customers and enterprises. For the Internet short video industry, through authenticity detection, forged videos can be discovered in time to avoid users being involved in subsequent infringement disputes.

[0093] Step 202: Perform frame splitting on the video to be uploaded to obtain the pictures after frame splitting.

[0094] By performing frame splitting on the video to be uploaded to obtain the pictures after frame splitting, the video authenticity detection is transformed into authenticity detection using pictures with relatively lower dimensions, reducing the difficulty of authenticity detection.

[0095] Step 203: Perform face detection on the pictures after frame splitting respectively to obtain face detection results.

[0096] Step 204: According to the face detection results, only retain the pictures containing faces to generate a verification picture set.

[0097] By performing face detection on the pictures after frame splitting respectively to obtain face detection results, and according to the face detection results, only retaining the pictures containing faces to generate a verification picture set, subsequent video authenticity recognition using the detected face features is realized. Compared with detecting features of reference objects, i.e., static entities, in pictures or videos, the confidence level of the video detection results is improved.

[0098] Step 205: Input the verification picture set into a pre-trained face detection model to obtain the detection classification results output by the face detection model for the verification picture set.

[0099] Step 206: Based on the detection classification results, determine the verification category of the video to be uploaded.

[0100] Step 207: Determine whether to perform the upload operation according to the verification category of the video to be uploaded.

[0101] Specifically, determining whether to perform the upload operation according to the verification category of the video to be uploaded improves the security and prudence of relevant services.

[0102] Specifically, in the online medical service scenario, since most of the involved medical service scenarios are relatively complex, such as multiple processes including online registration, queuing for consultation, drug distribution, and online payment. Throughout the process, through video shooting and face detection verification, it can ensure that the involved medical service scenarios are accurately served to the corresponding users, enabling the hospital to promptly remedy and investigate when errors occur in combination with online medical services.

[0103] In this embodiment, by obtaining the video to be uploaded; performing frame division processing on the video to be uploaded to obtain pictures after frame division processing; respectively performing face detection on the pictures after frame division processing to obtain face detection results; according to the face detection results, only retaining the pictures containing faces to generate a verification picture set; inputting the verification picture set into a pre-trained face detection model to obtain the detection classification results output by the face detection model for the verification picture set; based on the detection classification results, determining the verification category of the video to be uploaded; and determining whether to perform the upload operation according to the verification category of the video to be uploaded. By combining face detection technology to train a video authenticity detection model, it fully utilizes face features. According to the face features and the feature changes of face features in the entire video, it predicts the authenticity of the user's video to be uploaded, so as to be able to identify forged videos before the video is uploaded. Applying the method to the face recognition scenario in the financial business can not only automatically identify forged videos more intelligently, but also reduce the workload of manual identification and improve the efficiency of relevant services. The method can also be applied to the online medical service scenario. By performing face recognition detection in the online medical service scenario and combining the actual medical service results, it can promptly correct and process when errors occur in medical services, such as drug distribution errors and preferential rights errors, to ensure the accuracy of the online medical service scenario.

[0104] Continue to refer to Figure 3 , Figure 3 Yes Figure 2 is

[0105] Step 301: Perform frame division processing on the video to be uploaded based on a preset frame division time interval;

[0106] Step 302: Mark the pictures obtained through frame segmentation with frame sequence values and frame segmentation time points in sequence according to the sequence of frame segmentation processing and the preset time point markers.

[0107] Step 303: Obtain all the pictures that have both frame sequence values and frame segmentation time point markers as the pictures after the frame segmentation processing.

[0108] By marking the pictures obtained through frame segmentation with frame sequence values and frame segmentation time points in sequence, it is convenient to generate the verification picture set and identify the dynamic changes of the face features over time in the entire video later.

[0109] In this embodiment, the step of respectively performing face detection on the pictures after frame segmentation to obtain face detection results specifically includes: adopting target detection technology to respectively perform target face detection on the pictures after frame segmentation, and making different marks on the pictures containing the target face, where the target detection technology includes the image target detection technology based on the DeepLab series.

[0110] Specifically, the image target detection technology based on the DeepLab series uses the ResNet convolutional neural network as the backbone network, such as ResNet-34, ResNet-50, ResNet-101, etc., and extracts image features through convolution, so as to be able to screen out the pictures containing faces from the images.

[0111] In this embodiment, the step of generating a verification picture set by only retaining the pictures containing faces according to the face detection results specifically includes: screening out all the pictures containing the target face according to the different marks; adding all the pictures containing the target face to a preset set in the order of frame sequence values or frame segmentation time point markers to generate the verification picture set.

[0112] Specifically, adding all the pictures containing the target face to a preset set in the order of frame sequence values or frame segmentation time point markers to generate the verification picture set. When using the verification picture set to verify the authenticity of the video later, since all the pictures in the verification picture set are added to the preset set in the order of frame sequence values or frame segmentation time point markers, the change situation of the face features in the video can be obtained after the face features in each acquired picture are obtained.

[0113] In this embodiment, the face detection model includes a feature extraction network based on the ResNet convolutional neural network and a verification classification network based on the LSTM recurrent neural network. Specifically, the feature extraction network based on the ResNet convolutional neural network is used to extract face image features, while the verification classification network based on the LSTM recurrent neural network essentially utilizes the frame time points or frame sequence values to obtain temporal features, and combines the temporal features with the face features to obtain the changes in the face features corresponding to the pictures containing faces in sequence in the video over time.

[0114] Continue to refer to Figure 4 , in some alternative implementation manners, before step 205, there is also a step of pre-training the face detection model. Figure 4 It is a flowchart of a specific embodiment of pre-training the face detection model in the video upload verification method described in this application, including the following steps:

[0115] Step 401, obtain a labeled video data set, where the video data set contains positive and negative video samples with a preset first proportional relationship. The positive video samples correspond to real videos, and the negative video samples correspond to forged videos. The "labeled" means that the positive and negative properties of the samples have been labeled.

[0116] Specifically, the preset first proportional relationship, for example: 1:1, that is, the video data set contains an equal number of positive and negative video samples. Here, the first proportional relationship is set by the trainer himself. The negative video samples can be highly realistic forged videos generated for the positive video samples using a generative adversarial network.

[0117] Step 402, perform frame splitting on all video samples in the video data set to obtain the pictures after frame splitting, perform face detection on the pictures after frame splitting respectively to obtain face detection results, and according to the face detection results, only retain the pictures containing faces to generate a training picture set corresponding to each video sample respectively.

[0118] Specifically, the processing method in step 402 can combine the processing methods in steps 202 to 204 to perform picture marking to facilitate identifying the training picture sets corresponding to each video sample respectively.

[0119] Step 403, determine the actual authenticity of each training picture set according to the annotation result of each video sample.

[0120] Step 404, input the training picture sets corresponding to all video samples respectively into the face detection model to be trained.

[0121] Step 405: Extract the image features of different images in all training picture sets through the feature extraction network based on the ResNet convolutional neural network;

[0122] Specifically, extract the image features of different images in all training picture sets through the feature extraction network based on the ResNet convolutional neural network. Since the images in the same training picture set have a sequential frame sequence (time series) relationship, not only can the image features of different images in all training picture sets be extracted, where the image features include the face features contained in different images; subsequently, according to the time series relationship of different images in the same training picture set, the change situation of the face features in the images of the same training picture set can be obtained as implicit features.

[0123] Step 406: Extract the time series features of different images in all training picture sets through the verification and classification network based on the LSTM recurrent neural network;

[0124] Step 407: Obtain the image feature change result corresponding to each training picture set according to the time series features and the image features;

[0125] Specifically, obtain the image feature change result corresponding to each training picture set according to the time series features and the image features, that is, not only use the directly obtainable image features, but also use the dynamically changing image face features, thus ensuring the accuracy of subsequent detection of the face detection model.

[0126] Step 408: Based on the actual authenticity of each training picture set, perform binary classification on the image feature change results corresponding to all training picture sets to obtain the image feature change results of all true training picture sets and the image feature change results of all false training picture sets;

[0127] Step 409: Fit the first image feature change function according to the image feature change results of all true training picture sets, and,

[0128] Step 410: Fit the second image feature change function according to the image feature change results of all false training picture sets;

[0129] Step 411: Deploy the first image feature change function and the second image feature change function as authenticity verification functions to the classification verification node of the verification and classification network to obtain the pre-trained face detection model.

[0130] By combining the ResNet convolutional neural network to extract features of faces in images and the LSTM recurrent neural network to obtain temporal features, it is possible to obtain the changes in face features over time, that is, the above-mentioned image feature change results. Then, according to the authenticity of the video, binary classification is performed to fit the first change function of the image features and the second change function of the image features as the subsequent verification basis. This ensures the high availability of the pre-trained face detection model.

[0131] Continue to refer to Figure 5 , in some optional implementation manners, before step 411, there is also a step of testing the face detection model. Figure 5 is a flowchart of a specific embodiment of testing the face detection model in the video upload verification method described in this application, including the following steps:

[0132] Step 501, using a random sampling method, randomly obtain video samples with a preset second proportional relationship from the labeled video dataset as the test video set;

[0133] Specifically, the preset second proportional relationship, for example: randomly obtain 20% of the video samples from the labeled video dataset.

[0134] Step 502, perform frame splitting on all video samples in the test video set to obtain the pictures after frame splitting, perform face detection on the pictures after frame splitting, obtain the face detection results, and only retain the pictures containing faces according to the face detection results to generate a test picture set corresponding to each video sample;

[0135] Specifically, the processing method in step 502 can combine the processing methods of steps 202 to 204 to perform picture marking to facilitate identifying the test picture set corresponding to each video sample.

[0136] Step 503, determine the actual authenticity of each test picture set according to the annotation result of each video sample as the true result;

[0137] Step 504, input the test picture sets corresponding to all video samples into the face detection model to obtain the authenticity categories output by the face detection model for each test picture set as the test results;

[0138] Step 505, if the test result is consistent with the true result, the pre-training of the face detection model is completed; otherwise, adjust the processing parameters of the face detection model and re-perform training and testing until the test result is consistent with the true result.

[0139] In this embodiment, by performing tests during training, the high availability of the pre-trained face detection model is ensured.

[0140] Continue to refer to Figure 6 , Figure 6 is Figure 2 a flowchart of a specific embodiment of step 205 shown in

[0141] Step 601: Extract the image features of different pictures in the verification picture set through the feature extraction network based on the ResNet convolutional neural network;

[0142] Specifically, the extracting the image features of different pictures in the verification picture set through the feature extraction network based on the ResNet convolutional neural network includes extracting the face features in different pictures in the verification picture set.

[0143] Step 602: Extract the temporal features of different pictures in the verification picture set through the verification classification network based on the LSTM recurrent neural network;

[0144] Step 603: Obtain the image feature change result corresponding to the verification picture set according to the temporal features and the image features;

[0145] Step 604: Identify the image feature change function that the image feature change result conforms to through comparison;

[0146] Step 605: If the image feature change result conforms to the first image feature change function, the detection classification result is a real video;

[0147] Step 606: If the image feature change result conforms to the second image feature change function, the detection classification result is a forged video.

[0148] In this embodiment, the step of determining the verification category of the video to be uploaded based on the detection classification result specifically includes: if the detection classification result is a real video, the video to be uploaded is a real video; if the detection classification result is a forged video, the video to be uploaded is a forged video;

[0149] In this embodiment, the step of determining whether to perform an upload operation according to the verification category of the video to be uploaded specifically includes: if the video to be uploaded is a real video, perform the upload operation; if the video to be uploaded is a forged video, perform the upload blocking operation and send a feedback message of not uploading to the target upload end or the video provider end.

[0150] This application obtains the video to be uploaded; performs frame splitting on the video to be uploaded to obtain the pictures after frame splitting; performs face detection on the pictures after frame splitting respectively to obtain face detection results; according to the face detection results, only retains the pictures containing faces to generate a verification picture set; inputs the verification picture set into a pre-trained face detection model to obtain the detection classification results output by the face detection model for the verification picture set; determines the verification category of the video to be uploaded based on the detection classification results; and determines whether to perform an upload operation according to the verification category of the video to be uploaded. By training a video authenticity detection model in combination with face detection technology, the face features are fully utilized, and according to the face features and the feature changes of the face features in the whole video, the authenticity of the user's video to be uploaded is predicted, so that forged videos can be identified before the video is uploaded. Applying the method to the face recognition scenario in the financial business can not only automatically identify forged videos, making it more intelligent, but also reduce the workload of manual identification and improve the efficiency of related services. The method can also be applied to the online medical service scenario. By performing face recognition detection in the online medical service scenario and combining the actual medical service results, when errors occur in medical services, such as drug distribution errors and benefit preference errors, corrective measures can be taken in a timely manner according to the face recognition scenario to ensure the accuracy of the online medical service scenario.

[0151] Embodiments of this application can obtain and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0152] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0153] In the embodiments of the present application, by obtaining a video to be uploaded; performing frame division processing on the video to be uploaded to obtain pictures after frame division processing; performing face detection on the pictures after frame division processing respectively to obtain face detection results; according to the face detection results, only retaining the pictures containing faces to generate a verification picture set; inputting the verification picture set into a pre-trained face detection model to obtain the detection classification results output by the face detection model for the verification picture set; based on the detection classification results, determining the verification category of the video to be uploaded; and judging whether to perform an upload operation according to the verification category of the video to be uploaded. By training a video authenticity detection model by combining face detection technology, the face features are fully utilized, and according to the face features and the feature changes of the face features in the whole video, the authenticity of the video to be uploaded by the user is predicted, so that forged videos can be identified before the video is uploaded. Applying the method to the face recognition scenario in the financial business can not only automatically identify forged videos, be more intelligent, but also reduce the workload of manual identification and improve the efficiency of related services. The method can also be applied to the online medical service scenario. By performing face recognition detection in the online medical service scenario and combining the actual medical service results, when errors occur in medical services, such as drug distribution errors and preferential rights errors, corrective processing can be performed in a timely manner according to the face recognition scenario to ensure the accuracy of the online medical service scenario.

[0154] Further referring to Figure 7 as an implementation of the above Figure 2 shown method, an embodiment of a video upload verification device is provided in the present application. This device embodiment corresponds to the Figure 2 shown method embodiment and can be specifically applied to various electronic devices.

[0155] As Figure 7 shown, the video upload verification device 700 described in this embodiment includes: a to-be-uploaded video acquisition module 701, a video frame division processing module 702, a picture face detection module 703, a verification picture set generation module 704, a model detection classification module 705, a verification category determination module 706, and an upload operation judgment module 707. Among them:

[0156] The to-be-uploaded video acquisition module 701 is used to obtain a to-be-uploaded video;

[0157] The video frame division processing module 702 is used to perform frame division processing on the to-be-uploaded video to obtain pictures after frame division processing;

[0158] The picture face detection module 703 is used to perform face detection on the pictures after frame division processing respectively to obtain face detection results;

[0159] The verification image set generation module 704 is configured to generate a verification image set by retaining only the images containing faces according to the face detection results.

[0160] The model detection and classification module 705 is configured to input the verification image set into a pre-trained face detection model to obtain the detection and classification results output by the face detection model for the verification image set.

[0161] The verification category determination module 706 is configured to determine the verification category of the video to be uploaded based on the detection and classification results.

[0162] The upload operation judgment module 707 is configured to judge whether to perform an upload operation according to the verification category of the video to be uploaded.

[0163] In this application, a video to be uploaded is obtained; the video to be uploaded is frame-processed to obtain the frame-processed images; the frame-processed images are respectively subjected to face detection to obtain face detection results; according to the face detection results, only the images containing faces are retained to generate a verification image set; the verification image set is input into a pre-trained face detection model to obtain the detection and classification results output by the face detection model for the verification image set; based on the detection and classification results, the verification category of the video to be uploaded is determined; according to the verification category of the video to be uploaded, it is judged whether to perform an upload operation. By combining face detection technology to train a video authenticity detection model, the face features are fully utilized, and according to the face features and the feature changes of the face features in the whole video, the authenticity of the user's video to be uploaded is predicted, so that forged videos can be identified before the video is uploaded. Applying the method to the face recognition scenario in the financial business can not only automatically identify forged videos, making it more intelligent, but also reduce the workload of manual identification and improve the efficiency of related services. The method can also be applied to the online medical service scenario. By performing face recognition detection in the online medical service scenario and combining the actual medical service results, when errors occur in medical services, such as drug distribution errors and preferential right errors, corrective measures can be taken in a timely manner according to the face recognition scenario to ensure the accuracy of the online medical service scenario.

[0164] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through computer-readable instructions, and the computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the foregoing storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a Read-Only Memory (ROM), or a Random Access Memory (RAM).

[0165] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this document, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0166] To solve the above technical problems, an embodiment of the present application further provides a computer device. Specifically, please refer to Figure 8 , Figure 8 which is the basic structural block diagram of the computer device in this embodiment.

[0167] The computer device 8 includes a memory 8a, a processor 8b, and a network interface 8c that are communicatively connected to each other through a system bus. It should be noted that Figure 8 only the computer device 8 with components such as a memory 8a, a processor 8b, and a network interface 8c is shown in , but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0168] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through means such as a keyboard, a mouse, a remote control, a touchpad, or a voice control device.

[0169] The memory 8a includes at least one type of readable storage medium, which includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 8a may be an internal storage unit of the computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the memory 8a may also be an external storage device of the computer device 8, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, FlashCard, etc. equipped on the computer device 8. Of course, the memory 8a may also include both the internal storage unit and the external storage device of the computer device 8. In this embodiment, the memory 8a is generally used to store the operating system and various application software installed on the computer device 8, such as computer-readable instructions of a video upload verification method, etc. In addition, the memory 8a may also be used to temporarily store various types of data that have been output or will be output.

[0170] In some embodiments, the processor 8b may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. The processor 8b is generally used to control the overall operation of the computer device 8. In this embodiment, the processor 8b is used to run the computer-readable instructions stored in the memory 8a or process data, such as running the computer-readable instructions of the video upload verification method.

[0171] The network interface 8c may include a wireless network interface or a wired network interface, and this network interface 8c is generally used to establish a communication connection between the computer device 8 and other electronic devices.

[0172] The computer device proposed in this embodiment belongs to the field of image verification technology and is applied to the scenario of video authenticity identification. This application obtains the video to be uploaded; performs frame splitting on the video to be uploaded to obtain the pictures after frame splitting; performs face detection on the pictures after frame splitting respectively to obtain the face detection results; according to the face detection results, only retains the pictures containing faces to generate a verification picture set; inputs the verification picture set into a pre-trained face detection model to obtain the detection classification results output by the face detection model for the verification picture set; based on the detection classification results, determines the verification category of the video to be uploaded; and according to the verification category of the video to be uploaded, determines whether to perform the upload operation. By training a video authenticity detection model by combining face detection technology, it makes full use of face features and predicts the authenticity of the video to be uploaded by the user according to the face features and the feature changes of the face features in the entire video, so as to be able to identify forged videos before the video is uploaded. Applying the method to the face recognition scenario under financial business can not only automatically identify forged videos, be more intelligent, but also reduce the workload of manual identification and improve the efficiency of related services. The method can also be applied to the online medical service scenario. By performing face recognition detection in the online medical service scenario and combining the actual medical service results, it can correct and process in time according to the face recognition scenario when errors occur in medical services, such as drug distribution errors and preferential right errors, to ensure the accuracy of the online medical service scenario.

[0173] This application also provides another implementation manner, that is, to provide a computer-readable storage medium, which stores computer-readable instructions that can be executed by a processor to enable the processor to execute the steps of the video upload verification method as described above.

[0174] The computer-readable storage medium proposed in this embodiment belongs to the technical field of image verification and is applied to the scenario of video authenticity identification. This application obtains the video to be uploaded; performs frame splitting on the video to be uploaded to obtain the pictures after frame splitting; performs face detection on the pictures after frame splitting respectively to obtain face detection results; according to the face detection results, only retains the pictures containing faces to generate a verification picture set; inputs the verification picture set into a pre-trained face detection model to obtain the detection classification results output by the face detection model for the verification picture set; determines the verification category of the video to be uploaded based on the detection classification results; and determines whether to perform an upload operation according to the verification category of the video to be uploaded. By training a video authenticity detection model by combining face detection technology, it makes full use of face features, and predicts the authenticity of the user's video to be uploaded according to the face features and the feature changes of the face features in the whole video, so as to be able to identify forged videos before the video is uploaded. Applying the method to the face recognition scenario in the financial business can not only automatically identify forged videos, be more intelligent, but also reduce the workload of manual identification and improve the efficiency of related services. The method can also be applied to the online medical service scenario. By performing face recognition detection in the online medical service scenario and combining the actual medical service results, it can correct and process in time according to the face recognition scenario when errors occur in medical services, such as wrong drug distribution and wrong benefit discounts, to ensure the accuracy of the online medical service scenario.

[0175] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0176] Obviously, the embodiments described above are only a part of the embodiments of this application, rather than all the embodiments. The preferred embodiments of this application are given in the drawings, but they do not limit the patent scope of this application. This application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of this application more thorough and comprehensive. Although this application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments, or perform equivalent replacements on some of the technical features. Any equivalent structure that makes use of the content of the specification and drawings of this application, directly or indirectly applied in other related technical fields, is equally within the scope of patent protection of this application.

Claims

1. A video upload verification method, characterized in that: The steps include: Get the video to be uploaded; Performing frame processing on the video to be uploaded to obtain a frame-processed picture; Perform face detection on the frame-processed images to obtain face detection results; According to the face detection results, only the pictures containing faces are retained to generate a verification picture set; Inputting the verification picture set into a pre-trained face detection model, and obtaining a detection classification result output by the face detection model for the verification picture set; Based on the detection classification result, determining the verification category of the video to be uploaded; Determine whether to perform the upload operation according to the verification category of the video to be uploaded.

2. The video upload verification method according to claim 1, characterized in that: The step of performing frame processing on the video to be uploaded to obtain the frame-processed picture specifically includes: Performing frame processing on the video to be uploaded based on a preset frame time interval; According to the frame processing sequence and the preset time point marking component, the pictures obtained by the frame processing are marked with frame sequence values ​​and frame time points in sequence; All pictures having both frame sequence values ​​and frame time point marks are obtained as the pictures after the frame processing.

3. The video upload verification method according to claim 2, characterized in that: The step of performing face detection on the frame-processed pictures to obtain face detection results specifically includes: Using target detection technology, target faces are detected on the frame-processed pictures respectively, and pictures containing target faces are marked, wherein the target detection technology includes image target detection technology based on the DeepLab series; The step of retaining only pictures containing faces according to the face detection results to generate a verification picture set specifically includes: According to the distinguishing mark, all pictures containing the target face are screened out; All the pictures containing the target face are added to a preset set according to the order of frame sequence values ​​or frame time point marks to generate the verification picture set.

4. The video upload verification method according to claim 1, characterized in that: The face detection model includes a feature extraction network based on a ResNet convolutional neural network and a verification classification network based on an LSTM recurrent neural network. Before executing the step of inputting the verification picture set into the pre-trained face detection model and obtaining the detection and classification results output by the face detection model for the verification picture set, the method further includes: Acquire a labeled video data set, wherein the video data set includes positive and negative video samples in a preset first ratio relationship, the positive video samples correspond to real videos, and the negative video samples correspond to forged videos; Performing frame processing on all video samples in the video data set to obtain frame-processed pictures, performing face detection on the frame-processed pictures to obtain face detection results, and retaining only pictures containing faces according to the face detection results to generate training picture sets corresponding to all video samples; According to the annotation results of each video sample, the actual authenticity of each training picture set is determined; Inputting the training picture sets corresponding to all the video samples into the face detection model to be trained; Extracting image features of different pictures in all training picture sets through the feature extraction network based on the ResNet convolutional neural network; Extracting temporal features of different pictures in all training picture sets through the verification classification network based on the LSTM recurrent neural network; Obtaining image feature change results corresponding to each training picture set according to the time series features and the image features; Based on the actual authenticity of each training picture set, the image feature change results corresponding to all training picture sets are sorted into two categories to obtain the image feature change results of all true training picture sets and the image feature change results of all false training picture sets; According to the image feature change results of all true training pictures, the first image feature change function is fitted, and According to the image feature change results of all false training picture sets, a second image feature change function is fitted; The first change function of the image feature and the second change function of the image feature are deployed as authenticity verification functions to the classification verification node of the verification classification network to obtain the pre-trained face detection model.

5. The video upload verification method according to claim 4, characterized in that: Before executing the step of deploying the first change function of the image feature and the second change function of the image feature as authenticity verification functions to the classification verification node of the verification classification network to obtain the pre-trained face detection model, the method further includes: Using a random sampling method, randomly obtaining video samples with a preset second ratio relationship from the labeled video data set as a test video set; Performing frame processing on all video samples in the test video set to obtain frame-processed pictures, performing face detection on the frame-processed pictures to obtain face detection results, and retaining only pictures containing faces according to the face detection results to generate test picture sets corresponding to all video samples; According to the annotation results of each video sample, the actual authenticity of each test picture set is determined as the real result; Inputting the test picture sets corresponding to all the video samples into the face detection model, and obtaining the authenticity category output by the face detection model for each test picture set as the test result; If the test result is consistent with the true result, the pre-training of the face detection model is completed; otherwise, the processing parameters of the face detection model are adjusted, and training and testing are performed again until the test result is consistent with the true result.

6. The video upload verification method according to claim 4, characterized in that: The step of inputting the verification picture set into a pre-trained face detection model to obtain the detection and classification results output by the face detection model for the verification picture set specifically includes: Extracting image features of different pictures in the verification picture set through the feature extraction network based on the ResNet convolutional neural network; Extracting temporal features of different pictures in the verification picture set through the verification classification network based on the LSTM recurrent neural network; Obtaining image feature change results corresponding to the verification picture set according to the time series features and the image features; By comparison, identifying the image feature change function to which the image feature change result conforms; If the image feature change result conforms to the first image feature change function, then the detection and classification result is a real video; If the image feature change result conforms to the second image feature change function, the detection classification result is a forged video.

7. The video upload verification method according to claim 1 or 6, characterized in that: The step of determining the verification category of the video to be uploaded based on the detection classification result specifically includes: If the detection classification result is a real video, the video to be uploaded is a real video; If the detection classification result is a forged video, then the video to be uploaded is a forged video; The step of determining whether to perform the upload operation according to the verification category of the video to be uploaded specifically includes: If the video to be uploaded is a real video, perform the uploading operation; If the video to be uploaded is a forged video, an upload blocking operation is performed, and a feedback message of not uploading is sent to the target upload end or video provider.

8. A video upload verification device, characterized in that: include: The module for obtaining videos to be uploaded is used to obtain videos to be uploaded; A video frame processing module is used to perform frame processing on the video to be uploaded to obtain a frame-processed picture; The image face detection module is used to perform face detection on the frame-processed images to obtain face detection results; A verification picture set generation module is used to retain only pictures containing faces according to the face detection results to generate a verification picture set; A model detection and classification module, used to input the verification picture set into a pre-trained face detection model, and obtain the detection and classification results output by the face detection model for the verification picture set; A verification category determination module, used to determine the verification category of the video to be uploaded based on the detection classification result; The upload operation judgment module is used to judge whether to perform the upload operation according to the verification category of the video to be uploaded.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the video upload verification method according to any one of claims 1 to 7 when executing the computer-readable instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the video upload verification method according to any one of claims 1 to 7 are implemented.