Scene Classification Method, Device, Electronic Device and Storage Medium for In-Person Interview Video
Through the application of preprocessing and classification model for face-to-face video review, the problem of low efficiency and accuracy of face-to-face video scenario classification in the existing technology is solved, and more efficient and accurate scene classification is achieved.
Patent Information
- Application Number
- CN202210544228.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-18
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-05-18
AI Technical Summary
The existing technology cannot effectively classify scenes of face-to-face videos, resulting in low efficiency and accuracy.
By preprocessing the video frames in the video frame sequence set, similar video frames are deleted, and the scene image sample set is input to the pre-trained scene classification model for classification.
It improves the efficiency and accuracy of scene classification, reduces redundant information, and shortens the time for judging video scenes.
Smart Images

Figure CN114998782B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method, device, electronic device and storage medium for scene classification of face review videos. Background Art
[0002] Scene classification is becoming increasingly important in real life. For example, in the fields of anti-fraud and risk prevention and control, it is necessary to classify the scene of face review videos to ensure that customers conduct face review videos in a safe environment. However, the image content of face review videos is often that customers say something to the camera device, and the proportion of the face in the picture is relatively large, and the available scene information is relatively small. The prior art cannot perform effective scene classification based on this scene information, resulting in low efficiency and accuracy of scene classification of face review videos. Summary of the Invention
[0003] In view of the above, it is necessary to propose a method, device, electronic device and storage medium for scene classification of face review videos. After preprocessing the video frames in the video frame sequence set, a pre-trained scene classification model is used for scene classification, which improves the efficiency and accuracy of scene classification.
[0004] The first aspect of the present invention provides a method for scene classification of face review videos, and the method includes:
[0005] Responding to the received scene classification request, obtaining a face review video;
[0006] Converting the face review video into a video frame sequence set, and judging whether the face review video meets the scene classification environment according to the video frame sequence set;
[0007] When the face review video meets the scene classification environment, preprocessing the video frames in the video frame sequence set to obtain a set of scene image samples;
[0008] Inputting the set of scene image samples into a pre-trained scene classification model to obtain the classification result of each scene image;
[0009] Determining the scene classification result of the face review video according to the classification result of each scene image in the set of scene samples.
[0010] Optionally, the preprocessing the video frames in the video frame sequence set to obtain a set of scene image samples includes:
[0011] Determining the first video frame in the video frame sequence set as the current video frame;
[0012] Calculating the similarity between the current video frame and the next video frame of the current video frame;
[0013] When the similarity between the current video frame and the next video frame of the current video frame is greater than or equal to a preset similarity threshold, delete the next video frame of the current video frame from the video frame sequence set to obtain a new video frame sequence set, determine the first video frame in the new video frame sequence set as the current video frame, and repeatedly calculate the similarity between the current video frame in the new video frame sequence set and the next video frame of the current video frame until the similarity between the first video frame and the last video frame in the new video sequence set is calculated to obtain a set of scene image samples;
[0014] When the similarity between the current video frame and the next video frame of the current video frame is less than the preset similarity threshold, determine the next video frame of the current video frame as the new current video frame, and repeatedly calculate the similarity between the new current video frame and the next video frame of the new current video frame until the similarity between the new current video frame and the last video frame in the video sequence set is calculated to obtain a set of scene image samples.
[0015] Optionally, calculating the similarity between the current video frame and the next video frame of the current video frame includes:
[0016] Calculate the difference between the pixels at each position of the current video frame and the pixels at the corresponding position of the next video frame of the current video frame, and average the differences of the pixels at all positions to obtain the target mean of the pixels;
[0017] Calculate the quotient of the target mean of the pixels and the total number of pixels of the current video frame, and determine the calculated quotient as the similarity between the current video frame and the next video frame of the current video frame.
[0018] Optionally, the training process of the scene classification model includes:
[0019] Obtain multiple scenes and the first video set and the second video set corresponding to the scenes;
[0020] Decode the first video set to obtain a first set of sample images, and decode the second video set to obtain a second set of sample images;
[0021] Use a face detection algorithm to remove the face images of each sample image in the first set of sample images, and determine the multiple sample images after removing the face images as the third set of sample images;
[0022] Divide the second set of sample images into a training set and a test set;
[0023] Input the training set into a preset neural network for training to obtain a pre-trained model;
[0024] Input the test set into the pre-trained model for testing, and calculate the test passing rate.
[0025] Compare the test passing rate with a preset passing rate threshold.
[0026] When the test passing rate is greater than or equal to the preset passing rate threshold, determine that the training of the pre-trained model is completed, and based on the third sample image set, fine-tune the pre-trained model using a preset fine-tuning model to obtain a scene classification model.
[0027] Optionally, the determining the scene classification result of the face review video according to the classification results of each scene image in the scene sample set includes:
[0028] Obtain the highest confidence of each scene image from the classification results of each scene image, and compare the highest confidence with a preset confidence threshold.
[0029] Retain multiple scene images corresponding to the highest confidence greater than or equal to the preset confidence threshold, and determine the category corresponding to the highest confidence of each scene image as the target scene category of the corresponding scene image.
[0030] Classify the retained multiple scene images according to the target scene category to obtain a target scene sample set for each target scene category.
[0031] Based on the target scene sample set of each target scene category, count the total votes of each target scene category.
[0032] Calculate the quotient of the difference between the highest total votes and the second highest total votes among the multiple total votes of the multiple target scene categories and the sum of the multiple total votes to obtain the target scene category result.
[0033] Compare the target scene category result with a preset scene category threshold.
[0034] When the target scene category result is greater than or equal to the preset scene category threshold, determine the scene category with the highest total votes as the scene classification result of the face review video.
[0035] Optionally, the determining whether the face review video meets the scene classification environment according to the video frame sequence set includes:
[0036] Convert each video frame in the video frame sequence set into an HSV image to obtain an HSV image set; remove pixels in the face area of each HSV image in the HSV image set, retain pixels in the non-face area of each HSV image, and calculate a target brightness value of each HSV image based on the retained pixels in the non-face area of each HSV image;
[0037] Compare the target brightness value of each HSV image with the preset brightness threshold;
[0038] Counting the total number of HSV images whose target brightness value is less than the preset brightness threshold, to obtain a first total number;
[0039] Obtaining a target total number threshold based on a second total number of images in the HSV image set;
[0040] comparing the first total to the target total threshold;
[0041] When the first total number is less than the target total number threshold, it is determined that the face-to-face review video meets the scene classification environment.
[0042] Optionally, the method further comprises:
[0043] When the face-to-face review video does not meet the scene classification environment, switch to the manual scene classification review system.
[0044] A second aspect of the present invention provides a scene classification device for face-to-face video, the device comprising:
[0045] An acquisition module, used to acquire the face-to-face review video in response to the received scene classification request;
[0046] A judgment module, used for converting the face-to-face review video into a video frame sequence set, and judging whether the face-to-face review video satisfies a scene classification environment according to the video frame sequence set;
[0047] A preprocessing module, used for preprocessing the video frames in the video frame sequence set to obtain a scene image sample set when the face-to-face review video meets the scene classification environment;
[0048] An input module, used to input the scene image sample set into a pre-trained scene classification model to obtain a classification result for each scene image;
[0049] A determination module is used to determine the scene classification result of the face-to-face review video according to the classification result of each scene image in the scene sample set.
[0050] A third aspect of the present invention provides an electronic device, which includes a processor and a memory. When the processor executes a computer program stored in the memory, the method for classifying the scenarios of the face review video is implemented.
[0051] A fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for classifying the scenarios of the face review video is implemented.
[0052] In summary, for the method, device, electronic device, and storage medium for classifying the scenarios of the face review video according to the present invention, it is determined whether the face review video meets the scenario classification environment through a video frame sequence set. When the face review video meets the scenario classification environment, the video frames in the video frame sequence set are preprocessed, and similar video frames are deleted, reasonably filtering out a large number of redundant video frames. On the one hand, redundant information is reduced, and on the other hand, the calculation time of the scenario classification model can be greatly improved, shortening the time for judging the video scenario and improving the efficiency of classifying the scenarios of the face review video. The scenario image sample set is input into a pre-trained scenario classification model to obtain the classification results of each scenario image. During the training process of the scenario classification model, considerations are made from two aspects: open-source data and face review data, avoiding the phenomenon of overfitting of a small amount of data when using face review data to train the scenario classification model in the prior art, improving the accuracy of the trained scenario classification model, and thus improving the accuracy of scenario classification. According to the classification results of each scenario image in the scenario sample set, the scenario classification result of the face review video is determined, and by discarding some scenario images whose scenario classification cannot be confirmed, the efficiency and accuracy of scenario classification are improved. Description of the Drawings
[0053] Figure 1 is a flowchart of the method for classifying the scenarios of the face review video provided in the first embodiment of the present invention.
[0054] Figure 2 is a structural diagram of the device for classifying the scenarios of the face review video provided in the second embodiment of the present invention.
[0055] Figure 3 is a schematic structural diagram of the electronic device provided in the third embodiment of the present invention. Detailed Embodiments
[0056] In order to more clearly understand the above objects, features, and advantages of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs. The terms used in the specification of this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention.
[0058] Embodiment 1
[0059] Figure 1 It is a flowchart of the scenario classification method for the in-person review video provided by Embodiment 1 of the present invention.
[0060] In this embodiment, the scenario classification method for the in-person review video can be applied to an electronic device. For an electronic device that needs to classify the scenario of the in-person review video, the function of classifying the scenario of the in-person review video provided by the method of the present invention can be directly integrated on the electronic device, or run on the electronic device in the form of a Software Development Kit (SDK).
[0061] Embodiments of the present invention can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0062] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning, deep learning.
[0063] As Figure 1 shown, the scenario classification method for the in-person review video specifically includes the following steps. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.
[0064] S11, in response to the received scenario classification request, obtain the in-person review video.
[0065] In this embodiment, in the fields of anti-fraud and risk prevention and control, it is necessary to classify the scenario of the in-person review video. The scenario classification request is initiated from the client to the server. Specifically, the client can be a smart phone, an IPAD, or other existing intelligent devices. During the scenario classification process of the in-person review video, receive the scenario classification request sent by the client, and in response to the scenario classification request, obtain the in-person review video.
[0066] S12. Convert the in-person review video into a set of video frame sequences, and determine whether the in-person review video meets the scene classification environment according to the set of video frame sequences.
[0067] In this embodiment, the scene classification environment is preset. Determining whether the in-person review video meets the scene classification environment means determining whether the in-person review video is in a very dark environment. If it is in a very dark environment, it is difficult to determine which scene the in-person review is in, and the in-person review video does not meet the scene classification environment.
[0068] In this embodiment, when determining whether the in-person review video meets the scene classification environment according to the set of video frame sequences, the video frames in the set of video frame sequences are converted into the HSV space, and the brightness degree is judged by using the brightness value of each pixel. The larger the brightness value, the higher the brightness of the pixel.
[0069] In this embodiment, when conducting an in-person review video, if the user is in a dark environment, the flashlight of the mobile phone will be turned on, resulting in a relatively bright face image of the user in the in-person review video, which will increase the brightness value of each video frame in the in-person review video. If the average value of the brightness values of all pixels of each video frame is directly used, it is difficult to distinguish the video frames in the dark environment. Therefore, when determining whether the in-person review video meets the scene classification environment according to the set of video frame sequences, the pixels in the face area of each video frame in the in-person review video need to be removed, and the scene classification environment is judged according to the brightness values of the pixels in the non-face area of each video frame, which improves the accuracy of the scene classification environment judgment. When conducting the scene classification of the in-person review video subsequently, the judgment result of the scene classification environment is considered, which improves the accuracy and efficiency of the scene classification of the in-person review video.
[0070] In an alternative embodiment, determining whether the in-person review video meets the scene classification environment according to the set of video frame sequences includes:
[0071] Convert each video frame in the set of video frame sequences into an HSV image to obtain an HSV image set; remove the face area pixels in each HSV image in the HSV image set, retain the pixels in the non-face area of each HSV image, and calculate the target brightness value of each HSV image based on the pixels in the non-face area of the retained HSV images.
[0072] Compare the target brightness value of each HSV image with a preset brightness threshold.
[0073] Count the total sum of the number of HSV images whose target brightness value is less than the preset brightness threshold to obtain a first total sum.
[0074] Obtain a target total sum threshold based on the second total sum of the images in the HSV image set.
[0075] Compare the first total number with the target total number threshold;
[0076] When the first total number is less than the target total number threshold, determine that the in-person review video meets the scenario classification environment.
[0077] Further, the comparing the first total number with the target total number threshold further includes:
[0078] When the first total number is greater than or equal to the target total number threshold, determine that the in-person review video does not meet the scenario classification environment.
[0079] In this embodiment, a brightness threshold can be preset, and the target brightness value of each HSV image is compared with the preset brightness threshold. When the target brightness value of each HSV image is less than the preset brightness threshold, it is determined that this HSV image is in a dark environment, and the first total number of HSV images in the HSV image set that are in the dark environment is counted.
[0080] In this embodiment, the second total number refers to the sum of the total number of images in the HSV image set in the in-person review video, and different target total number thresholds are set for different second total numbers.
[0081] Exemplarily, if the second total number of the HSV image set in the in-person review video is 100 images, and the corresponding target total number threshold is 60 images, when the first total number of HSV images in the dark environment is greater than or equal to 60 images, it is determined that the in-person review video is in a dark environment, that is, it does not meet the scenario classification environment.
[0082] Further, the calculating the target brightness value of each HSV image based on the pixels of the non-face region of each retained HSV image includes:
[0083] Obtain the brightness value of each pixel in the non-face region of each retained HSV image, and average the multiple brightness values of the non-face region of each HSV image, and determine the average value as the target brightness value of the non-face region of the corresponding HSV image.
[0084] S13. When the in-person review video meets the scenario classification environment, preprocess the video frames in the video frame sequence set to obtain a set of scene image samples.
[0085] In an alternative embodiment, the preprocessing the video frames in the video frame sequence set to obtain a set of scene image samples includes:
[0086] Determine the first video frame in the video frame sequence set as the current video frame;
[0087] Calculate the similarity between the current video frame and the next video frame of the current video frame;
[0088] When the similarity between the current video frame and the next video frame of the current video frame is greater than or equal to a preset similarity threshold, delete the next video frame of the current video frame from the video frame sequence set to obtain a new video frame sequence set, determine the first video frame in the new video frame sequence set as the current video frame, and repeat calculating the similarity between the current video frame in the new video frame sequence set and the next video frame of the current video frame until the calculation of the similarity between the first video frame and the last video frame in the new video sequence set is completed to obtain a set of scene image samples;
[0089] When the similarity between the current video frame and the next video frame of the current video frame is less than the preset similarity threshold, determine the next video frame of the current video frame as the new current video frame, and repeat calculating the similarity between the new current video frame and the next video frame of the new current video frame until the calculation of the similarity between the new current video frame and the last video frame in the video sequence set is completed to obtain a set of scene image samples.
[0090] Further, the calculating the similarity between the current video frame and the next video frame of the current video frame includes:
[0091] Calculate the difference between the pixels at each position of the current video frame and the pixels at the corresponding position of the next video frame of the current video frame, and average the differences of the pixels at all positions to obtain the target mean of the pixels;
[0092] Calculate the quotient of the target mean of the pixels and the total number of pixels of the current video frame, and determine the calculated quotient as the similarity between the current video frame and the next video frame of the current video frame.
[0093] In this embodiment, a similarity threshold can be preset, and compare the calculated similarity between the current video frame and the next video frame of the current video frame with the preset similarity threshold. When the similarity between the current video frame and the next video frame of the current video frame is less than the preset similarity threshold, it indicates that the current video frame is similar to the next video frame of the current video frame, and delete the next video frame of the current video frame from the video frame sequence set.
[0094] In this embodiment, by calculating the similarity between the current video frame and the next video frame of the current video frame, and using the differences between adjacent video frames, similar video frames are deleted, a large number of redundant video frames are reasonably filtered out. On the one hand, redundant information is reduced, and on the other hand, the calculation time of the scene classification model can be greatly improved, the time for judging the video scene is shortened, and the scene classification efficiency of the face review video is improved.
[0095] S14. Input the scene image sample set into a pre-trained scene classification model to obtain the classification result of each scene image.
[0096] In this embodiment, when determining the scene category of a certain video frame in the face review video, the pre-trained scene classification model will give the confidence of each video frame belonging to each scene category, and the category corresponding to the highest confidence is determined as the category of the video frame. Among them, the scene categories may include indoor, outdoor, in-vehicle, public places, etc.
[0097] Specifically, the training process of the scene classification model includes:
[0098] Obtain multiple scenes and the first video set and the second video set corresponding to the scenes;
[0099] Decode the first video set to obtain a first sample image set, and decode the second video set to obtain a second sample image set;
[0100] Use a face detection algorithm to remove the face images of each sample image in the first sample image set, and determine the multiple sample images after removing the face images as a third sample image set;
[0101] Divide the training set and the test set from the second sample image set;
[0102] Input the training set into a preset neural network for training to obtain a pre-trained model;
[0103] Input the test set into the pre-trained model for testing, and calculate the test pass rate;
[0104] Compare the test pass rate with a preset pass rate threshold;
[0105] When the test pass rate is greater than or equal to the preset pass rate threshold, determine that the training of the pre-trained model is completed, and based on the third sample image set, use a preset fine-tuning model to fine-tune the pre-trained model to obtain a scene classification model.
[0106] In this embodiment, the preset fine-tuning model can be a Fine tuning model. The process of fine-tuning the pre-trained model using the Fine tuning model is a prior art, and will not be elaborated in this embodiment.
[0107] Further, the comparison of the test pass rate with the preset pass rate threshold further includes:
[0108] When the test pass rate is less than the preset pass rate threshold, increase the number of the training set and retrain the pre-trained scenario classification model.
[0109] In this embodiment, since the face review data in the face review video is relatively sensitive, it is difficult for institutions such as banks to provide a large amount of training data for each scenario during the training process of the scenario classification model. If a small amount of data is used for the training of the scenario classification model, data overfitting is likely to occur.
[0110] In this embodiment, the first video set is a face review video set of multiple scenarios. Since the proportion of face images in the video frames of the face review video set is relatively large, in order to prevent the face images from affecting the subsequent scenario classification results, the face detection algorithm is used to remove the face images in all the images in the first video set, and the multiple images after removal are determined as the third sample image set; the second video set is a large amount of scenario data provided by Place365. Specifically, the large amount of scenario data provided by Place365 is open-source data.
[0111] In this embodiment, when training the scenario classification model, a large amount of scenario data provided by Place365, that is, the second sample image set corresponding to the second video set, is used to pre-train the classification model to obtain a pre-trained model. The pre-trained model is fine-tuned through the face review data, that is, the third sample image set, to obtain a scenario classification model. During the training process of the scenario classification model, considerations are made from two aspects of open-source data and face review data, and the pre-trained model is fine-tuned based on the face review data to obtain a scenario classification model, ensuring the accuracy of the trained scenario classification model. At the same time, the phenomenon of overfitting of a small amount of data when only using face review data to train the scenario classification model in the prior art is avoided, improving the accuracy of the trained scenario classification model, and further improving the accuracy of scenario classification.
[0112] S15. According to the classification results of each scenario image in the scenario sample set, determine the scenario classification result of the face review video.
[0113] In this embodiment, when confirming the scenario classification result of the face review video, the classification results of each scenario image are considered.
[0114] In an optional embodiment, determining the scene classification result of the in-person interview video according to the classification results of each scene image in the scene sample set includes:
[0115] Obtain the highest confidence of each scene image from the classification results of each scene image, and compare the highest confidence with a preset confidence threshold;
[0116] Retain multiple scene images corresponding to the highest confidence greater than or equal to the preset confidence threshold, and determine the category corresponding to the highest confidence of each scene image as the target scene category of the corresponding scene image;
[0117] Classify the retained multiple scene images according to the target scene category to obtain a target scene sample set for each target scene category;
[0118] Based on the target scene sample set of each target scene category, count the total votes of each target scene category;
[0119] Calculate the quotient of the difference between the highest total votes and the second highest total votes among the multiple total votes of the multiple target scene categories and the sum of the multiple total votes to obtain the target scene category result;
[0120] Compare the target scene category result with a preset scene category threshold;
[0121] When the target scene category result is greater than or equal to the preset scene category threshold, determine the scene category with the highest total votes as the scene classification result of the in-person interview video.
[0122] Furthermore, comparing the target scene category result with a preset scene category threshold further includes:
[0123] When the target scene category result is less than the preset scene category threshold, switch to the manual scene classification review system.
[0124] In this embodiment, when judging the scene category of each scene image of the in-person interview video, the scene classification model will give the specific confidence of each scene image belonging to each type of scene, and determine the category corresponding to the highest confidence as the target scene category of each scene image.
[0125] In this embodiment, if there are multiple scenes in each scene image, it is necessary to determine which specific scene category each scene image belongs to.
[0126] Exemplarily, the preset confidence threshold is 0.4. For a three-classification scenario, it is determined which specific category each scene image belongs to among the in-vehicle, indoor, and outdoor scenes. For example: the confidence given by the scene classification model for the in-vehicle category is 0.82, the confidence for the indoor category is 0.10, and the confidence for the outdoor category is 0.08. Since the highest confidence 0.82 is greater than the preset confidence threshold 0.4, it is determined that the target scene category of the corresponding scene image is in-vehicle; the confidence given by the scene classification model for the in-vehicle category is 0.31, the confidence for the indoor category is 0.36, and the confidence for the outdoor category is 0.33. Since the highest confidence 0.36 is less than the preset confidence threshold 0.4, it indicates that discriminative information cannot be given in this scene image, and the scene classification model cannot accurately determine the scene category of this scene image, so this scene image is discarded.
[0127] In this embodiment, if there are a total of 100 scene images in the target scene sample sets of multiple target scene categories, the first scene classification result: 80 scene images are in-vehicle, 10 scene images are indoor, and 10 scene images are outdoor; the second scene classification result: 45 scene images are in-vehicle, 40 scene images are outdoor, and 15 scene images are indoor. For the second scene classification result, with 45 scene images in-vehicle, 40 scene images outdoor, and 15 scene images indoor, the scene classification model cannot determine which scene the face review video belongs to.
[0128] In this embodiment, the preset scene category threshold is 0.4. The target scene classification result is determined by calculating the difference between the highest total votes and the second-highest total votes divided by the total votes. The scene classification result of the face review video is determined according to the comparison between the target scene classification result and the preset scene category threshold. For example, the target scene classification result corresponding to the first scene classification result is: (80 - 10) / 100 = 0.7, and the target scene classification result corresponding to the second scene classification result is: (45 - 40) / 100 = 0.05. It can be seen that 0.7 is greater than 0.4, so it can be determined that the scene classification result of the face review video corresponding to the first scene classification result is the in-vehicle scene; while 0.05 is less than 0.4, indicating that the scene classification result of the face review video cannot be determined, and it is switched to manual review.
[0129] In this embodiment, based on the classification results of each scene image in the scene sample set, the scene classification result of the face review video is determined. By discarding some scene images whose scene classification cannot be confirmed, the efficiency and accuracy of scene classification are improved.
[0130] S16. When the face review video does not meet the scene classification environment, switch to the manual scene classification review system.
[0131] In this embodiment, when the face review video does not meet the scene classification environment, it is determined that the face review video belongs to a dark environment. In a dark environment, it is difficult for the scene classification subsystem of the face review video to determine the environment where the user is located from the face review video. Then, it switches to the manual scene classification review system for manual review, ensuring the accuracy rate of the scene classification of the face review video.
[0132] In summary, for the scene classification method of the face review video described in this embodiment, it is determined whether the face review video meets the scene classification environment through the video frame sequence set. When the face review video meets the scene classification environment, the video frames in the video frame sequence set are preprocessed to delete similar video frames, reasonably filtering out a large number of redundant video frames. On the one hand, it reduces redundant information, and on the other hand, it can greatly improve the calculation time of the scene classification model, shorten the time for judging the video scene, and improve the scene classification efficiency of the face review video. The scene image sample set is input into the pre-trained scene classification model to obtain the classification result of each scene image. During the training process of the scene classification model, considerations are made from two aspects: open-source data and face review data, avoiding the phenomenon of overfitting of a small amount of data when using face review data to train the scene classification model in the prior art, improving the accuracy rate of the trained scene classification model, and thus improving the accuracy rate of scene classification. According to the classification results of each scene image in the scene sample set, the scene classification result of the face review video is determined, and by discarding some scene images whose scene classification cannot be confirmed, the efficiency and accuracy rate of scene classification are improved.
[0133] Embodiment 2
[0134] Figure 2 It is the structural diagram of the scene classification device for the face review video provided in Embodiment 2 of the present invention.
[0135] In some embodiments, the scene classification device 20 for the face review video may include multiple functional modules composed of program code segments. The program codes of each program segment in the scene classification device 20 for the face review video can be stored in the memory of the electronic device and executed by the at least one processor to execute (see the detailed Figure 1 description) the functions of the scene classification of the face review video.
[0136] In this embodiment, the scenario classification device 20 of the in-person review video can be divided into multiple functional modules according to the functions it performs. The functional modules may include: an acquisition module 201, a judgment module 202, a preprocessing module 203, an input module 204, a determination module 205, and a switching module 206. As used in the present invention, a module refers to a series of computer-readable instruction segments that can be executed by at least one processor and can complete a fixed function, and is stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.
[0137] The acquisition module 201 is configured to acquire the in-person review video in response to a received scenario classification request.
[0138] In this embodiment, in the fields of anti-fraud, risk prevention and control, etc., it is necessary to classify the scenario of the in-person review video. A scenario classification request is initiated from the client to the server. Specifically, the client may be a smart phone, an IPAD or other existing smart devices, and the server may be a scenario classification subsystem for the in-person review video. During the scenario classification process of the in-person review video, for example, the client may send a scenario classification request to the scenario classification subsystem of the in-person review video, and the scenario classification subsystem of the in-person review video is configured to receive the scenario classification request sent by the client and acquire the in-person review video in response to the scenario classification request.
[0139] The judgment module 202 is configured to convert the in-person review video into a video frame sequence set and judge whether the in-person review video meets the scenario classification environment according to the video frame sequence set.
[0140] In this embodiment, the scenario classification environment is preset. Judging whether the in-person review video meets the scenario classification environment means judging whether the in-person review video is in a very dark environment. If it is in a very dark environment, it is difficult to judge which scenario the in-person review is in, and the in-person review video does not meet the scenario classification environment.
[0141] In this embodiment, when judging whether the in-person review video meets the scenario classification environment according to the video frame sequence set, the video frames in the video frame sequence set are converted into the HSV space, and the brightness degree is judged by using the brightness value of each pixel. The larger the brightness value, the higher the brightness of the pixel.
[0142] In the present embodiment, during the face review video, if the user is in a dark environment, the user will turn on the flashlight of the mobile phone, resulting in a brighter facial image of the user in the face review video, which will increase the brightness value of each video frame in the face review video. If the average brightness value of all pixels in each video frame is directly used, it will be difficult to distinguish the video frames in a dark environment. Therefore, when judging whether the face review video meets the scene classification environment according to the video frame sequence set, it is necessary to eliminate the pixels in the face area of each video frame in the face review video, and judge the scene classification environment according to the brightness values of the pixels in the non-face area of each video frame, thereby improving the accuracy of the scene classification environment judgment. When the scene classification of the face review video is subsequently performed, the judgment result of the scene classification environment is taken into consideration, thereby improving the accuracy and efficiency of the scene classification of the face review video.
[0143] In an optional embodiment, the judging module 202 judges whether the face-to-face review video satisfies the scene classification environment according to the video frame sequence set, including:
[0144] Convert each video frame in the video frame sequence set into an HSV image to obtain an HSV image set; remove pixels in the face area of each HSV image in the HSV image set, retain pixels in the non-face area of each HSV image, and calculate a target brightness value of each HSV image based on the retained pixels in the non-face area of each HSV image;
[0145] Compare the target brightness value of each HSV image with the preset brightness threshold;
[0146] Counting the total number of HSV images whose target brightness value is less than the preset brightness threshold, to obtain a first total number;
[0147] Obtaining a target total number threshold based on a second total number of images in the HSV image set;
[0148] comparing the first total to the target total threshold;
[0149] When the first total number is less than the target total number threshold, it is determined that the face-to-face review video meets the scene classification environment.
[0150] Further, comparing the first total with the target total threshold also includes:
[0151] When the first total is greater than or equal to the target total threshold, it is determined that the face-to-face review video does not meet the scene classification environment.
[0152] In this embodiment, a brightness threshold can be preset, and the target brightness value of each HSV image is compared with the preset brightness threshold. When the target brightness value of each HSV image is less than the preset brightness threshold, it is determined that this HSV image is in a dark environment, and the first total number of HSV images in the HSV image set that are in the dark environment is counted.
[0153] In this embodiment, the second total number refers to the sum of the total number of images in the HSV image set in the face review video, and different target total number thresholds are set for different second total numbers.
[0154] Exemplarily, if the second total number of the HSV image set in the face review video is 100 images, and the corresponding target total number threshold is 60 images, when the first total number of HSV images in the dark environment is greater than or equal to 60 images, it is determined that the face review video is in a dark environment, that is, it does not meet the scene classification environment.
[0155] Further, calculating the target brightness value of each HSV image based on the pixels in the non-face region of each retained HSV image includes:
[0156] Obtain the brightness value of each pixel in the non-face region of each retained HSV image, and average the multiple brightness values in the non-face region of each HSV image, and determine the average value as the target brightness value of the non-face region of the corresponding HSV image.
[0157] The preprocessing module 203 is configured to preprocess the video frames in the video frame sequence set to obtain a set of scene image samples when the face review video meets the scene classification environment.
[0158] In an alternative embodiment, the preprocessing module 203 preprocessing the video frames in the video frame sequence set to obtain a set of scene image samples includes:
[0159] Determine the first video frame in the video frame sequence set as the current video frame;
[0160] Calculate the similarity between the current video frame and the next video frame of the current video frame;
[0161] When the similarity between the current video frame and the next video frame of the current video frame is greater than or equal to a preset similarity threshold, delete the next video frame of the current video frame from the video frame sequence set to obtain a new video frame sequence set, determine the first video frame in the new video frame sequence set as the current video frame, and repeatedly calculate the similarity between the current video frame in the new video frame sequence set and the next video frame of the current video frame until the similarity between the first video frame and the last video frame in the new video sequence set is calculated to obtain a set of scene image samples;
[0162] When the similarity between the current video frame and the next video frame of the current video frame is less than the preset similarity threshold, determine the next video frame of the current video frame as the new current video frame, and repeatedly calculate the similarity between the new current video frame and the next video frame of the new current video frame until the similarity between the new current video frame and the last video frame in the video sequence set is calculated to obtain a set of scene image samples.
[0163] Further, the calculation of the similarity between the current video frame and the next video frame of the current video frame includes:
[0164] Calculate the difference between the pixels at each position of the current video frame and the pixels at the corresponding positions of the next video frame of the current video frame, and average the differences of the pixels at all positions to obtain the target mean value of the pixels;
[0165] Calculate the quotient of the target mean value of the pixels and the total number of pixels of the current video frame, and determine the calculated quotient as the similarity between the current video frame and the next video frame of the current video frame.
[0166] In this embodiment, the similarity threshold can be preset, and the calculated similarity between the current video frame and the next video frame of the current video frame is compared with the preset similarity threshold. When the similarity between the current video frame and the next video frame of the current video frame is less than the preset similarity threshold, it indicates that the current video frame is similar to the next video frame of the current video frame, and the next video frame of the current video frame is deleted from the video frame sequence set.
[0167] In this embodiment, by calculating the similarity between the current video frame and the next video frame of the current video frame, using the differences between adjacent video frames, deleting similar video frames, a large number of redundant video frames are reasonably filtered out. On the one hand, redundant information is reduced, and on the other hand, the calculation time of the scene classification model can be greatly improved, the time for judging the video scene is shortened, and the scene classification efficiency of the face review video is improved.
[0168] An input module 204 for inputting the set of scene image samples into a pre-trained scene classification model to obtain the classification results of each scene image.
[0169] In this embodiment, when determining the scene category of a certain video frame in the in-person review video, the pre-trained scene classification model will give the confidence of each video frame belonging to each type of scene, and the category corresponding to the highest confidence is determined as the category of this video frame. Among them, the scene categories may include indoor, outdoor, in-vehicle, public places, etc.
[0170] Specifically, the training process of the scene classification model includes:
[0171] Obtain multiple scenes and the first video set and the second video set corresponding to the scenes;
[0172] Decode the first video set to obtain a first sample image set, and decode the second video set to obtain a second sample image set;
[0173] Use a face detection algorithm to remove the face images of each sample image in the first sample image set, and determine the multiple sample images after removing the face images as a third sample image set;
[0174] Divide the training set and the test set from the second sample image set;
[0175] Input the training set into a preset neural network for training to obtain a pre-trained model;
[0176] Input the test set into the pre-trained model for testing, and calculate the test pass rate;
[0177] Compare the test pass rate with a preset pass rate threshold;
[0178] When the test pass rate is greater than or equal to the preset pass rate threshold, it is determined that the training of the pre-trained model is completed, and based on the third sample image set, the pre-trained model is fine-tuned using a preset fine-tuning model to obtain a scene classification model.
[0179] In this embodiment, the preset fine-tuning model may be a Fine tuning model. The process of fine-tuning the pre-trained model using the Fine tuning model is a prior art and will not be elaborated in this embodiment.
[0180] Further, the comparison of the test pass rate with the preset pass rate threshold further includes:
[0181] When the test pass rate is less than the preset pass rate threshold, increase the number of the training set and re-train the pre-trained scene classification model;
[0182] In this embodiment, since the face review data in the face review video is relatively sensitive, it is difficult for institutions such as banks to provide a large amount of training data for each scenario during the training process of the scenario classification model. If a small amount of data is used for training the scenario classification model, data overfitting is likely to occur.
[0183] In this embodiment, the first video set is a face review video set of multiple scenarios. Since the face images in the video frames of the face review video set account for a relatively large proportion, in order to prevent the face images from affecting the subsequent scenario classification results, the face detection algorithm is used to remove the face images in all the images in the first video set, and the multiple images after removal are accurately determined as the third sample image set; the second video set is a large amount of scenario data provided by Place365. Specifically, the large amount of scenario data provided by Place365 is open-source data.
[0184] In this embodiment, when training the scenario classification model, a large amount of scenario data provided by Place365, that is, the second sample image set corresponding to the second video set, is used to pre-train the classification model to obtain a pre-trained model. The pre-trained model is fine-tuned through the face review data, that is, the third sample image set, to obtain a scenario classification model. During the training process of the scenario classification model, considerations are made from two aspects of open-source data and face review data, and the pre-trained model is fine-tuned based on the face review data to obtain a scenario classification model, ensuring the accuracy of the trained scenario classification model. At the same time, it avoids the phenomenon of overfitting of a small amount of data when only using face review data to train the scenario classification model in the prior art, improves the accuracy of the trained scenario classification model, and further improves the accuracy of scenario classification.
[0185] The determination module 205 is configured to determine the scenario classification result of the face review video according to the classification result of each scenario image in the scenario sample set.
[0186] In this embodiment, when confirming the scenario classification result of the face review video, the classification result of each scenario image is considered.
[0187] In an alternative embodiment, the determination module 205 determines the scenario classification result of the face review video according to the classification result of each scenario image in the scenario sample set, including:
[0188] Obtain the highest confidence of each scenario image from the classification results of each scenario image, and compare the highest confidence with a preset confidence threshold;
[0189] Retain multiple scenario images whose highest confidence is greater than or equal to the preset confidence threshold, and determine the category corresponding to the highest confidence of each scenario image as the target scenario category of the corresponding scenario image.
[0190] Classify the multiple retained scene images according to the target scene category to obtain a target scene sample set for each target scene category;
[0191] Based on the target scene sample set for each target scene category, count the total votes for each target scene category;
[0192] Calculate the quotient of the difference between the highest total votes and the second highest total votes among the multiple total votes of the multiple target scene categories and the sum of the multiple total votes to obtain the target scene category result;
[0193] Compare the target scene category result with a preset scene category threshold;
[0194] When the target scene category result is greater than or equal to the preset scene category threshold, determine the scene category with the highest total votes as the scene classification result of the face-to-face review video.
[0195] Further, the comparing the target scene category result with the preset scene category threshold further includes:
[0196] When the target scene category result is less than the preset scene category threshold, switch to the manual scene classification review system.
[0197] In this embodiment, when judging the scene category of each scene image of the face-to-face review video, the scene classification model will give the specific confidence of each scene image belonging to each type of scene, and determine the category corresponding to the highest confidence as the target scene category of each scene image.
[0198] In this embodiment, if there are multiple scenes in each scene image, it is necessary to judge which specific scene category each scene image belongs to.
[0199] Exemplarily, the preset confidence threshold is 0.4. For a three-classification scene, judge which specific category of the in-vehicle, indoor, and outdoor scenes each scene image belongs to. For example: the scene classification model gives a confidence of 0.82 for belonging to the in-vehicle, a confidence of 0.10 for belonging to the indoor, and a confidence of 0.08 for belonging to the outdoor. The highest confidence of 0.82 is greater than the preset confidence threshold of 0.4, so it is determined that the target scene category of the corresponding scene image is in-vehicle; the scene classification model gives a confidence of 0.31 for belonging to the in-vehicle, a confidence of 0.36 for belonging to the indoor, and a confidence of 0.33 for belonging to the outdoor. The highest confidence of 0.36 is less than the preset confidence threshold of 0.4, indicating that discriminative information cannot be given in this scene image, and the scene classification model cannot accurately judge the scene category of this scene image, so this scene image is discarded.
[0200] In this embodiment, if there are a total of 100 scene images in the target scene sample sets of multiple target scene categories, the first scene classification result is as follows: 80 scene images are of the in-vehicle type, 10 scene images are of the indoor type, and 10 scene images are of the outdoor type; the second scene classification result is as follows: 45 scene images are of the in-vehicle type, 40 scene images are of the outdoor type, and 15 scene images are of the indoor type. For the second scene classification result, for the 45 scene images of the in-vehicle type, 40 scene images of the outdoor type, and 15 scene images of the indoor type, the scene classification model cannot determine which scene the face review video belongs to.
[0201] In this embodiment, the preset scene category threshold is 0.4. The target scene classification result is determined by calculating the difference between the highest total votes and the second-highest total votes and then dividing by the total number of votes. The target scene classification result is compared with the preset scene category threshold, and the scene classification result of the face review video is determined according to the comparison result. For example, the target scene classification result corresponding to the first scene classification result is: (80 - 10) / 100 = 0.7, and the target scene classification result corresponding to the second scene classification result is: (45 - 40) / 100 = 0.05. It can be seen that 0.7 is greater than 0.4, so it can be determined that the scene classification result of the face review video corresponding to the first scene classification result is the in-vehicle scene; while 0.05 is less than 0.4, indicating that the scene classification result of the face review video cannot be determined, and it is switched to manual review.
[0202] In this embodiment, based on the classification results of each scene image in the scene sample set, the scene classification result of the face review video is determined. By discarding some scene images for which the scene classification cannot be confirmed, the efficiency and accuracy of scene classification are improved.
[0203] The switching module 206 is used to switch to the manual scene classification review system when the face review video does not meet the scene classification environment.
[0204] In this embodiment, when the face review video does not meet the scene classification environment, it is determined that the face review video belongs to a dark environment. In a dark environment, it is difficult for the scene classification subsystem of the face review video to determine the environment where the user is located from the face review video, and it is switched to the manual scene classification review system for manual review, ensuring the accuracy of the scene classification of the face review video.
[0205] In summary, the scenario classification device for the face review video described in this embodiment determines whether the face review video meets the scenario classification environment through the video frame sequence set. When the face review video meets the scenario classification environment, it preprocesses the video frames in the video frame sequence set, deletes similar video frames, and reasonably filters out a large number of redundant video frames. On the one hand, it reduces redundant information, and on the other hand, it can greatly improve the calculation time of the scenario classification model, shorten the time for judging the video scenario, and improve the scenario classification efficiency of the face review video. The scenario image sample set is input into a pre-trained scenario classification model to obtain the classification result of each scenario image. During the training process of the scenario classification model, considerations are made from two aspects: open-source data and face review data, avoiding the phenomenon of overfitting of a small amount of data when using face review data to train the scenario classification model in the prior art, improving the accuracy of the trained scenario classification model, and thus improving the accuracy of scenario classification. According to the classification results of each scenario image in the scenario sample set, the scenario classification result of the face review video is determined, and by discarding some scenario images for which the scenario classification cannot be confirmed, the efficiency and accuracy of scenario classification are improved.
[0206] Embodiment 3
[0207] Refer to Figure 3 As shown, it is a schematic structural diagram of an electronic device provided by Embodiment 3 of the present invention. In a preferred embodiment of the present invention, the electronic device 3 includes a memory 31, at least one processor 32, at least one communication bus 33, and a transceiver 34.
[0208] Those skilled in the art should understand that Figure 3 The structure of the electronic device shown does not constitute a limitation of the embodiments of the present invention. It can be a bus structure or a star structure. The electronic device 3 may further include more or fewer other hardware or software than shown, or different component arrangements.
[0209] In some embodiments, the electronic device 3 is an electronic device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, a microprocessor, an application-specific integrated circuit, a programmable gate array, a digital processor, and an embedded device, etc. The electronic device 3 may further include a client device, and the client device includes, but is not limited to, any electronic product that can perform human-computer interaction with the client through a keyboard, a mouse, a remote control, a touchpad, or a voice control device, etc. For example, a personal computer, a tablet computer, a smart phone, a digital camera, etc.
[0210] It should be noted that the electronic device 3 is only an example, and other existing or future possible electronic products that can be adapted to the present invention should also be included in the protection scope of the present invention and are included herein by reference.
[0211] In some embodiments, the memory 31 is used to store program codes and various data, such as the scene classification device 20 of the face review video installed in the electronic device 3, and can achieve high-speed and automatic access to programs or data during the operation of the electronic device 3. The memory 31 includes a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc memories, a magnetic disk memory, a magnetic tape memory, or any other computer-readable medium capable of carrying or storing data.
[0212] In some embodiments, the at least one processor 32 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions packaged, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The at least one processor 32 is the control core (Control Unit) of the electronic device 3, connects various components of the entire electronic device 3 through various interfaces and lines, and executes various functions of the electronic device 3 and processes data by running or executing programs or modules stored in the memory 31 and calling data stored in the memory 31.
[0213] In some embodiments, the at least one communication bus 33 is configured to enable connection communication between the memory 31 and the at least one processor 32, etc.
[0214] Although not shown, the electronic device 3 may further include a power source (such as a battery) for supplying power to each component. Optionally, the power source may be logically connected to the at least one processor 32 through a power management device, so as to manage functions such as charging, discharging, and power consumption management through the power management device. The power source may further include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device 3 may further include a variety of sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0215] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0216] The integrated units implemented in the form of software function modules as described above may be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium and include several instructions for causing a computer device (which may be a personal computer, an electronic device, or a network device, etc.) or a processor to execute a part of the methods described in various embodiments of the present invention.
[0217] In a further embodiment, in combination with Figure 2 , the at least one processor 32 can execute the operating device of the electronic device 3 and various installed application programs (such as the scene classification device 20 for the face review video), program codes, etc. For example, the above-mentioned various modules.
[0218] Program codes are stored in the memory 31, and the at least one processor 32 can call the program codes stored in the memory 31 to execute related functions. For example, Figure 2 the various modules described in
[0219] are program codes stored in the memory 31 and are executed by the at least one processor 32, so as to implement the functions of the various modules to achieve the purpose of scene classification of the face review video.
[0220] In one embodiment of the present invention, the memory 31 stores a plurality of computer-readable instructions, and the at least one processor 32 executes the plurality of computer-readable instructions to implement the function of scene classification of the face review video.
[0221] Specifically, for the specific implementation method of the at least one processor 32 for the above instructions, reference may be made to Figure 1 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.
[0222] In several embodiments provided by the present invention, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0223] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0224] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a hardware plus software functional module.
[0225] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights. In addition, it is obvious that the word "comprising" does not exclude other units or, the singular does not exclude the plural. The multiple units or devices described in the present invention can also be implemented by one unit or device through software or hardware. First, second, etc. are used to represent names and do not represent any specific order.
[0226] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for scene classification of in-person review videos, characterized in that, the method includes: Upon receiving a scene classification request, obtain an in-person review video; Convert the in-person review video into a set of video frame sequences, and determine whether the in-person review video meets the scene classification environment according to the set of video frame sequences; When the in-person review video meets the scene classification environment, preprocess the video frames in the set of video frame sequences to obtain a set of scene image samples; Input the set of scene image samples into a pre-trained scene classification model to obtain the classification result of each scene image, wherein the training process of the scene classification model includes: obtaining multiple scenes and the corresponding first video set and second video set for each scene; decoding the first video set to obtain a first set of sample images, and decoding the second video set to obtain a second set of sample images; using a face detection algorithm to remove the face images from each sample image in the first set of sample images, and determining the multiple sample images after removing the face images as a third set of sample images; dividing the second set of sample images into a training set and a test set; inputting the training set into a preset neural network for training to obtain a pre-trained model; inputting the test set into the pre-trained model for testing, and calculating the test passing rate; comparing the test passing rate with a preset passing rate threshold; when the test passing rate is greater than or equal to the preset passing rate threshold, determine that the training of the pre-trained model is completed, and based on the third set of sample images, use a preset fine-tuning model to fine-tune the pre-trained model to obtain a scene classification model; According to the classification results of each scene image in the set of scene samples, determine the scene classification result of the in-person review video.
2. The method for scene classification of in-person review videos according to claim 1, characterized in that, the preprocessing of the video frames in the set of video frame sequences to obtain a set of scene image samples includes: Determine the first video frame in the set of video frame sequences as the current video frame; Calculate the similarity between the current video frame and the next video frame of the current video frame; When the similarity between the current video frame and the next video frame of the current video frame is greater than or equal to a preset similarity threshold, delete the next video frame of the current video frame from the set of video frame sequences to obtain a new set of video frame sequences, determine the first video frame in the new set of video frame sequences as the current video frame, and repeat calculating the similarity between the current video frame and the next video frame of the current video frame in the new set of video frame sequences until the similarity between the first video frame and the last video frame in the new set of video sequences is calculated to obtain a set of scene image samples; When the similarity between the current video frame and the next video frame of the current video frame is less than the preset similarity threshold, the next video frame of the current video frame is determined as the new current video frame, and the similarity between the new current video frame and the next video frame of the new current video frame is recalculated until the similarity between the new current video frame and the last video frame in the video sequence set is calculated to obtain a set of scene image samples.
3. The method for classifying the scene of the face review video according to claim 2, wherein, the calculation of the similarity between the current video frame and the next video frame of the current video frame includes: calculating the difference between the pixels at each position of the current video frame and the pixels at the corresponding positions of the next video frame of the current video frame, averaging the differences of the pixels at all positions to obtain the target mean value of the pixels; calculating the quotient of the target mean value of the pixels and the total number of pixels of the current video frame, and determining the calculated quotient as the similarity between the current video frame and the next video frame of the current video frame.
4. The method for classifying the scene of the face review video according to claim 1, wherein, the determination of the scene classification result of the face review video according to the classification results of each scene image in the scene sample set includes: obtaining the highest confidence level of each scene image from the classification results of each scene image, and comparing the highest confidence level with a preset confidence level threshold; retaining multiple scene images corresponding to the highest confidence level greater than or equal to the preset confidence level threshold, and determining the category corresponding to the highest confidence level of each scene image as the target scene category of the corresponding scene image; classifying the retained multiple scene images according to the target scene category to obtain a target scene sample set for each target scene category; statistically calculating the total votes of each target scene category based on the target scene sample set of each target scene category; calculating the quotient of the difference between the highest total votes and the second highest total votes among the multiple total votes of the multiple target scene categories and the sum of the multiple total votes to obtain a target scene category result; comparing the target scene category result with a preset scene category threshold; when the target scene category result is greater than or equal to the preset scene category threshold, determining the scene category with the highest total votes as the scene classification result of the face review video.
5. The method for classifying the scene of the face review video according to claim 1, wherein, the determination of whether the face review video meets the scene classification environment according to the video frame sequence set includes: converting each video frame in the video frame sequence set into an HSV image to obtain an HSV image set; removing the pixel regions of human faces in each HSV image in the HSV image set, retaining the pixels in the non-human face regions of each HSV image, and calculating the target brightness value of each HSV image based on the retained pixels in the non-human face regions of each HSV image; comparing the target brightness value of each HSV image with a preset brightness threshold; Count the total number of HSV images whose target brightness value is less than the preset brightness threshold to obtain the first total count; Obtain a target total count threshold based on the second total count of the images in the HSV image set; Compare the first total count with the target total count threshold; When the first total count is less than the target total count threshold, determine that the in-person review video meets the scene classification environment.
6. The scene classification method for in-person review videos according to any one of claims 1 to 5, wherein, the method further includes: When the in-person review video does not meet the scene classification environment, switch to the manual scene classification review system.
7. A scene classification device for in-person review videos, wherein, the device is used to implement the scene classification method for in-person review videos according to any one of claims 1 to 6, and the device includes: An acquisition module, configured to acquire an in-person review video in response to a received scene classification request; A judgment module, configured to convert the in-person review video into a set of video frame sequences, and judge whether the in-person review video meets the scene classification environment according to the set of video frame sequences; A preprocessing module, configured to preprocess the video frames in the set of video frame sequences to obtain a set of scene image samples when the in-person review video meets the scene classification environment; An input module, configured to input the set of scene image samples into a pre-trained scene classification model to obtain the classification result of each scene image; A determination module, configured to determine the scene classification result of the in-person review video according to the classification results of each scene image in the set of scene samples.
8. An electronic device, wherein, the electronic device includes a processor and a memory, and the processor is configured to implement the scene classification method for in-person review videos according to any one of claims 1 to 6 when executing a computer program stored in the memory.
9. A computer-readable storage medium, on which a computer program is stored, wherein, the computer program, when executed by a processor, implements the scene classification method for in-person review videos according to any one of claims 1 to 6.
Citation Information
Patent Citations
Scene classification method and device for face-to-face video,equipment and storage medium
CN113221835A