Method, System and Storage Medium for Detecting the Appearance of Bystanders in In-Person Interview Videos
By using pre-trained object detection model and feature point offset vector calculation methods in the face-to-face review video, the situation of people entering the camera is accurately identified, and the problems of difficulty and untimely detection in the prior art are solved, and efficient face-to-face review auxiliary information is achieved.
Patent Information
- Application Number
- CN202210826137.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-14
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-07-14
AI Technical Summary
There are problems in existing face-to-face review videos that are difficult to detect and are not timely, making it difficult to accurately identify the situation where others are in the camera.
By obtaining the face-to-face video frame, the target object in the video frame is detected using the pre-trained object detection model, and the target box is output. At the same time, calculate the characteristic point offset vector of the target box to obtain the confidence. When the confidence is less than the preset threshold, it is determined that there is someone else.
It realizes accurate identification of the situation of others entering the camera, provides effective face-to-face audit auxiliary information, and reduces face-to-face audit risk.
Smart Images

Figure CN115100571B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, system, and computer-readable storage medium for detecting the entry of bystanders into a video based on an in-person interview video. Background Art
[0002] In the field of computer vision, video image processing technology plays a very important role in the in-person interview for bank loan business. An in-person interview means that a loan applicant needs to answer questions from a staff member via a real-time video. At the end of the in-person interview, the in-person interview assistance system will predict whether the applicant has fraudulent behavior based on the applicant's performance in the video. During the in-person interview, for risk prevention and control purposes, it is necessary to ensure that the in-person interview object is alone, and the risk of a bystander answering questions on behalf of the in-person interview object should be eliminated. When a second person other than the in-person interview object is detected, the in-person interview process needs to be terminated. However, the existing in-person interview procedures have the following drawbacks in detecting the entry of bystanders:
[0003] 1) When a bystander enters the frame, the in-person interview object generally occupies the main area of the frame, while the bystander only shows a small part of the body, and the body of the bystander may be largely blocked by the in-person interview object; this makes the detection difficult;
[0004] 2) The entry of a bystander does not occur throughout the entire in-person interview process, nor do the bystander and the in-person interview object appear in the in-person interview camera at the same time; instead, there is only a frame switch at certain times, and the bystander replaces the in-person interview object in the in-person interview camera. Just detecting the number of people in the camera cannot timely discover that the in-person interview object has been replaced;
[0005] Therefore, there is an urgent need for a highly practical method for detecting the entry of bystanders into a video based on an in-person interview video. Summary of the Invention
[0006] The present invention provides a method, system, electronic device, and storage medium for detecting the entry of bystanders into a video based on an in-person interview video, and its main purpose is to solve at least one problem existing in the prior art.
[0007] To achieve the above object, a method for detecting the entry of bystanders into a video based on an in-person interview video provided by the present invention, which is applied to an electronic device, includes:
[0008] Obtain the in-person interview video of the target object, and extract multiple video frames in which the target object answers the in-person interview questions from the in-person interview video;
[0009] Detect the target object in the video frame through a pre-trained target detection model, and output the target box of the target object corresponding to each video frame; wherein, the detection target of the target box is a person;
[0010] Simultaneously obtain the feature points of each target box, pair the target boxes in pairs, and calculate the feature point offset vectors of each pair of target boxes;
[0011] Obtain the confidence according to the feature point offset vector. When the confidence is less than the preset confidence threshold, it is determined that there are bystanders in the face review video of the target object.
[0012] Furthermore, preferably, the target detection model is obtained by training with a training set enhanced by masoic data. Among them, the method for enhancing the masoic data of the training set includes:
[0013] Randomly obtain 4 training pictures with a width of W and a height of H in the training set, and create a blank picture with a width of 2W and a height of 2H;
[0014] Randomly sample 1 separation point in the central area of the blank picture; according to the separation point, divide the blank picture into four areas: upper left, lower right, lower left, and upper right;
[0015] Insert the 4 training pictures into the four areas of the blank picture;
[0016] Scale the blank picture to a size of width W and height H;
[0017] Traverse all the training pictures in the training set to complete the masoic data enhancement of the training set.
[0018] Furthermore, preferably, the method for training the target detection model with a training set enhanced by masoic data includes:
[0019] Perform face annotation and hand annotation on the training pictures in the training set enhanced by masoic data;
[0020] Input the training pictures into the neural network; the neural network respectively extracts face features and hand features from the training pictures;
[0021] Respectively obtain the predicted confidence for the extracted face features and hand features; and respectively obtain the first loss function and the second loss function according to the predicted confidence and the true confidence;
[0022] Adjust the parameters of the neural network according to the first loss function and the second loss function, so that the trained target detection model can simultaneously identify face feature labels and hand feature labels.
[0023] Furthermore, preferably, after obtaining the target box of the target object corresponding to each video frame, it also includes;
[0024] Detect the target box and the sample box of the target object obtained in advance, and determine the confidence level; wherein, the sample box of the target object obtained in advance is obtained by detecting a face image containing the target object using a pre-trained target detection model.
[0025] When the confidence level is less than a preset confidence threshold, it is determined that there are other people in the face review video of the target object.
[0026] Furthermore, preferably, the training set further includes training pictures containing in-vehicle scenes and training pictures containing human hands.
[0027] Furthermore, preferably, obtain video frames that determine that there are other people in the face review video of the target object.
[0028] Form a video segment of other people appearing on camera with the video frames where there are other people according to the frame numbers of the video frames.
[0029] Furthermore, preferably, the target detection model is any one of the DenseBox algorithm, YOLO algorithm, CornerNet algorithm, ExtremeNet algorithm, FCOS algorithm, and FoveaBox algorithm.
[0030] To solve the above problems, the present invention also provides a system for detecting the entry of other people into the frame based on a face review video, including:
[0031] A data acquisition unit for acquiring a face review video of a target object and extracting multiple video frames in which the target object answers face review questions from the face review video.
[0032] A target box acquisition unit for detecting the target object in the video frame through a pre-trained target detection model and outputting the target box of the target object corresponding to each video frame; wherein, the detection target of the target box is a person.
[0033] A determination unit for simultaneously acquiring the feature points of each target box, pairing the target boxes in pairs, and calculating the feature point offset vectors of each pair of target boxes; obtaining the confidence level according to the feature point offset vectors, and when the confidence level is less than a preset confidence threshold, determining that there are other people in the face review video of the target object.
[0034] To solve the above problems, the present invention also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps in the foregoing method for detecting the entry of other people into the frame based on a face review video.
[0035] To solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is executed by a processor in an electronic device to implement the above-mentioned method for detecting the intrusion of bystanders based on the face review video.
[0036] The above-mentioned method, system and storage medium for detecting the intrusion of bystanders based on the face review video provided by the present invention obtain the face review video of the target object, and intercept multiple video frames of the target object answering the face review questions from the face review video; detect the target object in the video frames through a pre-trained target detection model, and output the target box of the target object corresponding to each video frame; wherein, the detection target of the target box is a person; at the same time, obtain the feature points of each target box, pair the target boxes in pairs, and calculate the feature point offset vector of each pair of target boxes; obtain the confidence according to the feature point offset vector, and when the confidence is less than the preset confidence threshold, it is determined that there are bystanders in the face review video of the target object. The present invention accurately identifies the situation where a bystander replaces the target object or a bystander and the target object enter the frame together, and then provides effective face review auxiliary information for the face review personnel, reducing the face review risk. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 FIG. is a schematic flow chart of a method for detecting the intrusion of bystanders based on the face review video according to an embodiment of the present invention;
[0038] Figure 2 FIG. is a logical structure block diagram of a system for detecting the intrusion of bystanders based on the face review video according to an embodiment of the present invention;
[0039] Figure 3 FIG. is an internal structure schematic diagram of an electronic device for implementing the method for detecting the intrusion of bystanders based on the face review video according to an embodiment of the present invention.
[0040] The realization, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0041] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0042] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is a theory, method, technology and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0043] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, mechatronics, etc. The artificial intelligence software technology in this application is a machine learning technology based on convolutional neural networks. Based on convolutional neural networks, it can be applied to a variety of different fields, such as speech recognition, medical diagnosis, testing of application programs, etc.
[0044] By obtaining the face review video of the target object, multiple video frames in which the target object answers the face review questions are intercepted from the face review video; the target object in the video frames is detected by a pre-trained target detection model, and the target box of the target object corresponding to each video frame is output; wherein, the detection target of the target box is a person; at the same time, the feature points of each target box are obtained, the target boxes are paired in pairs, and the feature point offset vector of each pair of target boxes is calculated; a confidence level is obtained according to the feature point offset vector, and when the confidence level is less than a preset confidence level threshold, it is determined that there is a bystander in the face review video of the target object. The present invention accurately identifies the situation where a bystander replaces the target object to enter the camera or the bystander enters the camera together with the target object, and further provides effective face review auxiliary information for the face review personnel, reducing the face review risk.
[0045] Specifically, as an example, Figure 1 is a schematic flowchart of a method for detecting bystander entry into the camera based on a face review video provided by an embodiment of the present invention. Refer to Figure 1 As shown, the present invention provides a method for detecting bystander entry into the camera based on a face review video. This method can be executed by a device, and the device can be implemented by software and / or hardware.
[0046] In this embodiment, the method for detecting bystander entry into the camera based on a face review video includes: steps S110 to S150.
[0047] S110. By obtaining the face review video of the target object, multiple video frames in which the target object answers the face review questions are intercepted from the face review video.
[0048] Obtaining the face review video can be realized by computer vision (image) (Computer Vision, CV) technology. Computer vision is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to perform machine vision such as identifying, tracking, and measuring targets, and further performing graphic processing to make the computer process into an image that is more suitable for human eyes to observe or transmit to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes technologies such as image processing, image recognition, image semantic understanding, and image retrieval.
[0049] Specifically, a face review video of a user answering face review questions online is obtained by using a camera, a computer, a PAD, or the like.
[0050] S120. Detect the target object in the video frame through a pre-trained object detection model, and output the target box of the target object corresponding to each video frame; wherein, the detection target of the target box is a person.
[0051] In a specific embodiment, the object detection model is obtained by training with a training set after masoic data augmentation. The method for masoic data augmentation of the training set includes: S1211. Randomly obtain 4 training pictures with a width of W and a height of H in the training set, and create a blank picture with a width of 2W and a height of 2H; S1212. Randomly sample 1 separation point in the central area of the blank picture; according to the separation point, divide the blank picture into four areas: upper left, lower right, lower left, and upper right; S1213. Insert the 4 training pictures into the four areas of the blank picture; S1214. Scale the blank picture to a size of width W and height H; S1215. Traverse all the training pictures in the training set to complete the masoic data augmentation of the training set.
[0052] It should be noted that the conventional Mosaic data augmentation method simply means splicing 4 pictures by means of random scaling, random cropping, and random arrangement. However, there are some problems with the conventional masoic operation. This operation will splice 4 pictures together, and each picture will be cropped before splicing. This cropping will bring some problems. Generally, only part of the body of the bystander in the picture will be shown. Once cropped, the exposed body part may be even less, and it will no longer show that it is a person, but the detection box is still there, so it will lead to mislabeling. Therefore, during training, the cropping amplitude of the masoic operation is extremely small, and the cropping operation can even be removed. Specifically, the above masoic operation refers to splicing multiple different pictures into one picture. The width of the final picture is W, and the height is H. The process is as follows: First, randomly sample 4 pictures from the training pictures and create a blank picture to place the four pictures. The width of the blank picture is 2*W, and the height is 2*H. Then, randomly sample a dividing point in the blank picture to insert these four pictures. In order to make the cropping amplitude of the masoic extremely small, the dividing point needs to be as close to the center of the picture as possible, so the sampling area of the dividing point is limited to the central area of the blank picture. Secondly, after sampling the dividing point, the dividing point will divide the blank picture into four areas: upper left, upper right, lower left, and lower right. Insert the 4 training pictures randomly sampled before into the four areas of the blank picture. For each picture, scale the picture to the size of the corresponding blank area while maintaining the aspect ratio of the picture, and then paste it in the blank area. Finally, scale the blank picture to the size of height H and width W. By performing Mosaic data augmentation on the training set, randomly using 4 pictures, randomly scaling, and then randomly distributing and splicing, the detection data set is greatly enriched. In particular, the random scaling adds many small targets, making the network more robust. And it indirectly reduces the hardware requirements for the GPU.
[0053] It should be noted that for the preprocessing of the training set, in addition to Mosaic data augmentation, it can also include adaptive image padding and the integration of adaptive anchor box calculation at the Input input end to automatically set the initial anchor box size when replacing the data set; the specific preprocessing means are selected according to the actual application scenario and will not be specifically limited here.
[0054] In a specific embodiment, the method for training the object detection model through the training set after Mosaic data augmentation includes: S1210. Perform face annotation and hand annotation on the training pictures in the training set after Mosaic data augmentation.
[0055] It should be noted that image data is prepared, the image data is labeled to mark the target data, and the data set is divided into a training set, a validation set, and a test set. The training set is used to train the model, the validation set is used to verify the performance of the model after each Epoch, the data results of each training are recorded and plotted, and information such as accuracy, recall rate, and overfitting is obtained from the chart. The test set is used to test the performance of the final model. The image data here can be obtained from web crawling. The type of the image data is mainly that one person occupies the main picture of the image, and another person is beside this person, and only a part of the body may be exposed. As an improvement of this embodiment, in order to adapt to the scenario where the target object is in the vehicle and the vehicle seat is easily misjudged as a bystander, and the hand at a certain edge position of the picture is easily misjudged as a bystander, the training pictures containing the vehicle interior scene and the training pictures containing the hand are included in the training set.
[0056] S1220. Input the training pictures into the neural network; the neural network extracts face features and hand features from the training pictures respectively; S1230. Obtain the prediction confidence levels for the extracted face features and hand features respectively; and obtain the first loss function and the second loss function according to the prediction confidence levels and the true confidence levels respectively; S1240. Adjust the parameters of the neural network according to the first loss function and the second loss function so that the trained target detection model can identify face feature labels and hand feature labels simultaneously. Input the training data into the network to train the model. Supervise the parameters during the training process until the network reaches the optimum, and record the network parameters at this time.
[0057] That is to say, for the situation where the hand is misjudged as a bystander, in addition to the category of people, another detection category, that is, the hand, is added. So that the trained target detection model can identify the hand well and will not misdetect the hand and people. The specific implementation method of adding the detection category of the hand includes, first, in the existing training pictures, mark the hands in the pictures as training labels. Then, the detection category is increased by the category of the hand, people are marked as category 0, and hands are marked as category 1.
[0058] In a specific implementation process, the implementation processes of the two detection categories can be as follows: in the depth convolution and point convolution such as the Mobilenet network can be used to detect features, and then input to the output layer to obtain the first prediction confidence of the specified image category to which the image scene classification belongs. Then, the difference between the first prediction confidence and the first true confidence is calculated to obtain the first loss function; in the target detection network layer, the SSD network can be used. After the convolutional layers of the first 5 layers of VGG16, a convolutional feature layer is cascaded. In the convolutional feature layer, a set of convolutional filters are used to predict the offset parameters of the preselected default bounding box corresponding to the specified target category relative to the true bounding box and the second prediction confidence corresponding to the specified target category. The region of interest is the region of the preselected default bounding box. A position loss function is constructed according to the offset parameters, and a second loss function is obtained according to the difference between the second prediction confidence and the second true confidence. The first loss function, the second loss function, and the position loss function are weighted and summed to obtain the target loss function. According to the target loss function, the parameters of the neural network are adjusted by using the backpropagation algorithm to train the neural network.
[0059] S130. At the same time, obtain the feature points of each target box, pair the target boxes in pairs, and calculate the feature point offset vectors of each pair of target boxes.
[0060] It should be noted that the feature points of each video frame in the video frame sequence are obtained by analyzing the feature points of the acquired target boxes. The feature points refer to the points in the video frame that have distinct characteristics, can effectively reflect the essential characteristics of the video frame, and can identify the target object (person) in the video frame. Through the matching of the feature points, the matching of the target object can be completed, that is, the target object is identified and classified.
[0061] Specifically, after the model is trained, a relatively good target detection model (human body detection model) is obtained. Input an image to obtain the detection results of the people in the image and give the confidence of each detection result. If the confidence is greater than a threshold, it is considered that this detection result is a person. Since the present invention hopes to identify the situation where a bystander enters the camera during the switching of characters in the video screen, it is necessary to retain the features corresponding to this detection result during detection and judge whether it is the same person in the screen based on the features. The specific process is as follows: the trained feature extraction network extracts the features of the image, and the features F of the extracted image, the size of the features F is W*H*C. After the inference process of the detection head, a detection result O with a size of W*H*K can be obtained. W and H are the width and height of the features and the detection result respectively, and K represents the dimension of the detection result at each position. If one of the final obtained object boxes is given by the w and h positions of the detection result O, where (0≤w<W, 0≤h<H), then the feature f corresponding to the detection box is F[w,h], and the dimension of the feature is C. All the features corresponding to the detected object boxes are retained.
[0062] In a specific embodiment, in order to further improve the accuracy of in-person review detection, a comparison detection between the target box and the sample box is added; it is implemented through the following steps: S131. Detect the target box and the sample box of the pre-acquired target object, and determine the confidence level; wherein, the sample box of the pre-acquired target object is obtained after detecting a face image containing the target object by using a pre-trained target detection model; S132. When the confidence level is less than a preset confidence threshold, it is determined that there is someone else in the in-person review video of the target object. That is to say, it is necessary to determine the facial features of the target object under normal conditions, that is, before the in-person review, it is necessary to collect standard video information of the target user. In a specific embodiment, a video information collection stage with a fixed duration can be set before the in-person review program starts to collect static facial feature information of the target object.
[0063] S140. Obtain the confidence level according to the feature point offset vector. When the confidence level is less than a preset confidence threshold, it is determined that there is someone else in the in-person review video of the target object.
[0064] Specifically, the features corresponding to each detection box need to be used. Since the features of the same person are generally similar, this can be used to distinguish whether different people appear in the video. Retain the specific object box and the corresponding features for all frames. Here, the specific object box refers to the object detected by the box being a person, and the area of the box is the largest in this frame of the picture and exceeds half of the picture area. Traverse all frames in sequence. For example, for the i-th frame, assume the retained feature is fi. Calculate the cosine distance between this feature and the features retained in all other frames. If it is greater than a pre-set threshold T, it is considered that a new person appears in the camera.
[0065] It should be noted that the above object box and detection box are both target boxes.
[0066] In the specific implementation process, the appearance set of the person other than the target can also be obtained, which is implemented through the following steps: S141. Obtain the video frames that determine that there is someone else in the in-person review video of the target object; S142. Compose the video frames with someone else into a segment of the appearance of the person other than the target according to the frame numbers of the video frames. Specifically, according to the detection process of the entry of the person other than the target, it can be obtained whether someone else appears in each frame of a video, and the frame numbers of all frames in the video where the person other than the target appears are recorded. For example, assume a video has 1000 frames, and record that [15, 16, 17, 299, 300, 301, 555, 556, 557, 558] have someone else appear. Based on the recorded frame numbers, the final segment of the appearance of the person other than the target is obtained.
[0067] In a specific implementation process, the object detection model is any one of the DenseBox algorithm, the YOLO algorithm, the CornerNet algorithm, the ExtremeNet algorithm, the FCOS algorithm, and the FoveaBox algorithm.
[0068] In a scenario where the detection categories include two categories of human figures and human hands, the classification network layer includes two output layers. In the first output layer, the confidence levels of each specified scenario category to which the human figure belongs are output through a softmax classifier, and the human figure with the highest confidence level exceeding the confidence level threshold is selected as the human figure classification label for the face review video of the target object. The features of the extracted human figure are input into the object detection network layer for human figure detection. In the second output layer, the confidence level and corresponding position of the specified target category to which the human hand belongs are output through a softmax classifier, and the target category with the highest confidence level exceeding the confidence level threshold is selected as the human hand classification label for the face review video of the target object, and the position corresponding to the target classification label is output. The obtained human figure classification label and human hand classification label are used as the classification labels for object detection.
[0069] In summary, the method for detecting the intrusion of bystanders based on a face review video according to the present invention obtains the face review video of the target object, and extracts multiple video frames in which the target object answers the face review questions from the face review video; detects the target object in the video frames through a pre-trained object detection model, and outputs the target box of the target object corresponding to each video frame; wherein the detection target of the target box is a person; at the same time, the feature points of each target box are obtained, the target boxes are paired in pairs, and the feature point offset vector of each pair of target boxes is calculated; a confidence level is obtained according to the feature point offset vector, and when the confidence level is less than a preset confidence level threshold, it is determined that there is a bystander in the face review video of the target object. The present invention accurately identifies the situation where a bystander replaces the target object to enter the frame or a bystander enters the frame together with the target object, and further provides effective face review auxiliary information for the face review personnel, reducing the face review risk.
[0070] Corresponding to the above method for detecting the intrusion of bystanders based on a face review video, the present invention also provides a method for detecting the intrusion of bystanders based on a face review video. Figure 2 The functional modules of the system for detecting the intrusion of bystanders based on a face review video according to an embodiment of the present invention are shown.
[0071] As Figure 2As shown, the bystander intrusion detection system 200 based on the face review video provided by the present invention can be installed in an electronic device. According to the implemented functions, the bystander intrusion detection system 200 based on the face review video may include a data acquisition unit 210, a target box acquisition unit 220, and a determination unit 230. The units in the present invention may also be referred to as modules, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete a certain fixed function, and are stored in the memory of the electronic device.
[0072] In this embodiment, the functions of each module / unit are as follows:
[0073] The data acquisition unit 210 is configured to acquire the face review video of the target object, and intercept multiple video frames of the target object answering the face review questions from the face review video;
[0074] The target box acquisition unit 220 is configured to detect the target object in the video frame through a pre-trained target detection model, and output the target box of the target object corresponding to each video frame; wherein, the detection target of the target box is a person;
[0075] The determination unit 230 is configured to simultaneously acquire the feature points of each target box, pair the target boxes in pairs, and calculate the feature point offset vector of each pair of target boxes; obtain the confidence according to the feature point offset vector, and when the confidence is less than a preset confidence threshold, determine that there is a bystander in the face review video of the target object.
[0076] The target box acquisition unit 220 includes a target detection model training module. The target detection model training module is configured to perform face annotation and hand annotation on the training pictures in the training set after masoic data augmentation; input the training pictures into a neural network; the neural network performs face feature extraction and hand feature extraction on the training pictures respectively; obtain the prediction confidence of the extracted face features and hand features respectively; and obtain the first loss function and the second loss function according to the prediction confidence and the true confidence respectively; adjust the parameters of the neural network according to the first loss function and the second loss function, so that the trained target detection model can simultaneously identify the face feature label and the hand feature label.
[0077] The target detection model training module includes a training set data augmentation sub-module, which is used to randomly obtain 4 training pictures with a width of W and a height of H in the training set, and create a blank picture with a width of 2W and a height of 2H; randomly sample a dividing point in the central area of the blank picture; according to the dividing point, divide the blank picture into four regions: upper left, lower right, lower left, and upper right; insert the 4 training pictures into the four regions of the blank picture; scale the blank picture to a size of width W and height H; traverse all the training pictures in the training set to complete the masoic data augmentation of the training set.
[0078] For more specific implementation manners of the above-mentioned method for detecting the intrusion of a bystander into a face review video provided by the present invention, reference can be made to the description of the embodiments of the method for detecting the intrusion of a bystander into a face review video above, and they will not be enumerated one by one here.
[0079] As can be seen from the above embodiments, the system for detecting the intrusion of a bystander into a face review video proposed by the present invention obtains the face review video of the target object, and extracts multiple video frames in which the target object answers the face review questions from the face review video; detects the target object in the video frames through a pre-trained target detection model, and outputs the target box of the target object corresponding to each video frame; wherein, the detection target of the target box is a person; at the same time, obtains the feature points of each target box, pairs the target boxes in pairs, and calculates the feature point offset vector of each pair of target boxes; obtains the confidence according to the feature point offset vector, and when the confidence is less than a preset confidence threshold, determines that there is a bystander in the face review video of the target object. The present invention accurately identifies the situation where a bystander replaces the target object to enter the frame or the bystander and the target object enter the frame together, thereby providing effective face review auxiliary information for the face review personnel and reducing the face review risk.
[0080] As Figure 3 shown, the present invention provides an electronic device 3 for a method of detecting the intrusion of a bystander into a face review video.
[0081] The electronic device 3 may include a processor 30, a memory 31, and a bus, and may further include a computer program stored in the memory 31 and executable on the processor 30, such as a program 32 for detecting the intrusion of a bystander into a face review video.
[0082] Among them, the memory 31 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 31 can be an internal storage unit of the electronic device 3 in some embodiments, such as the mobile hard disk of the electronic device 3. The memory 31 can also be an external storage device of the electronic device 3 in some other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 3. Further, the memory 31 can also include both the internal storage unit and the external storage device of the electronic device 3. The memory 31 can be used not only to store application software installed in the electronic device 3 and various types of data, such as the code of the bystander intrusion detection program based on the face review video, etc., but also to temporarily store the data that has been output or will be output.
[0083] The processor 30 can be composed of integrated circuits in some embodiments. For example, it can be composed of a single packaged integrated circuit, or can also be composed of multiple integrated circuits with the same or different functions packaged together, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 30 is the control core (Control Unit) of the electronic device, connecting all components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules stored in the memory 31 (such as the bystander intrusion detection program based on the face review video, etc.), and calling the data stored in the memory 31, to execute various functions of the electronic device 3 and process data.
[0084] The bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is set to achieve connection and communication between the memory 31 and at least one processor 30, etc.
[0085] Figure 3 Only the electronic device with components is shown. Those skilled in the art can understand that Figure 3The structures shown do not constitute a limitation on the electronic device 3, and it may include fewer or more components than those shown, or combine certain components, or have different component arrangements.
[0086] For example, although not shown, the electronic device 3 may further include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to the at least one processor 30 through a power management device, so as to implement functions such as charge management, discharge management, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device 3 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0087] Furthermore, the electronic device 3 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 3 and other electronic devices.
[0088] Optionally, the electronic device 3 may further include a user interface. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 3 and to display a visual user interface.
[0089] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.
[0090] The in-person interview video-based bystander intrusion detection program 32 stored in the memory 31 of the electronic device 3 is a combination of multiple instructions. When running in the processor 30, it can achieve: obtaining the in-person interview video of the target object, and extracting multiple video frames of the target object answering the in-person interview questions from the in-person interview video; detecting the target object in the video frames through a pre-trained target detection model, and outputting the target box of the target object corresponding to each video frame; wherein, the detection target of the target box is a person; at the same time, obtaining the feature points of each target box, pairing the target boxes in pairs, and calculating the feature point offset vector of each pair of target boxes; obtaining a confidence level according to the feature point offset vector, and when the confidence level is less than a preset confidence level threshold, determining that there is a bystander in the in-person interview video of the target object.
[0091] Specifically, for the specific implementation method of the above instructions by the processor 30, reference can be made to Figure 1 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here. It should be emphasized that to further ensure the privacy and security of the above in-person interview video-based bystander intrusion detection program, the above in-person interview video-based bystander intrusion detection program is stored in the nodes of the blockchain where this server cluster is located.
[0092] Furthermore, if the module / unit integrated in the electronic device 3 is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).
[0093] The embodiment of the present invention also provides a computer-readable storage medium. The storage medium may be non-volatile or volatile. The storage medium stores a computer program, and when the computer program is executed by a processor, it can achieve: obtaining the in-person interview video of the target object, and extracting multiple video frames of the target object answering the in-person interview questions from the in-person interview video; detecting the target object in the video frames through a pre-trained target detection model, and outputting the target box of the target object corresponding to each video frame; wherein, the detection target of the target box is a person; at the same time, obtaining the feature points of each target box, pairing the target boxes in pairs, and calculating the feature point offset vector of each pair of target boxes; obtaining a confidence level according to the feature point offset vector, and when the confidence level is less than a preset confidence level threshold, determining that there is a bystander in the in-person interview video of the target object.
[0094] Specifically, for the specific implementation method when the computer program is executed by the processor, reference can be made to the description of the relevant steps in the embodiment of the in-person interview video-based bystander intrusion detection method, which will not be elaborated here.
[0095] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0096] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0097] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional modules.
[0098] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.
[0099] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any associated drawing marks in the claims should not be regarded as limiting the claimed rights.
[0100] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc. The blockchain can store medical data, such as personal health records, kitchen, inspection reports, etc.
[0101] In addition, obviously, the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or apparatuses stated in the system claims can also be implemented by one unit or apparatus through software or hardware. Words such as second are used to represent names and do not represent any specific order.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for detecting the entry of bystanders into the camera in a face-to-face interview video, which is applied to an electronic device, characterized in that, the method includes: Obtain the face-to-face interview video of the target object, and extract multiple video frames of the target object answering the face-to-face interview questions from the face-to-face interview video; Detect the target object in the video frame through a pre-trained object detection model, and output the target box of the target object corresponding to each video frame; wherein, the detection target of the target box is a person; the object detection model is obtained by training with a training set enhanced by masoic data augmentation; the training set also includes training pictures containing in-vehicle scenes and training pictures containing human hands; after the training set is enhanced by masoic data augmentation, it also undergoes preprocessing of adaptive image filling and adaptive anchor box calculation; At the same time, obtain the feature points of each target box, pair the target boxes in pairs, and calculate the feature point offset vector of each pair of target boxes; Obtain the confidence according to the feature point offset vector, and when the confidence is less than the preset confidence threshold, it is determined that there are bystanders in the face-to-face interview video of the target object.
2. The method for detecting the entry of bystanders into the camera in a face-to-face interview video according to claim 1, characterized in that, The method for augmenting the masoic data of the training set includes: Randomly obtain 4 training pictures with a width of W and a height of H in the training set, and create a blank picture with a width of 2W and a height of 2H; Randomly sample 1 separation point in the central area of the blank picture; according to the separation point, divide the blank picture into four areas: upper left, lower right, lower left, and upper right; Insert the 4 training pictures into the four areas of the blank picture; Scale the blank picture to a size of width W and height H; Traverse all the training pictures in the training set to complete the masoic data augmentation of the training set.
3. The method for detecting the entry of bystanders into the camera in a face-to-face interview video according to claim 2, characterized in that, The method for training the object detection model with the training set enhanced by masoic data augmentation includes: Perform face annotation and human hand annotation on the training pictures in the training set enhanced by masoic data augmentation; Input the training pictures into the neural network; the neural network extracts face features and human hand features from the training pictures respectively; Obtain the predicted confidence for the extracted face features and human hand features respectively; and obtain the first loss function and the second loss function according to the predicted confidence and the true confidence respectively; Adjust the parameters of the neural network according to the first loss function and the second loss function, so that the trained object detection model can identify face feature labels and human hand feature labels simultaneously.
4. The method for detecting the entry of bystanders into the camera in a face-to-face interview video according to claim 1, characterized in that, After obtaining the target box of the target object corresponding to each video frame, it further includes; Detect the target box with the pre-obtained sample box of the target object, and determine the confidence; wherein, the pre-obtained sample box of the target object is obtained by detecting the face picture containing the target object with a pre-trained object detection model. When the confidence level is less than a preset confidence threshold, it is determined that there are bystanders in the face review video of the target object.
5. The method for detecting bystanders entering the frame based on the face review video according to claim 1, characterized in that, obtain video frames in the face review video of the determined target object where there are bystanders; compose the video frames with bystanders into a bystander appearance segment according to the frame numbers of the video frames.
6. The method for detecting bystanders entering the frame based on the face review video according to claim 5, characterized in that, the target detection model is any one of the DenseBox algorithm, YOLO algorithm, CornerNet algorithm, ExtremeNet algorithm, FCOS algorithm and FoveaBox algorithm.
7. A system for detecting bystanders entering the frame based on the face review video, characterized in that, comprising: a data acquisition unit, configured to acquire the face review video of the target object, and intercept multiple video frames of the target object answering the face review questions from the face review video; a target box acquisition unit, configured to detect the target object in the video frame through a pre-trained target detection model, and output the target box of the target object corresponding to each video frame; wherein, the detection target of the target box is a person; the target detection model is obtained by training with a training set after masoic data augmentation; the training set also includes training pictures containing in-vehicle scenes and training pictures containing human hands; after the training set is subjected to masoic data augmentation, it also undergoes preprocessing of adaptive image filling and adaptive anchor box calculation; a determination unit, configured to simultaneously obtain the feature points of each target box, pair the target boxes in pairs, and calculate the feature point offset vector of each pair of target boxes; obtain the confidence level according to the feature point offset vector, and when the confidence level is less than a preset confidence threshold, determine that there are bystanders in the face review video of the target object.
8. An electronic device, characterized in that, the electronic device includes: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps in the method for detecting bystanders entering the frame based on the face review video according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, the computer program, when executed by a processor, implements the method for detecting bystanders entering the frame based on the face review video according to any one of claims 1 to 6.
Citation Information
Patent Citations
Video face examination auxiliary method, device and equipment and storage medium
CN112541474A
KR20220053670A