Method for detecting the presence of bystanders based on in-person interview videos and related devices
By applying human body detection model and motion direction analysis in the face-to-face video, the problem of being unable to accurately detect the presence of people in the face-to-face video in the previous technology is solved, and more efficient and accurate detection results are achieved.
Patent Information
- Application Number
- CN202210956394.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-10
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-08-10
AI Technical Summary
The prior art cannot accurately detect whether there are people present in the face-to-face video, especially when user A and user B do not appear on a certain frame at the same time.
Multiple video frames of the face-to-face video review based on the human body detection model are used to detect, and the number of human body detection boxes in each video frame is counted. If the number of frames is less than the preset number, a specific number of image sets are extracted, the movement information of the detection frame is calculated, the direction of motion is identified, and the video detection results are generated to determine whether there are people.
It improves the accuracy of the on-site detection of people in the face-to-face video, and can identify the phenomenon of multiple different users in different video frames, which enhances the efficiency and accuracy of the generation of detection results.
Smart Images

Figure CN115331144B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method for detecting the presence of bystanders based on an in-person interview video and related devices. Background Art
[0002] In the fields of anti-fraud, risk prevention and control, etc., in order to ensure the validity of the answers of the interviewees, it is usually required that the interviewees be in an independent and undisturbed space during the in-person interview process.
[0003] Currently, in the solutions for detecting the presence of bystanders, when it is detected that there are more than one person in a certain frame of the in-person interview video, it is determined that there are bystanders present in this in-person interview session. However, this solution cannot detect some situations where there are bystanders present. For example, user A and user B do not appear in a certain frame of the in-person interview video at the same time, and the current solution can only detect one person in any frame of the in-person interview video, resulting in an inability to accurately detect whether there are bystanders present in the in-person interview video. Summary of the Invention
[0004] In view of the above, it is necessary to provide a method for detecting the presence of bystanders based on an in-person interview video and related devices, which can solve the technical problem of being unable to accurately detect whether there are bystanders present in the in-person interview video.
[0005] On the one hand, the present invention proposes a method for detecting the presence of bystanders based on an in-person interview video, and the method for detecting the presence of bystanders based on an in-person interview video includes:
[0006] Obtain the to-be-detected in-person interview video, where the to-be-detected in-person interview video includes multiple video frames;
[0007] Detect the multiple video frames based on a pre-trained human body detection model to obtain the human body detection results of each video frame;
[0008] Count the number of bounding boxes in each human body detection result;
[0009] If the number of bounding boxes in multiple ones of the above is less than a first preset number, extract an image set with a number of bounding boxes equal to a second preset number from the multiple video frames, where the image set includes a first image and a second image;
[0010] Calculate the movement information of the second bounding box according to the first bounding box corresponding to the first image, the second bounding box corresponding to the second image, and the image size of the second image;
[0011] Obtain the first movement direction of the first bounding box in the first image, and identify the second movement direction of the second bounding box in the second image according to the movement information;
[0012] Generate the video detection result of the to-be-detected face review video according to the first movement direction, the second movement direction, and the movement information.
[0013] According to a preferred embodiment of the present invention, after counting the number of frames of the human detection frames in each human detection result, the method for detecting the presence of bystanders based on the face review video further includes:
[0014] If there is a frame number greater than or equal to the first preset number among the multiple frame numbers, it is determined that the to-be-detected face review video includes bystanders.
[0015] According to a preferred embodiment of the present invention, the extracting the image set with the frame number of the second preset number from the multiple video frames includes:
[0016] Extract the video frames with the frame number of the second preset number from the multiple video frames as the preliminary screening images;
[0017] Identify the frame numbers of the preliminary screening images in the multiple video frames;
[0018] Generate the image set according to the preliminary screening images corresponding to consecutive frame numbers.
[0019] According to a preferred embodiment of the present invention, the calculating the movement information of the second detection frame according to the first detection frame corresponding to the first image, the second detection frame corresponding to the second image, and the image size of the second image includes:
[0020] Generate the first center coordinate information of the first detection frame according to the first frame coordinate information of the first detection frame in the first image, and generate the second center coordinate information of the second detection frame according to the second frame coordinate information of the second detection frame in the second image;
[0021] Calculate the movement information according to the first center coordinate information, the second center coordinate information, and the image size.
[0022] According to a preferred embodiment of the present invention, the image size includes the image width and image height of the second image, the movement information includes the first movement information and the second movement information, and the calculation formula of the movement information is:
[0023] x_move = (this_xcenter - pre_xcenter) / Width;
[0024] y_move = (this_ycenter - pre_ycenter) / Height;
[0025] Wherein, x_move represents the first movement information, this_xcenter represents the information corresponding to the x direction in the second center coordinate information, pre_xcenter represents the information corresponding to the x direction in the first center coordinate information, Width represents the image width, y_move represents the second movement information of the second detection frame in the y direction, this_ycenter represents the information corresponding to the y direction in the second center coordinate information, pre_ycenter represents the information corresponding to the y direction in the first center coordinate information, and Height represents the image height.
[0026] According to a preferred embodiment of the present invention, the identifying the second movement direction of the second detection frame in the second image according to the movement information includes:
[0027] If the first movement information is greater than the first horizontal axis threshold, it is determined that the second movement direction moves right in the x direction; or
[0028] If the first movement information is less than the second horizontal axis threshold, it is determined that the second movement direction moves left in the x direction; or
[0029] If the first movement information is within the threshold interval constructed by the second horizontal axis threshold and the first horizontal axis threshold, it is determined that the second movement direction is the preset direction in the x direction; or
[0030] If the second movement information is greater than the first vertical axis threshold, it is determined that the second movement direction moves down in the y direction; or
[0031] If the second movement information is less than the second vertical axis threshold, it is determined that the second movement direction moves up in the y direction; or
[0032] If the second movement information is within the threshold interval constructed by the second vertical axis threshold and the first vertical axis threshold, it is determined that the second movement direction is the preset direction in the y direction.
[0033] According to a preferred embodiment of the present invention, the generating the video detection result of the to-be-detected face review video according to the first movement direction, the second movement direction, and the movement information includes:
[0034] Detect whether the first movement direction is the same as the second movement direction in the x direction, and detect whether the first movement direction is the same as the second movement direction in the y direction;
[0035] If the first movement direction is opposite to the second movement direction in the x direction, and the first movement information is greater than a first movement threshold, then determine that the video detection result is that the to-be-detected face review video includes a bystander; or
[0036] If the first movement direction is opposite to the second movement direction in the y direction, and the second movement information is greater than a second movement threshold, then determine that the video detection result is that the to-be-detected face review video includes a bystander.
[0037] On the other hand, the present invention also proposes a bystander presence detection device based on a face review video. The bystander presence detection device based on a face review video includes:
[0038] An acquisition unit, configured to acquire a to-be-detected face review video, where the to-be-detected face review video includes a plurality of video frames;
[0039] A detection unit, configured to detect the plurality of video frames based on a pre-trained human body detection model to obtain a human body detection result for each video frame;
[0040] A statistics unit, configured to count the number of bounding boxes in each human body detection result;
[0041] An extraction unit, configured to, if the number of bounding boxes is less than a first preset number, extract an image set with the number of bounding boxes being a second preset number from the plurality of video frames, where the image set includes a first image and a second image;
[0042] A calculation unit, configured to calculate movement information of the second bounding box according to a first bounding box corresponding to the first image, a second bounding box corresponding to the second image, and the image size of the second image;
[0043] An identification unit, configured to obtain a first movement direction of the first bounding box in the first image, and identify a second movement direction of the second bounding box in the second image according to the movement information;
[0044] A generation unit, configured to generate a video detection result of the to-be-detected face review video according to the first movement direction, the second movement direction, and the movement information.
[0045] On the other hand, the present invention also proposes an electronic device, where the electronic device includes:
[0046] A memory, storing computer-readable instructions; and
[0047] A processor, configured to execute the computer-readable instructions stored in the memory to implement the bystander presence detection method based on a face review video.
[0048] On the other hand, the present invention also provides a computer-readable storage medium storing computer-readable instructions that are executed by a processor in an electronic device to implement the method for detecting the presence of bystanders based on the face-to-face review video.
[0049] As can be seen from the above technical solutions, the present invention can quickly detect the multiple video frames through the human body detection model to improve the generation efficiency of the human body detection results, thereby improving the recognition efficiency of the video detection results. By comparing the number of bounding boxes of the human body detection with the first preset number and the second preset number, it is possible to identify whether there are multiple different users in different video frames in the to-be-detected face-to-face review video. Furthermore, by combining the first movement direction, the second movement direction, and the movement information, the accuracy of the video detection results can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flowchart of a preferred embodiment of the method for detecting the presence of bystanders based on the face-to-face review video of the present invention.
[0051] Figure 2 is a functional block diagram of a preferred embodiment of the device for detecting the presence of bystanders based on the face-to-face review video of the present invention.
[0052] Figure 3 is a schematic structural diagram of an electronic device of a preferred embodiment for implementing the method for detecting the presence of bystanders based on the face-to-face review video of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0054] As Figure 1 shown, it is a flowchart of a preferred embodiment of the method for detecting the presence of bystanders based on the face-to-face review video of the present invention. According to different requirements, the order of the steps in this flowchart can be changed, and some steps can be omitted.
[0055] The method for detecting the presence of bystanders based on the face-to-face review video can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, sense the environment, acquire knowledge, and use the knowledge to obtain the best results of theory, method, technology, and application systems.
[0056] The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0057] The method for detecting the presence of bystanders based on the face-to-face interview video is applied to one or more electronic devices. The electronic device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored computer-readable instructions. Its hardware includes, but is not limited to, microprocessors, application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0058] The electronic device can be any electronic product that can perform human-computer interaction with users. For example, personal computers, tablet computers, smart phones, personal digital assistants (PDAs), game consoles, Internet Protocol Televisions (IPTVs), smart wearable devices, etc.
[0059] The electronic device may include network devices and / or user devices. Among them, the network device includes, but is not limited to, a single network electronic device, a group of electronic devices composed of multiple network electronic devices, or a cloud composed of a large number of hosts or network electronic devices based on cloud computing.
[0060] The network where the electronic device is located includes, but is not limited to: the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.
[0061] 101. Obtain the face-to-face interview video to be detected, where the face-to-face interview video to be detected includes multiple video frames.
[0062] In at least one embodiment of the present invention, the face-to-face interview video to be detected refers to a video recorded when a bank or a securities company conducts a face-to-face interview with the person being interviewed. The face-to-face interview video to be detected can be extracted from a video recording device bound and associated with the bank or the securities company.
[0063] 102. Detect the multiple video frames based on a pre-trained human detection model to obtain the human detection result of each video frame.
[0064] In at least one embodiment of the present invention, the human detection model is used to identify the humans appearing in each video frame. The human detection model is a model generated by training and testing using the crawled pictures related to the face review scenario. The human detection model includes a first feature extraction network and a second feature extraction network.
[0065] The number of bounding boxes in the human detection result is the same as the number of people in the corresponding video frame.
[0066] In at least one embodiment of the present invention, the electronic device detecting the multiple video frames based on a pre-trained human detection model to obtain the human detection result of each video frame includes:
[0067] For any video frame, extract the feature of the any video frame based on the first feature extraction network to obtain the first feature information;
[0068] Extract the first feature information based on the second feature extraction network to obtain the second feature information;
[0069] Perform upsampling processing on the second feature information to obtain the sampling information;
[0070] Fuse the first feature information and the sampling information to obtain the target information;
[0071] Perform prediction on the target information to obtain the human detection result.
[0072] Wherein, the target information may be generated by splicing the first feature information and the sampling information.
[0073] By fusing the feature information of different scales, the accuracy of the human detection result can be improved.
[0074] 103. Count the number of bounding boxes in each human detection result.
[0075] In at least one embodiment of the present invention, the number of bounding boxes refers to the total number of bounding boxes in each human detection result.
[0076] In at least one embodiment of the present invention, after counting the number of bounding boxes in each human detection result, the method for detecting the presence of bystanders based on the face review video further includes:
[0077] If there is a number of bounding boxes greater than or equal to the first preset number among the multiple numbers of bounding boxes, determine that the face review video to be detected includes bystanders.
[0078] Among them, the first preset quantity can be set according to requirements. In an actual scenario, the first preset quantity can be 2.
[0079] Through the above embodiments, when it is detected that the number of boxes in any video frame is greater than or equal to the first preset quantity, the result that the to-be-detected face review video includes bystanders can be quickly determined.
[0080] In this embodiment, when it is detected that there is a number of boxes greater than or equal to the first preset quantity among the multiple numbers of boxes, the detection of the video frame by using the human body detection model is ended.
[0081] 104. If the multiple numbers of boxes are all less than the first preset quantity, an image set with the number of boxes being the second preset quantity is extracted from the multiple video frames. The image set includes a first image and a second image.
[0082] In at least one embodiment of the present invention, the first preset quantity is greater than the second preset quantity. In an actual scenario, the second preset quantity can be 1.
[0083] The image set includes video frames with consecutive frame numbers among the multiple video frames.
[0084] In at least one embodiment of the present invention, the electronic device extracts the image set with the number of boxes being the second preset quantity from the multiple video frames, including:
[0085] Extracting the video frames with the number of boxes being the second preset quantity from the multiple video frames as preliminary screening images;
[0086] Identifying the frame sequence numbers of the preliminary screening images in the multiple video frames;
[0087] Generating the image set according to the preliminary screening images corresponding to consecutive frame sequence numbers.
[0088] By extracting the preliminary screening images corresponding to consecutive frame sequence numbers as the image set for analysis, the accuracy of generating subsequent video detection results can be improved.
[0089] 105. Calculate the movement information of the second detection box according to the first detection box corresponding to the first image, the second detection box corresponding to the second image, and the image size of the second image.
[0090] In at least one embodiment of the present invention, the first detection box refers to the detection box output by the human body detection model after detecting the first image, and the second detection box refers to the detection box output by the human body detection model after detecting the second image. The image size includes the image width and image height of the second image. The movement information includes first movement information and second movement information.
[0091] In at least one embodiment of the present invention, the electronic device calculates the movement information of the second detection box according to the first detection box corresponding to the first image, the second detection box corresponding to the second image, and the image size of the second image, including:
[0092] Generate the first center coordinate information of the first detection box according to the first box coordinate information of the first detection box in the first image, and generate the second center coordinate information of the second detection box according to the second box coordinate information of the second detection box in the second image;
[0093] Calculate the movement information according to the first center coordinate information, the second center coordinate information, and the image size.
[0094] By identifying the first center coordinate information and the second center coordinate information and calculating the movement information, it is possible to avoid inaccurate calculation of the movement information due to the different sizes of the first detection box and the second detection box, thereby improving the accuracy of the movement information.
[0095] Specifically, the calculation formula for the first center coordinate information is:
[0096] pre_xcenter = (pre_x0 + pre_x1) / 2;
[0097] pre_ycenter = (pre_y0 + pre_y1) / 2;
[0098] Wherein, pre_xcenter represents the information corresponding to the x direction in the first center coordinate information, pre_ycenter represents the information corresponding to the y direction in the first center coordinate information, the coordinate information of the upper left corner point in the first box coordinate information is (pre_x0, pre_y0), and the coordinate information of the lower right corner point in the first box coordinate information is (pre_x1, pre_y1).
[0099] In this embodiment, the calculation method of the second center coordinate information is similar to that of the first center coordinate information, and details are not described herein again in this application.
[0100] Specifically, the calculation formula for the movement information is:
[0101] x_move = (this_xcenter - pre_xcenter) / Width;
[0102] y_move = (this_ycenter - pre_ycenter) / Height;
[0103] Wherein, x_move represents the first movement information, this_xcenter represents the information corresponding to the x - direction in the second center coordinate information, pre_xcenter represents the information corresponding to the x - direction in the first center coordinate information, Width represents the image width, y_move represents the second movement information of the second detection frame in the y - direction, this_ycenter represents the information corresponding to the y - direction in the second center coordinate information, pre_ycenter represents the information corresponding to the y - direction in the first center coordinate information, and Height represents the image height.
[0104] By combining the analysis of the second detection frame with the image size, the accuracy of the movement information can be improved.
[0105] 106. Obtain the first movement direction of the first detection frame in the first image, and identify the second movement direction of the second detection frame in the second image according to the movement information.
[0106] In at least one embodiment of the present invention, the first movement direction refers to the movement condition of the first detection frame in the first image relative to the previous video frame of the plurality of video frames, wherein the previous video frame refers to the adjacent video frame in front of the first image among the plurality of video frames.
[0107] The second movement direction refers to the movement condition of the second detection frame in the second image relative to the previous video frame of the plurality of video frames, wherein the previous video frame refers to the adjacent video frame in front of the second image among the plurality of video frames.
[0108] In at least one embodiment of the present invention, the electronic device identifying the second movement direction of the second detection frame in the second image according to the movement information includes:
[0109] If the first movement information is greater than the first horizontal axis threshold, it is determined that the second movement direction moves to the right in the x - direction; or
[0110] If the first movement information is less than the second horizontal axis threshold, it is determined that the second movement direction moves to the left in the x - direction; or
[0111] If the first movement information is within the threshold range constructed by the second horizontal axis threshold and the first horizontal axis threshold, determine that the second movement direction is the preset direction in the x direction; or
[0112] If the second movement information is greater than the first vertical axis threshold, determine that the second movement direction is moving downward in the y direction; or
[0113] If the second movement information is less than the second vertical axis threshold, determine that the second movement direction is moving upward in the y direction; or
[0114] If the second movement information is within the threshold range constructed by the second vertical axis threshold and the first vertical axis threshold, determine that the second movement direction is the preset direction in the y direction.
[0115] Wherein, the first horizontal axis threshold and the second horizontal axis threshold are opposite numbers to each other, the first vertical axis threshold and the second vertical axis threshold are opposite numbers to each other. Generally speaking, the first horizontal axis threshold is greater than the second horizontal axis threshold, and the first vertical axis threshold is greater than the second vertical axis threshold.
[0116] The preset direction means no movement.
[0117] In other embodiments, the recognition method of the second movement direction is similar to the recognition direction of the first movement direction, and this application will not elaborate on this.
[0118] In this embodiment, when the first image is the first frame in the image set, determine the first movement direction corresponding to the first image as the preset direction in both the x direction and the y direction.
[0119] 107. Generate a video detection result of the to-be-detected face review video according to the first movement direction, the second movement direction, and the movement information.
[0120] In at least one embodiment of the present invention, the video detection results include: there are bystanders in the to-be-detected face review video and there are no bystanders in the to-be-detected face review video.
[0121] In at least one embodiment of the present invention, the electronic device generating the video detection result of the to-be-detected face review video according to the first movement direction, the second movement direction, and the movement information includes:
[0122] Detect whether the first movement direction is the same as the second movement direction in the x direction, and detect whether the first movement direction is the same as the second movement direction in the y direction;
[0123] If the first movement direction is opposite to the second movement direction in the x direction, and the first movement information is greater than the first movement threshold, it is determined that the video detection result is that the to-be-detected face review video includes bystanders; or
[0124] If the first movement direction is opposite to the second movement direction in the y direction, and the second movement information is greater than the second movement threshold, it is determined that the video detection result is that the to-be-detected face review video includes bystanders.
[0125] Through the above embodiments, it is possible to identify whether there are multiple different users in different video frames in the to-be-detected face review video, thereby improving the accuracy of generating the video detection result.
[0126] It should be emphasized that to further ensure the privacy and security of the above video detection result, the above video detection result can also be stored in a node of a blockchain.
[0127] It can be seen from the above technical solutions that the present invention can quickly detect the multiple video frames through the human body detection model to improve the generation efficiency of the human body detection result, thereby improving the recognition efficiency of the video detection result. By comparing the number of frames of the human body detection frame with the first preset number and the second preset number, it is possible to identify whether there are multiple different users in different video frames in the to-be-detected face review video. Furthermore, by combining the first movement direction, the second movement direction, and the movement information, the accuracy of the video detection result can be improved.
[0128] As Figure 2 shown, it is a functional module diagram of a preferred embodiment of the bystander presence detection device based on the face review video of the present invention. The bystander presence detection device 11 based on the face review video includes an acquisition unit 110, a detection unit 111, a statistics unit 112, an extraction unit 113, a calculation unit 114, an identification unit 115, a generation unit 116, and a determination unit 117. The module / unit referred to in the present invention refers to a series of computer-readable instruction segments that can be acquired by a processor 13 and can complete fixed functions, and are stored in a memory 12. In this embodiment, the functions of each module / unit will be described in detail in subsequent embodiments.
[0129] The acquisition unit 110 acquires the to-be-detected face review video, and the to-be-detected face review video includes multiple video frames.
[0130] In at least one embodiment of the present invention, the to-be-detected face review video refers to a video generated by a bank or a securities company when conducting a face review on a person to be face-reviewed. The to-be-detected face review video can be extracted from a video recording device bound and associated with the bank or the securities company.
[0131] The detection unit 111 detects the multiple video frames based on a pre-trained human body detection model, and obtains the human body detection result of each video frame.
[0132] In at least one embodiment of the present invention, the human body detection model is used to identify the human bodies appearing in each video frame. The human body detection model is a model generated by training and testing with the pictures crawled related to the face review scenario. The human body detection model includes a first feature extraction network and a second feature extraction network.
[0133] The number of bounding boxes in the human body detection result is the same as the number of people in the corresponding video frame.
[0134] In at least one embodiment of the present invention, the detection unit 111 detects the multiple video frames based on a pre-trained human body detection model, and the human body detection result obtained for each video frame includes:
[0135] For any video frame, feature extraction is performed on the any video frame based on the first feature extraction network to obtain first feature information;
[0136] Based on the second feature extraction network, the first feature information is extracted to obtain second feature information;
[0137] Upsampling processing is performed on the second feature information to obtain sampling information;
[0138] The first feature information and the sampling information are fused to obtain target information;
[0139] Prediction is performed on the target information to obtain the human body detection result.
[0140] Among them, the target information may be generated by splicing the first feature information and the sampling information.
[0141] By fusing feature information of different scales, the accuracy of the human body detection result can be improved.
[0142] The statistical unit 112 counts the number of bounding boxes in each human body detection result.
[0143] In at least one embodiment of the present invention, the number of bounding boxes refers to the total number of bounding boxes in each human body detection result.
[0144] In at least one embodiment of the present invention, after counting the number of bounding boxes in each human body detection result, if there is a number of bounding boxes greater than or equal to a first preset number among the multiple numbers of bounding boxes, the determination unit 117 determines that the to-be-detected face review video includes bystanders.
[0145] Among them, the first preset quantity can be set according to requirements. In an actual scenario, the first preset quantity can be 2.
[0146] Through the above implementation manner, when it is detected that the number of boxes in any video frame is greater than or equal to the first preset quantity, the result that the to-be-detected face review video includes bystanders can be quickly determined.
[0147] In this embodiment, when it is detected that there is a number of boxes greater than or equal to the first preset quantity among the multiple numbers of boxes, the detection of the video frame by using the human body detection model is ended.
[0148] If the multiple numbers of boxes are all less than the first preset quantity, the extraction unit 113 extracts an image set with the number of boxes being the second preset quantity from the multiple video frames. The image set includes a first image and a second image.
[0149] In at least one embodiment of the present invention, the first preset quantity is greater than the second preset quantity. In an actual scenario, the second preset quantity can be 1.
[0150] The image set includes video frames with consecutive frame numbers among the multiple video frames.
[0151] In at least one embodiment of the present invention, the extraction unit 113 extracts an image set with the number of boxes being the second preset quantity from the multiple video frames, including:
[0152] Extracting video frames with the number of boxes being the second preset quantity from the multiple video frames as preliminary screening images;
[0153] Identifying the frame sequence numbers of the preliminary screening images in the multiple video frames;
[0154] Generating the image set according to the preliminary screening images corresponding to consecutive frame sequence numbers.
[0155] By extracting the preliminary screening images corresponding to consecutive frame sequence numbers as the image set for analysis, the accuracy of generating subsequent video detection results can be improved.
[0156] The calculation unit 114 calculates the movement information of the second detection box according to the first detection box corresponding to the first image, the second detection box corresponding to the second image, and the image size of the second image.
[0157] In at least one embodiment of the present invention, the first detection box refers to the detection box output by the human body detection model after detecting the first image, and the second detection box refers to the detection box output by the human body detection model after detecting the second image. The image size includes the image width and image height of the second image. The movement information includes first movement information and second movement information.
[0158] In at least one embodiment of the present invention, the calculation unit 114 calculates the movement information of the second detection box according to the first detection box corresponding to the first image, the second detection box corresponding to the second image, and the image size of the second image, including:
[0159] Generate the first center coordinate information of the first detection box according to the first box coordinate information of the first detection box in the first image, and generate the second center coordinate information of the second detection box according to the second box coordinate information of the second detection box in the second image;
[0160] Calculate the movement information according to the first center coordinate information, the second center coordinate information, and the image size.
[0161] By identifying the first center coordinate information and the second center coordinate information, the calculation of the movement information can avoid the inaccurate calculation of the movement information caused by the different sizes of the first detection box and the second detection box, thereby improving the accuracy of the movement information.
[0162] Specifically, the calculation formula of the first center coordinate information is:
[0163] pre_xcenter = (pre_x0 + pre_x1) / 2;
[0164] pre_ycenter = (pre_y0 + pre_y1) / 2;
[0165] Wherein, pre_xcenter represents the information corresponding to the x direction in the first center coordinate information, pre_ycenter represents the information corresponding to the y direction in the first center coordinate information, the coordinate information of the upper left corner point in the first box coordinate information is (pre_x0, pre_y0), and the coordinate information of the lower right corner point in the first box coordinate information is (pre_x1, pre_y1).
[0166] In this embodiment, the calculation method of the second center coordinate information is similar to that of the first center coordinate information, and the present application will not elaborate on this.
[0167] Specifically, the calculation formula of the movement information is:
[0168] x_move = (this_xcenter - pre_xcenter) / Width;
[0169] y_move = (this_ycenter - pre_ycenter) / Height;
[0170] Wherein, x_move represents the first movement information, this_xcenter represents the information corresponding to the x - direction in the second center coordinate information, pre_xcenter represents the information corresponding to the x - direction in the first center coordinate information, Width represents the image width, y_move represents the second movement information of the second detection box in the y - direction, this_ycenter represents the information corresponding to the y - direction in the second center coordinate information, pre_ycenter represents the information corresponding to the y - direction in the first center coordinate information, and Height represents the image height.
[0171] By analyzing the second detection box in combination with the image size, the accuracy of the movement information can be improved.
[0172] The recognition unit 115 obtains the first movement direction of the first detection box in the first image, and recognizes the second movement direction of the second detection box in the second image according to the movement information.
[0173] In at least one embodiment of the present invention, the first movement direction refers to the movement of the first detection box in the first image relative to the previous video frame of the plurality of video frames, wherein the previous video frame refers to the adjacent video frame in front of the first image among the plurality of video frames.
[0174] The second movement direction refers to the movement of the second detection box in the second image relative to the previous video frame of the plurality of video frames, wherein the previous video frame refers to the adjacent video frame in front of the second image among the plurality of video frames.
[0175] In at least one embodiment of the present invention, the recognition unit 115 recognizes the second movement direction of the second detection box in the second image according to the movement information, including:
[0176] If the first movement information is greater than the first horizontal axis threshold, it is determined that the second movement direction moves to the right in the x - direction; or
[0177] If the first movement information is less than the second horizontal axis threshold, it is determined that the second movement direction moves to the left in the x - direction; or
[0178] If the first movement information is within the threshold interval constructed by the second horizontal axis threshold and the first horizontal axis threshold, determine that the second movement direction is the preset direction in the x direction; or
[0179] If the second movement information is greater than the first vertical axis threshold, determine that the second movement direction is moving downward in the y direction; or
[0180] If the second movement information is less than the second vertical axis threshold, determine that the second movement direction is moving upward in the y direction; or
[0181] If the second movement information is within the threshold interval constructed by the second vertical axis threshold and the first vertical axis threshold, determine that the second movement direction is the preset direction in the y direction.
[0182] Wherein, the first horizontal axis threshold and the second horizontal axis threshold are opposite numbers to each other, and the first vertical axis threshold and the second vertical axis threshold are opposite numbers to each other. Generally speaking, the first horizontal axis threshold is greater than the second horizontal axis threshold, and the first vertical axis threshold is greater than the second vertical axis threshold.
[0183] The preset direction means no movement.
[0184] In other embodiments, the recognition method of the second movement direction is similar to the recognition direction of the first movement direction, and this application will not elaborate on this.
[0185] In this embodiment, when the first image is the first frame in the image set, determine the first movement direction corresponding to the first image as the preset direction in both the x direction and the y direction.
[0186] The generating unit 116 generates a video detection result of the to-be-detected face review video according to the first movement direction, the second movement direction, and the movement information.
[0187] In at least one embodiment of the present invention, the video detection results are: there are bystanders in the to-be-detected face review video and there are no bystanders in the to-be-detected face review video.
[0188] In at least one embodiment of the present invention, the generating unit 116 generating the video detection result of the to-be-detected face review video according to the first movement direction, the second movement direction, and the movement information includes:
[0189] Detect whether the first movement direction is the same as the second movement direction in the x direction, and detect whether the first movement direction is the same as the second movement direction in the y direction;
[0190] If the first movement direction is opposite to the second movement direction in the x direction, and the first movement information is greater than the first movement threshold, it is determined that the video detection result is that the video to be detected for face review includes bystanders; or
[0191] If the first movement direction is opposite to the second movement direction in the y direction, and the second movement information is greater than the second movement threshold, it is determined that the video detection result is that the video to be detected for face review includes bystanders.
[0192] Through the above embodiments, it is possible to identify whether there are multiple different users in different video frames in the video to be detected for face review, thereby improving the accuracy of generating the video detection result.
[0193] It should be emphasized that, to further ensure the privacy and security of the above video detection result, the above video detection result can also be stored in a node of a blockchain.
[0194] From the above technical solutions, it can be seen that the present invention can quickly detect the multiple video frames through the human body detection model to improve the generation efficiency of the human body detection result, thereby improving the recognition efficiency of the video detection result. By comparing the number of frames of the human body detection frame with the first preset number and the second preset number, it is possible to identify whether there are multiple different users in different video frames in the video to be detected for face review. Furthermore, by combining the first movement direction, the second movement direction, and the movement information, the accuracy of the video detection result can be improved.
[0195] As Figure 3 shown, it is a schematic structural diagram of an electronic device according to a preferred embodiment of the method for detecting the presence of bystanders based on a face review video of the present invention.
[0196] In an embodiment of the present invention, the electronic device 1 includes, but is not limited to, a memory 12, a processor 13, and computer-readable instructions stored in the memory 12 and executable on the processor 13, such as a program for detecting the presence of bystanders based on a face review video.
[0197] Those skilled in the art can understand that the schematic diagram is only an example of the electronic device 1, and does not constitute a limitation on the electronic device 1. It may include more or fewer components than shown, or combine some components, or different components. For example, the electronic device 1 may also include input / output devices, network access devices, buses, etc.
[0198] The processor 13 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor 13 is the operation core and control center of the electronic device 1, connecting various parts of the entire electronic device 1 through various interfaces and lines, and executing the operating system of the electronic device 1 and various installed application programs, program codes, etc.
[0199] Exemplarily, the computer-readable instructions may be divided into one or more modules / units, and the one or more modules / units are stored in the memory 12 and executed by the processor 13 to complete the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and the computer-readable instruction segments are used to describe the execution process of the computer-readable instructions in the electronic device 1. For example, the computer-readable instructions may be divided into an acquisition unit 110, a detection unit 111, a statistics unit 112, an extraction unit 113, a calculation unit 114, an identification unit 115, a generation unit 116, and a determination unit 117.
[0200] The memory 12 may be used to store the computer-readable instructions and / or modules. The processor 13 realizes various functions of the electronic device 1 by running or executing the computer-readable instructions and / or modules stored in the memory 12, and calling the data stored in the memory 12. The memory 12 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the electronic device. The memory 12 may include non-volatile and volatile memories, such as: hard disks, memories, plug-in hard disks, SmartMedia Cards (SMCs), Secure Digital (SD) cards, Flash Cards, at least one magnetic disk storage device, flash memory device, or other storage devices.
[0201] The memory 12 can be an external memory and / or an internal memory of the electronic device 1. Further, the memory 12 can be a memory in a physical form, such as a memory stick, a TF card (Trans-flash Card), and so on.
[0202] If the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present invention, it can also be completed by computer-readable instructions to instruct relevant hardware. The computer-readable instructions can be stored in a computer-readable storage medium. When the computer-readable instructions are executed by a processor, the steps of the above-mentioned various method embodiments can be implemented.
[0203] Among them, the computer-readable instructions include computer-readable instruction codes, and the computer-readable instruction codes can be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium can include: any entity or device capable of carrying the computer-readable instruction codes, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROM, Read-Only Memory), and random access memories (RAM, Random Access Memory).
[0204] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, essentially a decentralized database, is a string of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, and an application service layer, etc.
[0205] Combined Figure 1 , the memory 12 in the electronic device 1 stores computer-readable instructions to implement a bystander presence detection method based on an in-person interview video, and the processor 13 can execute the computer-readable instructions to achieve:
[0206] Obtain the in-person interview video to be detected, where the in-person interview video to be detected includes multiple video frames;
[0207] Detect the multiple video frames based on a pre-trained human detection model to obtain the human detection results of each video frame;
[0208] Count the number of bounding boxes in the human detection results for each person;
[0209] If the number of boxes in multiple ones of them is less than a first preset number, an image set with the number of boxes being a second preset number is extracted from the multiple video frames, and the image set includes a first image and a second image;
[0210] Calculate movement information of the second detection box according to a first detection box corresponding to the first image, a second detection box corresponding to the second image, and the image size of the second image;
[0211] Obtain a first movement direction of the first detection box in the first image, and identify a second movement direction of the second detection box in the second image according to the movement information;
[0212] Generate a video detection result of the to-be-detected face review video according to the first movement direction, the second movement direction, and the movement information.
[0213] Specifically, for the specific implementation method of the above computer-readable instructions by the processor 13, reference may be made to Figure 1 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.
[0214] In several embodiments provided by the present invention, it should be understood that the disclosed system, device, and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0215] Computer-readable instructions are stored on the computer-readable storage medium, wherein when the computer-readable instructions are executed by the processor 13, the following steps are implemented:
[0216] Obtain a to-be-detected face review video, where the to-be-detected face review video includes multiple video frames;
[0217] Detect the multiple video frames based on a pre-trained human body detection model to obtain a human body detection result for each video frame;
[0218] Count the number of boxes of the human body detection box in each human body detection result;
[0219] If the number of boxes in multiple ones of them is less than a first preset number, an image set with the number of boxes being a second preset number is extracted from the multiple video frames, and the image set includes a first image and a second image;
[0220] Calculate movement information of the second detection box according to a first detection box corresponding to the first image, a second detection box corresponding to the second image, and the image size of the second image;
[0221] Obtain the first movement direction of the first detection frame in the first image, and identify the second movement direction of the second detection frame in the second image according to the movement information;
[0222] Generate a video detection result of the to-be-detected face review video according to the first movement direction, the second movement direction, and the movement information.
[0223] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical unit, that is, it may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0224] In addition, in each embodiment of the present invention, each functional module can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.
[0225] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
[0226] In addition, it is obvious that the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices can also be implemented by one unit or device through software or hardware. Words such as first and second are used to represent names and do not indicate any specific order.
[0227] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for detecting the presence of bystanders based on face review videos, characterized in that, The method for detecting the presence of bystanders based on the face-to-face interview video includes: Obtain the face-to-face interview video to be detected, where the face-to-face interview video to be detected includes multiple video frames; Detect the multiple video frames based on a pre-trained human detection model to obtain the human detection result of each video frame; Count the number of bounding boxes in each human detection result; If the number of bounding boxes in multiple ones of the above is less than a first preset number, extract an image set with the number of bounding boxes being a second preset number from the multiple video frames, where the image set includes a first image and a second image; Calculate the movement information of the second bounding box according to the first bounding box corresponding to the first image, the second bounding box corresponding to the second image, and the image size of the second image, including: generating first center coordinate information of the first bounding box according to the first bounding box coordinate information of the first bounding box in the first image, and generating second center coordinate information of the second bounding box according to the second bounding box coordinate information of the second bounding box in the second image; calculating the movement information according to the first center coordinate information, the second center coordinate information, and the image size; Obtain the first movement direction of the first bounding box in the first image, and identify the second movement direction of the second bounding box in the second image according to the movement information; Generate a video detection result of the face-to-face interview video to be detected according to the first movement direction, the second movement direction, and the movement information; If there is a number of bounding boxes greater than or equal to the first preset number among the multiple numbers of bounding boxes, determine that the face-to-face interview video to be detected includes bystanders.
2. The method for detecting the presence of bystanders based on the face-to-face interview video according to claim 1, characterized in that, The extracting the image set with the number of bounding boxes being a second preset number from the multiple video frames includes: Extract video frames with the number of bounding boxes being the second preset number from the multiple video frames as preliminary screening images; Identify the frame numbers of the preliminary screening images in the multiple video frames; Generate the image set according to the preliminary screening images corresponding to consecutive frame numbers.
3. The method for detecting the presence of bystanders based on the face-to-face review video according to claim 1, wherein The image size includes the image width and image height of the second image, the movement information includes first movement information and second movement information, and the calculation formula of the movement information is: Among them, represents the first movement information, represents the information corresponding to the direction in the second center coordinate information, represents the information corresponding to the direction in the first center coordinate information, represents the image width, represents the second movement information of the second detection frame in the direction, represents the information corresponding to the direction in the second center coordinate information, represents the information corresponding to the direction in the first center coordinate information, represents the image height.
4. The method for detecting the presence of bystanders based on the face-to-face interview video according to claim 3, wherein, The identifying the second movement direction of the second bounding box in the second image according to the movement information includes: If the first movement information is greater than the first horizontal axis threshold, it is determined that the second movement direction is moving rightward in the direction; or If the first movement information is less than the second horizontal axis threshold, it is determined that the second movement direction is moving leftward in the direction; or If the first movement information is within the threshold interval constructed by the second horizontal axis threshold and the first horizontal axis threshold, it is determined that the second movement direction is a preset direction in the direction; or If the second movement information is greater than the first vertical axis threshold, it is determined that the second movement direction is moving downward in the direction; or If the second movement information is less than the second vertical axis threshold, it is determined that the second movement direction is upward movement in the direction; or If the second movement information is within the threshold range constructed by the second vertical axis threshold and the first vertical axis threshold, it is determined that the second movement direction is the preset direction in the direction.
5. The method for detecting the presence of bystanders based on the face-to-face review video according to claim 3, characterized in that, The generating the video detection result of the face-to-face interview video to be detected according to the first movement direction, the second movement direction, and the movement information includes: Detect whether the first movement direction and the second movement direction are the same in the direction, and detect whether the first movement direction and the second movement direction are the same in the direction; If the first movement direction is opposite to the second movement direction in the direction, and the first movement information is greater than the first movement threshold, it is determined that the video detection result is that there are bystanders in the video to be detected for face review; or If the first movement direction and the second movement direction are opposite in the direction, and the second movement information is greater than the second movement threshold, it is determined that the video detection result is that there are bystanders in the video to be detected for face review.
6. A bystander presence detection device based on in-person review videos, characterized in that, The device for detecting the presence of bystanders based on the face-to-face interview video includes: An acquisition unit for acquiring the face-to-face interview video to be detected, where the face-to-face interview video to be detected includes multiple video frames; A detection unit for detecting the multiple video frames based on a pre-trained human detection model to obtain the human detection result of each video frame; A statistics unit for counting the number of bounding boxes in each human detection result; An extraction unit for, if the number of bounding boxes in multiple ones of the above is less than a first preset number, extracting an image set with the number of bounding boxes being a second preset number from the multiple video frames, where the image set includes a first image and a second image; A calculation unit, configured to calculate movement information of the second detection box according to a first detection box corresponding to the first image, a second detection box corresponding to the second image, and an image size of the second image, including: generating first center coordinate information of the first detection box according to first box coordinate information of the first detection box in the first image, and generating second center coordinate information of the second detection box according to second box coordinate information of the second detection box in the second image; calculating the movement information according to the first center coordinate information, the second center coordinate information, and the image size; An identification unit, configured to obtain a first movement direction of the first detection box in the first image, and identify a second movement direction of the second detection box in the second image according to the movement information; A generation unit, configured to generate a video detection result of the to-be-detected face review video according to the first movement direction, the second movement direction, and the movement information; A determination unit, configured to determine that there are bystanders in the to-be-detected face review video if there is a box quantity greater than or equal to the first preset quantity among the multiple box quantities.
7. An electronic device, characterized in that, The electronic device includes: A memory, storing computer-readable instructions; and A processor, executing the computer-readable instructions stored in the memory to implement the method for detecting the presence of bystanders in a face review video according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and the computer-readable instructions are executed by a processor in the electronic device to implement the method for detecting the presence of bystanders in a face review video according to any one of claims 1 to 5.
Citation Information
Patent Citations
Tail detection method and device, electronic equipment and storage medium
CN110633636A
Countercurrent detection method and device, electronic equipment and storage medium
CN111582243A