Face detection and protection method, device, electronic device and storage medium

By using neural networked face detection and protection models in the video compression domain, the problems of large amount of face detection and poor modeling accuracy in the prior art are solved, and efficient face detection and privacy protection are achieved, taking into account the detection effects of static and moving targets.

CN113591681BActive Publication Date: 2025-05-02司马达拓(深圳)智能系统有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110860024.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-28
Publication Date
2025-05-02
Estimated Expiration
2041-07-28

AI Technical Summary

Technical Problem

The existing face detection methods have large calculation volume, poor modeling accuracy and poor detection effect, making it difficult to efficiently realize face detection and privacy protection in video surveillance.

Method used

A neural networked face detection model and face protection model are used to detect the target face in the video compression domain, and the detected target face is occluded in the image, and the convolutional neural network and recurrent neural network are trained.

Benefits of technology

It realizes efficient detection and privacy protection of faces in video images, improves face detection accuracy and detection speed, and can take into account the detection of static targets and moving targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113591681B_ABST
    Figure CN113591681B_ABST
Patent Text Reader

Abstract

The present invention provides a face detection and protection method, device, electronic device and storage medium, the method comprising: obtaining a screen image after decomposition of a video to be detected; inputting the screen image into a pre-trained face detection model to obtain position information of a face region on the screen image, and obtaining a face image with a target face based on the position information of the face region on the screen image; inputting the face image into a pre-trained face protection model to obtain a face occlusion image after occlusion processing. The method can effectively realize the dual functions of detection and privacy protection, and can detect both static targets in the face region and moving targets in the video image, achieving a balance between the two detection targets. In addition, the method has high face detection accuracy and greatly improves the speed of face detection and occlusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video data processing, and in particular to a face detection and protection method, device, electronic equipment and storage medium. Background Art

[0002] In recent years, the application of video surveillance technology has become more and more popular and widespread, and many scenarios involve monitoring images of people, such as internal monitoring of enterprises, park property monitoring, public place monitoring of hospitals and shopping malls, etc. However, in some scenarios, out of consideration for personal privacy protection, it is necessary to block or conceal the facial features of the protected person in the monitoring image, which includes the process of detecting the face in the monitoring image and the process of blocking the detected face.

[0003] Conventionally, face detection methods such as RetinaFace and DSFD are used. Such detection methods detect the complete RGB image in the surveillance video. Although they can detect moving targets and stationary targets, they require complete decoding of the surveillance video, which results in a very large consumption of overall computing resources.

[0004] There are also face target detection methods based on vector fields and face target detection methods based on DCT residual coefficients. Both of these detection methods are based on the calculation results of the similarity or residual of macroblocks to judge the moving target. They are not only for detecting the static target in the face target area, but can only detect the moving target. Therefore, these two methods will also detect redundant areas, which will also cause the problem of large amount of calculation and laboriousness. In addition, although the face target detection method based on vector fields has a faster detection speed, it is easily affected by the performance of the MV classification modeling model. If the split modeling is not well modeled, the detection effect of this method will also be very poor. While the face target detection method based on DCT residual coefficients has improved accuracy, it has a large amount of calculation and is easily affected by coding noise. Summary of the invention

[0005] The present invention provides a face detection and protection method, device, electronic device and storage medium, which are used to overcome the defects of the face detection method in the prior art, such as large amount of calculation, poor modeling accuracy and poor detection effect, and realize efficient detection and privacy protection of faces in video images.

[0006] The present invention provides a face detection and protection method, comprising:

[0007] Obtaining the decomposed image of the video to be detected;

[0008] Inputting the screen image into a pre-trained face detection model to obtain position information of a face region on the screen image, and obtaining a face image with a target face based on the position information of the face region on the screen image;

[0009] The face image is input into a pre-trained face protection model to obtain a face occlusion image that has been occluded.

[0010] The face detection and protection method provided by the present invention also includes:

[0011] The face occlusion image is input into the face protection model again to obtain a face restoration image after restoration processing.

[0012] According to the face detection and protection method provided by the present invention, the pre-training process of the face detection model includes:

[0013] The face area is marked on each frame of the decomposed video.

[0014] Divide the marked frames into a plurality of image groups, each group including a full-image frame and a plurality of differential-image frames;

[0015] For each screen image group, the full image is decoded to obtain an RGB image, and the RGB image is used as a training data set and combined with a first neural network for training to obtain a first detection output image;

[0016] For each screen image group, after fusing the current frame residual image, the current frame motion vector image and the corresponding accumulated residual image and accumulated motion vector image of each of the difference images, superimposing them with the full image in the same screen image group to obtain a composite image, using the composite image as a training data set, combined with the second neural network for training, to obtain a second detection output image;

[0017] The face detection model is constructed based on the first detection output image and the second detection output image obtained in each screen image group.

[0018] According to the face detection and protection method provided by the present invention, the face area marking refers to marking the position of the face area on the screen image, and the marking value at least includes the horizontal coordinate, vertical coordinate and width and height of the face area reference point.

[0019] According to the face detection and protection method provided by the present invention, the first neural network and the second neural network adopt any one of a convolutional neural network, a recurrent neural network and a combination thereof.

[0020] According to the face detection and protection method provided by the present invention, the pre-training process of the face protection model includes:

[0021] Inputting the face image into a variational autoencoder for encoding to obtain a face occlusion image;

[0022] Inputting the face occluded image back into the variational autoencoder for decoding to obtain a face restoration image;

[0023] The face protection model is constructed based on the face occlusion image and the face restoration image.

[0024] The present invention provides a face detection and protection device, comprising:

[0025] The acquisition module is used to obtain the decomposed picture of the video to be detected;

[0026] A detection module, used to input the screen image into a pre-trained face detection model, obtain position information of the face area on the screen image, and obtain a face image with a target face based on the position information of the face area on the screen image;

[0027] The occlusion module is used to input the face image into a pre-trained face protection model to obtain a face occlusion image that has been occluded.

[0028] The face detection and protection device provided by the present invention also includes:

[0029] The restoration module is used to input the face occluded image into the face protection model again to obtain a face restoration image after restoration processing.

[0030] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, all or part of the steps of the face detection and protection method as described in any one of the above items are implemented.

[0031] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements all or part of the steps of the face detection and protection method as described in any one of the above items.

[0032] The present invention provides a face detection and protection method, device, electronic device and storage medium. The method utilizes a neural network face detection model and a face protection model to detect a target face in a video compression domain, and to mask the face area of ​​the detected target face in an image, thereby effectively realizing the dual functions of detection and privacy protection, and can detect both stationary targets in the face area and moving targets in the video image, thereby taking both detection targets into consideration. In addition, the method has high face detection accuracy and greatly improves the speed of face detection and masking. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0034] Figure 1 This is one of the flow charts of the face detection and protection method provided by the present invention;

[0035] Figure 2 This is the second flow chart of the face detection and protection method provided by the present invention;

[0036] Figure 3 It is a schematic diagram of the pre-training process of the face detection model in the face detection and protection method provided by the present invention;

[0037] Figure 4 It is a schematic diagram of the pre-training process of the face protection model in the face detection and protection method provided by the present invention;

[0038] Figure 5 It is a structural schematic diagram of the face detection and protection device provided by the present invention;

[0039] Figure 6 It is a structural schematic diagram of the electronic device provided by the present invention.

[0040] Reference numerals:

[0041] 510: acquisition module; 520: detection module; 530: shielding module; 610: processor; 620: communication interface; 630: memory; 640: communication bus. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described in detail below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0043] The function of protecting the privacy of faces in videos is mainly realized through two parts: face detection and face area shielding. The face detection function and the face area shielding function in each embodiment of the present invention are both realized through a neural network model.

[0044] The following is combined with Figure 1-6 The face detection and protection method, device, electronic device and storage medium provided by the present invention are described in detail.

[0045] The present invention provides a face detection and protection method. Figure 1 is one of the flow charts of the face detection and protection method provided by the present invention, such as Figure 1 As shown, the method includes:

[0046] 300. Obtain a decomposed screen image of the video to be detected;

[0047] 400. Input the screen image into a pre-trained face detection model to obtain position information of a face region on the screen image, and obtain a face image with a target face based on the position information of the face region on the screen image;

[0048] 500. Input the facial image into a pre-trained face protection model to obtain a face occlusion image that has been occluded.

[0049] For various surveillance videos, the video files need to be split and processed to decompose into multiple frames, and then each frame is processed accordingly.

[0050] Obtain each frame of the screen image after decoding and decomposition of the current video to be detected. For each frame of the screen image, detection and recognition based on the face detection model are performed to identify each face image with the target face from multiple frames of the screen image. At this time, there may be one or more face images. In addition, when the screen image is input into the face detection model, the position information of the face area on the screen image is first obtained, and then the face image with the target face can be obtained based on the position information of the face area on the screen image. The position information of the face area on the screen image can be determined based on the relative coordinate position information of the face area on the screen image, etc. It can be operated according to the actual situation and is not specifically limited here.

[0051] One or more facial images are input into a pre-trained face protection model in sequence to mask the facial area in the image of the target face that needs privacy protection, thereby obtaining a corresponding face masked image, such as a face masked image with a mosaic.

[0052] Among them, the face detection model and face protection model adopt neural network learning model.

[0053] The face detection and protection method provided by the present invention utilizes a neural network face detection model and a face protection model to detect a target face in a video compression domain, and to mask the face region of the detected target face in an image, thereby effectively realizing the dual functions of detection and privacy protection, and can detect both stationary targets in the face region and moving targets in the video image, thereby achieving a balance between the two detection targets. In addition, the face detection method of the present invention has high accuracy and greatly improves the speed of face detection and masking.

[0054] According to the face detection and protection method provided by the present invention, Figure 2 FIG. 2 is a flow chart of the face detection and protection method provided by the present invention. Figure 2 As shown, the method is Figure 1 The embodiment also includes:

[0055] 600. The face occluded image is input into the face protection model again to obtain a face restoration image after restoration processing.

[0056] In some application scenarios, it is also necessary to restore the occluded image that has been occluded by the face area to restore the face image. In this case, the face occluded image can be input into the face protection model again, and the face protection model and the occlusion processing are reversed to obtain the restored face restoration image.

[0057] According to the face detection and protection method provided by the present invention, Figure 3 FIG. 1 is a schematic diagram of the pre-training process of the face detection model in the face detection and protection method provided by the present invention, such as Figure 3 As shown, the method also includes a face detection model training step, and step 100, the pre-training process of the face detection model includes:

[0058] 110. Labeling the face region of each frame image after decomposition of the video to be detected;

[0059] 120. Divide the marked frames into a plurality of image groups, each group including a full image and a plurality of differential images;

[0060] 130. For each screen image group, decode the full image to obtain an RGB image, use the RGB image as a training data set, and perform training in combination with a first neural network to obtain a first detection output image;

[0061] 140. For each screen image group, after fusing the current frame residual image, the current frame motion vector image and the corresponding accumulated residual image and accumulated motion vector image of each of the difference images, superimpose them with the full image in the same screen image group to obtain a composite image, use the composite image as a training data set, and perform training in combination with a second neural network to obtain a second detection output image;

[0062] 150. Construct the face detection model based on the first detection output image and the second detection output image obtained in each screen image group.

[0063] Specifically, a large amount of collected surveillance videos are decoded to obtain each frame of image, and the processing process is described by taking the current video to be detected as an example. After the video to be detected is decomposed and processed, multiple frames of screen images are obtained, such as several I-frame screen images and several P-frame screen images. The face area is marked for each frame of the screen image. The marking method is to use a rectangular target box (bounding box, referred to as bbox) to mark the position range occupied by the face area in the screen image. The face area marking refers to marking the position of the face area on the screen image, and the marking value at least includes the horizontal coordinate and vertical coordinate of the reference point of the face area and the width and height of the face area. That is, a target box of a face area consists of four marking values, namely the horizontal coordinate and vertical coordinate (x, y) of the upper left corner point position of the target in the image or the horizontal coordinate and vertical coordinate (cx, cy) of the center coordinate point position, as well as the width w of the face area and the height h of the face area.

[0064] Note: The labeling operation forms a data set, which only labels the face area in each frame image, and does not label any other area or other target in the image. In this way, a face area labeling data set in the video compression domain is obtained, which can be used as training data for the face detection model.

[0065] The labeled frames are divided into several groups of images, and each group includes one full-frame image and several differential images. Since the labeled frames still include several I-frame images and several P-frame images. Moreover, the I-frame image refers to the full-frame image that can be decoded, and generally requires a larger neural network model for training and reasoning; while the P-frame image is the motion vector and prediction residual relative to the previous frame, or the differential image, which contains only less main information, so a smaller neural network model can be selected for training and reasoning, and the number of P-frame images in all the images after each video is decoded is much larger than the I-frame image. If the video is understood as consisting of several groups of images, in other words, all the labeled frames of the video are divided into several groups of images, and each group is set to include one full-frame image and several differential images, then each group of images includes one I-frame image and several P-frame images, that is, forming a composition such as IPPP or IPPPP.

[0066] The data of each group of screen images is used as a training data set for model training. In addition, the face detection model is divided into two networks according to the different screen images for training respectively, so as to obtain the corresponding face detection networks respectively, and the two networks are combined to form a face detection model.

[0067] The entire image (I frame image) in the picture image group is decoded to obtain an RGB image, which is used as training data and combined with the first neural network for training. After obtaining the first detection output image output by the first neural network, the loss value is calculated together with the corresponding label data, and then backpropagation is performed based on the obtained loss value, and the parameters of the first neural network are updated accordingly to end this training.

[0068] For the differential images (each P frame image) in the same picture image group, decoding is not required. The current frame residual image, current frame motion vector image and its corresponding accumulated residual image and accumulated motion vector image obtained from the video compression domain are fused, and then superimposed with the full image (the I frame image) in the same picture image group to obtain a composite image. The composite image is used as training data and combined with the second neural network for training. After obtaining the second detection output image output by the second neural network, the loss value is calculated together with the corresponding label data, and then back-propagation is performed based on the obtained loss value, and the parameters of the second neural network are updated accordingly. The above training process is performed for each P frame image.

[0069] Note: When training a P-frame image, if the P-frame image is the first P-frame image in the current image group, directly superimpose its motion vector image, residual image and the I-frame image in the current image group, and send the superimposed composite image to the second neural network for training. If the P-frame image is not the first P-frame image in the current image group, then the motion vector image and residual image obtained from the P-frame image are respectively accumulated with the motion vector images and residual images of the previous P-frame images, and the accumulated motion vector image and residual image are superimposed with the I-frame image in the current image group, and the superimposed composite image is sent to the second neural network for training, until all the P-frame images in the current image group are trained, and then the training ends.

[0070] The face detection model is then comprehensively constructed based on each of the first detection output images and each of the second detection output images respectively obtained from each screen image group.

[0071] Therefore, the face detection model constructed by the above method has high detection accuracy and good detection effect. And the first neural network and the second neural network in the constructed face detection model can respectively perform image detection on the I frame image and the P frame image of the current time frequency to be detected, so as to more quickly detect the face image with the target face and improve the detection efficiency. Its essence belongs to the method of directly detecting the face in the video compression domain, by only decoding the I frame image, but not decoding the P frame image, in other words, by not completely decoding the video image of the video to be detected, it is possible to realize the detection process of both detecting static targets and detecting moving targets, thereby greatly reducing the calculation amount of the overall detection, and a smaller second neural network can be used for face detection for the P frame image, which also reduces the consumption of computing resources. And because the deep neural network is used to automatically learn features, the distribution characteristics between the target and background macroblock vector fields can be better learned, and a better model can be established, so as to improve the accuracy of target face detection.

[0072] According to the face detection and protection method provided by the present invention, specifically, the first neural network and the second neural network can both adopt any one of a convolutional neural network, a recurrent neural network and a combination thereof. Generally, the second neural network can adopt a neural network that is smaller than the first neural network. Moreover, when any one of a convolutional neural network, a recurrent neural network, a combination of a convolutional neural network and a recurrent neural network is adopted, it can specifically use a machine learning algorithm such as a non-maximum suppression algorithm (NMS for short) to perform a model training process, which can be specifically set according to the actual training scenario.

[0073] According to the face detection and protection method provided by the present invention, Figure 4 FIG. 1 is a schematic diagram of the pre-training process of the face protection model in the face detection and protection method provided by the present invention, such as Figure 4 As shown, the method further includes a step of training a face protection model, and step 200, the pre-training process of the face protection model includes:

[0074] 210. Input the face image into a variational autoencoder for encoding to obtain a face occlusion image;

[0075] 220. Input the face occluded image back to the variational autoencoder for decoding to obtain a face restoration image;

[0076] 230. Construct the face protection model based on the face occlusion image and the face restoration image.

[0077] A large number of face images with target faces detected based on the above face detection model are obtained. For each original face image, occlusion preprocessing is performed based on the mean pixel replacement method: the face part is divided into small unit areas and the mean of each small unit area is calculated respectively, and then the pixel value of the original image of the small unit area is replaced by the mean of each small unit area, so as to obtain a pre-occluded image y1 with mosaic, and the pre-occluded image y1 and the original face image y2 are used as a pair of reference data.

[0078] Furthermore, for each face image, each face image is input into the encoding part of the variational autoencoder for encoding processing, thereby obtaining an output encoded image as the actual face occlusion image x1. The actual face occlusion image x1 is input back into the decoding part of the variational autoencoder for reverse decoding processing, thereby obtaining an output decoded image as the face restoration image x2 of the actual face occlusion image x1.

[0079] The network loss is calculated based on the mosaic pre-occluded image y1 and the actual face occluded image x1, as well as the face restoration image x2 and the original face image y2, and the calculation result is back-propagated to update the network parameters of the variational autoencoder computing network. The face protection model is constructed based on the face occluded image x1 and the face restoration image x2, as well as the pre-occluded image y1 and the original face image y2.

[0080] The face protection model constructed in this way can not only block and protect the target face in scenarios where privacy protection is required, but also restore the original information of the face hidden in the face-blocked image in scenarios where the specific face needs to be viewed.

[0081] A face detection and protection device provided by the present invention is introduced below. The face detection and protection device can be understood as a device for implementing the above-mentioned face detection and protection method. The principles of the two are consistent and can be referenced to each other, so they will not be described in detail here.

[0082] The present invention provides a face detection and protection device. Figure 5 FIG. 1 is a schematic diagram of the structure of the face detection and protection device provided by the present invention. Figure 5 As shown, the device includes: a collection module 510, a detection module 520 and a shielding module 530, wherein:

[0083] The acquisition module 510 is used to obtain the decomposed screen image of the video to be detected;

[0084] The detection module 520 is used to input the screen image into a pre-trained face detection model to obtain position information of the face area on the screen image, and obtain a face image with a target face based on the position information of the face area on the screen image;

[0085] The occlusion module 530 is used to input the face image into a pre-trained face protection model to obtain a face occlusion image that has been occluded.

[0086] The face detection and protection device provided by the present invention includes an acquisition module 510, a detection module 520 and a shielding module 530, and the modules work in coordination with each other, so that the hardware platform of the device can detect the target face in the video compression domain through the neural network face detection model and face protection model deployed on the platform, and shield the face area of ​​the detected target face in the image, thereby realizing the dual functions of detection and privacy protection, and the face detection accuracy is high, and the face detection and shielding speed are also greatly improved.

[0087] According to the face detection and protection device provided by the present invention, Figure 5 Based on the embodiment shown, the device further includes a restoration module, wherein:

[0088] The restoration module is used to input the face occluded image into the face protection model again to obtain a face restoration image after restoration processing.

[0089] The restoration module is applied to the scene where a specific face needs to be viewed, so as to restore the original information of the face hidden in the face occlusion image.

[0090] The present invention also provides an electronic device, Figure 6 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 6As shown, the electronic device may include: a processor 610, a communication interface 620, a memory 630 and a communication bus 640, wherein the processor 610, the communication interface 620 and the memory 630 communicate with each other through the communication bus 640. The processor 610 may call the logic instructions in the memory 630 to execute all or part of the steps of the face detection and protection method, which includes:

[0091] Obtaining the decomposed image of the video to be detected;

[0092] Inputting the screen image into a pre-trained face detection model to obtain position information of a face region on the screen image, and obtaining a face image with a target face based on the position information of the face region on the screen image;

[0093] The face image is input into a pre-trained face protection model to obtain a face occlusion image that has been occluded.

[0094] On the other hand, the present invention further provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, when the program instructions are executed by a computer, the computer can execute the face detection and protection method provided in the above embodiments, the method comprising:

[0095] Obtaining the decomposed image of the video to be detected;

[0096] Inputting the screen image into a pre-trained face detection model to obtain position information of a face region on the screen image, and obtaining a face image with a target face based on the position information of the face region on the screen image;

[0097] The face image is input into a pre-trained face protection model to obtain a face occlusion image that has been occluded.

[0098] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, all or part of the steps of the face detection and protection method described in the above embodiments are implemented, and the method includes:

[0099] Obtaining the decomposed image of the video to be detected;

[0100] Inputting the screen image into a pre-trained face detection model to obtain position information of a face region on the screen image, and obtaining a face image with a target face based on the position information of the face region on the screen image;

[0101] The face image is input into a pre-trained face protection model to obtain a face occlusion image that has been occluded.

[0102] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the embodiments of the present invention. Those of ordinary skill in the art may understand and implement them without creative effort.

[0103] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the face detection and protection method described in each embodiment or some parts of the embodiment.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A face detection and protection method, characterized in that: include: Obtain the decomposed screen image of the video to be detected; Inputting the screen image into a pre-trained face detection model to obtain position information of a face region on the screen image, and obtaining a face image with a target face based on the position information of the face region on the screen image; Inputting the face image into a pre-trained face protection model to obtain a face occlusion image that has been occluded; The pre-training process of the face detection model includes: The face area is marked on each frame of the decomposed video to be detected; Divide the marked frames into a plurality of image groups, each group including a full-image frame and a plurality of differential-image frames; For each screen image group, the full image is decoded to obtain an RGB image, and the RGB image is used as a training data set and combined with a first neural network for training to obtain a first detection output image; For each screen image group, after fusing the current frame residual image, the current frame motion vector image and the corresponding accumulated residual image and accumulated motion vector image of each of the difference images, superimposing them with the full image in the same screen image group to obtain a composite image, using the composite image as a training data set, combined with the second neural network for training, to obtain a second detection output image; The face detection model is constructed based on the first detection output image and the second detection output image obtained in each screen image group.

2. The face detection and protection method according to claim 1, characterized in that: Also includes: The face occlusion image is input into the face protection model again to obtain a face restoration image after restoration processing.

3. The face detection and protection method according to claim 1, characterized in that: The face region marking refers to marking the position of the face region on the screen image, and the marking value at least includes the horizontal coordinate and vertical coordinate of the face region reference point and the width and height of the face region.

4. The face detection and protection method according to claim 1, characterized in that: The first neural network and the second neural network adopt any one of a convolutional neural network, a recurrent neural network and a combination thereof.

5. The face detection and protection method according to any one of claims 1-2, characterized in that: The pre-training process of the face protection model includes: Inputting the face image into a variational autoencoder for encoding to obtain a face occlusion image; Inputting the face occluded image back into the variational autoencoder for decoding to obtain a face restoration image; The face protection model is constructed based on the face occlusion image and the face restoration image.

6. A face detection and protection device, characterized in that: include: The acquisition module is used to obtain the decomposed picture of the video to be detected; A detection module, used to input the screen image into a pre-trained face detection model, obtain position information of the face area on the screen image, and obtain a face image with a target face based on the position information of the face area on the screen image; An occlusion module is used to input the face image into a pre-trained face protection model to obtain a face occlusion image after occlusion processing; The pre-training process of the face detection model includes: The face area is marked on each frame of the decomposed video to be detected; Divide the marked frames into a plurality of image groups, each group including a full-image frame and a plurality of differential-image frames; For each screen image group, the full image is decoded to obtain an RGB image, and the RGB image is used as a training data set and combined with a first neural network for training to obtain a first detection output image; For each screen image group, after fusing the current frame residual image, the current frame motion vector image and the corresponding accumulated residual image and accumulated motion vector image of each of the difference images, superimposing them with the full image in the same screen image group to obtain a composite image, using the composite image as a training data set, combined with the second neural network for training, to obtain a second detection output image; The face detection model is constructed based on the first detection output image and the second detection output image obtained in each screen image group.

7. The face detection and protection device according to claim 6, characterized in that: Also includes: The restoration module is used to input the face occluded image into the face protection model again to obtain a face restoration image after restoration processing.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, all or part of the steps of the face detection and protection method according to any one of claims 1 to 5 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, all or part of the steps of the face detection and protection method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Privacy protection method and system

    CN108520184A

  • Privacy protection method and device, apparatus and storage medium

    CN110135195A