Privacy shielding method and device and computer equipment
By performing multi-objective frame detection and motion estimation of video frames, and selecting the preferred target frame for occlusion, the missed coding problem in the existing privacy occlusion method is solved and a more accurate occlusion effect is achieved.
Patent Information
- Application Number
- CN202510202577.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-30
AI Technical Summary
The existing privacy occlusion methods have the problem of missed coding and lack effective solutions.
By target detection of the video frame to be blocked, a variety of target boxes are obtained, including face frames, head frames, head and shoulder frames, upper body frames and human body frames. The motion estimation is performed based on the target detection results of the reference frames of the video frame, the target detection results are updated, and the preferred target frames are selected for occlusion.
Through the detection and motion estimation of a variety of target boxes, the accuracy of the occlusion range is ensured, and the problem of missed coding is solved, and the effect of privacy occlusion is improved.
Smart Images

Figure CN120075494A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of camera technologies, and particularly to a privacy occlusion method, apparatus, and computer device. Background Art
[0002] With the development of camera technologies, security cameras can be seen everywhere. With the popularization of security cameras, privacy issues have gradually attracted attention. In some scenarios, it is necessary to perform privacy occlusion on the images captured by cameras.
[0003] Existing privacy occlusion methods mainly use object detection algorithms to detect objects in the pictures captured by cameras, generate detection results of target objects and target bounding boxes corresponding to the target objects. When the detection results of the target objects meet the occlusion trigger conditions, the regions of the target bounding boxes corresponding to the target objects are occluded. Using this method, if the object detection algorithm is inaccurate or there are target objects that the detection algorithm cannot detect normally, there will be a problem of missed blurring.
[0004] However, for the existing privacy occlusion methods, there is a problem of missed blurring, and currently no effective solution has been proposed. Summary of the Invention
[0005] Based on this, it is necessary to provide a privacy occlusion method, apparatus, and computer device for the above technical problems.
[0006] In a first aspect, the present application provides a privacy occlusion method. The method includes the following steps:
[0007] Obtain a video frame to be occluded;
[0008] Perform object detection on the video frame to be occluded to obtain the object detection result of the video frame; the object detection result includes the detected objects and multiple object bounding boxes corresponding to each object; among the multiple object bounding boxes, the ranges of different types of object bounding boxes overlap and have different sizes;
[0009] Based on the object detection result of the reference frame of the video frame, perform motion estimation on each video frame, and update the object detection result of each video frame according to the motion estimation result to obtain the updated object detection result of each video frame;
[0010] Based on the updated object detection result of each video frame, determine the preferred object bounding box of each object in each video frame, and use the preferred object bounding box to occlude each object in the video frame to obtain the occluded video frame.
[0011] In one embodiment, the multiple object bounding boxes include at least two of face bounding boxes, head bounding boxes, head-and-shoulder bounding boxes, upper-body bounding boxes, and body bounding boxes.
[0012] In one embodiment, after performing object detection on the video frame to be occluded and obtaining the object detection result of the video frame, the following steps are included:
[0013] Cache the object detection results of multiple adjacent video frames.
[0014] In one embodiment, the caching of the object detection results of multiple adjacent video frames includes the following steps:
[0015] Configure corresponding IDs for each object in the object detection result of the video frame to obtain the configuration result of each object in the video frame;
[0016] Configure different identifiers for different types of object bounding boxes in the object detection result of the video frame to obtain the configuration result of each object bounding box;
[0017] Cache the object detection results of multiple adjacent video frames based on the configuration result of each object and the configuration result of each object bounding box in the video frame.
[0018] In one embodiment, before performing motion estimation on each video frame based on the object detection result of the reference frame of the video frame, the following steps are included:
[0019] For video frames in the video frame queue except for the first two frames and the last two frames, determine at least two video frames before each video frame and at least two video frames after each video frame as the reference frames of each video frame;
[0020] For the first two frames in the video frame queue, determine at least two video frames after each video frame as the reference frames of each video frame;
[0021] For the last two frames in the video frame queue, determine at least two video frames before each video frame as the reference frames of each video frame.
[0022] In one embodiment, the performing of motion estimation on each video frame based on the object detection result of the reference frame of the video frame includes the following steps:
[0023] Based on the respective object detection results of the reference frames of the video frames, perform motion estimation on each of the video frames to obtain the supplementary detection results of each of the video frames; the supplementary detection results include supplementary objects in the video frames and multiple object bounding boxes corresponding to each supplementary object; the supplementary objects are objects that belong to the first object but do not exist in the object detection results of the current video frame; the first object is an object that exists in all the object detection results of the reference frame of the current video frame.
[0024] In one embodiment, the performing motion estimation on each of the video frames based on the respective object detection results of the reference frames of the video frames to obtain the supplementary detection results of each of the video frames includes the following steps:
[0025] Determine the supplementary objects of the object detection results of the video frames according to the first object in the respective object detection results of the reference frames of the video frames;
[0026] Based on the types of multiple object bounding boxes corresponding to the supplementary objects in the respective object detection results of the reference frames of the video frames, determine the types of the object bounding boxes corresponding to the supplementary objects in each of the video frames;
[0027] Based on the inter-frame pixel differences between the object bounding boxes of the same type of the supplementary objects in the respective object detection results of the reference frames of the video frames, perform motion estimation on each of the video frames to determine the positions of the object bounding boxes corresponding to the supplementary objects in each of the video frames.
[0028] In one embodiment, the determining the types of the object bounding boxes corresponding to the supplementary objects in each of the video frames based on the types of multiple object bounding boxes corresponding to the supplementary objects in the respective object detection results of the reference frames of the video frames includes the following steps:
[0029] Determine the type of the object bounding box that exists in all the object detection results of the reference frames of the video frames of the supplementary objects as the type of the object bounding box corresponding to the supplementary objects in the video frames.
[0030] In one embodiment, the determining the preferred object bounding boxes of the objects in each of the video frames based on the updated object detection results of each of the video frames includes the following steps:
[0031] Based on the multiple object bounding boxes of the objects in the updated object detection results of each of the video frames, determine the preferred object bounding boxes of the objects in each of the video frames according to the preset preferred object bounding box selection rules.
[0032] In a second aspect, the present application also provides a privacy occlusion device. The device includes:
[0033] A video frame acquisition module, configured to acquire a video frame to be occluded;
[0034] A detection module, configured to perform object detection on the video frame to be occluded, and obtain an object detection result of the video frame; the object detection result includes detected objects and multiple object bounding boxes corresponding to each object; among the multiple object bounding boxes, the ranges of different types of object bounding boxes overlap and have different sizes;
[0035] A result update module, configured to perform motion estimation on each video frame based on the object detection result of the reference frame of the video frame, and update the object detection result of each video frame according to the motion estimation result, so as to obtain the updated object detection result of each video frame;
[0036] And an occlusion module, configured to determine a preferred object bounding box for each object in each video frame based on the updated object detection result of each video frame, and use the preferred object bounding box to occlude each object in the video frame, so as to obtain an occluded video frame.
[0037] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the privacy occlusion method described in the first aspect above is implemented.
[0038] The above privacy occlusion method, device, and computer device perform object detection on a video frame to be occluded, obtain an object detection result of the video frame, perform motion estimation on each video frame according to the object detection result of the reference frame of the video frame, and update the object detection result of each video frame according to the motion estimation result. By motion estimation, the object detection result is updated to make the object detection result more accurate. By detecting multiple object bounding boxes with different region ranges, the optimal object bounding box can be selected, so that the occlusion range is also optimal, solving the problem of missed blurring in the existing privacy occlusion methods.
[0039] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:
[0041] Figure 1 It is a hardware structure block diagram of a terminal for the privacy occlusion method provided by an embodiment of the present application;
[0042] Figure 2 The flowchart of the privacy shielding method provided by an embodiment of the present application;
[0043] Figure 3 The flowchart of the privacy shielding method provided by a preferred embodiment of the present application;
[0044] Figure 4 The structural block diagram of the privacy shielding device provided by an embodiment of the present application. Detailed implementation manners
[0045] To understand the purpose, technical solution and advantages of the present application more clearly, the present application will be described and illustrated below with reference to the accompanying drawings and embodiments.
[0046] Unless otherwise defined, the technical terms or scientific terms involved in the present application shall have the general meanings understood by those with ordinary skills in the technical field to which the present application belongs. In the present application, words such as "a", "one", "a kind of", "the", "these" and the like do not indicate a limitation in quantity, and they can be singular or plural. The terms "including", "comprising", "having" and any variations thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connected", "coupled" and the like involved in the present application are not limited to physical or mechanical connections, but may include electrical connections, whether directly or indirectly connected. The term "plurality" involved in the present application means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may mean: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the associated objects before and after are in an "or" relationship. The terms "first", "second", "third" and the like involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.
[0047] The method embodiment provided in this embodiment can be executed on a terminal, a computer or a similar computing device. For example, it runs on a terminal. Figure 1 It is the hardware structural block diagram of the terminal of the privacy shielding method in this embodiment. As Figure 1 shown, the terminal may include one or more ( Figure 1Only one processor 102 and a memory 104 for storing data are shown. Among them, the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA. The above terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 The structure shown is only schematic and does not limit the structure of the above terminal. For example, the terminal may further include Figure 1 more or fewer components than those shown Figure 1 or have a different configuration from that shown.
[0048] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the privacy shielding method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely provided with respect to the processor 102, and these remote memories can be connected to the terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0049] The transmission device 106 is used to receive or send data via a network. The above network includes a wireless network provided by the communication provider of the terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0050] In this embodiment, a privacy shielding method is provided. Figure 2 is a flowchart of the privacy shielding method of this embodiment, as Figure 2 shown, and the process includes the following steps:
[0051] Step S210, obtain a video frame to be shielded.
[0052] The above-mentioned video frame to be occluded can be a pre-stored video frame to be occluded read from a cache space or a storage space, or a video stream captured in real time through a webcam or a video capture card. The above-mentioned video frame to be occluded can be a video frame that needs to be subject to target privacy occlusion in the current scenario. The specific privacy occlusion area can be specifically set according to specific requirements, and this embodiment does not make specific limitations here.
[0053] Step S220: Perform target detection on the video frame to be occluded to obtain the target detection result of the video frame; the target detection result includes the detected targets and multiple target boxes corresponding to each target; among the multiple target boxes, the ranges of different types of target boxes overlap and have different sizes.
[0054] In this step, performing target detection on the video frame to be occluded to obtain the target detection result of the video frame can be based on a preset method to perform target detection on the video frame to be occluded to obtain the target detection result of the video frame. The above-mentioned preset method can be one or more of a preset depth detection model, a preset target detection algorithm, etc. The above-mentioned depth detection model can include one or more of the Fast R-CNN (Fast Regions with Convolutional Neural Networks) model, the SSD (Single Shot MultiBox Detector) network model, the CenterNet (Center Point Detection for Object Detection) model, etc. The above-mentioned preset target detection algorithm can include one or more of methods such as background subtraction, sliding window detection, and the YOLO (You Only Look Once) algorithm. The above-mentioned targets can be one or more of a person, an animal, or other preset targets, etc. It should be noted that the above-mentioned targets can be different according to specific scenarios and requirements, and this embodiment does not make specific limitations here.
[0055] Because if only one type of target box is set, there may be a problem of missing blurring due to unreasonable setting of the target box, or unstable or poor detection effect of a single target box. For example, in the case of only partially showing a side face, if only a face box is set and the face box cannot be recognized due to fewer feature points shown, there will be a problem of missed detection. In this embodiment, by setting multiple types of target boxes, there is an overlap between the ranges of different types of target boxes and the range sizes are different. Appropriate target boxes can be selected through the setting of target boxes of different sizes to solve the problems existing in a single target box. Specifically, when the face box is not accurately recognized, by setting a head box, a head-and-shoulders box, an upper-body box, or a body box that is larger than the range of the face area, and selecting other appropriate target boxes except the face box for privacy occlusion, it is possible to still blur the face area when the face box cannot be detected, solving the problem of missing blurring.
[0056] Step S230: Based on the target detection results of the reference frame of the video frame, perform motion estimation on each video frame, and update the target detection results of each video frame according to the motion estimation results to obtain the updated target detection results of each video frame.
[0057] In this step, the above-mentioned performing motion estimation on each video frame based on the target detection results of the reference frame of the video frame may be performing motion estimation on each video frame based on the target detection results of each reference frame of the video frame to obtain the supplementary detection results of each video frame, and updating the target detection results of each video frame according to the supplementary detection results of each video frame to obtain the updated target detection results of each video frame. Because there may be a situation of missed detection only through the detection results of the current video frame, the detection results of the front and rear video frames can be combined to complement the possible missed detection situation of the current video frame. The above-mentioned motion estimation results may be the supplementary detection results of each video frame, that is, the supplementary targets in the video frame, and multiple target boxes corresponding to each supplementary target.
[0058] Step S240: Based on the updated target detection results of each video frame, determine the preferred target box of each target in each video frame, and use the preferred target box to occlude each target in the video frame to obtain the occluded video frame.
[0059] Among them, the above-mentioned preferred target box may be the optimal target box for occluding the target among multiple target boxes of the target. The above-mentioned determining the preferred target box of each target in each video frame based on the updated target detection results of each video frame may be determining the preferred target box of each target in each video frame according to multiple target boxes of each target in the updated target detection results of each video frame according to a preset preferred target box selection rule.
[0060] The above-mentioned occlusion of each target in the video frame using the preferred target box can be based on the preferred target box and adopt a preset occlusion method to occlude the target information in the area of the preferred target box. The above-mentioned preset occlusion method can be specifically set according to specific requirements, and this embodiment does not make specific limitations here. For example, the mask function can be used to occlude the target information in the area of the preferred target box, and one or more methods such as pixel mask, pixel blur, pixel replacement, and custom occlusion algorithm can also be used to occlude the target information in the area of the preferred target box.
[0061] In the above steps S210 to S240, by performing target detection on the video frame to be occluded, the target detection result of the video frame is obtained, and motion estimation is performed on each video frame according to the target detection result of the reference frame of the video frame, and according to the motion estimation result, the target detection result of each video frame is updated. By motion estimation, the target detection result of the video frame is updated to make the target detection result more accurate. By detecting target boxes in a variety of different region ranges, the optimal target box can be selected, so that the occlusion range is also optimal, solving the problem of missing censor in the existing privacy occlusion method.
[0062] Among them, in one embodiment, multiple types of target boxes include at least two of face box, head box, head and shoulder box, upper body box, and body box.
[0063] It should be noted that the determination of the types of the above multiple targets can specifically set the types of target boxes based on specific requirements, and this embodiment does not make specific limitations here.
[0064] Specifically, in one embodiment, after step S220, it includes:
[0065] Step S1, caching the target detection results of multiple adjacent video frames.
[0066] The above-mentioned caching of the target detection results of multiple adjacent video frames can be to first configure corresponding IDs (Identifications, target identifiers) for each target in the target detection results of multiple adjacent video frames, and configure different identifiers for different types of target boxes in the target detection results, so that each target corresponds to a unique ID, and each type of target box has a corresponding identifier. Then, based on the configuration results of each target in the video frame and the configuration results of each target box, the target detection results of multiple adjacent video frames are cached to obtain a queue of video frames. The number of video frames in the above queue can be specifically set according to specific circumstances, and this embodiment does not make specific limitations here.
[0067] For example, the queue of video frames can be expressed as:
[0068] Seq 0 ,…,SeqM
[0069] Among them, Seq 0 is the first frame in the queue of video frames, and Seq M is the Mth frame in the queue of video frames. M is the length of the video frame queue, indicating that there are M + 1 video frames in the queue, and M is a positive integer.
[0070] It should be noted that since the video frames are obtained from front to back in chronological order, the cached video frames are also arranged in the order of the time when the video frames are obtained or generated.
[0071] In addition, in one embodiment, step S1, caching the object detection results of multiple adjacent video frames, includes:
[0072] Step S12, configuring corresponding IDs for each object in the object detection results of the video frames to obtain the configuration results of each object in the video frames.
[0073] The above-mentioned ID can be represented by letters, numbers, special characters, and combinations of the above situations. In this embodiment, the format of the ID is not specifically limited here, as long as the same object can be characterized by a unique ID. For example, ID 1 can be used to represent the first object in the object detection results of the video frame, and ID 2 can be used to represent the second object in the object detection results of the video frame.
[0074] Step S14, configuring different identifiers for different types of object boxes in the object detection results of the video frames to obtain the configuration results of each object box.
[0075] If the types of object boxes include face boxes, head boxes, head and shoulder boxes, upper body boxes, and body boxes, Rect face can be used as the identifier for the face box, Rect head can be used as the identifier for the head box, Rect shoulder can be used as the identifier for the head and shoulder box, Rect upper can be used as the identifier for the upper body box, and Rect body can be used as the identifier for the body box.
[0076] Then, a single object ID and the corresponding multiple object boxes can be represented as:
[0077] (ID, {Rect face , Rect head , Rect shoulder , Rect upper , Rect body})
[0078] It should be noted that the five types of target bounding boxes in this target do not necessarily all exist, and are specifically determined according to the actual detection results.
[0079] For example, if the face bounding box of this target cannot be detected, this target and the corresponding multiple target bounding boxes can be expressed as:
[0080] (ID, {Rect head , Rect shoulder , Rect upper , Rect body )
[0081] Step S16: Based on the configuration results of each target and the configuration results of each target bounding box in the video frame, cache the target detection results of multiple adjacent video frames.
[0082] Specifically, each video frame can be represented in the form of a set, specifically the set Result i can be expressed as:
[0083]
[0084] where i is the index number of the current frame, and the index number of each frame is unique. ID ij is the ID representation of the target with the serial number j in the current frame, j ∈ [0, N], and j is a positive integer. j represents the serial number of the target ID in the current frame. Among them, the number of targets in the current frame is N + 1.
[0085] The target detection results of multiple adjacent video frames can be cached in the form of a set.
[0086] In the above steps S12 to S16, for each target in the target detection results of the video frame, a corresponding ID is configured, and different identifiers are configured for different types of target bounding boxes in the target detection results of the video frame. Then, based on the configuration results of each target and the configuration results of each target bounding box in the video frame, the target detection results of multiple adjacent video frames are cached. By configuring and caching the target detection results of each video frame, it is convenient to determine the reference frame of each video frame based on the cached results. Furthermore, based on the target detection results of the reference frame of the video frame, motion estimation is performed on each video frame, and according to the motion estimation results, the target detection results of each video frame are updated to obtain the updated target detection results of each video frame.
[0087] In one embodiment, before performing motion estimation on each video frame based on the target detection results of the reference frame of the video frame in step S230, it includes:
[0088] Step S2. For the video frames in the queue excluding the first two frames and the last two frames, at least two video frames before each video frame and at least two video frames after each video frame are determined as the reference frames of each video frame.
[0089] The above-mentioned first two frames are the first frame and the second frame in the queue of video frames arranged in chronological order from front to back. The above-mentioned last two frames are the last two frames in the queue of video frames arranged in chronological order from front to back. When the current video frame does not belong to the first two frames or the last two frames in the queue of video frames, then at least two video frames before the current video frame and at least two video frames after the current video frame are determined as the reference frames of the current video frame. The at least two video frames before each video frame can be adjacent video frames or non-adjacent video frames, which can be selected according to specific requirements, and are not specifically limited in this embodiment. The at least two video frames after each video frame can be adjacent video frames or non-adjacent video frames, which can be selected according to specific requirements, and are not specifically limited in this embodiment either.
[0090] Step S3. For the first two frames in the queue of video frames, at least two video frames after each video frame are determined as the reference frames of each video frame.
[0091] When the current video frame belongs to the first two frames in the queue of video frames, at this time, there are less than two video frames before the current video frame. Therefore, at least two video frames after the current video frame are determined as the reference frames of the current video frame.
[0092] Step S4. For the last two frames in the queue of video frames, at least two video frames before each video frame are determined as the reference frames of each video frame.
[0093] When the current video frame belongs to the last two frames in the queue of video frames, at this time, there are less than two video frames after the current video frame. Therefore, at least two video frames before the current video frame are determined as the reference frames of the current video frame.
[0094] In the above steps S2 to S4, by determining the reference frames of each video frame, it is convenient to perform motion estimation on each video frame according to the object detection results of the reference frames of each video frame, and update the object detection results of each video frame according to the motion estimation results, so as to obtain the updated object detection results of each video frame, making the object detection results of each video frame more accurate.
[0095] In addition, in one embodiment, step S230, performing motion estimation on each video frame based on the object detection results of the reference frames of the video frames, includes:
[0096] Step S232: Based on the respective object detection results of the reference frames of the video frames, perform motion estimation on each video frame to obtain the supplementary detection results of each video frame; the supplementary detection results include supplementary objects in the video frame and multiple object bounding boxes corresponding to each supplementary object; the supplementary objects are objects that belong to the first object but do not exist in the object detection results of the current video frame; the first object is an object that exists in all the object detection results of the reference frame of the current video frame.
[0097] In this step, the above-mentioned motion estimation of each video frame based on the object detection results of the reference frame of the video frame to obtain the supplementary detection results of each video frame can be to perform motion estimation on the object detection results of the current video frame according to the object detection results of the video frames before the current video frame, or / and, the object detection results of the video frames after the current video frame, so as to obtain the information that may be missed in the object detection results of the current video frame. Since the objects in the video frames may be moving, and since the acquisition frequency of each video frame is fixed, therefore, according to the movement of the objects between adjacent frames, the positions of the objects in the next frame or several frames, or, the previous frame or several frames can be estimated. By using this method, in some cases, such as when a face bounding box cannot be detected, the area to be occluded can be estimated according to the object bounding box of the previous frame or the next frame, and occlusion can be performed according to the estimation result.
[0098] Further, in one embodiment, step S232: Based on the respective object detection results of the reference frames of the video frames, perform motion estimation on each video frame to obtain the supplementary detection results of each video frame, including:
[0099] Step S2322: Determine the supplementary objects of the object detection results of the video frame according to the first object in the respective object detection results of the reference frame of the video frame.
[0100] Step S2324: Based on the types of multiple object bounding boxes corresponding to the supplementary objects in the respective object detection results of the reference frame of the video frame, determine the types of the object bounding boxes corresponding to the supplementary objects of each video frame.
[0101] Step S2326: Based on the inter-frame pixel differences between the object bounding boxes of the same type of the supplementary objects in the respective object detection results of the reference frame of the video frame, perform motion estimation on each video frame to determine the positions of the object bounding boxes corresponding to the supplementary objects of each video frame.
[0102] The above-mentioned inter-frame pixel difference can be the pixel difference in the positions in different video frames, that is, the distance difference in terms of pixel values. Since the frequency of obtaining image frames is the same, that is, based on the position change of the same object's same type of object bounding box in the previous video frame or the next video frame, the position of the object bounding box of the current video frame can be determined.
[0103] For example, in a queue of video frames, the reference frames of the first frame are the second frame and the third frame. The position of the head-and-shoulders target box of the third frame is 5 pixels different from that of the head-and-shoulders target box of the second frame and moves to the left. Then, it can be estimated that the position of the head-and-shoulders target box of the second frame is also 5 pixels different from that of the head-and-shoulders target box of the first frame and moves to the left. Based on this, the position of the head-and-shoulders target box of the first frame can be determined. In another example, in a queue of video frames, the reference frames of the first frame are the second frame and the fourth frame. The position of the head-and-shoulders target box of the fourth frame is 6 pixels different from that of the head-and-shoulders target box of the second frame and moves to the left. Then, it can be estimated that the position of the head-and-shoulders target box of the second frame may be 3 pixels different from that of the head-and-shoulders target box of the first frame and moves to the left. Based on this, the position of the head-and-shoulders target box of the first frame can be determined.
[0104] In the above steps S2322 to S2326, the supplementary targets of the object detection results of the video frames are determined through the first objects in the object detection results of the reference frames of the video frames. Based on the types of various target boxes corresponding to the supplementary targets in the object detection results of the reference frames of the video frames, the types of target boxes corresponding to the supplementary targets of each video frame are determined. And based on the inter-frame pixel differences between the target boxes of the same type of the supplementary targets in the object detection results of the reference frames of the video frames, motion estimation is performed on each video frame to determine the positions of the target boxes corresponding to the supplementary targets of each video frame. By determining the supplementary targets and the target boxes corresponding to the supplementary targets, the supplementary detection results of the video frames are updated. Furthermore, based on the updated object detection results of each video frame, occlusion is performed on each object of the video frame, making the detection of the objects and target boxes of the video frames more accurate and avoiding the situation of missed blurring.
[0105] In one embodiment, step S2324, based on the types of various target boxes corresponding to the supplementary targets in the object detection results of the reference frames of the video frames, determining the types of target boxes corresponding to the supplementary targets of each video frame includes:
[0106] Determine the type of the target box corresponding to the supplementary target of the video frame as the type of the target box that exists in all the object detection results of the reference frames of the video frame for the supplementary target.
[0107] For example, the reference frames include a first reference frame and a second reference frame. The types of target boxes that exist in the object detection result of the first reference frame for the supplementary target are a head box, a head-and-shoulders box, and an upper body box. The types of target boxes that exist in the object detection result of the second reference frame for the supplementary target are a head-and-shoulders box, an upper body box, and a human body box. Then, the types of target boxes that exist in both are a head-and-shoulders box and an upper body box. That is, the types of target boxes corresponding to the supplementary target of the obtained video frame are a head-and-shoulders box and an upper body box.
[0108] Among them, in one embodiment, the above step S240, based on the updated object detection results of each video frame, determines the preferred object bounding boxes of each object in each video frame, including:
[0109] Step S242, based on the multiple object bounding boxes of each object in the updated object detection results of each video frame, determines the preferred object bounding boxes of each object in each video frame according to a preset preferred object bounding box selection rule.
[0110] The above preset preferred object bounding box selection rule can be specifically set according to specific requirements, and this embodiment does not make specific limitations here. For example, if the types of object bounding boxes include face bounding boxes, head bounding boxes, head-and-shoulder bounding boxes, upper body bounding boxes, and body bounding boxes, the optimal object bounding box of each object can be selected according to the principle that the priority gradually decreases from the face bounding box to the body bounding box. That is, for each object, when there is a face bounding box in the detection bounding box corresponding to the object, the face bounding box is preferred; in the case where there is no face bounding box, the head bounding box is preferred, and so on, until the optimal object bounding box of the object is determined.
[0111] In one embodiment, if classified by object ID value (i.e., classified by object), the trajectory set of each ID under multiple video frames can be obtained. Furthermore, the position of the object bounding box of the target video frame can be estimated according to the trajectory sets of some video frames.
[0112] For example, object ID 0 's trajectory set can be expressed as:
[0113]
[0114] Among them, represents the object bounding box set of object ID 0 in the video frame with frame number i.
[0115] Each element in the above object ID 0 's trajectory set represents each trajectory bounding box of object ID 0 in the queue of cached video frames.
[0116] Based on the above trajectories for subsequent motion estimation, assuming that the length of the queue of cached video frames is 5, the trajectory set of object ID 0 can be expressed as:
[0117]
[0118] Then, calculate separately according to the object category. Taking the face bounding box as an example, the set of face bounding boxes in the trajectory is: Based on the inter-frame pixel differences of the face bounding boxes in this set, estimate the face coordinates before frame number 0. Assuming the number of frames predicted forward is 5, the prediction result is: The target boxes of other categories are all calculated in this way.
[0119] Final ID 0 The prediction result can be expressed as:
[0120]
[0121] It should be noted that, in order to reduce the situation of incorrect or redundant coding, this embodiment is only taken as an example. To ensure the effectiveness and accuracy of the estimation, usually, only one or two video frames need to be estimated forward or backward. The specific number of video frames to be estimated can be specifically set based on specific requirements.
[0122] The following describes and illustrates this embodiment through preferred embodiments.
[0123] Figure 3 is a flowchart of a privacy occlusion method provided by a preferred embodiment of the present application. As Figure 3 shown, the privacy occlusion method includes the following steps:
[0124] Step S301, obtain the video frame to be occluded.
[0125] Step S302, perform object detection on the video frame to be occluded to obtain the object detection result of the video frame.
[0126] Step S303, cache the object detection results of multiple adjacent video frames to obtain a queue of video frames.
[0127] Step S304, determine the reference frame of each video frame in the queue of video frames.
[0128] For the above determination of the reference frame of each video frame in the queue of video frames, for the video frames in the queue of video frames except for the first two frames and the last two frames, at least two video frames before each video frame and at least two video frames after each video frame can be determined as the reference frame of each video frame; for the first two frames in the queue of video frames, at least two video frames after each video frame can be determined as the reference frame of each video frame; for the last two frames in the queue of video frames, at least two video frames before each video frame can be determined as the reference frame of each video frame.
[0129] Step S305, determine the supplementary object of the object detection result of the video frame according to the first object in the object detection results of the reference frames of the video frame.
[0130] Step S306, determine the type of the target box corresponding to the supplementary object of each video frame based on the types of multiple target boxes corresponding to the supplementary object in the object detection results of the reference frames of the video frame.
[0131] Step S307: Based on the inter-frame pixel differences between the target boxes of the same type in the target detection results of each reference frame of the video frames for each supplementary target, perform motion estimation on each video frame to determine the positions of the target boxes corresponding to the supplementary targets of each video frame.
[0132] Step S308: Based on the supplementary targets in the target detection results of the video frames, the types of the target boxes corresponding to the supplementary targets, and the positions of the target boxes corresponding to the supplementary targets, determine the updated target detection results of each video frame.
[0133] Step S309: Based on the updated target detection results of each video frame, determine the preferred target boxes for each target in each video frame, and use the preferred target boxes to occlude each target in the video frame to obtain the occluded video frame.
[0134] In the above steps S301 to S309, by performing target detection on the video frames to be occluded, the target detection results of the video frames are obtained, and motion estimation is performed on each video frame according to the target detection results of the reference frame of the video frame. According to the motion estimation results, the target detection results of each video frame are updated. By motion estimation, the target detection results of the video frames are updated to make the target detection results more accurate. By detecting target boxes in multiple different regional ranges, the optimal target box can be selected, so that the occlusion range is also optimal, solving the problem of missed blurring in the existing privacy occlusion method.
[0135] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0136] Based on the same inventive concept, in this embodiment, a privacy occlusion device is further provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated here. The following terms "module", "unit", "sub-unit", etc. can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0137] In one embodiment, Figure 4 is a structural block diagram of a privacy shielding device provided by an embodiment of the present application. As Figure 4 shown, the privacy shielding device includes:
[0138] A video frame acquisition module 41, configured to acquire a video frame to be shielded;
[0139] A detection module 42, configured to perform object detection on the video frame to be shielded to obtain an object detection result of the video frame; the object detection result includes detected objects and multiple object bounding boxes corresponding to each object; among the multiple object bounding boxes, the ranges of different types of object bounding boxes overlap and have different sizes;
[0140] A result update module 43, configured to perform motion estimation on each video frame based on the object detection result of the reference frame of the video frame, and update the object detection result of each video frame according to the motion estimation result to obtain the updated object detection result of each video frame;
[0141] And a shielding module 44, configured to determine a preferred object bounding box for each object in each video frame based on the updated object detection result of each video frame, and use the preferred object bounding box to shield each object in the video frame to obtain a shielded video frame.
[0142] The above privacy shielding device performs object detection on the video frame to be shielded to obtain the object detection result of the video frame, performs motion estimation on each video frame according to the object detection result of the reference frame of the video frame, and updates the object detection result of each video frame according to the motion estimation result. It updates the object detection result of the video frame through motion estimation, making the object detection result more accurate. By detecting multiple object bounding boxes with different area ranges, it can select the optimal object bounding box, making the shielding range also optimal, and solving the problem of missed blurring in the existing privacy shielding method.
[0143] It should be noted that the above-mentioned each module can be a functional module or a program module, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned each module can be located in the same processor; or the above-mentioned each module can also be located in different processors in any combination form.
[0144] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements any one of the privacy shielding methods in the above embodiments.
[0145] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0146] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0147] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A privacy shielding method, characterized in that: The method comprises: Get the video frame to be blocked; Performing target detection on the video frame to be occluded to obtain a target detection result of the video frame; the target detection result includes the detected target and a plurality of target frames corresponding to each target; in the plurality of target frames, ranges of different types of target frames overlap and have different range sizes; Based on the target detection result of the reference frame of the video frame, motion estimation is performed on each of the video frames, and according to the motion estimation result, the target detection result of each of the video frames is updated to obtain an updated target detection result of each of the video frames; Based on the updated target detection results of each of the video frames, a preferred target frame of each target in each of the video frames is determined, and each target in the video frame is occluded using the preferred target frame to obtain a video frame after occlusion.
2. The privacy shielding method according to claim 1, characterized in that: The multiple target frames include at least two of a face frame, a head frame, a head-shoulder frame, an upper body frame, and a human body frame.
3. The privacy shielding method according to claim 1, characterized in that: After performing target detection on the video frame to be blocked and obtaining the target detection result of the video frame, the method includes: Cache the object detection results of multiple adjacent video frames.
4. The privacy shielding method according to claim 3, characterized in that: The cached target detection results of a plurality of adjacent video frames include: For each target in the target detection result of the video frame, a corresponding ID is configured to obtain a configuration result of each target in the video frame; For different types of target frames in the target detection result of the video frame, different identifiers are configured to obtain configuration results of each target frame; Based on the configuration results of each target in the video frame and the configuration results of each target frame, the target detection results of multiple adjacent video frames are cached.
5. The privacy shielding method according to claim 1, characterized in that: Before performing motion estimation on each of the video frames based on the target detection result of the reference frame of the video frame, the method includes: For the video frame queue, except for the first two frames and the last two frames, at least two video frames before each of the video frames and at least two video frames after each of the video frames are determined as reference frames for each of the video frames; For the first two frames in the queue of the video frames, at least two video frames located after each of the video frames are determined as reference frames for each of the video frames; For the last two frames in the queue of the video frames, at least two video frames located before each of the video frames are determined as reference frames for each of the video frames.
6. The privacy shielding method according to claim 5, characterized in that: The performing motion estimation on each of the video frames based on the target detection result of the reference frame of the video frame comprises: Based on the target detection results of the reference frame of the video frame, motion estimation is performed on each of the video frames to obtain supplementary detection results of each of the video frames; the supplementary detection results include supplementary targets in the video frame and multiple target frames corresponding to each supplementary target; the supplementary targets are targets that belong to the first target but do not exist in the target detection results of the current video frame; the first target is a target that exists in all target detection results of the reference frame of the current video frame.
7. The privacy shielding method according to claim 6, characterized in that: The method of performing motion estimation on each of the video frames based on each target detection result of the reference frame of the video frame to obtain a supplementary detection result of each of the video frames includes: Determining a supplementary target of the target detection result of the video frame according to the first target in each target detection result of the reference frame of the video frame; Determining the type of target frame corresponding to the supplementary target of each of the video frames based on the types of multiple target frames corresponding to the supplementary target in each target detection result of the reference frame of the video frame; Based on the inter-frame pixel differences between target frames of the same type in the target detection results of each supplementary target in the reference frame of the video frame, motion estimation is performed on each of the video frames to determine the position of the target frame corresponding to the supplementary target of each of the video frames.
8. The privacy shielding method according to claim 7, characterized in that: The determining the type of target frame corresponding to the supplementary target of each video frame based on the types of multiple target frames corresponding to the respective target detection results of the supplementary target in the reference frame of the video frame comprises: The target frame type of the supplementary target that exists in each target detection result of the reference frame of the video frame is determined as the target frame type corresponding to the supplementary target of the video frame.
9. The privacy shielding method according to claim 1, characterized in that: The determining, based on the updated target detection results of each of the video frames, a preferred target frame of each target in each of the video frames comprises: Based on the multiple target frames of each target in the updated target detection results of each of the video frames, the preferred target frame of each target in each of the video frames is determined according to a preset preferred target frame selection rule.
10. A privacy shielding device, characterized in that: The device comprises: A video frame acquisition module, used to acquire the video frame to be blocked; A detection module is used to perform target detection on the video frame to be occluded to obtain a target detection result of the video frame; the target detection result includes the detected target and a plurality of target frames corresponding to each target; in the plurality of target frames, the ranges of different types of target frames overlap and have different range sizes; A result updating module, configured to perform motion estimation on each of the video frames based on the target detection result of the reference frame of the video frame, and update the target detection result of each of the video frames according to the motion estimation result to obtain an updated target detection result of each of the video frames; And an occlusion module is used to determine the preferred target frame of each target in each video frame based on the updated target detection results of each video frame, and use the preferred target frame to occlude each target in the video frame to obtain the occluded video frame.
11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the privacy masking method according to any one of claims 1 to 9 are implemented.