Method and apparatus for applying privacy mask in image on basis of optical flow

The method addresses privacy protection by using object recognition and optical flow to ensure continuous and secure application of privacy masks, even when objects are temporarily unrecognized, enhancing data security and user satisfaction.

WO2025211538A1PCT designated stage Publication Date: 2025-10-09HANWHA VISION CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/021404
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-04
Filing Date
2024-12-30
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing privacy protection methods in images struggle to meet detailed user requirements and pose vulnerabilities in terms of data security, particularly when objects temporarily go unrecognized in frames.

Method used

A method using object recognition and optical flow to apply privacy masks continuously by tracking object movement across frames, applying masks only to designated objects, and terminating them when not recognized for a preset number of frames.

Benefits of technology

Maintains privacy mask continuity and enhances data security by ensuring privacy masks are applied only to intended objects and removed when not detected, reducing storage of sensitive information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024021404_09102025_PF_FP_ABST
    Figure KR2024021404_09102025_PF_FP_ABST
Patent Text Reader

Abstract

A method for applying a privacy mask on the basis of object recognition, according to one embodiment of the present invention, comprises the steps of: receiving an image including a first frame, a second frame and a third frame; performing object recognition on the first frame and, if an object to which a privacy mask is to be applied is recognized, determining region information of the object; if the object to which the privacy mask is to be applied is recognized in the second frame and the object to which the privacy mask is to be applied is not recognized in the third frame, calculating an optical flow-based motion field by using the first frame and the second frame; determining, in the third frame, on the basis of the motion field, an estimation region of the object to which the privacy mask is to be applied; and applying the privacy mask to the estimation region.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for applying privacy masks in images based on optical flow

[0001] The present invention relates to computer vision and image processing technology, and more particularly, to a method and device for applying a privacy mask based on optical flow and object recognition to protect personal privacy in an image.

[0002] With the recent increase in the ease of capturing and sharing videos, concerns about personal privacy are also growing. Existing privacy protection methods primarily involve manually applying privacy masks to specific objects within videos or automated processing using simple algorithms. However, these approaches struggle to meet the detailed user requirements and pose vulnerabilities in terms of data security.

[0003] The background technology described above is technical information that the inventor possessed for the purpose of deriving the present invention or acquired during the process of deriving the present invention, and cannot necessarily be said to be technology disclosed to the general public prior to the application for the present invention.

[0004] The various embodiments described herein have been proposed to solve the above-described problems, and specifically, provide a method for applying a privacy mask with continuity even when the object is temporarily unrecognized within a frame by identifying an object to which a privacy mask is to be applied using object recognition technology in an image and tracking the object's movement in consecutive frames using an optical flow-based motion field. In addition, the method provides a function for applying a privacy mask only to objects designated by the user, thereby satisfying the detailed requirements of the user.

[0005] A method for applying a privacy mask based on object recognition according to an embodiment of the present specification for achieving the above-described task may include: receiving an image including a first frame, a second frame, and a third frame; performing object recognition on the first frame to determine region information of an object to which a privacy mask is to be applied when the object to which the privacy mask is to be applied is recognized; calculating an optical flow-based motion field using the first frame and the second frame when the object to which the privacy mask is to be applied is recognized in the second frame and the object to which the privacy mask is to be applied is not recognized in the third frame; determining an estimated region of the object to which the privacy mask is to be applied in the third frame based on the motion field; and applying a privacy mask to the estimated region.

[0006] The object to which the above privacy mask is applied may include one or more objects specified by the user.

[0007] The method may further include a step of applying a privacy mask to an area of ​​the recognized object when the object to which the privacy mask is applied is recognized in the first frame.

[0008] After the third frame, if the object to which the privacy mask is applied is not recognized in a preset number or more of consecutive frames, a step of terminating the application of the privacy mask may be further included.

[0009] The above motion field can be calculated based on the vertices of the area information of the object.

[0010] The step of determining an estimated area of ​​an object to which a privacy mask is applied based on the motion field may include calculating an estimated area in the third frame by applying movement information indicated by the motion field to the vertices from the object areas recognized in the first frame and the second frame.

[0011] The step of applying a privacy mask to the above-mentioned estimated area may include setting a minimum-sized rectangular area including the estimated area, and applying any one of blurring, mosaic processing, and monochrome filling to the rectangular area to generate a privacy mask.

[0012] The method may include: a step of calculating an optical flow-based motion field using the first frame and the third frame when the object to which the privacy mask is applied is recognized in the third frame and the object to which the privacy mask is applied is not recognized in the second frame; a step of determining an estimated area of ​​the object to which the privacy mask is applied in the second frame based on the motion field; and a step of applying a privacy mask to the estimated area.

[0013] According to one embodiment of the present specification for achieving the above-described object, an image capturing device may include a capturing unit for capturing an image; a control unit for receiving an image including a first frame, a second frame, and a third frame from the image sensor, performing object recognition on the first frame to determine area information of the object when an object to be applied with a privacy mask is recognized, calculating an optical flow-based motion field using the first frame and the second frame when the object to be applied with the privacy mask is recognized in the second frame and the object to be applied with the privacy mask is not recognized in the third frame, determining an estimated area of ​​the object to be applied with the privacy mask in the third frame based on the motion field, and applying a privacy mask to the estimated area.

[0014] According to one embodiment of the present invention, an object to which a privacy mask is applied can be identified and a privacy mask can be applied.

[0015] According to one embodiment of the present invention, by tracking the movement of an object, continuity of privacy mask application can be maintained even in situations where the object is temporarily unrecognized.

[0016] According to one embodiment of the present invention, by distinguishing between the first memory and the second memory, image data containing personal information requiring privacy mask processing can be prevented from being stored in the second memory.

[0017] The effects of the present invention are not limited to the effects mentioned above.

[0018] FIG. 1 illustrates an imaging system according to one embodiment of the present invention.

[0019] FIG. 2 is a flowchart illustrating a masking operation of an image capturing device according to one embodiment of the present invention.

[0020] FIG. 3 is a flowchart illustrating a masking operation of an imaging device according to another embodiment of the present invention.

[0021] FIG. 4 is a flowchart illustrating a masking operation of an imaging device according to another embodiment of the present invention.

[0022] FIG. 5 is a drawing for explaining optical flow in one embodiment of the present invention.

[0023] FIG. 6 and FIG. 7 are drawings for explaining an example of a privacy mask application method according to one embodiment of the present invention.

[0024] FIG. 8 is a drawing for explaining an example of a privacy mask application method according to another embodiment of the present invention.

[0025] FIG. 9 is a block diagram showing the block configuration of an image capturing device according to one embodiment of the present invention.

[0026] The terms used in this invention are used only to describe specific embodiments and may not be intended to limit the scope of other embodiments. Singular expressions may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as commonly understood by those of ordinary skill in the art described in this invention. Terms defined in general dictionaries among the terms used in this invention may be interpreted as having the same or similar meaning in the context of the related technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined in this invention. In some cases, even if a term is defined in this invention, it cannot be interpreted to exclude embodiments of the present invention.

[0027] Hereinafter, various embodiments of the present invention will be described in detail with reference to the attached drawings so that those skilled in the art can easily practice them. However, the technical idea of ​​the present invention can be modified and implemented in various forms and is therefore not limited to the embodiments described in this specification. In describing the embodiments disclosed in this specification, if it is determined that a detailed description of a related known technology may obscure the gist of the technical idea of ​​the present invention, a detailed description of the known technology will be omitted. Identical or similar components will be given the same reference numerals, and redundant descriptions thereof will be omitted.

[0028] Here, the term '~ part' used in this embodiment refers to a component that performs a specific function performed by software or hardware such as an FPGA (Field Programmable Gate Array) or ASIC (Application Specific Integrated Circuit). However, the '~ part' is not limited to being performed by software or hardware. The '~ part' may exist in the form of data stored in an addressable storage medium, or may be implemented by instructions so that one or more processors are configured to perform a specific function.

[0029] Software may include a computer program, code, instructions, or a combination of one or more of these, and may configure a processing device to perform a desired operation or, independently or collectively, command the processing device. The software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage media or devices, or transmitted signal waves, for interpretation by the processing device or for providing instructions or data to the processing device. The software may be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media. The software may be read into main memory from another computer-readable medium, such as a data storage device, or from another device via a communications interface. Software instructions stored in main memory may cause the processor to perform processes or steps, which will be described in detail below. Alternatively, processes consistent with the principles of the present invention can be implemented using hardwired circuitry instead of, or in combination with, software instructions. Therefore, embodiments consistent with the principles of the present invention are not limited to any specific combination of hardware circuitry and software.

[0030] The terminology used in this application is only used to describe specific embodiments and is not intended to limit the present invention. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this application, it should be understood that the terms "comprises" or "has" indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but do not preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. Terms such as first, second, etc. may be used to describe various components, but the components should not be limited by the terms. The terms are used solely for the purpose of distinguishing one component from another.

[0031] 'Object recognition' according to the present invention may refer to identifying and classifying a specific object within a digital image or video. Such object recognition may be a concept encompassing object detection for identifying the location of an object within an image or video, object classification for assigning a detected object to a specific category, and object tracking for tracking the path of an object moving over time within a video, and may be implemented by models such as R-CNN (Regions with Convolutional Neural Network), Fast R-CNN, Faster R-CNN, YOLO (You Only Look Once), SSD (Single Shot Detector), CNN (Convolutional Neural Networks), AlexNet, VGGNet, ResNet, KCF (Kernelized Correlation Filters), TLD (Tracking, Learning and Detection), and SORT (Simple Online and Realtime Tracking).

[0032] The term "optical flow" according to the present invention may refer to a technology for analyzing the movement of an object or scene within a video. Specifically, optical flow represents the movement pattern of each pixel within an image sequence over time, and through this, the speed and direction of an object and dynamic changes in a scene can be estimated. Optical flow represents the movement of an image object between two consecutive frames in the form of a vector, and this vector can indicate the direction and distance (speed) that an object or image patch has moved, and may include a feature-based method such as the Lucas-Kanade method and a deep learning-based method such as FlowNet and PWC-Net.

[0033] FIG. 1 illustrates an imaging system according to one embodiment of the present invention.

[0034] Referring to Fig. 1, the video recording system (100) includes a network (103), a server (105), a user terminal (101), and a video recording device (107). Fig. 1 is an example for explaining the present invention, and the number of devices connected to the network (103) in the system according to one embodiment of the present invention is not limited.

[0035] The video recording system according to the present invention is a system that records a video using a video recording device (107) and provides the video to a user terminal (101). More specifically, the video recording device (107) or server (105) can perform masking on an object to which a privacy mask is applied and instruct a system to process the video.

[0036] The network (103) refers to a network that connects communication between multiple devices, either wired or wirelessly. According to one embodiment of the present invention, the network (103) may include wired networks such as local area networks (LANs), wide area networks (WANs), metropolitan area networks (MANs), and integrated service digital networks (ISDNs), and wireless networks such as wireless LANs, CDMA, Bluetooth, and satellite communication. The network (103) may be a closed network without any contact points or nodes that are connected to an external network. In other words, the network (103) may be a communication line that connects only predetermined components. According to one embodiment of the present invention, the network (103) may include a communication line that connects a server (105), a user terminal (101), and an image capturing device (107).

[0037] The server (105) may be implemented as a computer device or multiple computer devices that communicate with the user terminal (101) through a network to provide commands, codes, files, content, and services. The server (105) may provide content to the user terminal (101). The user terminal (101) may access the server (105) under the control of at least one program and receive services or content provided by the server (105).

[0038] The user terminal (101) refers to an electronic device that acquires information through a network (103) and provides the acquired information to a user. The user terminal (101) includes a fixed user terminal implemented as a computer device or a mobile user terminal. According to one embodiment of the present invention, the user terminal (101) may include a smart phone, a mobile phone, a navigation device, a computer, a laptop, a user terminal for digital broadcasting, a personal digital assistant (PDA), a portable multimedia player (PMP), or a tablet PC. The user terminal (101) may communicate with the server (105) through the network (103) using a wireless or wired communication method. The user terminal (101) may also receive a program or command from the server (105) and perform an operation according to the command.

[0039] According to one embodiment of the present invention, the user terminal (101) can receive user inputs such as touch, tap, drag, and click from the user.

[0040] The video capturing device (107) is a device that captures images by capturing a preset area for purposes such as surveillance or security. The video capturing device (107) can transmit the captured images to a server (105) or a user terminal (101). The video capturing device (107) according to one embodiment of the present invention can apply a privacy mask based on object recognition, and can dynamically apply the privacy mask based on area information of a recognized object.

[0041] The shape and type of the video recording device (107) illustrated in FIG. 1 are exemplary, and the spirit of the present invention is not limited thereto. Any device that acquires an image and transmits the acquired image through a connected network may correspond to the video recording device (107). The operation and configuration of such a video recording device will be described in detail below.

[0042] An image capturing system (100) according to one embodiment of the present invention can display a captured image on the screen of a user terminal (101).

[0043] FIG. 2 is a flowchart illustrating a masking operation of an image capturing device according to one embodiment of the present invention.

[0044] Referring to FIG. 2, in step S210, the video recording device can transmit an image captured by the camera unit to the control unit. The control unit can receive the image captured by the camera unit. The camera unit generates continuous frames through shooting, and the frames generated by the camera unit can be transmitted to a buffer of the camera unit. The frames stored in the buffer are transmitted to the control unit of the video recording device, and the control unit can store the transmitted frames in the first memory. At this time, the frames can be stored in a continuous memory space or stored using a data structure such as a ring buffer. The control unit selects and reads frames at a specific point in time, for example, the first frame and the second frame, from among the frames stored in the first memory. The frames can be selected in chronological order.

[0045] In one embodiment, the video recording device performs object recognition on the first frame at step S220, and if an object to which a privacy mask is applied is recognized, the area information of the object can be determined. The video recording device can use the first frame stored in the first memory area as input to a pre-trained object recognition algorithm, and output the type and location information of the object recognized in the first frame.

[0046] In one embodiment, if the object recognized in the first frame corresponds to an object to which a privacy mask is applied set by the user in step S230, the video capturing device may proceed to step S240 to determine area information of the object. For example, the area information may generally be expressed in the form of a bounding box, and the bounding box may be a rectangular area of ​​the minimum size that includes the object, and may be defined by upper left coordinates (x, y) and width and height values. The object to which the privacy mask is applied is an object that is identified and tracked through object recognition technology and to which at least one of blurring, mosaic, and solid color filling, which is a specific process for privacy protection, is applied. For example, the user may designate a person's face, a vehicle license plate, a body, a document, a specific object, and a symbol as a privacy mask. The object to which the privacy mask is applied may be designated in advance by the user, or may be dynamically identified and determined during an object recognition process using a deep learning model. The video capturing device may store the object area information determined in step S240 in the first memory.

[0047] In one embodiment, the video recording device may determine, at step S230, whether an object recognized in the first frame corresponds to an object to which a privacy mask is applied set by a user. For example, if a privacy mask is applied to a person, the object in the current frame may be recognized as an animal (e.g., a bear). In this case, the video recording device may also determine a probability that the object is not a person. In this case, the video recording device may consider whether the object is terminated only when the probability that the object is not a person is greater than a certain a%. The video recording device may not simply determine the type of object, but may also determine a probability that the object is not a designated object, and consider this when applying the privacy mask.

[0048] In one embodiment, the video recording device may determine, at step S230, whether an object recognized in the first frame corresponds to an object to which a privacy mask is applied set by a user. For example, as another example, if a privacy mask is applied to a person, the person may be recognized as an animal (e.g., a bear) in the current frame. At this time, the video recording device may also determine the probability that the person is an animal. In this case, the video recording device may consider whether the object is terminated only when the probability that the object is an animal is less than a certain a%. The video recording device may not only determine the type of object, but also determine the probability that the object is a designated object, and consider this when applying the privacy mask.

[0049] In one embodiment, the video capturing device can perform object recognition for the second frame as well, in step S250. If the object to which the privacy mask is applied, recognized in the first frame, is also recognized in the second frame, the optical flow can be calculated using the first frame and the second frame. For example, the video capturing device can detect feature points using a corner detector or a blob detector, and estimate the motion between frames by using the extracted feature points for matching corresponding points between frames. The video capturing device can calculate the optical flow from the first frame to the second frame using the matched feature point pairs and extract the motion field corresponding to the area of ​​the object to which the privacy mask is applied, recognized in the second frame.

[0050] In one embodiment, the video capturing device may determine an estimated region for estimating an object position in a subsequent frame based on the optical flow calculated using the first frame and the second frame at step S260. The video capturing device may estimate the object position in a third frame and calculate the vertex positions in the third frame by applying motion vectors corresponding to each vertex of the object region, i.e., the upper left, upper right, lower left, and lower right.

[0051] In one embodiment, the video recording device may, at step S260, apply a privacy mask to the estimated region in the third frame without a separate object recognition process for the third frame. The privacy mask may be applied to a minimum-sized rectangular region containing the estimated region, and for example, the privacy mask may correspond to any one of blurring, mosaic processing, and monochrome filling. The video recording device may repeat this process for the next frame and store the privacy-masked image in the second memory.

[0052] FIG. 3 is a flowchart illustrating a masking operation of an imaging device according to another embodiment of the present invention.

[0053] Referring to FIG. 3, in step S310, the video recording device can transmit an image captured by the camera unit to the control unit. The control unit can receive the image captured by the camera unit. The camera unit generates continuous frames through capturing, and the frames generated by the camera unit can be transmitted to a buffer of the camera unit. The frames stored in the buffer are transmitted to the control unit of the video recording device, and the control unit can store the transmitted frames in the first memory. At this time, the frames can be stored in a continuous memory space or stored using a data structure such as a ring buffer. The control unit selects and reads frames at a specific point in time, such as the first frame, the second frame, and the third frame, from among the frames stored in the first memory. The frames can be selected in chronological order.

[0054] In one embodiment, the video recording device performs object recognition on the first frame at step S320, and if an object to which a privacy mask is applied is recognized, the area information of the object can be determined. The video recording device can use the first frame stored in the first memory area as input to a pre-trained object recognition algorithm, and output the type and location information of the object recognized in the first frame.

[0055] According to one embodiment, if an object recognized in the first frame corresponds to an object to which a privacy mask is applied set by a user in step S330, the video capturing device may proceed to step S340 to determine area information of the object. For example, the area information may generally be expressed in the form of a bounding box, and the bounding box may be a rectangular area of ​​the minimum size that includes the object, and may be defined by upper left coordinates (x, y) and width and height values. The object to which the privacy mask is applied is an object that is identified and tracked through object recognition technology and to which at least one of blurring, mosaic, and solid color filling, which is a specific process for privacy protection, is applied. For example, the user may designate a person's face, a vehicle license plate, a body, a document, a specific object, and a symbol as a privacy mask. The object to which the privacy mask is applied may be designated in advance by the user, or may be dynamically identified and determined during an object recognition process using a deep learning model. The video capturing device may store the object area information determined in step S340 in the first memory.

[0056] In one embodiment, the video capturing device can perform object recognition for the second frame and the third frame at step S350. If the object to which the privacy mask is applied recognized in the first frame is also recognized in the second frame but is not recognized in the third frame, the optical flow can be calculated using the first frame and the second frame. For example, the video capturing device can detect feature points using a corner detector or a blob detector, and estimate the motion between frames by using the extracted feature points for matching corresponding points between frames. The video capturing device can calculate the optical flow from the first frame to the second frame using the matched feature point pairs and extract the motion field corresponding to the area of ​​the object to which the privacy mask is applied recognized in the second frame.

[0057] In one embodiment, the video capturing device may determine an estimated region for estimating an object position in a subsequent frame based on the optical flow calculated using the first frame and the second frame in step S360. The video capturing device may estimate the object position in a third frame and calculate the vertex positions in the third frame by applying motion vectors corresponding to each vertex of the object region, i.e., the upper left, upper right, lower left, and lower right.

[0058] In one embodiment, the video capturing device may apply a privacy mask to the estimated area in the third frame at step S360. The privacy mask may be applied to a minimum-sized rectangular area including the estimated area, and for example, the privacy mask may correspond to any one of blurring, mosaic processing, and monochrome filling. The video capturing device may repeat this process for the next frame. In addition, the video capturing device may end the application of the privacy mask if the object to which the privacy mask is applied is not recognized in a preset number or more of consecutive frames after the third frame. The video capturing device may store the privacy-masked image in the second memory.

[0059] FIG. 4 is a flowchart illustrating a masking operation of an imaging device according to another embodiment of the present invention.

[0060] Referring to FIG. 4, in step S410, the video recording device can transmit the video captured by the recording unit to the control unit. The control unit can receive the video captured by the recording unit. The recording unit generates continuous frames through recording, and the frames generated by the recording unit can be transmitted to a buffer of the recording unit. The frames stored in the buffer are transmitted to the control unit of the video recording device, and the control unit can store the received frames in the first memory. At this time, the frames can be stored in a continuous memory space or stored using a data structure such as a ring buffer. The control unit selects and reads frames at a specific point in time, such as the first frame, the second frame, and the third frame, from among the frames stored in the first memory. The frames can be selected in chronological order.

[0061] In one embodiment, the video recording device performs object recognition on the first frame at step S420, and if an object to which a privacy mask is applied is recognized, the area information of the object can be determined. The video recording device can use the first frame stored in the first memory area as input to a pre-trained object recognition algorithm, and output the type and location information of the object recognized in the first frame.

[0062] In one embodiment, if the object recognized in the first frame corresponds to an object to which a privacy mask is applied set by the user in step S430, the video capturing device may proceed to step S440 to determine area information of the object. For example, the area information may generally be expressed in the form of a bounding box, and the bounding box may be a rectangular area of ​​the minimum size that includes the object, and may be defined by upper left coordinates (x, y) and width and height values. The object to which the privacy mask is applied is an object that is identified and tracked through object recognition technology and to which at least one of blurring, mosaic, and solid color filling, which is a specific process for privacy protection, is applied. For example, the user may designate a person's face, a vehicle license plate, a body, a document, a specific object, and a symbol as a privacy mask. The object to which the privacy mask is applied may be designated in advance by the user, or may be identified and determined dynamically during an object recognition process using a deep learning model. The video capturing device may store the object area information determined in step S440 in the first memory.

[0063] In one embodiment, the video capturing device can perform object recognition for the second frame and the third frame at step S450. If the object to which the privacy mask is applied recognized in the first frame is also recognized in the third frame but is not recognized in the second frame, the first frame and the third frame can be used to calculate the optical flow. For example, the video capturing device can detect feature points using a corner detector or a blob detector, and estimate the motion between frames by using the extracted feature points for matching corresponding points between frames. The video capturing device can calculate the optical flow from the first frame to the third frame using the matched feature point pairs and extract the motion field corresponding to the area of ​​the object to which the privacy mask is applied recognized in the second frame.

[0064] In one embodiment, the video capturing device can determine an estimated region for estimating an object position in the next frame based on the optical flow calculated using the first frame and the third frame at step S460. The video capturing device can estimate the object position in the second frame and calculate the vertex positions in the second frame by applying motion vectors corresponding to each vertex of the object region, which is the upper left, upper right, lower left, and lower right.

[0065] In one embodiment, the video recording device may apply a privacy mask to the estimated region in the second frame at step S460. The privacy mask may be applied to a minimum rectangular region containing the estimated region, and for example, the privacy mask may correspond to any one of blurring, mosaic processing, and monochrome filling. The video recording device may repeat this process for the next frame. Additionally, the video recording device may store the privacy-masked image in the second memory.

[0066] Figure 5 is a diagram illustrating optical flow in one embodiment of the present invention. Specifically, for the first, second, and third frames, the description focuses on cases where the privacy masking target object is recognized in the first and second frames, and the privacy masking target object is not recognized in the third frame. This is for convenience of explanation, and it will be readily apparent to those skilled in the art that other embodiments can be implemented using the same method.

[0067] The first area (510) illustrated in FIG. 5 can indicate area information in the first frame, and the second area (520) can indicate area information for the second frame.

[0068] As illustrated in Fig. 5, optical flow can be calculated based on the movement of an object between the first region (510) and the second region (520). Optical flow is a technology for estimating information on the movement of an object between consecutive frames, i.e., a motion vector, and an imaging device can predict the position of an object to which a privacy mask is applied.

[0069] For example, when the coordinates of the upper left vertex of the first region (510) are (x1, y1), the coordinates of the lower right vertex are (x2, y2), the coordinates of the upper left vertex of the second region (520) are (x3, y3), and the coordinates of the lower right vertex are (x4, y4), the video recording device can calculate the movement vector between each vertex as follows.

[0070] Top left corner translation vector: (x3 - x1, y3 - y1)

[0071] Bottom right vertex translation vector: (x4 - x2, y4 - y2)

[0072] Based on the vertex motion vectors generated in this way, the video camera can more precisely predict the object's direction and speed of movement. Furthermore, it can indirectly determine information such as the object's size change and rotation. The video camera can estimate the object's position estimation area in the third frame as follows:

[0073] Predicted upper left corner coordinates: (x3 + (x3 - x1), y3 + (y3 - y1))

[0074] Predicted bottom-right corner coordinates: (x4 + (x4 - x2), y4 + (y4 - y2))

[0075] The video capture device determines the area to which the privacy mask is applied in the third frame based on the predicted vertex coordinates. At this time, the video capture device sets a rectangular area that includes the predicted area and is larger than the object area in the first and second frames, and applies the privacy mask to the area using at least one of blurring, mosaic processing, and / or monochrome filling.

[0076] FIG. 6 and FIG. 7 are drawings for explaining an example of a privacy mask application method according to one embodiment of the present invention.

[0077] Referring to FIGS. 6 and 7, drawings and flowcharts illustrating an example of a privacy mask termination condition and a privacy mask application method according to one embodiment of the present invention are shown.

[0078] According to one embodiment of the present invention, a privacy mask termination condition may be such that the privacy mask may be terminated if an object recognized in the previous n frames does not correspond to an object to which the privacy mask is applied. In this case, n is an arbitrary integer. For example, if n is set to 3, the privacy mask may be deactivated from the current frame if an object recognized in three consecutive frames does not correspond to an object to which the privacy mask is applied.

[0079] For example, as illustrated in FIG. 6, in a privacy mask termination condition according to another embodiment of the present invention, the coordinates of the area of ​​a recognized object may be represented as A and B, and the width and height of the object may be expressed as W and H, respectively. Next, if the x-axis coordinate of the object in the next frame is not consecutively within the range of Aa <= area (x-coordinate) of the object <= A+W+a, and the y-axis coordinate of the object is not consecutively within the range of Bb <= area (y-coordinate) of the object <= B+H+b for n frames, the mask application may be terminated. At this time, a, b, and n are different arbitrary integers.

[0080] For example, as illustrated in FIG. 7, in the mask termination condition according to one embodiment of the present invention, when a privacy mask is applied only to a specific recognized object, the privacy mask may be terminated due to recognition of another object. Therefore, rather than terminating the privacy mask solely by the type of object, the termination condition may be set using the object's recognized x, y coordinates together with the width and height. For example, assuming that the privacy mask is applied only to a car, in the next frame, the object may be mistakenly recognized as a motorcycle, and the privacy mask application may be canceled for the object recognized as a motorcycle.

[0081] In steps S510, S520, and S530, the privacy mask termination condition may be determined not only by the type of object, but also by determining whether the width or height of the current frame object is similar to the width or height of the previous frame object, and whether the current frame x or y coordinate is similar to the x or y coordinate of the previous frame object, so that the privacy mask application may be released.

[0082] FIG. 8 is a drawing for explaining an example of a privacy mask application method according to another embodiment of the present invention.

[0083] Referring to Figure 8, if an object gradually disappears from the image, object recognition may be interrupted. In this case, if the upper right (A+W, B), lower right (A+W, B+H), lower left (A, B+H), or upper left (A, B) of the object is close to the start or end point of the image, the object's movement can be predicted using optical flow. Therefore, even if object recognition is interrupted, the privacy mask can be continuously applied until the next n frames.

[0084] In Fig. 8 (a) and (b), an object is recognized and the direction and intensity of the movement can be predicted. As shown in Fig. 8 (a) and (b), the direction of movement is from left to right, and assuming that the intensity is 10 pixels in the right direction per frame, the value of a is 10. In Fig. 8 (c), although an object is not recognized, a privacy mask can be applied due to the direction and intensity of the movement in the previous frame. Therefore, a privacy mask can be predicted and applied as in Fig. 8 (d). Fig. 8 is an example of a direction from left to right, but a privacy mask can be applied in any direction as long as the direction and intensity of the movement are known.

[0085] FIG. 9 is a block diagram schematically illustrating the block configuration of an image capturing device according to one embodiment of the present invention.

[0086] A third image of the tolerance area can be created based on the parking management image captured by the camera unit (610) of the video capture device (107) to enable parking space management.

[0087] Referring to FIG. 9, the video recording device (107) may include a shooting unit (610), a communication unit (620), a first memory (630), a second memory (640), and a control unit (650). However, the illustrated components are not essential components. The video recording device (107) may be implemented with more components than the illustrated components, or may be implemented with fewer components. The components will now be described.

[0088] The photographing unit (610) can continuously capture images based on an image sensor that converts light energy into an electrical signal. The photographing unit (610) can store image data generated based on the image sensor in a buffer (611). The buffer (611) is a configuration for high-speed storage and transmission of image data, and can sequentially store images and transmit them according to a request from the control unit, for example, as a FIFO structure.

[0089] The communication unit (620) can be connected to a network via a wired or wireless connection and communicate with an external device. Here, the external device may be a server (105) and a user terminal (101). The communication unit (620) may support wireless communication protocols such as Wi-Fi, Bluetooth, Zigbee, and 3G / 4G / 5G mobile communication networks, or may support wired communication protocols such as Ethernet, USB, HDMI, and RS-232 serial ports.

[0090] The first memory (630) may designate a memory for high-speed storage and access of frame data, various setting values, program codes, etc. within the video recording device. For example, the first memory (630) may be implemented with a DRAM (Dynamic Random Access Memory) and an SRAM (Static RAM). The first memory (630) is utilized as a memory space for buffering video data and inter-frame reference, and may also be used to store program codes and intermediate data for video processing. In addition, the first memory (630) may temporarily store data received through a network or load / save various setting values.

[0091] The second memory (640) may indicate a memory in which stored data is preserved even when power is not supplied. The second memory (640) may be one of flash memory, ROM (Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), magnetic disc storage device, compact disc ROM (CD-ROM), digital versatile discs (DVDs), or other forms of optical storage device, magnetic cassette.

[0092] The control unit (650) can execute a program stored in the memory, read data or files stored in the memory, or store new files in the memory. The control unit (650) can execute commands stored in the memory. The control unit (650) can control the imaging device (107) to perform the operations of FIGS. 2 to 4, which are the overall operations of the imaging device.

[0093] The control unit (650) according to one embodiment of the present invention may receive sequential frames from the photographing unit (610) and store them in the first memory (630), or may receive frames stored in the first memory (630). The control unit (650) may apply a pre-learned object recognition algorithm to the first frame to determine whether a privacy mask target object has been recognized.

[0094] A control unit (650) according to one embodiment of the present invention receives an image including a first frame, a second frame, and a third frame from an image sensor, performs object recognition on the first frame, and if an object to which a privacy mask is applied is recognized, determines area information of the object, if an object to which a privacy mask is applied is recognized in the second frame and the object to which a privacy mask is applied is not recognized in the third frame, calculates an optical flow-based motion field using the first frame and the second frame, determines an estimated area of ​​an object to which a privacy mask is applied in the third frame based on the motion field, and applies a privacy mask to the estimated area.

[0095] The control unit (650) can store the privacy-masked image in the second memory (640).

[0096] According to one embodiment of the present invention, the control unit (650) can receive an image from the photographing unit (610) and perform AI object recognition, privacy mask correction, and privacy mask application. At this time, since it is in a single process, latency is significantly reduced, so the efficiency of privacy mask application can be increased. The value of metadata obtained through AI recognition may have an error depending on the recognition accuracy of the AI ​​model. At this time, the error rate of privacy mask application can be reduced through privacy mask correction. In addition, if the metadata value of AI recognition is accurate enough that privacy mask correction is not necessary, the privacy mask can be applied immediately.

[0097] The control unit (650) may include a CPU, RAM, ROM, a system bus, etc. The control unit (650) may be implemented with a single CPU or multiple CPUs (or DSP, SoC). In one embodiment, the control unit (650) may be implemented as a digital signal processor (DSP), a microprocessor, or a time controller (TCON) that processes digital signals. However, the present invention is not limited thereto, and may include one or more of a central processing unit (CPU), a micro controller unit (MCU), a micro processing unit (MPU), a controller, an application processor (AP), a communication processor (CP), or an ARM processor, or may be defined by the relevant terms. In addition, the control unit (650) may be implemented as a system on chip (SoC) or large scale integration (LSI) having a built-in processing algorithm, or may be implemented in the form of a field programmable gate array (FPGA). Furthermore, the control unit (650) may include a neural processing unit (NPU), a graphics processing unit (GPU), and a tensor processing unit (TPU).

[0098] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0099] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.

Claims

1. In a method of applying a privacy mask based on object recognition, A step of receiving an image including a first frame, a second frame, and a third frame; A step of performing object recognition on the first frame and determining area information of the object when an object to which a privacy mask is applied is recognized; A step of calculating an optical flow-based motion field using the first frame and the second frame, when the object to which the privacy mask is applied is recognized in the second frame and the object to which the privacy mask is applied is not recognized in the third frame; A step of determining an estimated area of ​​an object to which the privacy mask is applied in the third frame based on the motion field; and A method comprising the step of applying a privacy mask to the above-mentioned estimated area.

2. In paragraph 1, A method in which the above privacy mask application target object includes one or more objects specified by the user.

3. In paragraph 1, A method further comprising the step of applying a privacy mask to an area of ​​the recognized object when the object to which the privacy mask is applied is recognized in the first frame.

4. In paragraph 1, A method further comprising a step of terminating the application of the privacy mask when the object to which the privacy mask is applied is not recognized in a preset number of consecutive frames after the third frame.

5. In paragraph 1, A method in which the above motion field is calculated based on the vertices of the area information of the above object.

6. In paragraph 5, A method in which the step of determining an estimated area of ​​an object to which a privacy mask is applied based on the motion field includes applying movement information indicated by the motion field to the vertices from the object areas recognized in the first frame and the second frame to derive an estimated area in the third frame.

7. In paragraph 1, The step of applying a privacy mask to the above estimated area is: A method comprising: setting a minimum size rectangular area including the above-mentioned estimated area, and generating a privacy mask by applying any one of blurring, mosaic processing, and monochrome filling to the rectangular area.

8. In paragraph 1, A step of calculating an optical flow-based motion field using the first frame and the third frame, when the object to which the privacy mask is applied is recognized in the third frame and the object to which the privacy mask is applied is not recognized in the second frame; A step of determining an estimated area of ​​an object to which the privacy mask is applied in the second frame based on the motion field; and A method comprising the step of applying a privacy mask to the above-mentioned estimated area.

9. In the video recording device, The camera crew that shoots the video; A first memory that temporarily stores and processes data; A second memory that retains stored data even when power is cut off; and An image capturing device comprising a control unit configured to receive an image including a first frame, a second frame, and a third frame from the image sensor from the photographing unit or the first memory, perform object recognition on the first frame, and determine area information of the object when an object to which a privacy mask is applied is recognized, calculate an optical flow-based motion field using the first frame and the second frame when the object to which the privacy mask is applied is recognized in the second frame and the object to which the privacy mask is applied is not recognized in the third frame, determine an estimated area of ​​the object to which the privacy mask is applied in the third frame based on the motion field, apply a privacy mask to the estimated area, and store an image with the privacy mask applied to the estimated area in the second memory.

Citation Information

Patent Citations

  • Electronic device changing radio frequency path based on SAR and method for operating thereof

    KR1020230106478A

  • System and method for providing real-time prescription sharing and dispensing information monitoring service based on blockchain network

    KR1020240036251A

  • Semiconductor device with gate isolation structure and method for forming the same

    KR102647010B1

  • Image processing method and device thereof

    WO2020171257A1

  • KR20220093548A