De-identification system for recognition object in CCTV image
By detecting and converting the feature points of the actual object in the edge device, forming virtual objects, and using these virtual objects as artificial intelligence learning data on the management server, the problem of personal information protection when facial recognition data is leaked in the prior art is solved, and secure facial recognition data use and privacy protection are achieved.
Patent Information
- Application Number
- JP2024192269
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-02
- Filing Date
- 2024-10-31
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-10-31
AI Technical Summary
The prior art is difficult to use facial recognition images as artificial intelligence learning data on the premise of ensuring the privacy of personal information, especially when facial recognition databases are leaked, personal information cannot be effectively protected.
By detecting unique feature points of actual objects in edge devices, converting them into virtual objects, and using these virtual objects as AI learning data on the management server. The system includes a facial recognition unit, a facial de-recognition execution unit and a data transmission unit, which can process and convert video data in real time.
It realizes safe use of facial recognition images while maintaining facial recognition features and tracking attributes, and effectively protects personal privacy and prevents unauthorized use and disclosure of sensitive personal information.
Smart Images

Figure 2025077027000001_ABST
Abstract
Description
[Technical field]
[0001] An embodiment of the present invention relates to a de-identification system for recognizable objects in CCTV footage.
[0002] This invention was supported by the following national research and development projects: [Project unique number]1781000008 [Subject number]RS-2023-00221181 [Department name] Personal Information Protection Commission [Name of issue management (specialized) agency] Korea Internet and Security Agency [Research project name] Research and development (R&D) on technology to strengthen personal information protection [Research Project Title] Real-time face de-identification technology that enables analysis of connections between the same subjects in face recognition CCTV [Contribution rate] 1 / 1 [Name of project executing organization] World Vertex Co., Ltd. [Research Period] 2023.04.01~2025.12.31 [Background technology]
[0003] Recently, the use of deep learning technology has been rapidly increasing in the field of intelligent video surveillance technology, and in some applications, successful technological developments have been made that exceed human recognition capabilities. A large-scale learning database is required to develop face recognition technology based on deep learning, but if the constructed learning database is exposed to the outside, there is a possibility that privacy issues may arise.
[0004] For this reason, it is difficult to even secure a learning database. As a result, there is a demand for protection technology that prevents personal information from being exposed and allows the video collected for use as a learning database to be used even if it is leaked to the outside.
[0005] In order to solve the problem of privacy violations that may occur when facial images are exposed, techniques such as masking of video information (mosaic, smiley icon processing, etc.) and encryption techniques are being used.
[0006] There is a problem in that CCTV footage processed in this way cannot be used for marketing or as data for artificial intelligence learning. Even if such existing technology is applied to a face learning database, it is not a fundamentally safe solution because the original face image must be used for face learning after going through an unmasking or decoding process.
[0007] In particular, while mosaic-processed or blurred video is an appropriate method for avoiding violation of portrait rights and privacy, it is virtually impossible to use it as learning data because it is impossible to recognize or detect features of human objects in the video. In addition, when a large number of people appear in a video, the video itself becomes distracting due to the numerous mosaics and blurs, making it difficult for viewers to concentrate, making it difficult to use for marketing purposes. [Prior art documents] [Patent documents]
[0008] [Patent Document 1] Korean Patent Publication No. 10-2023-0080111 (Publication date: June 7, 2023) [Patent Document 2] Korean Patent Publication No. 10-2020-0036656 (Publication date: April 7, 2020) [Patent Document 3] Korean Patent Publication No. 10-2515011 (Registration Date: March 23, 2023) [Patent Document 4] Korean Patent Publication No. 10-2022-0122457 (Publication date: September 2, 2022) [Patent Document 5] Korean Patent Publication No. 10-2420151 (Publication Date: July 7, 2022) Summary of the Invention [Problem to be solved by the invention]
[0009] An embodiment of the present invention provides a de-identification system for recognized objects in CCTV images that de-identifies faces in captured images by converting them into non-existent virtual faces in real time, while maintaining unique features and traceability so that they can be used for artificial intelligence learning and analysis, thereby enabling safe use of facial recognition images and protecting privacy, and preventing the unauthorized use and infringement of sensitive personal information. [Means for solving the problem]
[0010] A system for de-identifying recognized objects in CCTV images according to an embodiment of the present invention includes an edge device that detects unique feature points of a real object recognized from CCTV images, converts the real object whose unique feature points have been detected into a virtual object to de-identify the real object, and transmits the de-identified image and the unique feature points, and a management server that receives the de-identified image and the unique feature points from the edge device and uses them as data for artificial intelligence learning.
[0011] The edge device may also include a person object face recognition unit that detects unique facial feature points of a real person object from a CCTV video and recognizes the real person object, respectively; a person object face de-identification execution unit that converts the real person object into a virtual person object based on a change in external attributes of the face of the real person object including the unique facial feature points as metadata, thereby generating a de-identified video in which the real person object in the CCTV video has been de-identified; and a data transmission unit that compresses and transmits the de-identified video and the metadata to the management server.
[0012] Also, the human object face recognition unit may include a face detection area setting unit that sets a preset point using an alpha pose of a pose estimation model as a face detection area for the actual human object.
[0013] In addition, the human object face recognition unit includes a parameter adjustment unit that adjusts at least one parameter of a minimum size of a detection area for a real human object, a detection reliability, a maximum number of detections, and a non-detection option, and the parameter adjustment unit can adjust the parameter for the non-detection option by adjusting the number of unique facial feature points detected from the human object.
[0014] In addition, the human object face de-identification execution unit may convert the real human object into a virtual human object by changing at least one external attribute of the real human object among hair color, hairstyle, skin color, beard, aging, and facial expression.
[0015] In addition, the human object face de-identification execution unit divides the external attributes of the face into domains and adjusts the degree of change for each external attribute, and may adjust the degree of de-identification so that the results of the changes in the external attributes do not overlap with actual human objects in the CCTV video.
[0016] The edge device may further include a human object clothing feature detection unit that detects clothing features of the real human object from CCTV footage, and a human object clothing de-identification execution unit that de-identifies the clothing of the real human object based on changes in external attributes of the clothing of the real human object based on the clothing features.
[0017] The human object clothing non-identification executing unit may change any one of external attributes of the color, design pattern, style, and type of the clothing of the real human object.
[0018] In addition, the human object clothing de-identification execution unit divides the external attributes of clothing into domains and adjusts the degree of change for each external attribute, and may adjust the degree of de-identification so that changes in the external attributes do not overlap between actual human objects in the CCTV footage. Effect of the Invention
[0019] According to the present invention, a de-identification system for recognized objects in CCTV images can be provided, which performs de-identification processing by converting faces in captured images into non-existent virtual faces in real time, while maintaining unique features and traceability so that they can be used for artificial intelligence learning and analysis, thereby enabling safe use of facial recognition images and protecting privacy, and preventing the unauthorized use and infringement of sensitive personal information. [Brief description of the drawings]
[0020] [Figure 1] 1 is a schematic diagram showing the overall configuration of a system for de-identifying a recognized object in a CCTV video according to an embodiment of the present invention; [Diagram 2] FIG. 2 is a block diagram showing a configuration of an edge device according to an embodiment of the present invention. [Diagram 3] 2 is a block diagram showing a configuration of a person object face recognition unit according to the embodiment of the present invention. FIG. [Figure 4] 5 is a diagram illustrating functions of a face detection area setting unit according to the embodiment of the present invention. FIG. [Diagram 5] 4 is a diagram for explaining a function of a parameter adjustment unit according to the embodiment of the present invention. FIG. [Figure 6] FIG. 2 is a diagram illustrating an overall operation of an edge device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0021] The terms used in this specification will be briefly explained below, and the present invention will be described in detail. The terms used in the present invention are selected as widely used general terms as possible in consideration of the functions in the present invention, but they may change depending on the intentions or precedents of the engineers in the field, the emergence of new technologies, etc. In some cases, the applicant may arbitrarily select terms, and in this case, the meanings of the terms will be explained in detail in the description of the present invention. Therefore, the terms used in the present invention are defined based on the meanings of the terms and the overall content of the present invention, rather than simply the names of the terms. Throughout the specification, when a part "includes" a certain component, this does not mean that other components are excluded, but that other components may be included, unless otherwise specified. Furthermore, the terms "unit" and "module" described in the specification mean a unit that processes at least one function or operation, which may be realized as hardware or software, or a combination of hardware and software. Hereinafter, the embodiments of the present invention will be described in detail with reference to the accompanying drawings so that those skilled in the art to which the present invention pertains can easily carry out the embodiments. However, the present invention may be embodied in various different forms and is not limited to the embodiments described herein. In addition, in the drawings, parts that are not related to the description are omitted in order to clearly explain the present invention, and similar parts are designated by similar reference numerals throughout the specification.
[0022] Figure 1 is an overview diagram showing the overall configuration of a de-identification system for recognized objects in CCTV footage according to an embodiment of the present invention, Figure 2 is a block diagram showing the configuration of an edge device according to an embodiment of the present invention, Figure 3 is a block diagram showing the configuration of a human object face recognition unit according to an embodiment of the present invention, Figure 4 is a diagram showing the function of a face detection area setting unit according to an embodiment of the present invention, Figure 5 is a diagram shown to explain the function of a parameter adjustment unit according to an embodiment of the present invention, and Figure 6 is a diagram showing the overall operation of the edge device according to an embodiment of the present invention.
[0023] Referring to FIG. 1 , a de-identification system 1000 for a recognizable object in a CCTV video according to an embodiment of the present invention may include at least one of an edge device 100 and a management server 200.
[0024] The edge device 100 detects unique features of a real object recognized from a CCTV image, converts the real object whose unique features have been detected into a virtual object, and de-identifies the real object. The edge device 100 can then transmit the de-identified image and data (metadata) regarding the unique features to the management server 200 as artificial intelligence learning data.
[0025] In this embodiment, the object may be a person (person), but is not limited thereto, and may be an animal (pet such as a dog or cat). However, the following objects will be described with a person (person) as the object most suitable for the purpose of the invention.
[0026] Such an edge device 100 may include at least one of a human object face recognition unit 110, a human object face de-identification execution unit 120, a human object clothing feature detection unit 130, a human object clothing de-identification execution unit 140, and a data transmission unit 150, as shown in FIG. 2.
[0027] The human object face recognition unit 110 detects unique facial feature points of real human objects from CCTV images, and can recognize the real human objects.
[0028] Such a human object face recognition unit 110 can detect faces in real time from images and videos using the "Retina Face" model, which is a SOTA model, and extract unique facial feature points for each object from the detected face image. By changing the "Backbone" of the "Retina Face" model to a lightweight model, real-time inference is possible even in a CPU environment. The "Retina Face" model is a model that can flexibly extract from 5 to 1000 feature points, and is suitable for using feature points in object analysis.
[0029] In addition, the human object face recognition unit 110 can extract feature points for face images of various scales that are problematic in CCTV video by applying FPN (Feature Pyramid Network). In addition, the human object face recognition unit 110 can maintain high detection performance even for small face images in CCTV video by adding a "Context Module" to each "Pyramid" level and strengthening the storage area.
[0030] In addition, the human object face recognition unit 110 can realize the preprocessing function for the image of the object detection model using CUDA (Compute Unified Device Architecture), and can also realize high speed processing using the "TensorRT" model. The human object face recognition unit 110 may include at least one of a face detection area setting unit 111 and a parameter adjustment unit 112, as shown in FIG.
[0031] The face detection area setting unit 111 can set a preset point as a face detection area for a real human object using an alpha pose of a pose estimation model for face detection rate and continuous tracking. If only a face detection model is applied, the accuracy of the side of the face is poor and the back cannot be detected, so that the continuous tracking performance of the object is degraded. Therefore, by applying a pose estimation model, continuous tracking of the object is possible even if face detection is impossible.
[0032] For example, a face region may be set on a point indicating the neck by pose estimation, and a detection process may be performed based on the set region. Since it is possible to estimate that a face exists on a specific key point obtained by such pose estimation, it is possible to provide more reliable results by combining it with existing face recognition results.
[0033] The parameter adjustment unit 112 may adjust at least one parameter of a minimum size of a detection region for a real human object, a detection confidence level, a maximum number of detections, and a non-detection option.
[0034] More specifically, the parameter adjustment unit 112 can adjust parameters for the non-detection option by setting (adjusting) parameters for the minimum size of a face that can be detected from an image or video, setting (adjusting) a threshold for detection reliability, setting (adjusting) the maximum number of clustered faces that can be detected, and adjusting the number of unique face feature points. In the case of the non-detection option, since there is a high possibility that facial feature points will not be detected for an object with a low degree of face exposure due to wearing a mask or the like, by adjusting the number of feature points and adjusting the load on the edge device 100, it becomes possible to detect feature points for an object even in such a situation.
[0035] In this way, the human object face recognition unit 110 detects facial feature points in real time in a crowded area using a walk-through method, and in this walk-through method, it is important to recognize not only the front face but also the side and back, and for this reason, by applying a posture estimation model, it is possible to recognize faces not only from the front and side but also from the side and back. In the process of detecting a large number of facial feature points, the inference speed of the detection model can be improved by adjusting the number of feature points to be detected in order to reduce the delay speed, and in order to improve the real-time face detection speed, it is possible to reduce the weight through algorithms such as model structure improvement, weight pruning, and model compression.
[0036] The human object face de-identification execution unit 120 can convert a real human object into a virtual human object based on changes in external attributes of the face of the real human object, whose eigenface feature points are included as metadata, and generate a de-identified image in which the real human object in the CCTV image is de-identified.
[0037] For example, the human object face de-identification execution unit 120 can adjust the virtual face features (attributes) as intended by designating and learning the facial features as a domain with reference to Image to Image Translation and Face Attribute Manipulation (Editing) models such as StarGAN, StarGAN2, STGAN, and StyleGANEX. The "StarGAN" model is a model capable of performing multi-domain conversion with a single neural network, and is suitable for generating a virtual face image and identifying virtual facial features to convert a face. That is, a specific parameter (domain) can be set to convert into a virtual face having desired external features (attributes), and the degree of de-identification by a combination of various external features can also be adjusted.
[0038] The human object face de-identification performing unit 120 can convert a real human object into a virtual human object by changing at least one external attribute of the real human object among hair color, hairstyle, skin color, beard, aging, and facial expression. In this case, the human object face de-identification performing unit 120 divides the external attributes of the face into domains and adjusts the degree of change for each external attribute, and can adjust the degree of de-identification so that the change results of the external attributes do not overlap with real human objects in the CCTV video.
[0039] That is, the degree of change in external attributes can be adjusted so that virtual human objects (converted results) converted within one CCTV video are not identical to each other. For example, if real human objects A and B exist, the degree and type of change can be set so that A and B's hair colors are converted to be the same, or the same hairstyle and length, or the same skin color, but different combinations of external attributes are applied. By simultaneously applying many parameters, the degree of de-identification can be further increased, and it is also possible to select whether or not to stop tracking the same person between CCTV videos by setting random parameters.
[0040] In this way, even if the external attributes of a real person object are changed and the real person object is converted into a virtual person object, the eigenface feature points of the virtual person object can be maintained as they are as metadata of the virtual person object. In other words, even if a real person object is converted into a virtual person object, the eigenface feature points of the virtual person object are maintained to be the same as the eigenface feature points of the real person object (the features of the object are inherent), and the trackability based on the eigenface feature points can be maintained as they are when tracking the virtual person object in the video. The human object clothing feature detection unit 130 can detect clothing features of a real human object from a CCTV image.
[0041] For example, if there are two actual human objects A and B in a CCTV image, it is possible to detect the clothing features of each object (type of clothing, color, design pattern, material, etc.), such as A being wearing a red long dress and B being wearing a black short-sleeved shirt and shorts.
[0042] The human object clothing de-identification performing unit 140 may de-identify the clothing of the real human object according to changes in external attributes of the clothing of the real human object based on the clothing features.
[0043] More specifically, the human object clothing non-identification executing unit 140 may change any one of the external attributes of the color, design pattern, style, and type of the clothing of the actual human object.
[0044] For example, when a clothing feature is detected for real person object A that the object is wearing a red long dress, the color and style of the clothing can be changed to a yellow short two-piece suit. Also, when a clothing feature is detected for real person object B that the object is wearing a black short-sleeved shirt and short pants, the color and style of the clothing can be changed to a white long-sleeved shirt for the upper body and gray long pants for the lower body.
[0045] In this way, even if the external attributes of the clothes of the real person object are changed and the clothes are converted into other clothes, the data of the clothes features as metadata of the virtual person object can be maintained as is. In other words, even if the clothes of the real person object are converted into new virtual clothes, the clothes feature data of the virtual person object is maintained to be the same as the clothes feature data possessed by the real person object, and the trackability based on the clothes feature data can be maintained as is when tracking the virtual person object in the image.
[0046] The human object clothing de-identification execution unit 140 divides the external attributes of clothing into domains and adjusts the degree of change for each external attribute, and can adjust the degree of de-identification so that changes in external attributes do not overlap between actual human objects in CCTV footage.
[0047] That is, the degree of change in the external attributes related to clothing can be adjusted so that the clothes of virtual human objects (converted results) converted in one CCTV video are not identical to each other. For example, if there are real human objects A and B, the degree and type of change can be set so that A and B are not both converted into clothes of the same color and style, but rather, different combinations of external attributes are applied to them.
[0048] In this manner, the human object clothing anonymization performing unit 140 can further strengthen the personal information protection function by adding clothing anonymization for the actual human object. The data transmission unit 150 can compress the de-identified image and transmit it to the management server 200 together with metadata (eigenface feature points, clothing features). The management server 200 can receive the de-identified images and unique features from the edge device 100 and use them as data for artificial intelligence learning.
[0049] Conventionally, violations of portrait rights and privacy have been avoided by mosaic or blurring the faces of human objects in CCTV footage, but such image processing methods are unable to recognize human objects or detect their features, and therefore cannot be used as learning data.
[0050] The de-identified video of this embodiment contains inherent facial feature points for each person object as metadata. From the viewer's perspective, it is not possible to identify the individuals appearing in the video. However, an artificial intelligence learning algorithm can use the metadata to learn about the person objects in the video. This protects portrait rights and privacy and can be used as artificial intelligence learning data, and can also be used for various purposes such as marketing to CCTV footage.
[0051] The above description is merely one embodiment for implementing the system for de-identifying recognized objects in CCTV footage according to the present invention, and the present invention is not limited to the above embodiment. The technical spirit of the present invention is to the extent that a person having ordinary skill in the art to which the invention pertains can make various modifications without departing from the gist of the present invention, as claimed in the following claims. [Explanation of symbols]
[0052] 1000: A De-Identification System for Recognized Objects in CCTV Videos 100: Edge devices 110: Person object face recognition unit 111: face detection area setting unit 112: Parameter adjustment section 120: Person object face de-identification execution unit 130: Person object clothing feature detection unit 140: Person object clothing non-identification execution unit 150: Data transmission unit 200: Management server
Claims
1. an edge device that detects characteristic features of a real object recognized from a CCTV image, converts the real object from which the characteristic features are detected into a virtual object, thereby de-identifying the real object, and transmits the de-identified image and the characteristic features; a management server that receives the de-identified image and the unique feature points from the edge device and uses them as data for artificial intelligence learning; The edge device includes: a person object face recognition unit for detecting characteristic facial feature points of a real person object from a CCTV image and recognizing the real person object; a person object face de-identification execution unit that converts the real person object into a virtual person object according to a change in an external attribute of the face of the real person object including the eigenface feature points as metadata, thereby generating a de-identified image in which the real person object in the CCTV image is de-identified, and adjusts a degree of change for each external attribute by dividing the external attribute of the face into a domain, and adjusts the degree of de-identification so that the change in the external attribute does not overlap between real person objects in the CCTV image; a person object clothing feature detection unit for detecting clothing features of an actual person object from a CCTV image; The clothing of the real person object is de-identified according to the change in the external attributes of the clothing of the real person object based on the clothing feature, and the external attributes of the clothing are divided into domains to adjust the degree of change for each external attribute, and the degree of de-identification is adjusted so that the change in the external attributes does not overlap between real person objects in the CCTV video. A system for de-identifying recognizable objects in CCTV video comprising:
2. The edge device includes: a data transmission unit that compresses and transmits the de-identified image and the metadata to the management server.
2. The system for de-identifying recognizable objects in CCTV video as recited in claim 1.
3. The person object face recognition unit The face detection area setting unit sets a preset point as a face detection area for a real person object using an alpha pose of a pose estimation model.
2. The system for de-identifying recognizable objects in CCTV video as recited in claim 1.
4. The person object face recognition unit a parameter adjustment unit for adjusting at least one parameter of a minimum size of a detection region for a real person object, a detection confidence level, a maximum number of detections, and a non-detection option; The parameter adjustment unit is Adjusting the parameters for the non-detection option by adjusting the number of unique facial feature points detected from the real person object; The non-detection option is an option for reducing and setting the number of eigenface feature points detected from the real person object when it is predicted that the exposure level of the face of the real person object is relatively low due to wearing a mask compared to the real person object not wearing a mask, and a parameter for the non-detection option is the eigenface feature points detected from the real person object.
2. The system for de-identifying recognizable objects in CCTV video as recited in claim 1.
5. The person object face de-identification execution unit, To convert a real person object into a virtual person object by changing at least one external attribute of the real person object among hair color, hairstyle, skin color, beard, aging, and facial expression.
2. The system for de-identifying recognizable objects in CCTV video as recited in claim 1.
6. The person object clothing non-identification execution unit is The external attributes of the clothing of the actual person object are changed, among which color, design pattern, style, and type.
2. The system for de-identifying recognizable objects in CCTV video as recited in claim 1.
Citation Information
Patent Citations
Monitoring system and monitoring method
CN118038350A
Egg food products and methods for producing egg food products
JP2020513858A
Systems and methods for image de-identification
JP2020522829A
Face image de-identification apparatus and method
KR1020200036656A
Method for de-identifying face and body contained in video data, and device performing the same
KR1020220122457A