Personal information de-identification apparatus for utilizing unstructured data, and operation method thereof

The personal information de-identification device addresses the limitations of existing technologies by effectively extracting and anonymizing personal information from unstructured data in documents, images, and voices, improving data usability and compliance.

WO2025143978A1PCT designated stage expired Publication Date: 2025-07-03ONE DATA TECH CO LTD

Patent Information

Application Number
PCT/KR2024/096501
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-27
Filing Date
2024-11-13
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Current personal information de-identification technologies are inadequate for unstructured data such as images and voices, lagging behind global standards and limiting the usability and compliance with data restrictions.

Method used

A personal information de-identification device that includes an input unit, text extraction unit, personal information detection unit, and de-identification unit, capable of extracting and anonymizing text from electronic documents, images, and voices using OCR and voice recognition, and performing pseudonymization, categorization, and masking based on the detected information type.

Benefits of technology

Enhances the usability of unstructured data by de-identifying personal information, reducing the risk of misuse, and ensuring compliance with data protection laws.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024096501_03072025_PF_FP_ABST
    Figure KR2024096501_03072025_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a personal information de-identification apparatus for utilizing unstructured data, and to an operation method thereof, the personal information de-identification apparatus for utilizing unstructured data comprising: an input unit for receiving unstructured data for personal information detection and de-identification; a text extraction unit for extracting text from the unstructured data received from the input unit; a personal information detection unit for detecting personal information included in the text extracted by the text extraction unit; a personal information de-identification unit for performing de-identification according to the type of the detected personal information; and an output unit for generating unstructured data comprising the de-identified personal information.
Need to check novelty before this filing date? Find Prior Art

Description

Personal information de-identification device for utilizing unstructured data and its operating method

[0001] The present invention relates to a personal information de-identification device for utilizing unstructured data and a method for operating the device, and more particularly, to a personal information de-identification device for utilizing unstructured data, such as electronic documents, videos, images, and voices (files), which detects and de-identifies personal information contained in such unstructured data, and a method for operating the device.

[0002] Big data analysis is rapidly expanding beyond simple text analysis to include analysis based on video and audio information. However, the personal information contained within this data often limits its practical use, leading to growing technological and market demand for solutions.

[0003] To revitalize the big data industry, which is often called the driving force of the Fourth Industrial Revolution, a de-identification solution is needed that de-identifies personal information, such as identification and sensitive information contained in images, voices, and documents, and allows them to be freely utilized by both the public and private sectors.

[0004] Currently, personal information de-identification technology has been developed primarily for structured data. For unstructured data like text, the focus is on masking specific patterns, such as names, place names, and resident registration numbers. In particular, domestic technology for de-identifying personal information for video and audio lags behind global standards.

[0005] Therefore, the existing personal information de-identification technology has the problems described above, and there is a need to develop a technology to de-identify personal information from unstructured data to complement the above problems.

[0006] The matters described as background technology above are only intended to enhance understanding of the background of the present invention, and should not be taken as an admission that they correspond to prior art already known to those of ordinary skill in the art.

[0007] Accordingly, the present invention has been proposed to solve the above-mentioned conventional problems, and the purpose of the present invention is to provide a device for extracting text from unstructured data and detecting and de-identifying personal information contained in unstructured data such as electronic documents, videos, images, and voices (files).

[0008] A personal information de-identification device for utilizing unstructured data according to an embodiment of the present invention may be characterized by including: an input unit for receiving unstructured data to detect and de-identify personal information; a text extraction unit for extracting text from the unstructured data input from the input unit; a personal information detection unit characterized by detecting the personal information included in the text extracted from the text extraction unit; a personal information de-identification unit for de-identifying the detected personal information according to the type of the detected personal information; and an output unit for generating unstructured data including the de-identified personal information.

[0009] In addition, the text extraction unit may further include an electronic document text extraction module that extracts text through document parsing if the input unstructured data is an electronic document file; a video and image text extraction module that extracts text through optical character recognition if the input unstructured data is a video and image file; and a voice text extraction module that extracts text through voice recognition if the input unstructured data is a voice file.

[0010] In order to solve the above problem, a personal information de-identification device for utilizing unstructured data according to an embodiment of the present invention includes a communication interface unit that receives first unstructured data of a document, an image, a video, or an audio file, and a control unit that detects personal information from the received first unstructured data, de-identifies the detected personal information in different forms according to the type of the detected personal information, and generates and outputs second unstructured data including the personal information de-identified in different forms.

[0011] The above control unit can detect and de-identify personal information by extracting text from the first unstructured data, analyzing the extracted text, and grouping objects expressed differently for the same object into the same object.

[0012] The control unit may extract text through document parsing when the first unstructured data is an electronic document file, extract text through optical character recognition (OCR) when the first unstructured data is a video or image file, and extract text through voice recognition when the first unstructured data is a voice file.

[0013] The above control unit can pseudonymize personal information of proper nouns, categorize personal information of numerical data, and mask personal information of unique numbers.

[0014] A method for operating a personal information de-identification device for utilizing unstructured data according to an embodiment of the present invention for solving the above problem includes a step in which a communication interface unit receives first unstructured data of a document, an image, a video, or an audio file, and a step in which a control unit detects personal information from the received first unstructured data, de-identifies the detected personal information in different forms according to the type of the detected personal information, and generates and outputs second unstructured data including the personal information de-identified in different forms.

[0015] The step of generating the second unstructured data may detect and de-identify personal information by extracting text from the first unstructured data, analyzing the extracted text, and grouping objects expressed differently for the same object into the same object.

[0016] The step of generating the second unstructured data may extract text through document parsing if the first unstructured data is an electronic document file, extract text through optical character recognition (OCR) if the first unstructured data is a video or image file, and extract text through voice recognition if the first unstructured data is a voice file.

[0017] The step of generating the above second non-standard data may include pseudonymizing personal information of proper nouns, categorizing personal information of numerical data, and masking personal information of unique numbers.

[0018] According to the personal information de-identification device for utilizing unstructured data as described above, the following effects are achieved.

[0019] First, anonymization of personal information for unstructured data can have the effect of increasing the usability of unstructured data.

[0020] Second, it has the effect of reducing the damage from misuse of personal information by anonymizing personal information for unstructured data.

[0021] The effects that can be obtained from the present invention are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood by a person having ordinary skill in the art to which the present invention pertains from the description below.

[0022] FIG. 1 is a configuration diagram of a personal information de-identification device for utilizing non-standard data according to one embodiment of the present invention.

[0023] FIG. 2 is an exemplary diagram showing a process of de-identifying personal information included in an electronic document file according to one embodiment of the present invention.

[0024] FIG. 3 is an exemplary diagram showing a process of de-identifying personal information contained in video and image files according to one embodiment of the present invention.

[0025] FIG. 4 is an exemplary diagram showing a process of de-identifying personal information included in a voice file according to one embodiment of the present invention.

[0026] FIG. 5 is a diagram illustrating a personal information de-identification system for utilizing unstructured data according to an embodiment of the present invention.

[0027] Figure 6a is a diagram illustrating the process of analyzing personal information risks and de-identifying unstructured data for safe use of unstructured data.

[0028] FIG. 6b is a diagram for explaining the personal information de-identification operation in text data based on the neuro-symbolic model of the personal information de-identification device of FIG. 5.

[0029] Figures 7a, 7b and 7c are diagrams for explaining the operation of entity detection based on a neuro-symbolic model.

[0030] Figures 8a and 8b are diagrams for explaining the operation of personal information anonymization linked to entity name detection technology.

[0031] Figure 9 is a diagram for explaining the operation of anonymizing personal information to prevent content corruption.

[0032] Fig. 10 is a block diagram illustrating the detailed structure of the personal information de-identification device of Fig. 5.

[0033] Figure 11 is a flowchart showing the operation process of a personal information de-identification device according to an embodiment of the present invention.

[0034] Hereinafter, a personal information de-identification device (100) for utilizing unstructured data according to embodiments of the present invention will be described in detail with reference to the attached drawings. The present invention may have various modifications and various forms, and specific embodiments will be illustrated in the drawings and described in detail in the text. However, this is not intended to limit the present invention to a specific disclosed form, but should be understood to include all modifications, equivalents, or substitutes included in the spirit and technical scope of the present invention. In describing each drawing, similar reference numerals are used to refer to similar components. In the attached drawings, the dimensions of structures are illustrated in an enlarged form compared to actual size to ensure clarity of the present invention, or in a reduced form compared to actual size to understand the schematic configuration.

[0035] Furthermore, while terms such as "first" and "second" may be used to describe various components, these components should not be limited by these terms. These terms are used solely to distinguish one component from another. For example, without departing from the scope of the present invention, the first component may be referred to as the "second component," and similarly, the second component may also be referred to as the "first component." Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by one of ordinary skill in the art to which the present invention pertains. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined herein.

[0036] The present invention relates to a personal information de-identification device (100) for utilizing unstructured data, and more specifically, to a device for detecting and de-identifying personal information contained in unstructured data such as electronic documents, videos, images, and voices.

[0037] FIG. 1 is a block diagram of a personal information de-identification device for utilizing unstructured data according to one embodiment of the present invention, FIG. 2 is an exemplary diagram showing a process of de-identifying personal information included in an electronic document file according to one embodiment of the present invention, FIG. 3 is an exemplary diagram showing a process of de-identifying personal information included in a video and image file according to one embodiment of the present invention, and FIG. 4 is an exemplary diagram showing a process of de-identifying personal information included in a voice file according to one embodiment of the present invention.

[0038] As illustrated in FIG. 1, the present invention includes an input unit (101) for receiving unstructured data to detect and de-identify personal information, a text extraction unit (102) for extracting text from the unstructured data input from the input unit (101), a personal information detection unit (103) characterized by detecting the personal information included in the text extracted from the text extraction unit (102), a personal information de-identification unit (104) for de-identifying the detected personal information according to the type of the detected personal information, and an output unit (105) for generating unstructured data including the de-identified personal information.

[0039] The non-standard data input into the above input unit (101) is characterized by being any one of an electronic document file, a video and image file, and a voice file.

[0040] In addition, the text extraction unit (102) performs the task of extracting text from the unstructured data input through the input unit (101), and is characterized by further including an electronic document text extraction module that extracts text through document parsing if the input unstructured data is an electronic document file, a video and image text extraction module that extracts text through optical character recognition if the input unstructured data is a video and image file, and a voice text extraction module that extracts text through voice recognition if the input unstructured data is a voice file.

[0041] The type of the above electronic document file is characterized by being one of hwp / hwpx, docx, pptx, xlsx, and pdf.

[0042] And the above personal information detection unit (103) derives analysis results for 12 entity names related to the personal information through entity name recognition technology, and detects personal information through combination of the entity name recognition results and a personal information pattern knowledge base, and the 12 entity names related to the personal information are characterized as person, school, company, certificate, major, work department, club, occupation, position, country, region, and period.

[0043] Lastly, the personal information de-identification unit (104) performs pseudonymization processing on proper nouns such as names of people and organizations and categorization processing on numerical data such as height and age according to the type of personal information detected by the personal information detection unit (103), and performs masking processing on unique numbers such as resident registration numbers, thereby performing personal information de-identification on the detected personal information.

[0044] In addition to the above, the personal information de-identification device (100) for utilizing non-standard data of FIG. 1 can perform various operations, and since related contents will be continuously covered later, detailed contents will be replaced with those contents.

[0045] FIG. 5 is a diagram illustrating a personal information de-identification system for utilizing unstructured data according to an embodiment of the present invention, FIG. 6a is a diagram explaining a process of analyzing personal information risks and de-identifying unstructured data for safe utilization of unstructured data, FIG. 6b is a diagram explaining a neuro-symbolic model-based text data personal information de-identification operation of the personal information de-identification device of FIG. 5, FIGS. 7a, 7b, and 7c are diagrams explaining a neuro-symbolic model-based named entity detection operation, FIGS. 8a and 8b are diagrams explaining a named entity detection technology-linked personal information de-identification operation, and FIG. 9 is a diagram explaining a content corruption prevention personal information de-identification operation.

[0046] As illustrated in FIG. 5, a personal information de-identification system (490) for utilizing unstructured data according to an embodiment of the present invention may be configured to include an unstructured data provision device (500) and a personal information de-identification device (510), and may further include a communication network such as a WCDMA network, an e-Node network, a cloud communication network, a 5G network, or a 6G network operated by a communication company.

[0047] The unstructured data provision device (500) may include various types of devices that provide unstructured data to the personal information de-identification device (510). For example, a storage medium such as a USB may be a representative example. Furthermore, the device may include not only PC-based terminal devices such as computers used by users, but also mobile-based terminal devices such as smartphones. The terminal devices may have memory installed therein, and unstructured data may be stored therein. Furthermore, the unstructured data provision device (500) may include a server, such as a cloud server, that provides services online. Examples of such servers include those of companies that provide SNS services.

[0048] Unstructured data, unlike numerical data, refers to data with complex forms and structures, such as images, videos, and documents, that are not structured. Examples of unstructured data include traditional data like books, magazines, medical records, audio, and video, as well as data generated on mobile devices and online, such as emails, Twitter, and blogs. This unstructured data may contain personal information, and the de-identification of such personal information is considered crucial in Korea, particularly with legal revisions and technological advancements, such as the three data laws (e.g., the Personal Information Protection Act, the Information and Communications Network Act, and the Credit Information Act amendments) and big data analysis.

[0049] The personal information de-identification device (510) according to an embodiment of the present invention may refer to a PC-based terminal device, a mobile-based terminal device, or a server operated online. The personal information de-identification device (510) may be equipped with a program for de-identifying personal information according to an embodiment of the present invention, and through this, personal information in the unstructured data (or first unstructured data) provided by the unstructured data providing device (500) may be de-identified, and unstructured data (or second unstructured data) containing the de-identified personal information may be generated and output. For example, an engineer who wishes to utilize big data analysis technology may provide unstructured data to the personal information de-identification device (510) through his or her computer. In addition, the engineer may receive data in which personal information in the unstructured data, such as voice files or videos, provided by the engineer, is de-identified and may freely utilize the data.

[0050] Figures 6a to 9 illustrate various configurations and operations of a personal information de-identification device (510) according to an embodiment of the present invention. As previously described with respect to the personal information de-identification device (100) of Figure 1, the personal information de-identification device (510) extracts text data from the input unstructured data when unstructured data is input. To this end, the personal information de-identification device (510) may include a text extraction module. Here, the module may mean a hardware module, a software module, or a module by a combination thereof, and may mean software, a hardware group, or a unit for performing a specific operation. The text extraction module may include, as illustrated in Figure 6a, an electronic document text extraction module that parses an electronic document to extract text, an OCR-based text extraction module that performs an intelligent OCR operation to extract text from images, etc., and a voice recognition / transcription module that extracts voice signals from voice files into text. Of course, the intelligent OCR herein may include noise data that is unavailable when the extraction accuracy is low according to a general OCR program. Therefore, it can be seen as a program that can select available text data by applying an artificial intelligence program to filter out such noisy text data.

[0051] Since characters can appear not only in document areas such as the header, footer, and body of a document, but also in tables and images, performing character recognition without distinguishing between areas containing characters will not only result in a decrease in recognition accuracy, but also cause the recognized characters to be mixed up. Therefore, image segmentation must be performed first. Image segmentation is an object detection task that classifies multiple objects in a single image and estimates their locations, and can be predicted using bounding boxes that indicate the types and locations of the objects. In an embodiment of the present invention, training data is constructed by adding bounding boxes to raw data using open source image labeling tools such as VGG Image Annotator, Labellmg, OpenLabeler, and ImgLab, and an image segmentation model based on YOLO (You Only Look Once) can be trained.

[0052] Once the text region is recognized, text is recognized from the text region through image preprocessing, text detection, text recognition, and postprocessing. The image preprocessing step involves correcting the image to make it easier for the computer to recognize the text region. This can include converting a color image to grayscale, analyzing pixel values ​​to increase brightness and contrast, and then dividing the pixel values ​​into two ranges, classified as 0 and 1, through binarization. Various techniques can be used, such as spot removal, line removal, and layout analysis. Text detection is the step that extracts the text region from the entire image. It determines the text region and its rotation angle, and makes the text horizontal to increase the recognition rate. Postprocessing further improves accuracy by examining the content of the text and correcting any unnatural words or characters. For example, the word "후길동" can be modified to "홍길동" because it is contextually understood to be a first name and the surname "후" does not exist. In embodiments of the present invention, EasyOCR, currently the most popular open source, can be used for text detection and recognition. EasyOCR uses CRAFT, a text detection tool developed and released by Naver, and can also easily generate training data (or learning data) through font files for learning. In the case of table type processing, text extraction from table images is basically similar to processing text images, but the task of extracting structural information of the table (i.e., rows and columns) must be additionally performed. Therefore, in embodiments of the present invention, TableNet, which can distinguish rows and columns in the table area within the image, can be used to extract text data while maintaining the structural information of the table.

[0053] In addition, with regard to the personal information detection operation in voice data, the personal information de-identification device (510) must first perform a STT (Speech-to-Text) process in which a computer interprets the spoken language spoken by a person through speech recognition and converts the content into text data in order to detect personal information in the voice data. Most of the acoustic models applied to existing commercial services were based on the HMM (Hidden Markov Model), which is a probability statistics method, but in the embodiment of the present invention, KoSpeech, which is a model based on RNN (Recurrent Neural Networks) of the sequence-to-sequence method that has recently shown good performance, can be utilized. KoSpeech is capable of learning all features contained in voice data, such as grammar and pronunciation, and provides various acoustic models (e.g., Deep Speech2, LAS (Listen, Attend and Spell), Speech Transfomer, Joint CTC-Attention), so that the optimal model can be selected and used through comparative experiments.

[0054] The personal information de-identification device (510) can perform deep learning-based named entity recognition operation for personal information detection as shown in FIGS. 6A and 6B, and can also perform personal information detection operation based on a neuro-symbolic model. Here, the neuro-symbolic model may mean a model in which symbolic artificial intelligence (e.g., rule-based model) that is effective for inference and application and neuro-symbolic artificial intelligence (e.g., deep learning-based model) that is effective for pattern learning are combined. FIGS. 7A to 7C well illustrate the neuro-symbolic model-based named entity detection technology. Personal information-related named entity detection of text data based on the neuro-symbolic model can be performed by using natural language processing element technology (e.g., morphological analysis, syntax analysis) for named entity recognition in text data and a pre-trained language model (e.g., KoELECTRA-base). In relation to personal information detection, as mentioned above when explaining the personal information detection unit (103) of FIG. 1, the analysis results for 12 named entities related to personal information are derived through named entity recognition technology, and personal information is detected through the combination of the named entity recognition results and the personal information pattern knowledge base, and the 12 named entities related to personal information can detect people, schools, companies, certifications, majors, work departments, clubs, occupations, positions, countries, regions, periods, etc. FIGS. 7a and 7c show examples to help explain some of the steps of FIG. 7b, and the configuration of FIGS. 7a and 7c in which the same alphabets as the alphabets connected by dotted lines next to the steps of FIG. 7b are connected is related to the steps of FIG. 7b.

[0055] In addition, the personal information de-identification device (510) can perform a personal information de-identification operation as shown in FIGS. 6a and 6b when the personal information detection operation is completed. FIGS. 8a and 8b show an example of a personal information de-identification screen linked to a named entity detection technology. The de-identification processing targeting the named entity detection result can be applied with reference to the 'personal information de-identification action guideline' as shown in the table in FIG. 8 (b). Such guideline data can be pre-stored in memory and used. The personal information de-identification device (510) according to an embodiment of the present invention can perform different forms of de-identification processing operations depending on the type of personal information when de-identifying personal information. The personal information de-identification device (510) can perform various forms of de-identification operations, such as pseudonymizing proper nouns such as names of people and names of organizations, categorizing numerical data such as height and age, and masking unique numbers such as resident registration numbers.

[0056] The personal information de-identification device (510) according to an embodiment of the present invention can also perform a cross-reference operation for personal information in text data to prevent content corruption. In order to utilize data in which personal information has been de-identified as training data, identical entity names existing in the text data must be pseudonymized / anonymized into the same type to prevent data contamination. In an embodiment of the present invention, this problem is regarded as a coreference resolution in natural language processing and a program applying the corresponding technology can be executed. Here, coreference resolution is a natural language processing problem of finding and linking various noun phrases (mentions) expressing an arbitrary entity. In an embodiment of the present invention, this can be solved by enhancing the named entity recognition technology. This is well illustrated in Figure 9. As shown in Figure 9, in an embodiment of the present invention, by applying an identical entity pseudonymization / anonymization module, text data personal information detection, and a de-identification base module, personal information risk analysis / de-identification operations for each type of electronic document / text, image data, and voice data can be performed. In addition to anonymizing recognized entity names in unstructured data to the same type (e.g., Hong Gil-dong, etc.), the personal information de-identification device (510) can also determine, for example, that Representative Jang Je-won and Representative Jang in text data are the same type by executing a program (or algorithm) that performs a cross-reference resolution operation of natural language processing, as shown in FIG. 9. Afterwards, based on the results of grouping (or clustering) of identical entities, the identical entities in the text can be pseudonymized / anonymized. Of course, during this process, the personal information de-identification device (510) can also change the particles (e.g., Lee / Ga, Eul / Reul, etc.) when pseudonymizing / anonymizing. Anonymization can be determined based on the particles, or the particles can be changed based on the anonymization.This can be set and operated in various ways according to the intention of the system designer, and the decision on anonymization or investigation can be made based on the learning results of the data by using an artificial intelligence program, and the decision on pseudonymization / anonymization or investigation can be modified.

[0057] In addition to the above, the non-standard data provision device (500) and the personal information de-identification device (510) of FIG. 5 can perform various operations, and since the related contents have been previously explained through FIG. 1 or will be continuously discussed later, the detailed contents will be replaced with those contents.

[0058] Fig. 10 is a block diagram illustrating the detailed structure of the personal information de-identification device of Fig. 5.

[0059] As illustrated in FIG. 10, the personal information de-identification device (510) of FIG. 5 according to an embodiment of the present invention is a terminal device such as a computer or smartphone used by users, or a server device such as a cloud server, and includes part or all of a communication interface unit (1000), a control unit (1010), a personal information de-identification unit (1020), and a storage unit (1030).

[0060] Here, “including some or all” means that some components, such as the storage unit (1030), may be omitted to configure the personal information de-identification device (510), or some components, such as the personal information de-identification unit (1020), may be integrated into other components, such as the control unit (1010), etc. In order to help a sufficient understanding of the invention, it is described as including all.

[0061] The communication interface unit (1000) can receive unstructured data such as electronic documents, images, and videos by communicating or being connected to storage media such as USB or various types of terminal devices. In addition, the communication interface unit (1000) can transmit the received unstructured data to the control unit (1010). The communication interface unit (1000) can perform operations such as modulation / demodulation, muxing / demuxing, encoding / decoding, and scaling to convert resolution in the process of communicating or performing interface operations with external devices such as storage media, and since this is obvious to those skilled in the art, further description thereof will be omitted.

[0062] In addition, the communication interface unit (1000) can de-identify personal information in the received unstructured data under the control of the control unit (1010), receive the unstructured data containing the de-identified personal information, and provide it again to the terminal device of the user, etc. When the first unstructured data in which the personal information has not been de-identified is received, it is transmitted as the second unstructured data in which the personal information has been de-identified. Here, the de-identification of the personal information can be performed in the form of pseudonymization, categorization, or masking depending on the type of the personal information. The de-identification operation is also well shown in screen (a) of FIG. 8. For example, when the personal information de-identification unit (1020) provides a service in the form of an online platform through a cloud server, etc., a UI screen like that of FIG. 8 (a) can be provided on the screen of a user terminal device such as a computer or smartphone. When the de-identification button is selected on the screen, the original sentence (i.e., the first unstructured data) is displayed in the output space of the original sentence, and the de-identified unstructured data (i.e., the second unstructured data) is displayed in the output space of the de-identified sentence. The communication interface unit (1000) may be involved in this operation.

[0063] The control unit (1010) includes a processor such as a CPU, MPU, GPU, etc., and performs overall control operations of the communication interface unit (1000), personal information de-identification unit (1020), and storage unit (1030). When big data such as electronic documents, images, and videos are collected online through the communication interface unit (1000), the control unit (1010) can temporarily store the data in the storage unit (1030) and then retrieve the data to request data analysis or personal information de-identification processing from the personal information de-identification unit (1020).

[0064] Of course, the control unit (1010) according to the embodiment of the present invention may also perform an operation to build, or load, a program for de-identifying personal information according to the embodiment of the present invention into the personal information de-identification unit (1020). For example, the personal information de-identification unit (1020) may be loaded with a basic model of artificial intelligence, and when a training operation is requested by inputting learning data into the program from a developer's computer, etc., the learning operation of the program may be performed. In other words, when learning data related to personal information entity detection is provided, it may be stored in the storage unit (1030) and then provided to the personal information de-identification unit (1020) to enable learning.

[0065] The personal information de-identification unit (1020) can perform text extraction operations in a corresponding manner for unstructured data provided in various forms such as electronic documents, voice files, and images by distinguishing the format of the unstructured data, such as the file format. An operation to determine the file format may be performed first. For example, in the case of electronic documents, the personal information de-identification unit (1020) can perform text extraction operations after parsing the electronic document. In addition, in the case of PDFs or images, text can be extracted based on intelligent OCR. In the case of voice files, text can be extracted through voice recognition. Based on the sound volume and pattern information in the voice signal, the recognition result and the data in the dictionary related to the recognition result can be compared with the data in the dictionary, and a specific text can be extracted based on the comparison.

[0066] In addition, the personal information de-identification unit (1020) can perform an operation to detect personal information by analyzing the extracted text data following the text extraction operation. Of course, in addition to the existing rule-based detection method, various programs such as supervised learning, unsupervised learning, and semi-supervised learning that combines the two can be applied to pre-learn learning data for entity detection and then detect personal information based on the learning results. For example, since the LLM (Large Language Model) is a model specialized for natural language processing, using such a model can lead to superior natural language processing performance. In an embodiment of the present invention, personal information can be detected by applying a Korean dependent phrase analysis model or a Korean named entity recognition model based on KoELECTRA. Of course, since various program models can be applied to detect personal information, the embodiment of the present invention will not be particularly limited to any one form. However, the personal information de-identification unit (1020) according to an embodiment of the present invention can perform an identical entity grouping (or clustering) operation as described in FIG. 9. For example, since Representative Jang Je-won and Representative Jang belong to the same group, a program for cross-reference resolution using natural language processing according to an embodiment of the present invention can be executed to accurately detect their entities. Entities in the same group, expressed in different forms, undergo anonymization / pseudonymization processing in the same form.

[0067] In addition, if it is confirmed that there are multiple entity names (e.g., Representative Jang Je-won, Jang Ui-won) included in the same entity group within the text data, the personal information de-identification unit (1020) can generally perform anonymization / pseudonymization processing in the same form (e.g., Representative Hong Gil-dong), and at this time, the form (Representative Hong Gil-dong, Hong Ui-won) may be determined depending on the number of entity name forms included within the text data. For example, if the entity name 'Representative Jang Je-won' within the text data is greater than the entity name 'Jang Ui-won', the personal information de-identification unit (1020) can change the entity names included in the same entity group (Representative Jang Je-won, Jang Ui-won) to 'Representative Hong Gil-dong'. Of course, in the opposite case, the entity names can be changed to 'Hong Ui-won'.

[0068] In some cases, it may be processed in an anonymized / pseudonymized form (e.g. Rep. Jang Je-won -> Rep. Hong Gil-dong, Rep. Jang Ui-won -> Rep. Hong) while maintaining each original form.

[0069] Furthermore, the personal information de-identification unit (1020) can perform a de-identification operation on the detected object when the detection of the personal information object is completed in the extracted text data. In an embodiment of the present invention, different forms of de-identification operations can be performed depending on the type of personal information. For example, pseudonymization can be performed on proper nouns such as people's names, categorization can be performed on numerical data such as height or age, and masking can be performed on unique numbers such as resident registration numbers. For example, numerical data can be categorized into AB for all heights, and EF for all ages, and so on, in the form of the same number or the same letter.

[0070] In addition, the personal information de-identification unit (1020) can perform anonymization processing through entity recognition of '후길동' recognized by OCR as '홍길동', and through the natural language cross-reference resolution operation above, in the case of Representative Jang Je-won and Representative Jang, they can be determined to be the same group entities and the same form of de-identification operation can be performed. For example, in the case of the LLM model, natural language processing may be possible, and therefore, by analyzing the sentences of a specific paragraph, it can be determined through context analysis, etc. that Representative Jang Je-won and Representative Jang are the same entities as shown in FIG. 9, and thus they can be classified as the same group entities. In the case of Representative Kim Ki-hyun and Representative Kim, they can also be determined to be the same group entities and thus classified.

[0071] In addition, the personal information de-identification unit (1020) can perform anonymization considering the survey based on the de-identified entity name during the anonymization process, or can perform an operation to automatically change the survey (e.g., this / that, this / that) after anonymization (e.g., Hong Gil-dongga → Hong Gil-dongi-ga, Hong Gil-dongeun, Hong Gil-dongi, etc.). At this time, the decision on the specific survey can be made by processing sentences using an artificial intelligence LLM model, etc. to determine the survey that is natural to the context. With regard to the use of the survey, it is entirely possible to correct the survey after anonymization based on the learning results after applying an artificial intelligence program to learning data, and whether to select anonymity during anonymization considering the survey or change the survey after anonymization can be determined and performed in various ways depending on the intention of the service system or system designer, and therefore the embodiments of the present invention will not be particularly limited to any one form.

[0072] The storage unit (1030) can temporarily store various types of unstructured data processed under the control of the control unit (1010) and then provide the data to the personal information de-identification unit (1020) for data analysis or personal information de-identification processing.

[0073] In addition to the above, the communication interface unit (1000), control unit (1010), personal information de-identification unit (1020), and storage unit (1030) of FIG. 10 can perform various operations, and other detailed information has been sufficiently explained above, so those contents will be used instead.

[0074] According to an embodiment of the present invention, the communication interface unit (1000), the control unit (1010), the personal information de-identification unit (1020), and the storage unit (1030) of FIG. 10 are configured as physically separate hardware modules, but each module may store and execute software for performing the above operations therein. However, the software is a collection of software modules, and each module may be formed of hardware, so there is no particular limitation to the configuration, such as software or hardware. For example, the storage unit (1030) may be hardware, such as storage or memory. However, since it is also possible to store information (repository) in software, there is no particular limitation to the above.

[0075] Meanwhile, as another embodiment of the present invention, the control unit (1010) may include a CPU and a memory, and may be formed as a single chip. The CPU may include a control circuit, an operation unit (ALU), a command interpretation unit, and a registry, and the memory may include a RAM. The control circuit may perform a control operation, the operation unit may perform an operation of binary bit information, and the command interpretation unit may perform an operation of converting a high-level language into machine language and vice versa, including an interpreter or a compiler, and the registry may be involved in software data storage. According to the above configuration, for example, at the initial stage of the operation of the personal information de-identification device (510) of FIG. 5, the program stored in the personal information de-identification unit (1020) may be copied and loaded into the memory, i.e., RAM, and then executed, thereby rapidly increasing the data operation processing speed. In the case of a deep learning model, it may be loaded into the GPU memory instead of the RAM and executed by accelerating the execution speed using the GPU.

[0076] Figure 11 is a flowchart showing the operation process of a personal information de-identification device according to an embodiment of the present invention.

[0077] For convenience of explanation, referring to FIG. 11 together with FIG. 5, a personal information de-identification device (510) according to an embodiment of the present invention receives first unstructured data, such as a document, video, or audio file (S1100). The unstructured data may be provided offline via a storage medium, but it may also be provided online.

[0078] In addition, the personal information de-identification device (510) detects personal information from the received first unstructured data, de-identifies the personal information in different forms depending on the type of the detected personal information, and generates and outputs second unstructured data containing the de-identified personal information (S1110).

[0079] The personal information de-identification device (510) can perform operations to extract text from unstructured data prior to detecting personal information in the unstructured data. For example, in the case of images, text can be extracted using an intelligent OCR program. Conventional OCR may include text that is not available. For example, recognizing Hong Gil-dong as Hoot Gil-dong is a representative example. An intelligent OCR program can perform operations such as correcting Hoot Gil-dong to Hong Gil-dong. For example, it is also possible to select unnecessary noise data such as symbols (e.g., *).

[0080] The personal information de-identification device (510) performs different forms of de-identification operations depending on the type of personal information, but can also perform operations to de-identify individuals who are expressed differently in text data by recognizing them as the same entity group. As explained above, representative cases include Representative Kim Ki-hyun being expressed as Representative Kim or Representative Jang Je-won being expressed as Representative Jang. Of course, in addition, operations are required to recognize individuals who are later referred to by nicknames as the same entity. To this end, the embodiment of the present invention can prevent the problem of content corruption in which the same entity is recognized differently and de-identified in different forms in advance through the cross-reference resolution operation of natural language processing. For example, in the process of executing a program for cross-reference resolution of natural language processing, if it is not possible to determine whether Representative Jang and Representative Jang are the same person, even through operations such as artificial intelligence deep learning, the personal information de-identification device (510) can inquire of the program developer or service manager, and at this time, the artificial intelligence program can directly receive an answer through a recognizable prompt screen, thereby continuing the operation of grouping identical objects.

[0081] In addition to the above, the personal information de-identification device (510) of FIG. 5 can perform various operations, and other detailed information has been sufficiently explained above, so we will replace it with those contents.

[0082] Even though all components constituting the embodiments of the present invention have been described as being combined or operating in combination, the present invention is not necessarily limited to such embodiments. That is, within the scope of the present invention, all of the components may be selectively combined and operated one or more times. In addition, although all of the components may be implemented as individual hardware, some or all of the components may be selectively combined and implemented as a computer program having program modules that perform some or all of the functions of the combined hardware in one or more pieces. The codes and code segments constituting the computer program will be readily inferred by those skilled in the art. Such a computer program may be stored in a non-transitory computer-readable storage medium and read and executed by a computer, thereby implementing the embodiments of the present invention.

[0083] Here, the non-transitory readable storage medium refers to a medium that permanently stores data and can be read by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specifically, the above-described programs may be stored and provided on a non-transitory readable storage medium, such as a CD, DVD, hard disk, Blu-ray disc, USB, memory card, or ROM.

[0084] While the technical concepts of the present invention described above have been specifically described in preferred embodiments, it should be noted that the above-described embodiments are for illustrative purposes only and are not intended to be limiting. Furthermore, those skilled in the art will appreciate that various embodiments are possible within the scope of the technical concepts of the present invention. Therefore, the true scope of technical protection of the present invention should be defined by the technical concepts of the appended claims.

Claims

1. A communication interface unit for receiving first non-standard data of a document, image, video or audio file; and A control unit that detects personal information from the first non-standard data received above, de-identifies the detected personal information into different forms according to the type of the detected personal information, and generates and outputs second non-standard data containing the personal information de-identified into different forms; A device for anonymizing personal information for utilizing unstructured data, including:

2. In paragraph 1, The above control unit is a personal information de-identification device for utilizing unstructured data, which detects and de-identifies personal information by extracting text from the first unstructured data, analyzing the extracted text, and grouping objects expressed differently for the same object into the same object.

3. In paragraph 2, The above control unit extracts text through document parsing if the first unstructured data is an electronic document file, extracts text through optical character recognition (OCR) if the first unstructured data is a video or image file, and extracts text through voice recognition if the first unstructured data is a voice file. A personal information de-identification device for utilizing unstructured data.

4. In paragraph 1, The above control unit is a personal information anonymization device for utilizing non-standard data, which pseudonymsize personal information of proper nouns, categorizes personal information of numerical data, and masks personal information of unique numbers.

5. A step of receiving first non-standard data of a document, image, video or audio file by the communication interface unit; and A step in which the control unit detects personal information from the first non-standard data received, de-identifies the detected personal information into different forms according to the type of the detected personal information, and generates and outputs second non-standard data containing the personal information de-identified into different forms; A method for operating a personal information anonymization device for utilizing unstructured data, including:

6. In paragraph 5, The step of generating the second non-standard data is: A method for operating a personal information de-identification device for utilizing unstructured data, which detects and de-identifies personal information by extracting text from the first unstructured data, analyzing the extracted text, and grouping objects expressed differently for the same object into the same object.

7. In paragraph 6, The step of generating the second non-standard data is: A method for operating a personal information de-identification device for utilizing unstructured data, wherein if the first unstructured data is an electronic document file, text is extracted through document parsing, if the first unstructured data is a video or image file, text is extracted through optical character recognition (OCR), and if the first unstructured data is a voice file, text is extracted through voice recognition.

8. In paragraph 5, The step of generating the second non-standard data is: A method of operating a personal information anonymization device for utilizing non-standard data, in which personal information of proper nouns is pseudonymized, personal information of numerical data is categorized, and personal information of unique numbers is masked.

Citation Information

Patent Citations

  • Computer-implemented method, system, computer program, and storage medium for data anonymization

    JP2021504798A

  • Aerosol generating device and operation method thereof

    KR1020240164340A

  • Semiconductor memory device and manufacturing method of semiconductor memory device

    KR1020250000704A

  • Method for detecting private information and measuring data exposure possibility from unstructured data

    KR102533008B1

  • Process

    KR102602229B1

Cited By

  • Method and system for data processing based on multi-agent modules

    KR102994543B1