Masking system, masking method, and masking program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- MITSUBISHI ELECTRIC DIGITAL INNOVATION CORP
- Filing Date
- 2025-01-24
- Publication Date
- 2026-08-05
Smart Images

Figure 2026126717000001_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to addressing the risk of leakage of confidential information.
Background Art
[0002] Generative AI is used to generate things such as text. AI is an abbreviation for artificial intelligence. For example, generative AI can be used to generate summaries of conversation texts and the like.
[0003] However, when using generative AI, there is a risk that if a document containing confidential information such as personal information is input, it may lead to the leakage of confidential information.
[0004] Patent Document 1 discloses a method for generating text using a large language model (LLM). Also, Patent Document 1 describes that when selling a text or an LLM, it is necessary to remove personal information and confidential information.
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, if the confidential information in the document input to the generative AI is masked by means unrelated to the description of the confidential information, the text generated by the generative AI may become inaccurate. For example, if the confidential information in the input document is replaced with asterisks, the text generated by the generative AI may be in a language / style different from the language / style of the input document.
[0007] This disclosure aims to enable the masking of confidential information within documents using appropriate means in accordance with how the confidential information is described. [Means for solving the problem]
[0008] The masking system described in this disclosure is A mask string determination unit determines a mask string to mask the secret string, based on the character types of the secret string, which is a string representing confidential information among the strings shown in the document data. A masking unit that uses the mask string to mask the secret string in the document data, It is equipped with. [Effects of the Invention]
[0009] According to this disclosure, confidential information within a document is masked using a string of characters corresponding to the type of characters in the confidential information. In other words, according to this disclosure, confidential information within a document can be masked using appropriate means according to how the confidential information is written. [Brief explanation of the drawing]
[0010] [Figure 1] Configuration diagram of the masking system 200 in Embodiment 1. [Figure 2] Functional configuration diagram of the masking system 200 in Embodiment 1. [Figure 3] Flowchart of the masking method in Embodiment 1. [Figure 4] A diagram illustrating the overview of document generation using generative AI. [Figure 5] This figure shows an example of a summary generated using a generation AI based on conversational text as input. [Figure 6] A diagram illustrating an example of replacing personal names with asterisks. [Figure 7] A diagram illustrating an example of how inappropriate areas are masked. [Figure 8] This diagram shows an example of a summary generated when the names in the dialogue are not masked. [Figure 9]A diagram showing an example of a summary generated when masking the name part of a conversation sentence with an asterisk. [Figure 10] A diagram showing an example of a summary generated when masking the name part of a conversation sentence with a character string in the same language as the name notation.
Mode for Carrying Out the Invention
[0011] In the embodiments and the drawings, the same elements or corresponding elements are denoted by the same reference numerals. The description of elements denoted by the same reference numerals as the described elements will be omitted or simplified as appropriate. The arrows in the figures mainly indicate the flow of data or the flow of processing.
[0012] Embodiment 1. The masking system 200 will be described based on FIGS. 1 to 10.
[0013] ***Description of the Configuration*** Based on FIG. 1, the configuration of the masking system 200 will be described. The masking system 200 includes a masking device 100.
[0014] The masking device 100 is a computer including hardware such as a processor 101, a memory 102, an auxiliary storage device 103, a communication device 104, and an input / output interface 105. These hardware components are connected to each other via signal lines.
[0015] The processor 101 is an IC that performs arithmetic processing and controls other hardware. For example, the processor 101 is a CPU. IC is an abbreviation for Integrated Circuit. CPU is an abbreviation for Central Processing Unit.
[0016] Memory 102 is a volatile or non-volatile storage device. Memory 102 is also called main memory. For example, memory 102 is RAM. Data stored in memory 102 is saved to auxiliary storage device 103 as needed. RAM is an abbreviation for Random Access Memory.
[0017] The auxiliary storage device 103 is a non-volatile storage device. For example, the auxiliary storage device 103 is a ROM, HDD, flash memory, or a combination thereof. Data stored in the auxiliary storage device 103 is loaded into memory 102 as needed. ROM is an abbreviation for Read Only Memory. HDD is an abbreviation for Hard Disk Drive.
[0018] The communication device 104 is a receiver and transmitter. For example, the communication device 104 is a communication chip or NIC. Communication of the masking device 100 is performed using the communication device 104. NIC is an abbreviation for Network Interface Card.
[0019] The input / output interface 105 is a port to which input and output devices are connected. For example, the input / output interface 105 is a USB terminal, the input devices are a keyboard and mouse, and the output device is a display. Input and output of the masking device 100 are performed via the input / output interface 105. USB is an abbreviation for Universal Serial Bus.
[0020] The masking device 100 comprises elements such as a secret string extraction unit 110, a mask string determination unit 120, a masking unit 130, and a target document generation unit 140. These elements are implemented in software.
[0021] The auxiliary storage device 103 stores a masking program that allows the computer to function as a secret string extraction unit 110, a mask string determination unit 120, a masking unit 130, and a target document generation unit 140. The masking program is loaded into memory 102 and executed by the processor 101. The auxiliary storage device 103 also stores the operating system. At least a portion of the OS is loaded into memory 102 and executed by the processor 101. Processor 101 runs the masking program while simultaneously running the OS. OS is an abbreviation for Operating System.
[0022] The data for the masking program (input data, output data, etc.) is stored in the storage unit 190. The auxiliary storage device 103 functions as a storage unit 190. However, storage devices such as memory 102, registers in the processor 101, and cache memory in the processor 101 may function as a storage unit 190 instead of the auxiliary storage device 103, or together with the auxiliary storage device 103.
[0023] The masking program can be recorded (stored) in a computer-readable format on a non-volatile recording medium such as an optical disc or flash memory.
[0024] Figure 2 shows the functional configuration of the masking system 200. The mask string determination unit 120 includes a character type determination unit 121, a rule selection unit 122, and a rule application unit 123. Each element of the masking system 200 will be described later.
[0025] ***Explanation of operation*** The operating procedure of the masking system 200 corresponds to the masking method. Furthermore, the operating procedure of the masking system 200 corresponds to the processing procedure performed by the masking program.
[0026] The masking method will be explained based on Figure 3. In step S110, the secret string extraction unit 110 extracts the secret string from the original document data 191.
[0027] Original document data 191 is document data that represents a document (original text) containing confidential information. The original document data 191 is pre-stored in, for example, the storage unit 190 and read out from the storage unit 190.
[0028] Confidential information is information that should be kept secret. Examples of confidential information include personal information and proper nouns. Personal information includes names of people, etc. For example, the original text is a transcript of a conversation, and the names of the people who spoke are included in the original text.
[0029] A secret string is a string of characters contained in the original text that represents secret information (for example, a person's name).
[0030] For example, the secret string extraction unit 110 uses a part-of-speech analysis function to extract the secret string from the source document data 191. The part-of-speech analysis function analyzes the part of speech of each word contained in a text.
[0031] For example, the secret string extraction unit 110 uses a generation AI to extract the secret string from the source document data 191. The secret string extraction unit 110 provides the generating AI with the source document data 191 and an extraction instruction as input. The extraction instruction is a sentence that instructs the AI to extract the secret string. If the goal is to extract a name as the secret string, for example, the extraction instruction would be the sentence "Extract the name". The generating AI extracts the secret string from the source document data 191 according to the extraction instructions and outputs the extracted secret string.
[0032] For example, the secret string extraction unit 110 extracts the secret string from the source document data 191 using regular expressions. A regular expression representing the pattern of the secret string is defined in advance. The secret string extraction unit 110 extracts strings that match the regular expression. The extracted strings are the secret strings.
[0033] The secret string extraction unit 110 may extract the secret string using a combination of part-of-speech analysis, generation AI, and regular expressions.
[0034] In steps S121 to S123, the mask string determination unit 120 determines a mask string for each secret string extracted from the source document data 191 based on the character types of the secret string.
[0035] A mask string is a string used to mask a secret string.
[0036] In step S121, the character type determination unit 121 determines the character type of each secret string.
[0037] The character type corresponds to the type of language (Japanese, English, Korean, etc.). For example, the character types include the alphabet, kana / kanji, and Hangul characters. Kana includes hiragana and katakana. For example, if each character constituting the secret string is an alphabet character, the character type determination unit 121 determines that the character type of the secret string is English (alphabet).
[0038] The character set of a secret string can be determined based on its character encoding. For example, a character code table is used. The character code table shows the range of character codes for each character type and the character code for each character. The character type determination unit 121 checks the character codes of the characters included in the secret string, selects the character code range that includes the checked character codes, and determines that the character type corresponding to the selected character code range is the character type of the secret string.
[0039] In step S122, the rule selection unit 122 selects a mask rule for the character types of each secret string from the mask rule data 192.
[0040] Mask rule data 192 is data that shows mask rules associated with each character type. For example, mask rule data 192 is in tabular format. The mask rule data 192 is pre-stored in the memory unit 190.
[0041] A mask rule is a set of rules for mask strings. A mask rule defines a mask string as a string of the same character type as the character type it is associated with. An example of a mask rule for the character type "English (alphabet)" is "name<number>". <number> will be replaced with a number. An example of a mask rule for the character type "Japanese (Kana / Kanji)" is "Kana<Chinese numeral>". <Chinese numeral> will be replaced with Chinese numerals.
[0042] In step S123, the rule application unit 123 determines a mask string for each secret string according to the mask rules for the character types of the secret string.
[0043] For example, suppose the secret string's character type is "English (alphabet)" and the mask rule for "English (alphabet)" is "name<number>". In this case, a string like "name1" or "name2" would be the mask string for the secret string.
[0044] In step S130, the masking unit 130 uses a mask string to mask the secret strings in the source document data 191 for each secret string.
[0045] The secret string is masked by being replaced with a mask string.
[0046] In step S130, mask document data 193 is generated. Masked document data 193 is the original document data 191 with each secret string masked. In other words, the original document data 191 is the document data before masking, and the masked document data 193 is the document data after masking.
[0047] In step S140, the target document generation unit 140 generates target document data 194 using the mask document data 193.
[0048] Objective document data 194 is document data that shows the objective document without including confidential information. The target text is a new text based on the original text. An example of a target text is a summary of the original text.
[0049] For example, the target document generation unit 140 takes the mask document data 193 as input and generates the target document data 194 using a generation AI. Publicly available generation AIs can be used. If the target document is a summary of the original document, the target document generation unit 140 takes the masked document data 193 as input and instructs the summarization generation AI to generate a summary. The summarization generation AI then generates a summary of the text (masked document) shown in the masked document data 193 and outputs the generated summary. Since confidential information is masked in the masked document, the generated summary does not contain any confidential information. The data showing the output summary becomes the target document data 194.
[0050] ***Effects of Embodiment 1*** The masking system 200 features a two-stage process. The masking system 200 does not directly instruct the generating AI to "mask" and perform the masking process, but rather performs the masking process in two stages as follows. First, the masking system 200 extracts the string to be masked. The masking system 200 then replaces the extracted string with a predetermined string using a very ordinary string conversion. This helps to avoid situations where unintended areas are masked. Furthermore, the extracted string can be used for necessary tasks. For example, if the extracted string is a name, it can be used for identity verification. Also, if the extracted string is an ID (identifier), it can be used for initialization.
[0051] The masking system 200 features same-language code processing. The masking system 200 masks the original string with similar strings. Specifically, the masking system 200 masks the original string with strings from the same language. This makes it possible to avoid being affected by changes in character encoding (language) when using a generation AI after masking.
[0052] The effects of Embodiment 1 will be explained with specific examples. Figure 4 shows an overview of document generation using generative AI. (1) Inquiries / summaries may be generated using a generative AI language model specialized for a specific field / task. When summarizing a dialogue, a large amount of conversational text will be input into the generative AI for training the language model. (2) If the input document contains confidential information / personal information, the confidential information / personal information will be learned by the language model of the generating AI. Furthermore, the confidential information / personal information will be output according to the learning results. In other words, there is a risk of information leakage.
[0053] Figure 5 shows an example of a summary generated using a generation AI based on a conversational text as input. The dialogue contains the proper noun "Akira Equipment" and the personal name "Takano." The summary also contains this personal information.
[0054] Traditionally, the following methods have been used to avoid such risks. (1) 1. Mask confidential information within the document beforehand. 2. Do not use publicly available language models on the cloud. (2) In the masking process, personal names and other identifying information are replaced with strings based on predetermined specifications. Figure 6 shows an example where a personal name is replaced with "****". (3) Three types of techniques are known for performing masking. The conversion mechanism in Figure 6 is implemented using the following techniques. 1. Utilization of morphological analysis: A dictionary is used to identify the part of speech of words, and words in a text are compared with the dictionary to mask "proper nouns," "person names," or "organization names." 2. Use of regular expressions: Define regular expressions that apply to phone numbers or email addresses, and mask strings that match the specifications. 3. Utilizing Generative AI: Give the Generative AI instructions such as "Mask personal names in the following conversation with "****"" to mask personal names.
[0055] However, these methods have the following problems: (1) The generating AI may mask inappropriate parts. Figure 7 shows an example of how inappropriate areas are masked. For example, if you want to mask a personal name, the speaker label may also be masked. While speaker labels can sometimes contain personal names, masking the speaker label is inappropriate when the label is a speaker attribute (e.g., "customer").
[0056] (2) The content of the generated text may change significantly depending on how the mask is applied. Masking is a measure taken for confidentiality purposes. Therefore, in aspects other than protection (output content, writing style, etc.), we want the output to be as identical as possible whether it is masked or not. However, the output may change significantly depending on how the masking is done. Figure 8 shows an example of a summary generated when the names in the dialogue are not masked. If you instruct the system to summarize a conversation without masking the names "Takano" and "Kurata," it will generate a correct summary. Figure 9 shows an example of a summary generated when the names in a conversation are masked with asterisks. If you mask "Takano" and "Kurata" with "****" and instruct the system to summarize the conversation, the Japanese conversation may be summarized in English.
[0057] The masking system 200 solves these problems through two-stage processing and same-language code processing. Figure 10 shows an example of a summary generated when the name portion of a conversation is masked with a string in the same language (character code range of the same type) as the name. If you replace "Takano" with "Kana 1" and "Kurata" with "Kana 2" and instruct the system to summarize the conversation, a summary will be generated that is in the same language as the conversation and does not contradict the content.
[0058] ***Supplement to Embodiment 1*** By setting the mask string to an empty string, the secret string can be removed from the original text using the masking system 200.
[0059] If the language used in the source document is fixed to a single language, it is unnecessary to determine the character type of the secret string (S121) and select the mask rule (S122). In this case, the string of the language used in the source document is determined to be the mask string (S123). If the language used in the source document is specified, it is not necessary to determine the character type of the secret string (S121). In this case, a mask rule for the specified language is selected (S122), and the string of the specified language is determined to be the mask string (S123). For example, if the original text is in Japanese, personal names (including names of foreigners) will be masked with strings of kanji, hiragana, or katakana.
[0060] The input of source document data 191 and the output of target document data 194 may be performed using an API or using standard input / output. API is an abbreviation for Application Program Interface.
[0061] Embodiment 1 is an example of a preferred form and is not intended to limit the technical scope of this disclosure. Embodiment 1 may be implemented in part or in combination with other forms. The procedure described using flowcharts, etc., may be modified as appropriate.
[0062] The masking system 200 may be implemented using multiple devices. The masking device 100 may also be used as a server. Users access the masking device 100 via the network using a user terminal (client) and utilize the functions (services) of the masking device 100. Each element of the masking device 100 may be implemented using software, hardware, firmware, or a combination thereof. The "part" of each element of the masking device 100 may be read as "process," "step," "circuit," or "circuit."
[0063] The various aspects of this disclosure are described below as appendices. (Note 1) A mask string determination unit determines a mask string to mask the secret string, based on the character types of the secret string, which is a string representing confidential information among the strings shown in the document data. A masking unit that uses the mask string to mask the secret string in the document data, A masking system equipped with [a specific feature].
[0064] (Note 2) The mask string determination unit determines the character type of the secret string based on the character code of the secret string. The masking system described in Appendix 1.
[0065] (Note 3) The mask string determination unit determines the mask string to be a string of the same character type as the secret string. The masking system described in Appendix 1 or Appendix 2.
[0066] (Note 4) The mask string determination unit selects a mask rule from the mask rule data, which indicates mask rules associated with each character type, that corresponds to the same character type as the secret string, and determines the mask string according to the selected mask rule. A masking system as described in any one of the appendices 1 through 3.
[0067] (Note 5) The system includes a secret string extraction unit that extracts the secret string from the document data. A masking system described in any one of the appendices 1 through 4.
[0068] (Note 6) The secret string extraction unit extracts the secret string from the document data using a part-of-speech analysis function. The masking system described in Appendix 5.
[0069] (Note 7) The secret string extraction unit extracts the secret string from the document data using a generation AI. The masking system described in Appendix 5.
[0070] (Note 8) The secret string extraction unit extracts the secret string from the document data using a regular expression. The masking system described in Appendix 5.
[0071] (Note 9) The system includes a target document generation unit that generates target document data using the masked document data to generate new target document data that shows a new sentence based on the sentence shown in the unmasked document data, without including the confidential information. A masking system as described in any one of the appendices 1 through 8.
[0072] (Note 10) The aforementioned target document generation unit takes the masked document data as input and generates the target document data using a generation AI. The masking system described in Appendix 9.
[0073] (Note 11) The mask string determination unit determines a mask string to mask the secret string, based on the character types of the secret string, which is a string representing confidential information among the strings shown in the document data. The masking unit uses the mask string to mask the secret string in the document data. Masking method.
[0074] (Note 12) A mask string determination process that determines a mask string to mask the secret string, based on the character types of the secret string, which is a string representing confidential information among the strings shown in the document data, A masking process that masks the secret string in the document data using the mask string, A masking program to cause a computer to execute something. [Explanation of symbols]
[0075] 100 Masking device, 101 Processor, 102 Memory, 103 Auxiliary storage device, 104 Communication device, 105 Input / Output interface, 110 Secret string extraction unit, 120 Mask string determination unit, 121 Character type discrimination unit, 122 Rule selection unit, 123 Rule application unit, 130 Masking unit, 140 Target document generation unit, 190 Storage unit, 191 Original document data, 192 Mask rule data, 193 Mask document data, 194 Target document data, 200 Masking system.
Claims
1. A mask string determination unit determines a mask string to mask the secret string, based on the character types of the secret string, which is a string representing confidential information among the strings shown in the document data. A masking unit that uses the mask string to mask the secret string in the document data, A masking system equipped with [a specific feature].
2. The mask string determination unit determines the character type of the secret string based on the character code of the secret string. The masking system according to claim 1.
3. The mask string determination unit determines the mask string to be a string of the same character type as the secret string. The masking system according to claim 1.
4. The mask string determination unit selects a mask rule from the mask rule data, which indicates mask rules associated with each character type, that corresponds to the same character type as the secret string, and determines the mask string according to the selected mask rule. The masking system according to claim 1.
5. The system includes a secret string extraction unit that extracts the secret string from the document data. The masking system according to claim 1.
6. The secret string extraction unit extracts the secret string from the document data using a part-of-speech analysis function. The masking system according to claim 5.
7. The secret string extraction unit extracts the secret string from the document data using generating AI. The masking system according to claim 5.
8. The secret string extraction unit extracts the secret string from the document data using a regular expression. The masking system according to claim 5.
9. The system includes a target document generation unit that generates target document data using the masked document data to generate new target document data that shows a new sentence based on the sentence shown in the unmasked document data, without including the confidential information. A masking system according to any one of claims 1 to 8.
10. The aforementioned target document generation unit takes the masked document data as input and generates the target document data using generation AI. The masking system according to claim 9.
11. The mask string determination unit determines a mask string to mask the secret string, based on the character types of the secret string, which is a string representing confidential information among the strings shown in the document data. The masking unit uses the mask string to mask the secret string in the document data. Masking method.
12. A mask string determination process that determines a mask string to mask the secret string, based on the character types of the secret string, which is a string representing confidential information among the strings shown in the document data, A masking process that masks the secret string in the document data using the mask string, A masking program to cause a computer to execute something.