Evaluation device

WO2026203217A1PCT designated stage Publication Date: 2026-10-01NT T INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/012535
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2026-10-01

Smart Images

  • Figure JP2025012535_01102026_PF_FP_ABST
    Figure JP2025012535_01102026_PF_FP_ABST
Patent Text Reader

Abstract

This evaluation device extracts a continuous token sequence from teacher data of a language model, and divides the extracted token sequence into a prefix, which is a token sequence of a predetermined number of tokens from the beginning of the extracted token sequence, and a suffix, which is a token sequence following the prefix. Next, the evaluation device creates a first token sequence such that, when the first token sequence is input to the language model, the probability that the suffix is generated as a second token sequence following the first token sequence is increased. The evaluation device then inputs the created first token sequence to the language model, and causes the language model to generate the second token sequence following the first token sequence. Thereafter, the evaluation device evaluates whether the second token sequence matches the suffix, and outputs the result of the evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Evaluation Apparatus

[0001] The present invention relates to an evaluation apparatus for evaluating the risk of information leakage and copyright infringement of data used as training data for language models.

[0002] Conventionally, there are technologies for evaluating the risk of information leakage and copyright infringement of data used as training data for large language models (LLMs). For example, there is a technology for evaluating whether the same data as the training data is output from an LLM (see Non-Patent Documents 1 and 2).

[0003] The above technology is a technique that inputs a part of a sentence collected from the web or a part of a sentence included in training data as a prefix to the LLM, and checks whether the subsequent sentence (suffix) of the prefix output from the LLM is included in the training data.

[0004] N. Carlini, et al., "Extracting Training Data from Large Language Models," USENIX Security Symposium, 2021.M. Nasr, et al., "Scalable Extraction of Training Data from(Production)Language Models," arXiv preprint arXiv:2311.17035, 2023.

[0005] However, the above technology does not take into account the risk that a suffix included in the training data may be output from the LLM when a sentence other than the aforementioned prefix is input to the LLM. For this reason, there is a possibility that the risk of information leakage and copyright infringement of LLMs may be underestimated.

[0006] Accordingly, an object of the present invention is to solve the aforementioned problems and appropriately evaluate risks such as information leakage and copyright infringement of LLMs.

[0007] To solve the aforementioned problems, the present invention is characterized by comprising: a token sequence extraction unit that extracts a sequence of consecutive tokens from training data of a language model and divides the extracted token sequence into a prefix, which is a predetermined number of token sequences from the beginning, and a suffix, which is a token sequence that follows the prefix; a creation unit that creates a first token sequence such that, when the first token sequence is input to the language model, the probability that the suffix will be generated as a second token sequence following the first token sequence is increased; a generation unit that inputs the created first token sequence to the language model and generates a second token sequence following the first token sequence; and an evaluation unit that evaluates whether the generated second token sequence matches the suffix and outputs the result of the evaluation.

[0008] According to the present invention, the risks of information leakage and copyright infringement of LLMs can be appropriately evaluated.

[0009] Figure 1 is a diagram illustrating the overview of the evaluation device. Figure 2 is a diagram showing an example of the configuration of the evaluation device. Figure 3 is a flowchart showing an example of the processing procedure performed by the evaluation device. Figure 4 is a flowchart showing an example of the prefix optimization processing procedure performed by the evaluation device. Figure 5 is a flowchart showing an example of the prefix optimization processing procedure performed by the evaluation device. Figure 6 is a flowchart showing an example of the prefix optimization processing procedure performed by the evaluation device. Figure 7 is a flowchart showing an example of the prefix optimization processing procedure performed by the evaluation device. Figure 8 is a diagram showing an example of a computer that runs the evaluation program.

[0010] The following describes embodiments for carrying out the present invention with reference to the drawings. The present invention is not limited to these embodiments.

[0011] [Overview] The overview of the evaluation device of this embodiment will be explained using Figure 1. The evaluation device evaluates the risk of information leakage and copyright infringement of data used as training data for a language model, for example, in the following way. In the following explanation, the language model to be evaluated for risk is described as a large-scale language model (LLM) as an example, but is not limited to this.

[0012] First, the evaluation device extracts a sequence of consecutive tokens from the LLM training data, splits it into a prefix and a target suffix. Then, the evaluation device optimizes the prefix so that the target suffix is ​​more likely to be output from the LLM (details of prefix optimization are described later).

[0013] Subsequently, the evaluation device inputs the optimized prefix into the LLM, causing the LLM to output a suffix following that prefix. Finally, the evaluation device assesses whether the suffix output by the LLM matches the target suffix (risk assessment).

[0014] In this way, the evaluation device assesses the risk of information leakage from LLM by using a prefix that increases the probability that the target suffix will be output by LLM, thus preventing underestimation of the risk.

[0015] [Configuration Example] Next, an example of the configuration of the evaluation device 10 will be described using Figure 2. The evaluation device 10 includes, for example, an input / output unit 11, a communication unit 12, a storage unit 13, and a control unit 14.

[0016] The input / output unit 11 is an interface that handles the input and output of various types of data. For example, the input / output unit 11 receives input requests for LLM risk assessment and outputs the LLM risk assessment results from the control unit 14. The communication unit 12 is an interface for communicating with external devices via a network. For example, the communication unit 12 accesses the LLM to be assessed for risk and its training data via the network.

[0017] The memory unit 13 stores data, programs, etc., that are referenced when the control unit 14 performs various processes. The memory unit 13 is implemented by semiconductor memory elements such as RAM (Random Access Memory) and flash memory, or by storage devices such as hard disks and optical discs.

[0018] The control unit 14 is responsible for controlling the entire evaluation device 10. The functions of the control unit 14 are realized, for example, by the CPU (Central Processing Unit) executing a program stored in the memory unit 13.

[0019] [Control Unit] The control unit 14 includes a token sequence extraction unit 141, a prefix optimization unit (creation unit) 142, a suffix generation unit (generation unit) 143, and a risk assessment unit 144. The determination unit 145, shown by the dashed line, may or may not be equipped in the evaluation device 10; the case where it is equipped will be described later.

[0020] [Token Sequence Extraction Unit] The token sequence extraction unit 141 extracts consecutive token sequences from the LLM training data. The token sequence extraction unit 141 then divides the extracted token sequence into a prefix, which is a predetermined number of token sequences from the beginning, and a suffix, which is the token sequence that follows the prefix.

[0021] For example, the token sequence extraction unit 141 extracts a sequence of consecutive tokens [p,s t Extract ]. [p,s t ] of which p is prefix, s t The suffix is ​​used for the target. t The number of tokens is, for example, p is around 5 to 50, s t For example, it is around 50.

[0022] The token sequence extraction unit 141, for example, when extracting token sequences from training data, uses LLM to extract token sequences that are considered to pose a high risk if the same token sequence is generated repeatedly. For example, when the evaluation device 10 evaluates privacy-related risks, the token sequence extraction unit 141 extracts token sequences from emails and content created by individual users. Also, when the evaluation device 10 evaluates copyright infringement risks, the token sequence extraction unit 141 extracts token sequences from web articles, books, and other content. Regardless of which risk the evaluation device 10 is evaluating, it is preferable not to use data in which the same token sequence appears repeatedly, such as web page templates.

[0023] [Prefix Optimization Unit] The prefix optimization unit 142 optimizes the prefix so that the LLM can easily generate the target suffix. For example, if the prefix optimization unit 142 receives a token sequence as a prefix from the LLM, it creates a token sequence that increases the probability that the target suffix will be generated as the suffix following that token sequence. Examples of prefix optimization methods include the following:

[0024] [Method 1: In-context learning method] For example, the prefix optimization unit 142 extracts combinations of prefix and suffix from the training data and creates an in-context learning instruction for LLM using the extracted combinations. The prefix optimization unit 142 then outputs the created instruction as an optimized prefix.

[0025] For example, let [p1,s1] and [p2,s2] be combinations of prefix and suffix separately extracted from the above training data. The prefix optimization unit 142 may randomly extract the above [p1,s1] and [p2,s2] from the training data, or [p,s t You may also extract data similar to [ ]. The number of pairs to extract can be, for example, 1 to 10, but there is no limit to the number.

[0026] Then, the prefix optimization unit 142 uses the above prefixes p, [p1,s1], and [p2,s2] to create and output the following in-context learning instructions for LLM.

[0027] "Generate suffixes following prefixes. prefix:p1,suffix:s1, prefix:p2,suffix:s2, prefix:p,suffix:"

[0028] The prefix optimization unit 142 creates the above-mentioned in-context learning instruction statement for the LLM as a prefix, thereby increasing the probability that when the prefix is ​​input to the LLM, the above-mentioned target suffix will be generated (output) as a suffix following that prefix in the LLM.

[0029] [Method 2: Method to improve the probability of prefix generation] Alternatively, the prefix optimization unit 142 may optimize the prefix as follows. For example, the prefix optimization unit 142 divides the above prefix p into two parts, [p1, p2], and updates (changes) p1 so that the probability of p2 being generated in LLM (generation probability) is improved.

[0030] For example, when p1 is input to the LLM, the prefix optimization unit 142 performs a token substitution on p1 so that the probability of generating p2 following p1 is improved. The substituted token is denoted as p1'.

[0031] Furthermore, the substitution of p1 tokens that improves the generation probability of p2 can be performed using the technique described in Reference 1 below.

[0032] Reference 1: A. Zou, et al., Universal and Transferable Adversarial Attacks on Aligned Language Models, arXiv:2307.15043, 2023.

[0033] The prefix optimization unit 142 then outputs [p1',p2], which has undergone the above substitution, as the optimized prefix.

[0034] Furthermore, when the prefix optimization unit 142 performs the token replacement of p1 as described above, it may insert a token sequence a before p1 and then perform the replacement of the token sequence into which a has been inserted.

[0035] In this case, when [a,p1] is input to the LLM, the prefix optimization unit 142 performs a substitution of a and p1 to improve the probability that p2 will be generated as the token following [a,p1]. The substituted a and p1 are denoted as a' and p1'. The prefix optimization unit 142 then outputs [a',p1',p2], which has undergone the above substitution, as the optimized prefix. Note that the above substitution can also be performed using the technique described in the above-mentioned reference 1.

[0036] [Method 3: Method for eliciting a positive response from LLM] If, when a prefix is ​​input to LLM, there is a possibility that LLM may reject the output of the token sequence following that prefix, the prefix optimization unit 142 may optimize the prefix as follows. For example, the prefix optimization unit 142 optimizes the prefix so that the LLM response starts with a positive suffix like the following.

[0037] "Yes. We will generate a suffix to follow the prefix. prefix: p, suffix:

[0038] For example, if a prefix created by Method 1 or Method 2 is input to the LLM, but the LLM responds by refusing to output the token sequence following that prefix, the prefix optimization unit 142 will create a prefix that, when the prefix is ​​input to the LLM, increases the probability of generating a token sequence (suffix) that starts with a string with a positive meaning (for example, "Yes.") as described above.

[0039] For example, the prefix optimization unit 142 performs a substitution of the token p in LLM so as to improve the probability of generating such a positive suffix. The prefix optimization unit 142 then outputs the substituted p' as the optimized prefix.

[0040] Furthermore, when replacing the token p described above, the prefix optimization unit 142 may insert token strings before and after p. For example, let the inserted token strings be a and b, and when [a,p,b] is input to the LLM, replacement of a, p, and b is performed so as to increase the probability that the above-described suffix is generated as a token following [a,p,b]. Let a, p, and b after replacement be a', p', and b'. Then, the prefix optimization unit 142 outputs [a',p',b'] subjected to the above replacement as an optimized prefix. Note that the above replacement can also be performed using the technique described in Document 1 above.

[0041] [Suffix Generation Unit] When the suffix generation unit 143 acquires the prefix optimized by the prefix optimization unit 142, the suffix generation unit 143 inputs the prefix to the LLM, and causes the LLM to generate a suffix following the prefix. Then, the suffix generation unit 143 outputs the suffix generated by the LLM to the risk evaluation unit 144.

[0042] As a method for generating the suffix, there are various methods such as a method of imparting randomness to the output from the LLM. If this setting is possible, a method of selecting the suffix with the highest generation probability in the LLM can be considered. When the output from the LLM has randomness, the suffix generation unit 143 may input the optimized prefix to the LLM a plurality of times and acquire a plurality of suffixes from the LLM.

[0043] [Evaluation Unit] The risk evaluation unit 144 evaluates whether the suffix generated by the LLM matches the target suffix, and outputs the evaluation result. The risk evaluation unit 144 may evaluate whether the suffix acquired from the suffix generation unit 143 completely matches the target suffix, or may evaluate whether the Jaccard coefficient, edit distance, longest common subsequence, or the like of the tokens of the suffix acquired from the suffix generation unit 143 and the target suffix is equal to or greater than a threshold value.

[0044] Furthermore, when the risk assessment unit 144 acquires a plurality of suffixes from the suffix generation unit 143, it may be determined as a match if one or more of the suffixes match the target suffix.

[0045] Furthermore, for example, when the risk assessment unit 144 intends to assess the risk related to a specific token sequence for the LLM to be evaluated, the risk may be assessed using that token sequence. Further, when the risk assessment unit 144 intends to assess the risk of the entire dataset (training data) for the LLM to be evaluated, the assessment can be performed by extracting a plurality of token sequences from the training data and calculating the probability that the LLM outputs a matching suffix.

[0046] [Example of processing procedure] Next, an example of a processing procedure executed by the evaluation apparatus 10 will be described with reference to FIG. 3. First, when the evaluation apparatus 10 receives an LLM evaluation request, the token sequence extraction unit 141 extracts a continuous token sequence from training data of the LLM (S1), and splits the extracted token sequence into a prefix and a target suffix. Thereafter, the prefix optimization unit 142 optimizes the prefix so that the target suffix is easily output from the LLM (S2).

[0047] After S2, the suffix generation unit 143 inputs the prefix optimized in S2 to the LLM, and causes the LLM to generate a suffix following the prefix (S3). Then, the risk assessment unit 144 assesses whether the suffix generated in S3 matches the target suffix (S4: risk assessment). Then, the risk assessment unit 144 outputs the assessment result.

[0048] Next, the process of S2 in FIG. 3 will be described in detail with reference to FIG. 4. Here, an example of a processing procedure when the prefix optimization unit 142 executes prefix optimization by in-context learning (method 1) of the LLM in S2 of FIG. 3 will be described.

[0049] For example, the prefix optimization unit 142 extracts combinations of prefix and suffix from the training data (S21). Note that the above combinations of prefix and suffix are different from the token sequences extracted by the token sequence extraction unit 141.

[0050] After S21, the prefix optimization unit 142 uses the prefix created by the token sequence extraction unit 141 and the combination of prefix and suffix extracted in S21 to create an in-context learning instruction for LLM, and outputs the created instruction as an optimized prefix (S22: Creates and outputs an in-context learning instruction for LLM using the combination of prefix and suffix extracted).

[0051] Furthermore, using Figures 5 and 6, an example of the processing procedure when the prefix optimization unit 142 performs prefix optimization in S2 of Figure 3 by a method (Method 2) that improves the probability of prefix generation will be explained.

[0052] For example, the prefix optimization unit 142 divides the prefix p into [p1,p2] (S31 in Figure 5). Next, the prefix optimization unit 142 replaces p1 with p1' in the LLM so that the probability of generating p2 is improved (S32). After that, the prefix optimization unit 142 outputs the substituted [p1',p2] as the optimized prefix (S33).

[0053] Furthermore, when the prefix optimization unit 142 performs the token substitution of p1 as described above, if it inserts token sequence a before p1, it performs the substitution using, for example, the processing procedure shown in Figure 6.

[0054] The prefix optimization unit 142, similar to S31 in Figure 5, divides the prefix p into [p1,p2] (S41), and then, when [a,p1] is input to the LLM, it replaces a and p1 with a',p1' to improve the probability of generating p2 (S42). After that, the prefix optimization unit 142 outputs the replaced [a',p1',p2] as the optimized prefix (S43).

[0055] Furthermore, using Figure 7, an example of the processing procedure when the prefix optimization unit 142 performs prefix optimization in S2 of Figure 3 by a method (Method 3) that extracts a positive response from the LLM will be explained.

[0056] For example, the prefix optimization unit 142 replaces prefix p with p' in LLM so that the probability of generating a token sequence starting with a string with a positive meaning is improved (S51). Then, the prefix optimization unit 142 outputs the replaced p' as the optimized prefix (S52).

[0057] [Other Embodiments] The evaluation device 10 may further include a determination unit 145 (see Figure 2) that selects a prefix optimization method (method 1, method 2, method 3 described above) to be used by the prefix optimization unit 142.

[0058] For example, the determination unit 145 inputs the prefix p created by the token sequence extraction unit 141 to the LLM and obtains its response. If the determination unit 145 determines that the response from the LLM is not a token sequence following p (and is not a suffix of p), it selects "Method 1" as the prefix optimization method that the prefix optimization unit 142 should use. The prefix optimization unit 142 then optimizes the prefix using Method 1. As a result, the evaluation device 10 can cause the prefix optimization unit 142 to create a prefix that more reliably generates a suffix for the LLM.

[0059] On the other hand, if the determination unit 145 determines that the response from LLM is a sequence of tokens following p (it is a suffix of p), but is different from the target suffix, the determination unit 145 selects "Method 2" as the prefix optimization method that the prefix optimization unit 142 should use. The prefix optimization unit 142 then optimizes the prefix using Method 2. As a result, the evaluation device 10 can cause the prefix optimization unit 142 to create a prefix that improves the probability of generating the target suffix in LLM.

[0060] Furthermore, the determination unit 145 inputs the prefix p created by the token sequence extraction unit 141 to the LLM, and if it determines that the response matches the target suffix, it determines that prefix optimization by the prefix optimization unit 142 is unnecessary.

[0061] Furthermore, if the LLM receives a prefix p created by the token sequence extraction unit 141, or a prefix optimized by the prefix optimization unit 142 using method 1 or method 2, and the LLM responds by rejecting the output of the token sequence following that prefix, the determination unit 145 selects "method 3" as the prefix optimization method that the prefix optimization unit 142 should use. The prefix optimization unit 142 then optimizes the prefix using method 3. As a result, the evaluation device 10 can cause the prefix optimization unit 142 to create a prefix that is less likely to be rejected by the LLM.

[0062] Furthermore, if the determination unit 145 inputs the prefix created by method 3 described above to the LLM, but determines that the response does not match the target suffix, the prefix optimization unit 142 may select a prefix optimization method (method 1 or method 2) to use, depending on the content of the response from the LLM. In that case, the prefix optimization unit 142 recreates the prefix using the selected prefix optimization method (method 1 or method 2).

[0063] [System Configuration, etc.] Furthermore, the components of each part shown in the diagram are functional concepts and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown in the diagram, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions. In addition, all or any part of the processing functions performed by each device can be realized by a CPU and the program executed on that CPU, or by hardware using wired logic.

[0064] Furthermore, among the processes described in the embodiments described above, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above document and drawings can be arbitrarily changed unless otherwise specified.

[0065] [Program] The evaluation device 10 described above can be implemented by installing a program (evaluation program) as packaged software or online software on a desired computer. For example, by having the computer run the above program, the computer can be made to function as the evaluation device 10. The term "computer" here includes mobile communication terminals such as smartphones, mobile phones and PHS (Personal Handyphone System), and terminals such as PDA (Personal Digital Assistant).

[0066] Figure 8 shows an example of a computer running an evaluation program. Computer 1000 has, for example, memory 1010 and CPU 1020. Computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.

[0067] Memory 1010 includes ROM (Read Only Memory) 1011 and RAM (Random Access Memory) 1012. ROM 1011 stores, for example, a boot program such as BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to the hard disk drive 1090. The disk drive interface 1040 is connected to the disk drive 1100. For example, a removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is connected to, for example, a display 1130.

[0068] The hard disk drive 1090 stores, for example, the OS 1091, application program 1092, program module 1093, and program data 1094. That is, the program that defines each process executed by the evaluation device 10 is implemented as a program module 1093 in which executable code for a computer is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for executing a process similar to the functional configuration of the evaluation device 10 is stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).

[0069] Furthermore, the data used in the processing of the above-described embodiment is stored as program data 1094 in, for example, memory 1010 or hard disk drive 1090. The CPU 1020 then reads the program module 1093 and program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as needed and executes them.

[0070] Furthermore, the program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090; for example, they may be stored in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (LAN (Local Area Network), WAN (Wide Area Network), etc.). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via a network interface 1070.

[0071] 10 Evaluation device 11 Input / Output unit 12 Communication unit 13 Storage unit 14 Control unit 141 Token sequence extraction unit 142 Prefix optimization unit (creation unit) 143 Suffix generation unit (generation unit) 144 Risk assessment unit (evaluation unit) 145 Judgment unit

Claims

1. An evaluation device comprising: a token sequence extraction unit that extracts a sequence of consecutive tokens from training data of a language model and divides the extracted token sequence into a prefix, which is a predetermined number of token sequences from the beginning of the sequence, and a suffix, which is a token sequence that follows the prefix; a creation unit that creates a first token sequence such that, when the first token sequence is input to the language model, the probability that the suffix will be generated as a second token sequence following the first token sequence increases; a generation unit that inputs the created first token sequence to the language model and generates a second token sequence following the first token sequence; and an evaluation unit that evaluates whether the generated second token sequence matches the suffix and outputs the result of the evaluation.

2. The evaluation apparatus according to claim 1, characterized in that the creation unit creates an instruction statement that instructs the language model to generate a token sequence following the prefix obtained by the token sequence extraction unit, using in-context learning of the language model, which uses a combination of prefix and suffix separately extracted from the training data, as the first token sequence.

3. The evaluation apparatus according to claim 1, characterized in that the creation unit divides the prefix into a third token sequence which is a predetermined number of token sequences from the beginning and a fourth token sequence which follows the third token sequence as the first token sequence, modifies the third token sequence to increase the probability that the fourth token sequence is generated as the token sequence which follows the third token sequence when the third token sequence is input to the language model, and creates a token sequence in which the portion of the third token sequence in the prefix is ​​replaced with the modified third token sequence.

4. The evaluation device according to claim 1, wherein, when the token sequence created by the creation unit is input to the language model, but the language model responds by refusing to output a token sequence following the token sequence, the creation unit creates the first token sequence to the language model such that, when the first token sequence is input, the probability that a second token sequence beginning with a string of positive meaning will be generated following the first token sequence is increased.