Determination device, determination method, and determination program
Patent Information
- Application Number
- JP2025032068
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2026-09-09
AI Technical Summary
【0012】 本発明によれば、LLMの利用時におけるリスクを低減するとともに、リスク排除対象のカテゴリを柔軟にカスタマイズできる。
Smart Images

Figure 2026144648000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a determination apparatus, a determination method, and a determination program.
Background Art
[0002] Generative AI (Artificial Intelligence) that generates text, images, audio, and video using trained machine learning models has attracted attention (see, for example, Patent Documents 1 and 2).
[0003] For example, as a text-based generative AI, there is a Large Language Model (LLM), which is a natural language processing model trained using a large amount of text data. In particular, GPT (Generative Pre-trained Transformer) (chatGPT (registered trademark)) from OpenAI (registered trademark) has attracted much attention and has gained a large number of users.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
Problem to be Solved by the Invention
[0005] Here, it is difficult to prevent the leakage of personal information from the output of an LLM. This is because it is difficult to erase data once learned by an LLM, and it is known that email addresses of specific companies included in training data can be obtained even from ChatGPT (registered trademark) by using a special prompt.
[0006] Furthermore, because it is not possible to completely eliminate hallucination in LLMs at this time, LLMs may output incorrect expertise. In particular, LLMs may generate incorrect answers to questions in areas such as law, medicine, health, and finance, which could result in direct harm to users.
[0007] Thus, there was a problem in that it was difficult to eliminate all potential risks from the LLM output. Furthermore, LLM operators specifically requested that risks be removed from the LLM output in categories that they desired.
[0008] The present invention has been made in view of the above, and aims to provide a determination device, determination method, and determination program that reduce the risks when using LLM and allow for flexible customization of the categories to be targeted for risk elimination. [Means for solving the problem]
[0009] To solve the above-mentioned problems and achieve the objective, the determination device of the present invention is characterized by comprising: a linking unit that exchanges information with a processing device having a first natural language processing model that generates text; an evaluation unit that evaluates the safety of a first prompt that instructs the first natural language processing model to generate text, and / or the safety of text generated by the first natural language processing model, using a machine learning model that evaluates the safety of text for each of a plurality of safety categories; a determination unit that determines whether or not to input the first prompt to the first natural language processing model, and / or whether or not to output the text generated by the first natural language processing model to the source that requested the text, based on the safety of each category when the machine learning model evaluates the first prompt and / or the text generated by the first natural language processing model; a notification unit that notifies the processing device of the determination result by the determination unit; and an addition unit that, upon receiving a request to add a category from the processing device, adds the requested category as a custom category associated with the processing device to the categories to be evaluated by the machine learning model.
[0010] Furthermore, the determination method of the present invention is a determination method executed by a determination device, and is characterized by including the steps of: exchanging information with a processing device having a first natural language processing model that generates text; evaluating the safety of a first prompt that instructs the first natural language processing model to generate text, and / or the safety of text generated by the first natural language processing model, using a machine learning model that evaluates the safety of text for each of a plurality of safety categories; determining whether or not to input the first prompt to the first natural language processing model, and / or whether or not to output the text generated by the first natural language processing model to the source that requested the text, based on the safety of each category when the machine learning model evaluated the first prompt and / or the text generated by the first natural language processing model; notifying the processing device of the determination result in the determination step; and, when the processing device receives a request to add a category, adding the requested category as a custom category associated with the processing device to the categories to be evaluated by the machine learning model.
[0011] Furthermore, the determination program of the present invention causes a computer to perform the following steps: exchange information with a processing device having a first natural language processing model that generates text; evaluate the safety of a first prompt that instructs the first natural language processing model to generate text, and / or the safety of text generated by the first natural language processing model, using a machine learning model that evaluates the safety of text for each of several categories related to safety; determine whether or not to input the first prompt to the first natural language processing model, and / or whether or not to output the text generated by the first natural language processing model to the source that requested the text, based on the safety of each category when the machine learning model evaluated the first prompt and / or the text generated by the first natural language processing model; notify the processing device of the determination result in the determination step; and, when the processing device requests the addition of a category, add the requested category as a custom category associated with the processing device to the categories to be evaluated by the machine learning model. [Effects of the Invention]
[0012] According to the present invention, the risks associated with using LLM are reduced, and the categories of risks to be eliminated can be flexibly customized. [Brief explanation of the drawing]
[0013] [Figure 1] Figure 1 is a diagram showing an example of the configuration of a communication system according to an embodiment. [Figure 2] Figure 2 is a diagram illustrating the overview of the communication system processing in the embodiment. [Figure 3] Figure 3 shows an example of the notification content from the determination device shown in Figure 1 to the service provision device. [Figure 4] Figure 4 illustrates the process of adding categories in the embodiment. [Figure 5] Figure 5 shows an example of the notification content from the determination device shown in Figure 1 to the service provision device. [Figure 6] FIG. 6 is a diagram illustrating an example configuration of the determination apparatus shown in FIG. 1. [Figure 7] FIG. 7 is a diagram illustrating an example evaluation result of the evaluation model shown in FIG. 6. [Figure 8] FIG. 8 is a diagram for explaining custom category addition processing according to the embodiment. [Figure 9] FIG. 9 is a diagram for explaining custom category addition processing according to the embodiment. [Figure 10] FIG. 10 is a diagram for explaining custom category addition processing according to the embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of a text determination request transmitted from the service providing apparatus shown in FIG. 1 to the determination apparatus. [Figure 12] FIG. 12 is a diagram illustrating an example evaluation result of the evaluation model shown in FIG. 6. [Figure 13] FIG. 13 is a diagram illustrating an example of content notified from the determination apparatus shown in FIG. 1 to the service providing apparatus. [Figure 14] FIG. 14 is a diagram for explaining assist processing for creating category descriptions according to the embodiment. [Figure 15] FIG. 15 is a diagram illustrating an example of an assist prompt. [Figure 16] FIG. 16 is a diagram illustrating an example description of a category created by the LLM shown in FIG. 1. [Figure 17] FIG. 17 is a diagram illustrating an example description generated using the LLM shown in FIG. 1. [Figure 18] FIG. 18 is a diagram illustrating an example description generated using the LLM shown in FIG. 1. [Figure 19] FIG. 19 is a diagram illustrating an example evaluation result of the evaluation model shown in FIG. 6. [Figure 20] FIG. 20 is a diagram illustrating an example description generated using the LLM shown in FIG. 1. [Figure 21] FIG. 21 is a diagram illustrating an example evaluation result of the evaluation model shown in FIG. 6. [Figure 22]Figure 22 is a sequence diagram showing the processing procedure for communication processing in the embodiment. [Figure 23] Figure 23 is a sequence diagram showing the processing procedure for communication processing in the embodiment. [Figure 24] Figure 24 is a sequence diagram showing the processing procedure for communication processing in the embodiment. [Figure 25] Figure 25 is a sequence diagram showing the processing procedure for communication processing in the embodiment. [Figure 26] Figure 26 shows the evaluation model of the judgment device and the accuracy evaluation of other models. [Figure 27] Figure 27 illustrates an example of the application of the embodiment. [Figure 28] Figure 28 illustrates an example of the application of the embodiment. [Figure 29] Figure 29 illustrates an example of the application of the embodiment. [Figure 30] Figure 30 illustrates an example of the application of the embodiment. [Figure 31] Figure 31 shows an example of output provided to a user by a service provider. [Figure 32] Figure 32 shows an example of a service screen for a conventional help desk bot. [Figure 33] Figure 33 shows an example of a help desk bot service screen when the determination device according to the embodiment is applied for safety determination. [Figure 34] Figure 34 shows an example of an LLM response returned when the embodiment is not applied. [Figure 35] Figure 35 shows an example of an LLM response returned when the embodiment is not applied. [Figure 36] Figure 36 shows an example of a computer in which a determination device is realized when a program is executed. [Modes for carrying out the invention]
[0014] Hereinafter, one embodiment of the present invention will be described in detail with reference to the drawings. However, the present invention is not limited by this embodiment. Furthermore, in the drawings, the same parts are denoted by the same reference numerals.
[0015] [Embodiment] [Processing System] The configuration of the communication system according to Embodiment 1 will now be described. The communication system according to Embodiment has a determination device that evaluates the safety of input and output text of a Large-Language-Model (LLM) and determines whether input and output to the LLM are permitted based on the evaluation results. In the communication system according to Embodiment, the risk when using the LLM is reduced by restricting the input of LLMs with low safety prompts and / or outputs from LLMs with low safety.
[0016] Figure 1 shows an example of the configuration of a communication system according to an embodiment. As shown in Figure 1, the communication system according to the embodiment includes a user terminal 20 used by an end user, a service provision device 30 (processing device) that provides services to the end user by communicating with the user terminal 20, a generation AI (Artificial Intelligence) server 40, and a determination device 10 that communicates between the service provision device 30 and the generation AI server 40. Note that the number of user terminals 20 may be two or more.
[0017] The user terminal 20 is a terminal device capable of inputting text data and audio data, outputting text data and audio data, and communicating with the service provider 30. The user terminal 20 is, for example, a PC (Personal Computer), a notebook PC, a tablet terminal, a smartphone, etc. The user terminal 20 sends a prompt (first prompt) to the service provider 30 instructing the LLM 31 to generate a response text to the inquiry. When the user terminal 20 receives text generated by the LLM 31 (first natural language processing model) from the service provider 30, it displays and / or outputs the received text as audio.
[0018] The service provider 30 has an LLM 31. The LLM 31 generates text according to prompts entered from the user terminal 20. Text is entered into the LLM 31. The text entered into the LLM 31 may be text data converted from voice data. The service provider 30 uses the LLM 31 to provide a chatbot service, provide a web page for inquiries, and / or control connections to a customer center.
[0019] The determination device 10 evaluates the safety of the input and output text of LLM31 for each of several safety categories, and determines whether input and output to LLM31 are permitted based on the evaluation results. The determination device 10 restricts input of low-safety prompts to LLM31, and / or restricts low-safety outputs from LLM31. When the determination device 10 receives a request to add a category from the service provider 30, it adds the requested category to the evaluation categories of the evaluation model 134.
[0020] The generating AI server 40 has an LLM 41. The LLM 41 (second natural language processing model) generates text according to the prompt input from the judgment device 10 and returns the generated text to the judgment device 10.
[0021] Figure 2 is a diagram illustrating the overview of the communication system processing in the embodiment. As shown in Figure 2, normally, when a prompt is input from the user terminal 20 to the LLM 31 (arrow Y1), the text generated by the LLM 31 in response to the prompt is output to the user terminal 20 (arrow Y2).
[0022] Here, there may be an input (prompt) (arrow Y11) from the user terminal 20 that suggests an attempt to misuse the LLM 31. In this embodiment, before input to the LLM 31, the determination device 10 uses an evaluation model 134 (machine learning model) that evaluates the safety of the text to determine the safety of the text to be input to the LLM 31 (Figure 2 (1)). The determination device 10 then instructs that input of low-safety prompts be restricted to the LLM 31 (Figure 2 (2)). Following the instructions of the determination device 10, the service provider 30 does not input any restricted prompts (arrow Y11) to the LLM 31.
[0023] Furthermore, text generated by LLM31 in accordance with permitted input prompts may lead to the leakage of confidential information, the generation of malicious content, or the provision of false expertise (arrow Y12). In this embodiment, before outputting to the requesting user terminal 20, the determination device 10 uses an evaluation model 134 (machine learning model) to determine the security of the text generated by LLM31 (Figure 2(1)). The determination device 10 then instructs that the output of generated text with low security be restricted (Figure 2(3)). The service provider 30, in accordance with the instructions of the determination device 10, does not output text whose output has been restricted (arrow Y12) to the user terminal 20.
[0024] If the determination device 10 determines that the use of LLM31 should be restricted, it notifies the service provider device 30 of the restriction on the use of LLM31 and the category that was evaluated as having the lowest level of safety. Figure 3 shows an example of the notification content from the determination device 10 to the service provider device 30 shown in Figure 1. As shown in Figure 3, the determination device 10 notifies the service provider device 30 of "Unsafe" and the category "Violent Crime".
[0025] Upon receiving this notification, the service provider 30 restricts the use of LLM31. In this way, the determination device 10 can reduce the risks associated with the use of LLM31. Default categories include, for example, violent crime, nonviolent crime, verbal abuse, sexual offenses, warning, indiscriminate weapons, hatred, self-injury, sexual content, child exploitation, privacy (personal information), professional advice, and / or intellectual property.
[0026] Furthermore, the determination device 10 customizes the categories of risk to be eliminated in accordance with the services provided by the service provision device 30. Figure 4 is a diagram illustrating the category addition process in this embodiment.
[0027] If there are categories of text that the LLM31 should exclude from input and output in a service provided using the service provider 30 (such as a chatbot service), the determination device 10 will store them as a category set unique to the service provider 30.
[0028] For example, let's consider the case where the operator of the service provision device 30 wants to add the category "Competitor Comparison" as a category for which they want to restrict the use of LLM 31. In this case, the operator enters "Competitor Comparison" in the category name field C11 on the operator's management screen W11 (Figure 4), enters the description text D11 which is the description of the text content corresponding to this category, and presses the add button B11. The management screen W11 is a screen provided by the judgment device 10 via the service provision device 30.
[0029] In response, the determination device 10 adds "Competitor Comparison" as a custom category for the service provision device 30. Then, the determination device 10 adds "Competitor Comparison" as a custom category associated with the service provision device 30.
[0030] The determination device 10 uses the evaluation model 134 to determine the safety of the input text to LLM31 and / or the text generated by LLM31 for the category "Competitor Comparison". If the safety of the input text to LLM31 and / or the text generated by LLM31 is low for the category "Competitor Comparison", the determination device 10 instructs input restrictions to LLM31 and output restrictions from LLM31.
[0031] If the determination device 10 determines that the use of LLM31 should be restricted, it notifies the service provider device 30 of the restriction on the use of LLM31 and the category "Competitor Comparison". Figure 5 shows an example of the notification content from the determination device 10 to the service provider device 30 shown in Figure 1. As shown in Figure 5, the determination device 10 notifies the service provider device 30 of "Unsafe" and the category "Competitor Comparison". Upon receiving this notification, the service provider device 30 restricts the use of LLM31.
[0032] Thus, the determination device 10 allows for flexible customization of the categories of risks to be eliminated by linking them to the service provision device 30.
[0033] Furthermore, if the determination device 10 receives a request from the service provider 30 to add a custom category, it may generate a description of the category to be added and present it to the service provider 30 to assist in creating the category description. In addition, the determination device 10 may switch the categories to be evaluated in the evaluation model 134 to the default categories or custom categories associated with the service provider 30, in accordance with the services provided by the service provider 30.
[0034] [Judgment device] Next, the determination device 10 will be described. Figure 6 is a diagram showing an example of the configuration of the determination device 10 shown in Figure 1. As shown in Figure 6, the determination device 10 has a communication unit 11, a storage unit 12, and a control unit 13.
[0035] The communication unit 11 is a communication interface that sends and receives various types of information with other devices connected via a network or the like. The communication unit 11 is implemented using a NIC (Network Interface Card) or the like, and communicates between other devices (for example, a service provider 30) and the control unit 13 (described later) via telecommunication lines such as a LAN (Local Area Network) or the Internet.
[0036] The storage unit 12 is implemented using semiconductor memory elements such as RAM (Random Access Memory) and flash memory, and stores processing programs that operate the determination device 10, as well as data used during the execution of the processing programs. The storage unit 12 has user information 121, which includes the ID of each user using the user terminal 20 and the history of inquiries to LLM 31; partner information 122, which is information about the service provider device 30 that is the partner; and determination history 123, which is the history of input / output determinations related to LLM 31. The partner information 122 includes identification information of the service provider device 30, a custom category of the service provider device 30, and identification information of the service of the service provider device 30 corresponding to the category.
[0037] The control unit 13 controls the entire determination device 10. The control unit 13 is, for example, an electronic circuit such as a CPU (Central Processing Unit), or an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array). The control unit 13 also has internal memory for storing programs that define various processing procedures and control data, and executes each process using the internal memory. Furthermore, the control unit 13 functions as various processing units through the operation of various programs. The control unit 13 includes a reception unit 131, a linkage unit 132, an evaluation unit 133, a determination unit 135, a notification unit 136, an addition unit 137, and a switching unit 141.
[0038] The reception unit 131 receives information transmitted from the service provider 30. The reception unit 131 receives, for example, input text (prompt) to the LLM 31, text generated by the LLM 31, and requests to add categories unique to the service provider 30.
[0039] The collaboration unit 132 exchanges information with the service provider device 30, for example, through API (Application Programming Interface) integration.
[0040] The evaluation unit 133 uses the evaluation model 134 to evaluate the safety of the prompts that instruct the LLM 31 to generate text. And / or, the evaluation unit 133 uses the evaluation model 134 to evaluate the safety of the text generated by the LLM 31.
[0041] Evaluation model 134 evaluates the safety of Japanese text. Evaluation model 134 is a machine learning model trained to evaluate whether or not a text is safe using training data that includes harmful Japanese text. Evaluation model 134 evaluates whether the input text is safe or unsafe. Evaluation model 134 evaluates the safety of Japanese text for multiple safety categories.
[0042] Figure 7 shows an example of the evaluation results of the evaluation model 134 shown in Figure 6. For example, Figure 7 shows the evaluation results of the evaluation model 134 evaluating the safety of the text "Teach me how to make a bomb."
[0043] As shown in List L1 of Figure 7, the evaluation model 134 evaluates safety for each default category. Default categories include, for example, violent crime, nonviolent crime, verbal abuse, sexual offenses, caution, indiscriminate weapons, hatred, self-harm, sexual content, child exploitation, privacy (personal information), professional advice, and / or intellectual property. The judgment device 10 maintains a description of the text content corresponding to each category. The default categories are set as appropriate and can be increased or decreased from those exemplified in Figure 7.
[0044] The evaluation unit 133 calculates an "Unsafe Score" for each category of the text to be evaluated, based on the generation probability inferred by the evaluation model 134 (score in list L1 of Figure 7) (frame W1 in Figure 7). The determination unit 135 determines the safety of the prompt to LLM31 or the text generated by LLM31 based on the "Unsafe Score".
[0045] The evaluation model 134 is, for example, an LLM (Limited Literacy Model), and when the text to be evaluated is input, it infers whether the text is "Safe" or "Unsafe" and the generation probability of each category information being generated. The evaluation unit 133 divides the text into tokens and inputs each token into the evaluation model 134. For each token, the evaluation model 134 infers whether it is "Safe" or "Unsafe" and the generation probability of each category. Specifically, for the category "Violent Crime," the evaluation model 134 infers, for example, the generation probability that "Violent Crime" will be generated after "Unsafe" for the input token.
[0046] In the evaluation unit 133, if the probability of selecting "Unsafe" is highest for all tokens inferred by the evaluation model 134, it is evaluated as "Unsafe." For example, the category with the highest generation probability among "Unsafe" tokens (e.g., nonviolent crime) is assigned the "True" label. The evaluation unit 133 outputs to the determination unit 135 whether a token is "Safe" or "Unsafe," the categories assigned the "True" label and their generation probabilities, and the generation probabilities of other categories that were not assigned the "True" label. In Figure 7, the score corresponds to the generation probability.
[0047] The determination unit 135 determines whether or not to input a prompt to the LLM 31 based on the evaluation result by the evaluation unit 133. And / or, the determination unit 135 determines whether or not to output the text generated by the LLM 31 to the source of the text generation request (user terminal 20) based on the evaluation result by the evaluation unit 133.
[0048] If the evaluation unit 133 evaluates the prompt as unsafe, the determination unit 135 restricts input of the prompt to the LLM 31. On the other hand, if the evaluation unit 133 evaluates the prompt as safe, the determination unit 135 allows input of the prompt to the LLM 31.
[0049] Furthermore, if the evaluation unit 133 determines that the text generated by LLM31 is unsafe, the determination unit 135 restricts the output of the text generated by LLM31. On the other hand, if the evaluation unit 133 determines that the text generated by LLM31 is safe, the determination unit 135 allows the output of the text generated by LLM31.
[0050] Furthermore, the determination unit 135 determines whether or not to use the LLM31 based on the safety of each category when the evaluation model 134 evaluates the input text or output text to the LLM31.
[0051] For example, the determination unit 135 calculates an "Unsafe Score" for each category of the text to be evaluated, based on the generation probability inferred by the evaluation model 134 (score in list L1 of Figure 7) (frame W1 in Figure 7). Specifically, the determination unit 135 calculates a weighted sum of the scores of the top three categories judged as "Unsafe," and the result is the "Unsafe Score." The categories used in the calculation are not limited to the top three categories judged as "Unsafe," but are set appropriately depending on the score distribution, whether or not True / False labels are assigned, etc. Based on the "Unsafe Score," the determination unit 135 determines the safety of the prompt to LLM 31 or the text generated by LLM 31. The determination result by the determination unit 135 is stored in the storage unit 12 as a determination history 123.
[0052] The notification unit 136 notifies the service provider device 30 of the determination result made by the determination unit 135. If the determination unit 135 determines that the content shown in frame W1 of Figure 7 should restrict the use of LLM31, it notifies the service provider device 30 via the notification unit 136 of the restriction on the use of LLM31 and the category that was evaluated as having the lowest safety. The notification unit 136 may also notify the service provider device 30 of the "Unsafe Score".
[0053] When the add-on unit 137 receives a request to add a category from the service provision device 30, it adds the requested category to the categories to be evaluated in the evaluation model 134.
[0054] The additional unit 137 includes a category additional unit 138 (addition unit), a prompt creation unit 139 (creation unit), and an explanation presentation unit 140 (presentation unit).
[0055] The category addition unit 138 adds the requested category to the evaluation target categories of the evaluation model 134 in response to a category addition request from the service provision device 30. The category addition unit 138 adds the requested category as a custom category associated with the service provision device 30, separate from the default evaluation target categories.
[0056] This section explains the assistance function for creating category descriptions in the additional section 137.
[0057] When the prompt generation unit 139 receives a request to add a category and a request to generate a description of the category from the service provision device 30, it creates an assist prompt (second prompt) that instructs the LLM 41 to generate a description of the category to be added.
[0058] The explanation display unit 140 inputs an assist prompt to the LLM 41 and presents the explanation generated by the LLM 41 to the service provision device 30.
[0059] When the category addition unit 138 receives notification from the service provider 30 that it will adopt the explanation presented by the explanation presentation unit 140 as the explanation for the category to be added, it adds the category to be added as a custom category associated with the service provider 30, separate from the default categories, to the evaluation target of the text safety of the evaluation model 134.
[0060] If the prompt generation unit 139 receives notification from the service provider 30 that it will not adopt the explanation presented by the explanation presentation unit 140 as the explanation for the category to be added, it recreates the assist prompt. The explanation presentation unit 140 then inputs the recreated second prompt into the LLM 41, and the explanation generated by the LLM 41 is presented to the service provider 30 again.
[0061] The switching unit 141 switches the category to be evaluated by the evaluation model 134 to the default category or a custom category associated with the service provider 30, in accordance with the category selection instruction of the service provider 30. The evaluation model 134 evaluates the text safety for the category switched by the switching unit 141.
[0062] Furthermore, the switching unit 141 may switch the evaluation target category of the evaluation model 134 to the default category or a custom category associated with the service provider 30, in accordance with the service provided by the service provider 30. In this case, it is pre-registered whether the category associated with each service provided by the service provider 30 is the default category or a custom category associated with the service provider 30.
[0063] For example, when an end user enters text into the web page for connection control to the customer center provided by the service provider 30, the switching unit 141 sets the evaluation target category of the evaluation model 134 to the default category for the text entered into the connection control web page. On the other hand, when an end user enters text into the chatbot provided by the service provider 30, the evaluation target category of the evaluation model 134 to the custom category associated with the service provider 30 is set for the text entered into the chatbot.
[0064] [Adding a Category] Next, the process for adding custom categories in the determination device 10 will be described. Figures 8 to 10 illustrate the process for adding custom categories in the embodiment.
[0065] As shown in Figure 4, let's explain using the example of a case where the operator of the service provision device 30 wants to add the category "Competitor Comparison" as a category in which the use of LLM31 should be restricted. The operator enters "Competitor Comparison" in the category name field C11 on the operator's management screen W11 (Figure 4), enters a description for this category D11, and presses the add button B11.
[0066] After verifying the category description D11 in the playground, the operator can save and retain the category. Specifically, if the operator wants to retain the category "Competitor Comparison Set," they select the retain button B21 on the operator's management screen W21 (Figure 8). Note that each management screen, including management screens W11, W21, W31 (described later), and W41 (described later), is provided from the judgment device 10 to the terminal used by the operator via the service provision device 30.
[0067] Next, the category addition unit 138 assigns the category ID "ktg30a14" for the category "Competitor Comparison Set" to the service provision device 30 (Figure 9 (1)), and as shown in the management screen W31, the custom category "Competitor Comparison Detection Set" C31 becomes selectable. The category addition unit 138 stores the category ID "ktg30a14", the category set name "Competitor Comparison Detection Set", and the category description D31 verified by the operator as a custom category (code R31) of the service provision device 30.
[0068] Operators can edit the description of the custom category "Competitor Comparison Detection Set," delete categories, and add new categories.
[0069] For example, the operator selects the edit button B31 on the management screen W31 (Figure 9) for the custom category "Competitor Comparison Detection Set". In this case, the category addition section 138 accepts the operator's edits to the category name and category description D31.
[0070] The operator selects the delete button B32 on the management screen W31 (Figure 9) for the custom category "Competitor Comparison Detection Set". In this case, the category addition section 138 accepts the operator's request to delete the "Competitor Comparison Set".
[0071] Furthermore, the operator selects the "Add Category" button B33 on the management screen W31 (Figure 9). In this case, the category addition section 138 accepts the operator's addition of a new custom category.
[0072] The operator can select whether to evaluate the safety of LLM31's input and output text using the default category or the custom category "Competitor Comparison Detection Set" by pulling down the category set field C41 on the management screen W41 (Figure 10).
[0073] In this case, depending on the operator's category selection on the management screen W41, the switching unit 141 switches the evaluation target category of the evaluation model 134 to the default category or the custom category "Competitor Comparison Detection Set". The switching unit 141 may also automatically switch the evaluation target category of the evaluation model 134 to the default category or a custom category associated with the service provider 30, in accordance with the services provided by the service provider 30, regardless of the operator's category selection.
[0074] Figure 11 shows an example of a text judgment request sent from the service provider device 30 shown in Figure 1 to the judgment device 10. Figure 12 shows an example of the evaluation result of the evaluation model 134 shown in Figure 6. Figure 13 shows an example of the notification content from the judgment device 10 to the service provider device 30 shown in Figure 1.
[0075] Figures 11 to 13 illustrate the case where the end user of the service provider 30 sends the text T31, "ABC Company and XYZ Company offer similar services, but which one is better?", as input text to the LLM 31. The service provider 30 requests that this text T31 be evaluated in the custom category "Competitor Comparison Detection Set" (see Figure 11 "category_set_id").
[0076] In response, the judgment device 10 receives an evaluation request from the service provider 30 for text T31 in the custom category "Competitor Comparison Detection Set". The evaluation unit 133 calculates the "Unsafe Score" for the custom category "Competitor Comparison Detection Set" based on the generation probability that the evaluation model 134 inferred for text T31 using the description of the custom category "Competitor Comparison Detection Set". As shown in list L31 in Figure 12, for text T31, the evaluation model 134 evaluated the Unsafe Score of the competitor comparison category (frame R31) as 0.998 (frame P31), so the evaluation unit 133 assigns the True label to the competitor comparison category.
[0077] The determination unit 135 determines the safety of text T31 for the custom category "Competitor Comparison Detection Set" based on the "Unsafe Score" and the label. The determination unit 135 determines that there is a restriction on the use of LLM31 because the Unsafe Score for the competitor comparison category (frame R31 in Figure 12) is 0.998 (frame P31 in Figure 12).
[0078] As shown in Figure 13, the notification unit 136 notifies the service provider 30 of the text T31, including the usage restrictions of LLM31 ("Unsafe" and "true"), the category "Competitor Comparison" which was evaluated as having the lowest security, and the Unsafe Score "0.9977". The notification unit 136 may also send the list L31 shown in Figure 12 to the service provider 30.
[0079] [Assistance process for creating category descriptions] Next, we will explain the assistance process for creating category descriptions in the determination device 10.
[0080] When an operator adds a category, they need to create a descriptive text (category description) to exclude text that falls under that category. However, even operators may find it difficult to create category descriptions. Therefore, the judgment device 10 uses LLM41 to assist in creating category descriptions.
[0081] Figure 14 illustrates the assistance process for creating category descriptions in the embodiment.
[0082] This explanation will take the example of a case where the operator of the service provision device 30 wants to add the category "Competitor Comparison" as a category for which the use of LLM 31 should be restricted, and requests the automatic generation of a category description. In this case, the operator enters "Competitor Comparison" in the category name field C61 of the operator's management screen W61 (Figure 14) and selects the "Automatically generate description from category name" button B62. The management screen W61 is a screen provided by the judgment device 10 via the service provision device 30.
[0083] The determination device 10 then receives a request to add the category "Competitor Comparison" and a request to generate a description for this category.
[0084] In this case, the prompt creation unit 139 creates an assist prompt that instructs the generation of a description for the category to be added. The assist prompt created by the prompt creation unit 139 is set in LLM41. LLM41 is an LLM that can output in JSON format, such as ChatGPT® and Gemini.
[0085] Figure 15 shows an example of an assistance prompt. The assistance prompt P71 in Figure 15 includes an instruction R71 that instructs the generation of text (description) for a new category, and an instruction R73 that instructs the output format of the generated description. The prompt creation unit 139 inputs the category "Competitor Comparison," which was added by the operator, into the new category field in instruction R71. Then, the prompt creation unit 139 instructs instruction R73 to respond with the text for the new category as a JSON object.
[0086] Furthermore, the prompt generation unit 139 provides the assist prompt P71 with one or more default categories and corresponding descriptions as example sentences.
[0087] Specifically, the prompt generation unit 139 presents explanation example R72 as EXAMPLE in the assist prompt P71. As explanation example R72, the prompt generation unit 139 provides one or more category names and descriptions of the default categories, such as violent crime, nonviolent crime, abusive language, sexual offenses, caution, indiscriminate weapons, hatred, self-harm, sexual content, child exploitation, privacy, professional advice, and / or intellectual property, in the form of few-shots. Here, by increasing or decreasing the number of EXAMPLEs provided (two examples in Figure 15), randomness can be introduced into the explanation output results by LLM41.
[0088] The explanation presentation unit 140 provides the LLM41 with the assist prompt P71 created by the prompt creation unit 139. As a result, the LLM41 returns the explanation for the category "Competitor Comparison" to the determination device 10.
[0089] Figure 16 shows an example of a category description created by LLM41 shown in Figure 1. For example, LLM41 generates descriptions of Pattern 1 and Pattern 2 in Figure 16 according to the instructions of the assist prompt P71. The description presentation unit 140 presents the descriptions of Pattern 1 and Pattern 2 generated by LLM41 to the service provision device 30.
[0090] The operator uses the explanations of Pattern 1 and Pattern 2 presented to the determination device 10 to compare and verify them in the playground and select the optimal explanation.
[0091] The operator selects either Pattern 1 or Pattern 2 as the description for the category "Competitor Comparison" and selects the add button B61. As a result, the category addition unit 138 issues a category ID to the custom category "Competitor Comparison" of the service provision device 30, and stores the category ID, the category set "Competitor Comparison", and the category description D31 selected by the operator.
[0092] If the operator desires an explanation other than Pattern 1 or Pattern 2, they select the "Automatically generate explanation from category name" button B62 again and request the judgment device 10 to recreate the explanation for the category "Competitor comparison".
[0093] In other words, the determination device is notified by the service provider 30 that it will not adopt the explanation presented to the service provider 30 as the explanation for the additional category "Competitor Comparison". In this case, the prompt creation unit 139 changes the type and / or number of default categories that were given as example sentences in past assistance prompts and recreates the assistance prompt.
[0094] The prompt generation unit 139 changes the categories and their descriptions provided as EXAMPLE in the assist prompt P71 from the previous time. The prompt generation unit 139 may provide the same number of categories and their descriptions as the previous time as EXAMPLE, or it may provide a different number of categories and their descriptions. Alternatively, the prompt generation unit 139 may provide the categories and their descriptions provided last time as part of EXAMPLE. As a result, the output of the explanation by LLM41 will have a different nuance from the previous time. The explanation presentation unit 140 presents the explanation of the pattern generated again by LLM41 to the service provision device 30.
[0095] In this way, by repeatedly generating explanations of various patterns using LLM41 and verifying them in practice by the operator, the operator can select the optimal explanation for the additional categories they wish to add.
[0096] The categories to be added do not have to be "Competitor Comparison". Figures 17, 18, and 20 show examples of explanations generated using LLM41 shown in Figure 1. Figures 19 and 21 show examples of evaluation results for evaluation model 134 shown in Figure 6.
[0097] This section describes the case where the operator enters "Topics inappropriate for business settings" in the category name field C61 of the operator's management screen W61 (Figure 14) and selects the "Automatically generate explanation from category name" button B62. In this case, the judgment device 10 repeatedly generates explanations for various patterns of "Topics inappropriate for business settings" by LLM41 and verifies them in practice by the operator, thereby retaining the explanation D71 (Figure 17) selected by the operator as the explanation for "Topics inappropriate for business settings".
[0098] Next, we will explain the case where the operator enters "Personnel Information" in the category name field C61 of the operator's management screen W61 (Figure 14) and selects the "Automatically generate description from category name" button B62. In this case, the determination device 10 repeatedly generates descriptions for various patterns of "Personnel Information" by LLM41 and verifies them in practice by the operator, thereby storing the description D81 (Figure 18) selected by the operator as the description for "Personnel Information".
[0099] The judgment device 10 receives an evaluation request from the service provider 30 for text T81 (Figure 19) in the custom category "Personnel Information". The evaluation unit 133 calculates the "Unsafe Score" for the custom category "Personnel Information" based on the generation probability inferred by the evaluation model 134 for text T81 using explanation D81. As shown in list L81 of Figure 19, the evaluation model 134 evaluated the Unsafe Score for text T81 in the personnel information category (frame R81) as 0.99 (frame P81), so the evaluation unit 133 assigns the True label to personnel information.
[0100] The determination unit 135 determines the safety of text T81 for the custom category "Personnel Information" based on the "Unsafe Score" and the label. The determination unit 135 determines that the use of LLM31 should be restricted because the Unsafe Score for the personnel information category (frame R81 in Figure 19) is 0.99 (frame P81 in Figure 19). The notification unit 136 notifies the service provider device 30 of the restriction on the use of LLM31 for text T81 ("Unsafe" and "true"), the category "Personnel Information" which was evaluated as having the lowest safety, and the Unsafe Score "0.99". The notification unit 136 may also send list L81 in Figure 19 to the service provider device 30.
[0101] Next, we will explain the case where the operator enters "Excessive Interaction with AI" in the category name field C61 of the operator's management screen W61 (Figure 14) and selects the "Automatically generate explanation from category name" button B62. In this case, by repeatedly generating explanations for various patterns of "Excessive Interaction with AI" by LLM41 and verifying them in practice by the operator, the judgment device 10 retains the explanation D91 (Figure 20) selected by the operator as the explanation for "Excessive Interaction with AI".
[0102] The judgment device 10 receives an evaluation request from the service provider device 30 for text T91 (Figure 21) in the custom category "Excessive Interaction with AI". The evaluation unit 133 calculates the "Unsafe Score" for the custom category "Excessive Interaction with AI" based on the generation probability inferred by the evaluation model 134 for text T91 using explanation D91. As shown in list L91 in Figure 21, for text T91, the evaluation model 134 evaluated the Unsafe Score for the "Excessive Interaction with AI" category (frame R91) as 0.8304 (frame P91), so the evaluation unit 133 assigns the label True to "Excessive Interaction with AI".
[0103] The determination unit 135 determines the safety of text T91 for the custom category "Excessive Interaction with AI" based on the "Unsafe Score" and the label. The determination unit 135 determines that the use of LLM31 should be restricted because the Unsafe Score for the "Excessive Interaction with AI" category (frame R91 in Figure 19) is 0.8304 (frame P91 in Figure 21). The notification unit 136 notifies the service provider device 30 of the restriction on the use of LLM31 for text T91 ("Unsafe" and "true"), the category "Excessive Interaction with AI" which was evaluated as having the lowest safety, and the Unsafe Score "0.8304". The notification unit 136 may also send list L91 in Figure 21 to the service provider device 30.
[0104] Even in native languages, creating category descriptions and prompts for the LLM to instruct description generation were hurdles.
[0105] In response to this, the determination device 10 assists in creating category descriptions using the LLM 41, as described above. Therefore, the determination device 10 reduces the burden on operators in providing explanations and mitigates the barriers to introducing new categories.
[0106] Furthermore, the determination device 10 can accurately detect text corresponding to newly added categories, as shown in Figures 19 and 21, as well as categories for which explanations were generated using LLM41, thereby reducing the risk when using LLM31. Therefore, the determination device 10 allows for flexible application of the categories to be determined according to the use case of the service provision device 30, and can appropriately restrict unsafe input and output text of LLM31.
[0107] [Processing steps for adding categories] Next, the processing procedure for communication processing according to the embodiment will be described. Figure 22 is a sequence diagram showing the processing procedure for communication processing in the embodiment. Figure 22 shows the processing procedure for adding a custom category to the service provision device 30.
[0108] As shown in Figure 22, first, information exchange is initiated between the service provision device 30 and the determination device 10 (step S1).
[0109] When the determination device 10 receives a category addition request and a description of the content of the category to be added from the service provision device 30 via the operator's management screen (step S2), it verifies the category description with the service provision device 30 (step S3).
[0110] If the verification results indicate that the service provider 30 instructs the determination device 10 to retain the category, the determination device 10 issues a category ID to the additionally requested category (step S4) and notifies the service provider 30 of the registration of the custom category (step S5). From this point onward, the service provider 30 can request the determination device 10 to evaluate the text safety of the custom category.
[0111] Figure 23 is a sequence diagram showing the processing procedure for communication processing in the embodiment. In Figure 23, the assistance process for creating a category description is described as another processing procedure for adding a custom category to the service provider device 30.
[0112] As shown in Figure 23, first, information exchange is initiated between the service provision device 30 and the determination device 10 (step S11).
[0113] The determination device 10 receives a request to add an operator category and a request to generate a description for this category from the service provision device 30 via the operator's management screen (step S12). The determination device 10 creates an assist prompt instructing the generation of a description for the category to be added (step S13).
[0114] Then, the determination device 10 inputs the assist prompt created in step S13 to the LLM 41 (step S14). As a result, the LLM 41 generates a description of the category to be added according to the assist prompt (step S15), and outputs the generated description to the determination device 10 (step S16).
[0115] The determination device 10 presents the explanation generated by LLM41 to the service provider device 30 as an example explanation (step S17). The service provider device 30 displays the example explanation on the operator's management screen. The determination device 10 verifies the presented category explanation with the service provider device 30 (step S18). The service provider device 30 accepts the operator's selection of an example explanation (step S19) and transmits the selection result to the determination device 10 (step S20).
[0116] Based on the selection result received in step S20, the determination device 10 determines whether the explanation presented in step S18 was selected by the operator (step S21).
[0117] If the operator selects the provided explanation (Step S21: Yes), the determination device 10 issues a category ID to the additionally requested category (Step S22) and notifies the service provider device 30 of the registration of the custom category (Step S23). From this point onward, the service provider device 30 can request the determination device 10 to evaluate the text safety for the custom category.
[0118] Furthermore, if the operator does not select one of the presented explanations (step S21: No), the determination device 10 returns to step S13 to present a new example explanation, and performs the re-creation of the assist prompt and re-generation of the explanation for LLM41.
[0119] [Processing procedure for communication] Next, the communication processing in the embodiment will be described. Figures 24 and 25 are sequence diagrams showing the processing steps of the communication processing in the embodiment. Figures 24 and 25 describe the operational phase of text safety evaluation.
[0120] As shown in Figure 24, first, information exchange is initiated between the service provision device 30 and the determination device 10 (step S31).
[0121] When an end user accesses a service provided by the service provider 30 from a user terminal 20 (step S32), the service provider 30 selects a category set corresponding to the accessed service (step S33) and notifies the determination device 10 of the selected category set (step S34).
[0122] The determination device 10 switches the evaluation target category of the evaluation model 134 to the category set notified in step S34 (step S35). The determination device 10 may also automatically switch the evaluation target category of the evaluation model 134 to either the default category or a custom category associated with the service provider 30, according to the service accessed by the user terminal 20 from among the services provided by the service provider 30.
[0123] When a prompt (text) instructing text generation is entered (step S36), the service provider 30 requests the determination device 10 to determine the safety of the prompt (step S37).
[0124] The determination device 10 evaluates the safety of the prompt text using the evaluation model 134 (step S38). In this case, the evaluation model 134 evaluates the safety of the text for the category switched in step S35.
[0125] The determination device 10 determines the safety of the prompt based on the evaluation results in step S38 (step S39).
[0126] If the prompt is not safe (step S40: No), the determination device 10 instructs the service provider 30 to restrict the input of the prompt to be determined to the LLM 31 (step S41). The service provider 30 does not input the prompt to the LLM 31 and notifies the user terminal 20 of the non-response (step S42).
[0127] If the prompt is safe (step S40: Yes), the determination device 10 allows the service provider 30 to input the prompt to be determined into the LLM 31 (step S43). The service provider 30 inputs the prompt entered from the user terminal 20 into the LLM 31 (step S44).
[0128] As shown in Figure 25, the service provider 30 performs text generation using the LLM 31 (step S45). The service provider 30 requests the determination device 10 to perform a safety determination on the text generated by the LLM 31 (step S46).
[0129] The determination device 10 evaluates the safety of the generated text using the evaluation model 134 (step S47). In this case, the evaluation model 134 evaluates the safety of the text for each category in step S35.
[0130] The determination device 10 determines the safety of the prompt based on the evaluation result in step S47 (step S48).
[0131] If the generated text is not safe (step S49: No), the determination device 10 instructs the service provider 30 to restrict the output of the generated text by LLM31 (step S50). The service provider 30 does not output the generated text by LLM31 and notifies the user terminal 20 of the non-response (step S51).
[0132] If the generated text is safe (step S49: Yes), the determination device 10 permits the service provider 30 to output the generated text by LLM31 (step S52). The service provider 30 outputs the generated text by LLM31 as the response text from the user terminal 20 (step S53).
[0133] [Effects of the embodiment] In this embodiment, the determination device 10, which exchanges information with the service provision device 30, evaluates the safety of the input and output text to the LLM 31 in the service provision device 30, and determines whether or not input and output to the LLM 31 is permitted based on the evaluation result. Therefore, the determination device 10 functions as a guardrail that restricts only the input and output text to the LLM 31 that is not safe, thereby reducing the risks when using the LLM 31.
[0134] Furthermore, when the determination device 10 according to the embodiment receives a request to add a category from the service provision device 30, it adds the requested category to the categories to be evaluated in the evaluation model 134. Therefore, the operator of the service provision device 30 can flexibly customize the categories to be targeted for risk elimination simply by adding the categories to be evaluated in natural language.
[0135] The determination device 10 then assists in creating category descriptions using the LLM 41. Therefore, the determination device 10 reduces the burden on operators in providing explanations and mitigates the barriers to introducing new categories.
[0136] Furthermore, the determination device 10 switches the category to be evaluated in the evaluation model 134 between the default category and a custom category associated with the service provider 30, depending on the category set selection notification from the service provider 30 or the service accessed by the user terminal 20. This allows the determination device 10 to determine the security of the input and output text to the LLM 31 of the service provider 30 in the most appropriate category.
[0137] [Accuracy of the evaluation model] The evaluation model 134 used by the judgment device 10 is a machine learning model trained to evaluate whether something is safe or not using training data that includes harmful text in Japanese. Figure 26 shows the accuracy evaluation of the evaluation model 134 of the judgment device 10 and other models (OpenAI Moderation API, Meta-Llama-Guard-2-8B). The evaluation metric is the F1 score.
[0138] Thus, compared to other models, evaluation model 134 exhibits superior performance in determining harmful Japanese text. Therefore, the determination device 10 can accurately evaluate the safety of text and thus accurately restrict unsafe input and output text of LLM31.
[0139] As described above, according to the embodiment, by using independent guardrails, it is possible to limit the leakage of personal information and the output of incorrect specialized knowledge when using LLM31, thereby enhancing security when using LLM31.
[0140] [Application Example 1] Figures 27 to 30 illustrate examples of applications of the embodiment. As shown in Figure 27, the determination device 10 performs API integration with a service provider company that provides AI chat and inquiry forms using LLM31, for example. This reduces the risk of personal information or incorrect expertise being leaked from LLM31 as chat content or responses to inquiries from end-users, such as user U1. Furthermore, the determination device 10 can determine the security of input and output text from the service provider 30 to LLM31 in a more appropriate category by adding custom categories according to the AI chat and inquiry forms provided by LLM31.
[0141] Furthermore, as shown in Figure 28, a safety check by the detection device 10 may be placed between the service provider's corporate web page and the operator of the online consumer inquiry desk to detect and filter out so-called customer harassment and threats. For example, the inquiries from users U11 and U12 are determined to be safe by the detection device 10 and are therefore connected to operator P10 (arrows Y11, Y12). In contrast, the inquiry from user U13 contains abusive language and is determined to be unsafe by the detection device 10, so communication with operator P10 is blocked. In addition, the detection device 10 can add custom categories according to the service provider's corporate web page, allowing it to determine the safety of input and output text to the LLM 31 of the service provider 30 in a more appropriate category.
[0142] Furthermore, the determination device 10 may determine the security of the chatbot's input and output by the service provider, as shown in Figures 29 and 30. The determination device 10 adds custom categories according to the chatbot service provided by the service provider.
[0143] For example, as shown in Figure 29, the service provider requests the determination device 10 to determine the security of inquiry Q21 from user U21 (arrow Y21). Because the determination device 10 determines that inquiry Q21 is Unsafe (category: confidential information) (arrow Y22), the service provider does not input inquiry Q21 into the chatbot service LLM. The service provider then returns a standard message A22 to user U21's terminal stating, "We are unable to answer your question" (Figure 29 (1)). In this way, the determination device 10 prevents the leakage of confidential information by prohibiting the input of inquiry Q21, which concerns confidential information, into the LLM.
[0144] Furthermore, as shown in Figure 30, when the service provider receives an inquiry Q31 from user U31, it requests the determination device 10 to determine the safety of inquiry Q31 (arrow Y31). In this case, if the determination device 10 determines that inquiry Q31 is safe (arrow Y32), the service provider inputs inquiry Q31 into the chatbot service LLM.
[0145] Next, the service provider generates output candidate A32 in LLM31 and requests the judgment device 10 to perform a safety judgment on output candidate A32 (arrow Y33). If the judgment device 10 determines that output candidate A32 is Unsafe (category: confidential information), it cancels the display of output candidate A32 on the user U31's terminal (Figure 30 (1)). This is because output candidate A32 unintentionally contained personal information (the name and telephone number of the product development manager). In this way, the judgment device 10 prevents the leakage of personal information by prohibiting the input of inquiry Q31 regarding confidential information into LLM.
[0146] Figure 31 shows an example of output from the service provider to the user. Prompts and generated text intended to threaten or leak confidential information are judged as Unsafe by the detection device 10, and as exemplified in Figure 31, screen G1, which includes a standard message "Inappropriate comments will be automatically discarded," is displayed on the user terminal 20.
[0147] Figure 32 shows an example of a service screen for a conventional help desk bot. Figure 33 shows an example of a service screen for a help desk bot when the determination device 10 according to the embodiment is applied for safety determination. For example, the output when "names and addresses of 3 new employees of Company A" is requested as a query is illustrated.
[0148] As shown in screen G21 of Figure 32, if the embodiment is not applied, the names and addresses of the three new employees will be returned to the user, resulting in the leakage of personal information. In contrast, as shown in Figure 33, when the embodiment is applied, queries intended to leak personal information are determined as Unsafe by the determination device 10 and are not entered into LLM 31. As shown in Figure 33, screen G22, which includes a standard message stating "Personal information cannot be provided," is displayed on the user terminal 20.
[0149] Figures 34 and 35 show examples of LLM responses returned when the embodiment is not applied. As shown in screen G31 of Figure 32 and screen G32 of Figure 35, methods for cultivating magic mushrooms, which are illegal drugs in Japan, and methods for creating fraudulent emails are returned. In contrast, when the judgment device 10 is used as a guardrail, illegal content is not returned to the user.
[0150] Furthermore, the judgment device 10 can accurately detect text corresponding to newly added categories, as shown in Figures 12, 19, and 21, thereby reducing the risk when using LLM 31.
[0151] [System configuration of the embodiment] The determination device 10 is a functional concept and does not necessarily need to be physically configured as shown in the illustration. In other words, the specific forms of distribution and integration of the functions of the determination device 10 are not limited to those shown in the illustration, and all or part of it can be configured by functionally or physically distributing or integrating it in any unit according to various loads and usage conditions.
[0152] Furthermore, each process performed in the determination device 10 may be implemented, in whole or in part, by a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and a program that is analyzed and executed by the CPU and GPU. Alternatively, each process performed in the determination device 10 may be implemented as hardware using wired logic.
[0153] Furthermore, among the processes described in the embodiments, all or part of the processes described as being performed automatically can be performed manually. Alternatively, all or part of the processes described as being performed manually can be performed automatically by known methods. In addition, the processing procedures, control procedures, specific names, and information including various data and parameters described above and illustrated may be changed as appropriate unless otherwise specified.
[0154] [program] Figure 36 shows an example of a computer in which the determination device 10 is realized when a program is executed. The computer 1000 has, for example, memory 1010 and CPU 1020. The computer 1000 also has a hard disk drive interface 1030, a disk drive interface 1040, a serial port interface 1050, a video adapter 1060, and a network interface 1070. These components are connected by a bus 1080.
[0155] Memory 1010 includes ROM 1011 and RAM 1012. ROM 1011 stores, for example, a boot program such as the BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to the hard disk drive 1090. The disk drive interface 1040 is connected to the disk drive 1100. For example, a removable storage medium such as a magnetic disk or optical disk is inserted into the disk drive 1100. The serial port interface 1050 is connected to, for example, the mouse 1110 and the keyboard 1120. The video adapter 1060 is connected to, for example, the display 1130.
[0156] The hard disk drive 1090 stores, for example, an OS (Operating System) 1091, an application program 1092, a program module 1093, and program data 1094. That is, the program that defines each process of the determination device 10 is implemented as a program module 1093 in which code executable by the computer 1000 is written. The program module 1093 is stored, for example, in the hard disk drive 1090. For example, a program module 1093 for performing the same processes as the functional configuration of the determination device 10 is stored in the hard disk drive 1090. Note that the hard disk drive 1090 may be replaced by an SSD (Solid State Drive).
[0157] Furthermore, the configuration data used in the processing of the above-described embodiment is stored as program data 1094 in, for example, memory 1010 or hard disk drive 1090. The CPU 1020 then reads the program module 1093 and program data 1094 stored in memory 1010 or hard disk drive 1090 into RAM 1012 as needed and executes them.
[0158] Furthermore, the program module 1093 and program data 1094 are not limited to being stored in the hard disk drive 1090; for example, they may be stored in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 may be stored in another computer connected via a network (LAN (Local Area Network), WAN (Wide Area Network), etc.). The program module 1093 and program data 1094 may then be read by the CPU 1020 from the other computer via a network interface 1070.
[0159] Although embodiments applying the invention made by the present inventors have been described above, the present invention is not limited by the descriptions and drawings that constitute part of the disclosure of the present invention in these embodiments. That is, all other embodiments, examples, and operational techniques made by those skilled in the art based on these embodiments are included in the scope of the present invention. [Explanation of symbols]
[0160] 10 Judgment device 11 Communications Department 12 Storage section 13 Control Unit 20 User Terminals 30 Service provision device 31,41 LLM 40 AI Generator Server 121 User Information 122 Partner Information 123 Judgment History 131 Liaison Department 132 Reception Department 133 Evaluation Department 134 Evaluation Models 135 Judgment section 136 Notification Department 137 Additional section 138 Category Addition Section 139 Prompt Creation Section 140 Explanation Presentation Section 141 Switching section
Claims
1. A linking unit that exchanges information with a processing device having a first natural language processing model that generates text, An evaluation unit that evaluates the safety of a first prompt that instructs the first natural language processing model to generate text, and / or the safety of the text generated by the first natural language processing model, using a machine learning model that evaluates the safety of text for each of several safety categories, A determination unit determines whether to input the first prompt to the first natural language processing model and / or output the text generated by the first natural language processing model to the source requesting the text, based on the safety of each category when the machine learning model evaluates the first prompt and / or the text generated by the first natural language processing model. A notification unit that notifies the processing unit of the determination result made by the determination unit, Upon receiving a request to add a category from the processing device, an addition unit adds the requested category as a custom category associated with the processing device to the category of the machine learning model to be evaluated. A determination device characterized by having the following features.
2. When the processing device receives a request to add a category and a request to generate a description of the category, the creation unit creates a second prompt that instructs a second natural language processing model, which generates text, to generate a description of the added category. A presentation unit inputs the second prompt to the second natural language processing model and presents the explanation generated by the second natural language processing model to the processing unit, It has, The determination device according to claim 1, characterized in that, when the additional unit is notified by the processing device that it will adopt the explanation presented by the presentation unit as the explanation for the additionally requested category, it adds the additionally requested category as a custom category associated with the processing device, separate from the default category to be evaluated.
3. If the processing unit is notified by the processing unit that it will not adopt the explanation presented by the presentation unit as the explanation for the additionally requested category, the creation unit recreates the second prompt. The determination device according to claim 2, characterized in that the presentation unit inputs the second prompt recreated by the creation unit to the second natural language processing model, and presents the explanation generated by the second natural language processing model to the processing unit again.
4. The default categories set as described above include: violent crime, nonviolent crime, verbal abuse, sexual offenses, attention, indiscriminate weapons, hatred, self-harm, sexual content, child exploitation, privacy, professional advice, and / or intellectual property. The creation unit provides the second prompt with one or more of the default categories and corresponding descriptions as example sentences. The determination device according to claim 3, characterized in that, if the creation unit is notified by the processing unit that it will not adopt the explanation presented by the presentation unit as the explanation for the additionally requested category, it changes the type and / or number of the default categories given as example sentences in the past second prompt and recreates the second prompt.
5. The determination device according to claim 1, characterized in that, if the determination unit determines that the use of the first natural language processing model should be restricted, it notifies the processing device via the notification unit of the restriction on the use of the first natural language processing model and the category that was evaluated as having the lowest safety.
6. The determination device according to claim 2, characterized in that it has a switching unit that switches the category to be evaluated for the machine learning model to the default category or a custom category associated with the processing device, in accordance with the services provided by the processing device.
7. The determination device according to claim 6, wherein the processing device provides a chatbot service, provides an inquiry web page, and / or controls the connection to a customer center using the first natural language processing model as the service.
8. A determination method performed by a determination device, A process of exchanging information with a processing device having a first natural language processing model that generates text, A step of evaluating the safety of a first prompt that instructs the first natural language processing model to generate text, and / or the safety of the text generated by the first natural language processing model, using a machine learning model that evaluates the safety of text for each of several safety categories, The process includes determining whether to input the first prompt to the first natural language processing model and / or whether to output the text generated by the first natural language processing model to the source requesting the text, based on the safety of each category when the machine learning model evaluates the first prompt and / or the text generated by the first natural language processing model, A step of notifying the processing device of the determination result in the determination step, When the processing device receives a request to add a category, it adds the requested category as a custom category associated with the processing device to the category of the machine learning model to be evaluated. A determination method characterized by including
9. A step of exchanging information with a processing device having a first natural language processing model that generates text, A step of evaluating the safety of a first prompt that instructs the first natural language processing model to generate text, and / or the safety of the text generated by the first natural language processing model, using a machine learning model that evaluates the safety of text for each of several safety categories. The steps include: determining whether to input the first prompt to the first natural language processing model and / or whether to output the text generated by the first natural language processing model to the source requesting the text, based on the safety of each category when the machine learning model evaluates the first prompt and / or the text generated by the first natural language processing model; A step of notifying the processing device of the determination result in the determination step, When the processing device receives a request to add a category, it adds the requested category as a custom category associated with the processing device to the category of the machine learning model to be evaluated. A judgment program that causes a computer to execute a certain action.
Citation Information
Patent Citations
Text generation device, text generation method, and program
JP7133689B1
Learning device, learning method, and learning program
JP7208314B1