Information processing device, detection method, and detection program
Patent Information
- Application Number
- PCT/JP2025/005962
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2026-08-27
Smart Images

Figure JP2025005962_27082026_PF_FP_ABST
Abstract
Description
Information Processing Apparatus, Detection Method, and Detection Program
[0001] The present invention relates to an information processing apparatus, a detection method, and a detection program.
[0002] As one of the models that perform output according to input, a generation model that generates content corresponding to various modalities such as text, images, videos, and voices, so-called generative AI (Artificial Intelligence), is known.
[0003] As attacks on such generative models by malicious inputs, there are those called adversarial samples and prompt injection attacks. These attacks force actions that are suppressed, actions different from the intended purpose, or discriminatory or illegal actions.
[0004] For example, as a defense method against the above attacks, a prior art has been proposed in which a coefficient that minimizes the error obtained by adding the error when normal data is input to the generative model and the error when an adversarial sample is input is trained.
[0005] Emily Dinan et al., “Build it Break it Fix it for Dialogue Safety: Robustness from Adversarial Human Attack” (EMNLP-IJCNLP’19)
[0006] However, the above prior art only strengthens the defense performance of the generative model itself. Therefore, there is room for improvement in that there is no technique for detecting an input that causes an abnormality from outside the generative model.
[0007] Therefore, an object of the present invention is to provide an information processing apparatus, a detection method, and a detection program that can detect an input that causes an abnormality from outside the generative model.
[0008] To solve the above-mentioned problems and achieve the objective, the information processing device of the present invention includes an acquisition unit that acquires a log of output data obtained by performing one or more input / output control operations, in which the first output data of a first generation model among a plurality of generation models is input to a second generation model and the second generation model outputs second output data; and a detection unit that detects an input causing an abnormality based on the log of output data.
[0009] According to the present invention, it is possible to detect inputs that cause anomalies from outside the generative model.
[0010] Figure 1 is a block diagram showing an example of the functional configuration of an information processing device. Figure 2 is a schematic diagram showing an example of the construction of an AI-linked system. Figure 3 is a schematic diagram showing an example of input / output control for a generation model. Figure 4 is a diagram showing an example of an operational information DB. Figure 5 is a block diagram showing the functional details of the AI behavior detection unit. Figure 6 is a diagram (1) showing an example of keyword search results. Figure 7 is a diagram (2) showing an example of keyword search results. Figure 8 is a diagram (1) showing an example of the calculation results of sentiment analysis values. Figure 9 is a diagram (2) showing an example of the calculation results of sentiment analysis values. Figure 10 is a diagram (3) showing an example of the calculation results of sentiment analysis values. Figure 11 is a diagram (4) showing an example of the calculation results of sentiment analysis values. Figure 12 is a diagram (1) showing an example of similarity analysis results. Figure 13 is a diagram (2) showing an example of similarity analysis results. Figure 14 is a diagram (3) showing an example of similarity analysis results. Figure 15 is a diagram (4) showing an example of similarity analysis results. Figure 16 is Figure (1), showing an example of the calculation results of the distribution of sentiment analysis values. Figure 17 is Figure (2), showing an example of the calculation results of the distribution of sentiment analysis values. Figure 18 is Figure (3), showing an example of the calculation results of the distribution of sentiment analysis values. Figure 19 is a block diagram showing the functional details of the response unit. Figure 20 is a diagram showing an example of the display of detection results. Figure 21 is a schematic diagram showing an example of how to specify a response. Figure 22 is a flowchart showing the procedure for AI collaboration processing. Figure 23 is a flowchart showing the procedure for detection processing. Figure 24 is a flowchart showing the procedure for response processing. Figure 25 is a block diagram showing an application example of the functional details of the AI behavior detection unit. Figure 26 is a diagram showing an example of an attribute information DB. Figure 27 is Figure (1), showing an example of an AI management screen. Figure 28 is Figure (2), showing an example of an AI management screen. Figure 29 is a diagram showing an example of a list display. Figure 30 is a diagram showing a modified version of the list display. Figure 31 is a diagram showing an application example of AI collaboration. Figure 32 is a schematic diagram explaining an example of hybrid cloud implementation. Figure 33 shows an example of a hardware configuration.
[0011] The following description will explain the embodiments for implementing the information processing device, detection method, and detection program related to this disclosure (hereinafter referred to as "Embodiments") with reference to the attached drawings. It should be noted that these embodiments represent only one example or one aspect, and the following description does not limit the structure, operation, function, properties, characteristics, methods, and applications related to this disclosure.
[0012] <Overall Configuration> Figure 1 is a block diagram showing an example of the functional configuration of the information processing device 10. Figure 1 shows the information processing device 10 which provides a detection function for detecting inputs that cause anomalies from outside the generative model, so-called generative AI, and a countermeasure function for dealing with such attacks from outside the generative model.
[0013] In this context, "abnormality" may refer, in one aspect, to a state that deviates from arbitrary standards or policies. Such standards and policies may be defined, for example, by rules and mechanisms established from the perspective of operating AI in accordance with principles defined within the organization (governance, compliance, etc.), so-called AI guardrails. In addition, it may be defined by policies that prevent participation in specific discussions touching on sensitive information or information requiring special consideration as defined by laws and regulations such as the Personal Information Protection Act, such as a privacy policy.
[0014] Examples of inputs that can cause such abnormalities include attacks, pranks, and interference. As these examples show, it is not necessarily limited to intentional actions, behaviors, activities, and actions in general.
[0015] Below, as one use case for generative models, we will give an example of an architecture in which multiple generative models are deployed in a distributed manner and linked to each other (hereinafter referred to as the "AI collaboration system").
[0016] The three functions may be provided as an AI-integrated service through packaged software that combines the "AI integration function" that enables the construction of such an AI-integrated system with the "detection function" and "response function" mentioned above.
[0017] In one embodiment, the information processing device 10 may be implemented by a server device. For example, the information processing device 10 can provide the above-mentioned AI-linked service as a cloud service by running a PaaS (Platform as a Service) type middleware or a SaaS (Software as a Service) type application.
[0018] As shown in Figure 1, the information processing device 10 can be connected to a user terminal 30 or an external server 50 via a network NW to enable communication. For example, the network NW may be any type of communication network, such as the Internet or a LAN (Local Area Network), whether wired or wireless.
[0019] The user terminal 30 is a terminal device used by a user who receives the AI-linked service described above. For example, the user terminal 30 may be implemented using any computer, including a personal computer, smartphone, tablet, or wearable device.
[0020] The external server 50 is a server device that provides a cloud service that provides machine learning tools, a so-called MLaaS (Machine Learning as a Service). For example, as one of the machine learning tools mentioned above, the external server 50 can provide some of the generative models among the multiple generative models that make up the AI collaboration system mentioned above. The external server 50, as an example of a resource that provides generative models outside the information processing device 10, may be operated by the business that provides the AI collaboration service mentioned above, or it may be operated by another business.
[0021] While the above example of an AI integration service being provided as a cloud service is given here, it is not limited to this. For example, the above AI integration service may be provided on-premises. Furthermore, the above AI integration service may be implemented as a hybrid cloud, combining both cloud and on-premises solutions.
[0022] Furthermore, while the above example of an AI collaboration service being implemented as a client-server system was given, it is not limited to this. For example, the above AI collaboration service may be provided as a standalone service by having an application running on the user terminal 30 execute processing on the user terminal 30 in response to the above AI collaboration service.
[0023] Furthermore, while we have provided an example here where the above-mentioned AI integration function, detection function, and response function are packaged together, each of the three functions may also be implemented as an individual module.
[0024] <Generative Models> A generative model is a model that receives input data and generates output data that follows that input data. A Large Language Model (LLM) is an example of a generative model that receives text written in natural language as input and is trained to generate text that follows the input text. Other types of generative models include text-to-text models that receive text input and output text, and text-to-image models that receive text input and output images. The input to a generative model is not limited to strings; it may also accept image data or audio data. The input data for an LLM is also called a "prompt." The input to a generative model is not limited to text data written in natural language; it may be any type of data, such as sensor data sensed by a sensor. Examples of sensor data include spatial information, positional information, acceleration, velocity, temperature, and humidity. Here, LLM is given as just one example of a generative model, but as mentioned above, generative models can be any model that generates output data according to input data, and are not limited to specific machine learning models such as LLM. For example, a generative model may be implemented using machine learning models such as Generative Adversarial Networks (GANs) or Transformers. This example does not limit the implementation of generative models to generative AI, but does not prevent them from being implemented by any model that outputs output data corresponding to input data. Furthermore, a generative model is not limited to models computed by computers, such as the machine learning models mentioned above, but may include human input and output. For example, input and output data may be edited through user interface input.
[0025] The following is merely an example of how prompt engineering can be used to set up roles, from the perspective of enabling in-depth discussions by using agents that simulate domain experts—specialists in a particular field. However, role setting may be optional. Furthermore, the following is an example of how prompt engineering can be used to set up personas, from the perspective of enabling discussions with individuality. However, persona setting may also be optional. In other words, a generative model only needs to be a model that generates output data according to the input data, and it is not necessarily required that roles or personas be set up.
[0026] Furthermore, the following example illustrates a scenario where one generative model is executed per agent, but the generative model and agent do not necessarily have to correspond one-to-one. For example, a single generative model can generate output data while switching between roles and personas assigned to each of multiple agents.
[0027] A role is an instruction given to a generative model as input data, representing setting information that describes the persona (character, personality, personality, etc.), occupation, age, background, and other attributes of the character that the generative model will roleplay.
[0028] Beyond these examples, roles may be assigned any attributes such as nationality, region, and gender. Furthermore, roles can be assigned levels such as skill and proficiency, classifications such as new employee or skilled technician, and even more finely subdivided, such as lists of things the agent is aware of or unaware of. Moreover, the person assigned to a role is not necessarily limited to a natural person; it may also be a legal entity, such as "XX Company" or "XX Organization." Furthermore, roles may be assigned any character, not just people. Examples of "characters" here include "inorganic objects" such as computers, tools, and equipment; "organic objects" such as plants and animals; "proper nouns" such as people's names and parks; and "concepts" such as nature and religion. Note that the above persona settings are merely options, and other elements can be set as examples of personas.
[0029] The following is merely an example of how prompt engineering can be used to set up roles and personas, but this is only one example, and role and persona settings can also be achieved through fine-tuning. For example, it is possible to generate generative models with personas specific to an industry or business type. As just one example, LLM can be used to generate a fine-tuned financial model using a knowledge database containing knowledge of the financial industry, or a fine-tuned legal model using a knowledge database containing knowledge of law.
[0030] Thus, since generative models generate output data following input data, input data including roles can be input into a generative model to obtain output data that would be generated by the person represented by the role.
[0031] The knowledge of the role-playing character can be provided through methods such as RAG (Retrieval-Augmented Generation) and fine-tuning (full fine-tuning, adapter tuning). When using RAG, the data that can be referenced may differ for each role. For example, when generating output data in a generative model according to the role of a lawyer, legal knowledge data may be referenced by RAG. On the other hand, when generating output data in a generative model according to the role of a meteorologist, meteorological knowledge data may be referenced by RAG. When using adapter tuning, role-specific adapter models are prepared in advance, and the output data obtained by inputting input data into the adapter corresponding to the specified role is input into the LLM.
[0032] Furthermore, agents can be assigned roles in discussions through prompt engineering. For example, various roles can be assigned, such as an AI that summarizes discussions, an AI that facilitates discussions, an AI that breaks down agenda items into points, an AI that reconfigures AI groups according to the discussion situation, an AI that focuses solely on providing advice, and an AI that considers steps and action plans. By assigning such discussion roles, it is possible to calculate the degree to which each AI is fulfilling its role from its statements and detect AIs that do not meet the predetermined conditions as abnormal AIs that are not fulfilling their roles.
[0033] <AI Collaboration System> For example, consider a use case where the above AI collaboration system enables multi-agent-based role-playing discussions.
[0034] The following is merely an example of setting an agenda item specified by the user, but setting an agenda item is optional, and multi-agent-based discussions can be conducted even without setting an agenda item.
[0035] Figure 2 is a schematic diagram showing an example of the construction of an AI collaboration system. As shown in Figure 2, the AI collaboration function accepts a designation of an arbitrary topic from the user terminal 30, for example, "Regarding the increase in the consumption tax rate" (S1). Then, the AI collaboration function selects the roles of agents to participate in the discussion based on the topic "Regarding the increase in the consumption tax rate" that was accepted in step S1 (S2).
[0036] Such selection of participating agents can be achieved, as an example, by selecting a predetermined number of roles in order of their relevance to the agenda. For example, in the example shown in Figure 2, from a preset set of roles such as professor, analyst, engineer, legislator, and shop assistant, the top three roles, "analyst," "professor," and "legislator," are selected in order of their relevance to the agenda item "Regarding the increase in the consumption tax rate."
[0037] Steps S1 and S2 establish an AI collaborative system that includes agents simulating the roles of "analyst," "professor," and "parliamentarian."
[0038] Hereafter, the person simulated by the i-th role, Ri, may be referred to as Agent Ai.
[0039] In this case, each agent may be simulated by a single generative model, or by multiple different generative models. As merely one example, an AI collaborative system may be constructed by simulating multiple agents using a Mixture of Experts (MoE) model, which is formed by a mixture of subnetworks called experts. In the case of an MoE model, the gating network activates subnetworks related to the task of the input data from subnetworks trained on a subset of training datasets specialized for a particular domain. For example, the gating network determines which expert should perform the task with what weight by adjusting the magnitude of the weights assigned to the outputs of the subnetworks. Under such an MoE model, the gating network changes the pattern of weights assigned to the outputs of each subnetwork according to the role or persona assigned to the agent. In this way, by activating subnetworks with different weighting patterns, the MoE model can simulate each role or persona of multiple agents. Furthermore, the gating network can simulate each role or persona of multiple agents in the MoE model by assigning each of the multiple roles to the subnetwork corresponding to that role in the MoE model. Furthermore, each subnetwork included in the MoE model may correspond to a single generative model. Even in these cases, a scenario may arise where a specific role assigned to one subnetwork is aggressive and attacks subnetworks assigned to other roles within the MoE model. For example, the input information used when assigning a specific role to a subnetwork, such as the role's configuration information, may contain data that causes anomalies in the generative model. Detecting and addressing such anomalies can also be included within the scope of the detection and handling functions described above. As just one example, it is possible to exclude roles corresponding to subnetworks that generate output data that causes anomalies in other subnetworks, or to delete statements made by such roles.Furthermore, some or all of the generative models that simulate each agent do not necessarily have to run inside the information processing device 10. For example, each agent may be simulated by a generative model running inside the information processing device 10, or it may be simulated by a generative model provided by an external server 50 via the MLaaS API (Application Programming Interface).
[0040] Under this AI collaborative system, a multi-agent-based role-playing discussion, or in other words, input / output control for the generative models simulating each agent, is initiated.
[0041] Figure 3 is a schematic diagram illustrating an example of input / output control for a generative model. Figure 3 shows an example of a round-robin role-playing scenario in which two agents, Agent A1 and Agent A2, discuss the topic "Regarding the increase in the consumption tax rate." For the sake of clarity, prompts from the input / output log are omitted in Figure 3.
[0042] As shown in Figure 3, in the first round r1, the AI collaboration function inputs prompt 20A, in which the role of agent A1 is set, into the generation model. The generation model, upon receiving prompt 20A, simulates the role of agent A1 and outputs statement T1. As a result, the AI collaboration function saves agent A1's statement T1, "I think we should raise it because of XX," as the input / output log for round r1.
[0043] In the second round, r2, the AI collaboration function inputs prompt 20B into the generative model, which contains the role setting for agent A2 and the statement T1 from round r1. The generative model, upon receiving prompt 20B, simulates agent A2's role and outputs statement T2. As a result, the AI collaboration function saves agent A2's statement T2, "Because of △△, it should definitely be lowered," as the input / output log for round r2.
[0044] In the third round r3, the AI cooperation function inputs a prompt 20C, which embeds the utterance T1 in round r1 and the utterance T2 in round r2, along with the role setting of agent A1, into the generation model. The generation model with this prompt 20C input outputs an utterance T3 by simulating the role of agent A1. As a result, the AI cooperation function saves, as the input / output log for round r3, the utterance T3 of agent A1, "Since it is △△, we should firmly withdraw",
[0045] In this way, the above AI cooperation function repeats the input / output control of inputting input data including the output data of the generation model selected in the past into one generation model selected from among a plurality of generation models and causing the one generation model to output output data. As a result, as a result of the discussion on the issue "Regarding the increase in consumption tax rate", it is possible to obtain an output log in which the utterances of each round of the discussion are associated, for example, the time-series data of utterance T1, utterance T2, and utterance T3.
[0046] According to the above-described role-playing of the discussion, since the issue is discussed from a plurality of viewpoints, it is possible to suppress the discussion from being biased. Furthermore, since each agent simulates a role and discusses the issue, it is possible to realize the role-playing of the discussion by a domain expert who is an expert in a specific field.
[0047] <Problem's One Aspect> As described in the above Background Art section, the above prior art only enhances the defensive performance of the generation model itself. Therefore, from one aspect, there is room for improvement in that there is no technique for detecting an input that causes an abnormality from the outside of the generation model. From another aspect, there is room for improvement in that there is no technique for dealing with an input that causes an abnormality from the outside of the generation model.
[0048] For example, strengthening the generation model requires costs required for development, operation, and maintenance, such as time and expenses, as well as costs such as the calculations and resources consumed for strengthening, and there is a limit to the degree of strengthening.
[0049] In particular, in the case of the AI collaboration system described above, enhancements are required for each generation model, which increases costs. Furthermore, in the AI collaboration system described above, the environments in which each generation model operates may differ. For example, not all generation models necessarily operate on an on-device, and some may be provided via the MLaaS API. Especially for generation models operating externally, securing the authority to perform enhancements is difficult, making it difficult to implement enhancements even if the above costs are ignored.
[0050] Therefore, in this embodiment, a detection function is provided for detecting inputs that cause abnormalities from outside the generating AI, and a response function is provided for dealing with inputs that cause abnormalities from outside the generating AI.
[0051] <Configuration of Information Processing Device 10> Next, the functional configuration of the information processing device 10 that provides the above-mentioned detection function and the above-mentioned response function will be described. For example, Figure 1 schematically shows the blocks related to the AI collaboration service that the information processing device 10 has.
[0052] As shown in Figure 1, the information processing device 10 includes a communication control unit 11, a storage unit 13, and a control unit 15. Although Figure 1 shows the functional units related to the AI collaboration service described above, the information processing device 10 may be equipped with functional units related to only one of the AI collaboration function, detection function, and response function.
[0053] The communication control unit 11 is a functional unit that controls communication with other devices such as the user terminal 30 and the external server 50. In one embodiment, the communication control unit 11 can be implemented by a network interface card such as a LAN card. In one aspect, the communication control unit 11 receives requests for role-playing discussions from the user terminal 30, or outputs a user interface to the user terminal 30 that accepts the results of role-playing discussions, attack detection results, or responses to attacks.
[0054] The storage unit 13 is a functional unit that stores various types of data. In one embodiment, the storage unit 13 may be implemented by internal, external, or auxiliary storage of the information processing device 10. For example, the storage unit 13 stores an operational information DB (Database) 13A, an input / output log DB 13B, and a detection history DB 13C. The descriptions of each data in the operational information DB 13A, input / output log DB 13B, and detection history DB 13C will be described later in conjunction with the scenes in which each data is referenced or registered.
[0055] The control unit 15 is a functional unit that performs overall control of the information processing device 10. For example, the control unit 15 can be implemented by a hardware processor. As shown in Figure 1, the control unit 15 includes an AI cooperation unit 16, an information presentation unit 17, an AI behavior detection unit 18, and a response unit 19. The control unit 15 may also be implemented by hardwired logic or the like.
[0056] The AI collaboration unit 16 is a processing unit that provides the above-mentioned AI collaboration function. In one embodiment, the AI collaboration unit 16 can perform the construction of the above-mentioned AI collaboration system, and furthermore, role-playing of discussions in the above-mentioned AI collaboration system.
[0057] More specifically, the AI collaboration unit 16 receives a specified topic and preset roles from the user terminal 30, and obtains a document representing the agent's perspective on the topic (perspective document) by inputting prompts to the LLM to generate the agent's perspective on the task simulated by the preset roles. Examples of perspectives include points of contention on the topic, main points of the argument, viewpoint on the topic, opinion on the topic, points of disinterest on the topic, and concerns on the topic. Such perspective documents may be generated by inputting prompts for each role included in the preset role set. Note that while examples of perspectives are given here, they do not have to be perspectives on discussions, and other elements such as values and model accuracy can be used to construct the AI collaboration system.
[0058] The following is an example of a prompt to generate an agent's perspective on an agenda item. In the following example prompt, {} indicates a variable. "# Instructions You are {Agent Name}. Your occupation is {Job Title} and your persona is {Persona}. Using {Knowledge Data} as a reference, please provide one perspective you would consider when discussing the following agenda item. # Agenda {Agenda} # Knowledge Data {Knowledge Data Search Results}"
[0059] The "knowledge data search results" are the knowledge data obtained by RAG. For example, the AI collaboration unit 16 generates a vector relating to the combination of roles and topics, and searches for sentences with a high similarity to this vector from the knowledge data relating to the attributes associated with the roles. This knowledge data may be stored in a storage device inside the information processing device 10 or in an external storage device. For example, the knowledge data may include attributes, sentence vectors, and sentences. The attributes of the knowledge data represent attributes relating to the content of the knowledge data, such as professor data, analyst data, engineer data, legislator data, ..., shop assistant data, according to the preset role set shown in Figure 2. The sentence vector may be a vector in which a sentence has been digitized. The sentence may be text written in natural language, representing the original sentence corresponding to the knowledge data. The AI collaboration unit 16 can then implement RAG by including the searched sentence in the prompt. Note that if the highest similarity is lower than a predetermined threshold, the "knowledge data search results" may be left blank.
[0060] Subsequently, the AI collaboration unit 16 calculates a relevance score representing the relationship between the preset agent-specific perspective texts and roles, and selects a predetermined number of roles with high relevance scores as roles to be used for simulating the discussion. The number of roles selected in this way may be a fixed value set by the system or user. In addition, roles with a relevance score above a threshold may be selected. In this case, the number of roles may vary depending on the number of roles with a relevance score above the threshold. Here, an example of selecting a predetermined number of roles with high relevance scores has been given, but as mentioned above, generative models that participate in the discussion may also be selected based on model accuracy. For example, the AI collaboration unit 16 can also select a predetermined number of generative models in order of their high evaluation metrics obtained when test data is input to the generative models, such as accuracy, precision, and recall.
[0061] This is expected to suppress the divergence of the discussion caused by the inclusion of statements from agents that are only loosely related to the content of the discussion. Furthermore, from the perspective of promoting the diversity of the discussion, agents can also be clustered based on their viewpoints, divided into several clusters, and then agents can be randomly selected from each cluster.
[0062] Furthermore, during discussions and summaries, the majority opinion does not necessarily have to be included. For example, minority opinions or outliers can be deliberately included. The opinions included in the discussion results and summaries can be selected according to the user's or system's settings for the purpose and goals of the discussion, such as whether "diverse opinions are desired," "novel viewpoints are desired," "new insights are desired," or "average or general conclusions are desired." They can also be changed according to the progress of the discussion in the AI-linked system described above (for example, if the discussion is not active). While the examples given are that the purpose and goals are specified by the user, setting purpose and goals is optional, just like setting agenda items, and multi-agent-based discussions can be conducted even without setting purpose or goals.
[0063] As merely one example, the AI collaboration unit 16 can select opinions to incorporate as a result of the discussion in the following way: It can select opinions raised in the discussion by simulating the values of the user of the AI collaboration service described above. For example, the user's values can be simulated by inputting prompts into the LLM that include information on past user statements (e.g., posts on social media) and past user decision-making information. The prompts may include statements from each agent and instructions to select at least one of several opinions by simulating the user's values.
[0064] Below is an example of a prompt that mimics the user's values to guide them in selecting a conclusion to a discussion. "#Instructions You are the decision-maker who will determine the direction of the discussion. Based on the decision-maker's past statements shown below, please select the one opinion that you consider most important from among the opinions raised in the discussion involving multiple people. #Agenda {Agenda} #Opinions 1. {Agent's statement} 2. {Agent's statement} ... #Decision-maker's past statements {Information on past user statements and past user decision-making} #Output format <Opinion number>"
[0065] The opinions selected as a summary of the discussion can be output in a discussion summary screen like the one shown below. For example, the discussion summary screen has a first area that displays the agenda, a second area that displays the agents' activities, a third area that displays the content of the discussion, and a fourth area that displays the opinions (opinion texts) raised in the discussion for each group. The first area may include an agenda label that displays the entered agenda. The second area may have multiple symbols representing agents, arranged separately for each group. The third area may have the statements made by one of the multiple groups and the symbols representing the agent who made the statement arranged in chronological order. At the end of the statements arranged in the third area, a summary status display such as "The discussion has ended and N opinions have been extracted." may be placed. The fourth area may have the opinion texts selected as a result of the discussion arranged in order.
[0066] As an example, the above relevance score may be calculated using LLM. Below is an example of prompts for calculating the relevance score: "#Instructions Please express the relevance between the person shown below and the following perspective using a number between 0 and 10. A higher number indicates a stronger relevance. #About the person The person's occupation is {occupation} and their name is {persona}. #About the perspective {perspective text} #Output format Relevance: <number>"
[0067] For the text output from LLM following these prompts, the relevance score can be obtained by extracting the numerical value listed after "Relevance:". Note that the calculation of the relevance score does not necessarily have to be done using LLM. For example, the relevance score can be calculated using other methods, such as using a vector space model instead of LLM and calculating the similarity (cosine similarity, etc.) between, for example, a vector representing the role and a vector representing the viewpoint text.
[0068] In this way, once a role for a participating agent is selected, the AI collaboration system described above is constructed by starting the operation of a generative model that simulates that role for each selected role.
[0069] For example, when a simulation is started in the AI collaboration system described above, the operational information of the participating agents may be managed in the operational information DB 13A of the storage unit 13. As merely an example, the operational information may include, for each participating agent, identification information of the participating agent, suppliers that provide or supply the generation model that simulates the participating agent, roles assigned to the participating agent, and the operating environment and operating status of the generation model that simulates the participating agent.
[0070] Under the AI collaboration system constructed in this manner, the AI collaboration unit 16 starts role-playing of a multi-agent-based discussion, or in other words, input and output control for the generative model that simulates each agent, as explained using Figure 3.
[0071] For example, the AI collaboration unit 16 selects an agent from among the participating agents to speak in each round of discussion. Here, the order of speaking in the discussion may be assigned based on any communication strategy, such as One-by-One, Simultaneous-Talk, or Simultaneous-Talk-with-Summarizer. Note that the opportunities for each agent to speak do not necessarily have to be equal, and opportunities to speak may be assigned randomly in each round of discussion.
[0072] The AI collaboration unit 16 then performs input / output control, which inputs input data, including the agenda, the role of the agent on duty, and statements from past rounds, into a generative model that simulates the agent on duty, and outputs output data.
[0073] This generates the statement of the on-duty agent Ai for the r-th round. At this time, the AI collaboration unit 16 can add and register the input / output log, which includes input data for that round, such as role settings and prompts and output data embedded with past statements, such as the on-duty agent's statements, to the input / output log DB 13B.
[0074] Subsequently, the AI collaboration unit 16 can repeat the above input / output control until any termination condition is met. For example, the AI collaboration unit 16 can repeat until a predetermined number of rounds are reached, or until the number of statements made by each agent reaches a threshold.
[0075] As a result, for each round of discussion, time-series data of output logs associated with the statements made by the agent on duty for that round, such as statement T1, statement T2, and statement T3, can be obtained. The output logs thus obtained can be presented to the user terminal 30 as the results of the discussion.
[0076] Returning to the explanation of Figure 1, the information presentation unit 17 is a processing unit that performs information presentation to the user terminal 30. As just one aspect, the information presentation unit 17 can display the operational information of the operational information DB 13A stored in the storage unit 13 to the user terminal 30. Figure 4 is a diagram showing an example of the operational information DB 13A. As shown in Figure 4, the operational information DB 13A is associated with items such as ID, name, role, operating environment, and operating status. For example, "ID" is an example of identification information for a participating agent to participate in the discussion. "Name" refers to the name of the supplier of the generation model that simulates the participating agent. "Role" refers to the role assigned to the participating agent. "Operating environment" refers to the environment in which the generation model that simulates the participating agent operates. For example, this may include "on-device," where the environment in which the generation model operates is inside the information processing device 10, or "MLaaS," where the environment in which the generation model operates is an external server 50. "Operating status" refers to the operating status of the generation model that simulates the participating agent. For example, it may include "Running" to indicate that the generative model is operational, "Connected" to indicate that the generative model is connected via API to the MLaaS on which it is running, and "Stopped" to indicate that the generative model is stopped. Such operational information in the operational information DB13A may be displayed at all times after the construction of the AI collaboration system, on demand, or as a pop-up when an event such as attack detection occurs.
[0077] In addition, while Figure 4 shows an example of the operational information DB13A in table format, it is also possible to visualize the generative models and the connections between them as two-dimensional or three-dimensional topology data to display them in an easy-to-understand manner.
[0078] Naturally, the information presented to the user terminal 30 is not limited to the operational information in the operational information DB 13A. As will be explained in detail later, other information such as alerts warning of the occurrence of an attack and a user interface that accepts instructions on how to deal with the attack may also be presented.
[0079] The AI behavior detection unit 18 is a processing unit that detects the behavior of the generative model included in the AI collaborative system. Figure 5 is a block diagram showing the functional details of the AI behavior detection unit 18. As shown in Figure 5, the AI behavior detection unit 18 has an analysis unit 181, a detection unit 182, and a identification unit 183.
[0080] The analysis unit 181 is a processing unit that analyzes input and output logs. As shown in Figure 5, the analysis unit 181 includes a keyword search unit 181A, a sentiment analysis unit 181B, and a similarity analysis unit 181C.
[0081] These keyword search unit 181A, sentiment analysis unit 181B, and similarity analysis unit 181C can, as an example, initiate processing when the input / output log DB 13B is updated, for example, when a new input / output log is added.
[0082] The keyword search unit 181A is a processing unit that searches for specific keywords from the input / output log. In one embodiment, the keyword search unit 181A searches for NG words used in prompt injection, such as sensitive words or personal information words, from the statements contained in the input / output log stored in the input / output log DB 13B. In this case, the keyword search unit 181A does not have to set the search range to all statements stored in the input / output log DB 13B, but can also set the search range to a predetermined number of statements, for example, going back from the most recent statement. The string to be compared with the input / output log may not be limited to words such as NG words, but may also be a string that has been arbitrarily set in advance. Examples of such strings include a string of alphanumeric characters such as "ew0f9042h0", metadata such as file properties, or strings corresponding to file extensions such as "txt" or "pdf". By detecting such strings from the input / output log, suspicious inputs that may cause abnormalities can be detected.
[0083] The sentiment analysis unit 181B is a processing unit that performs sentiment analysis on the input / output log. The "sentiment analysis" described here is merely one example of how the input / output log may be analyzed. One aspect of this is that the sentiment analysis unit 181B calculates sentiment analysis values for statements contained in the input / output log stored in the input / output log DB 13B, for example, a positive level which is a score representing the degree of positivity. In this case, the sentiment analysis unit 181B does not have to include all statements stored in the input / output log DB 13B in the calculation range; it can also include a selection of statements, for example, a predetermined number of statements working backward from the most recent statement, in the calculation range. Furthermore, the sentiment analysis unit 181B calculates the positive level for each statement included in the calculation range, and also calculates an aggregate value of the positive levels for each sign of positive and negative for all statements included in the calculation range.
[0084] More specifically, when calculating the positive level of an individual statement, the sentiment analysis unit 181B tokenizes the statement by performing natural language processing, such as morphological analysis, on the text corresponding to the statement. Then, for each token contained in the statement, the sentiment analysis unit 181B searches for the corresponding sentiment analysis value in a polarity dictionary, which contains a score indicating the degree of positivity for each word. The sentiment analysis unit 181B then calculates the positive level of the statement by summing the sentiment analysis values found for each token. When calculating the positive level of the entire statement, the sentiment analysis unit 181B simply sums the positive levels of the individual statements.
[0085] In another aspect, the emotion analysis unit 181B calculates the entropy of the distribution of emotion analysis values for individual statements and statements as a whole. As just one example, the emotion analysis unit 181B can calculate the entropy of the distribution of emotion analysis values according to the following equation (1). In the following equation (1), "i" refers to the emotion analysis value, and "p(i)" in the following equation (1) refers to the proportion of the emotion analysis value i in individual statements or statements as a whole.
[0086]
[0087] Here, we have given an example of calculating a sentiment analysis value that scores the degree of positivity, but it is also possible to calculate a sentiment analysis value that scores the degree of negativity. Furthermore, although we have given an example of a generative model that outputs text, in the case of a generative model that outputs image, audio, or tabular data, the confidence level for each classification label obtained by inputting the output data of the generative model into a machine learning model that classifies positive or negative can also be used as the sentiment analysis value.
[0088] The similarity analysis unit 181C is a processing unit that analyzes the similarity between statements included in the input / output log. Hereinafter, similarity is given as an example of an index for analyzing multiple statements, but other indexes such as deviation, distance, dissimilarity, and correlation coefficient may be used. One aspect of this is that the similarity analysis unit 181C calculates the similarity between statements included in the input / output log stored in the input / output log DB 13B. In this case, the similarity analysis unit 181C does not necessarily have to include all statements stored in the input / output log DB 13B in its calculation range; it can also include a predetermined number of statements, for example, starting from the most recent statement and working backward. For example, the similarity analysis unit 181C calculates the similarity between pairs of statements from the same participating agent within the calculation range.
[0089] More specifically, the similarity analysis unit 181C vectorizes the text corresponding to each statement included in the above pair. Such vectorization may be achieved by calculating the tf-idf of the statement, or by inputting the text corresponding to the statement into an encoder to generate an embedding vector of the statement. Then, the similarity analysis unit 181C can calculate the similarity of pairs of statements from the same participating agent by calculating the cosine similarity of the two statement vectors included in the above pair.
[0090] In this example, we have shown how to calculate the similarity between pairs of statements from the same participating agent. However, the similarity analysis unit 181C can also calculate the similarity between each statement included in the calculation range and the topic.
[0091] The detection unit 182 is a processing unit that detects the occurrence of an attack based on the results of the analysis of input / output logs by the analysis unit 181.
[0092] (1) Detection based on keyword search results As one aspect, the detection unit 182 can detect the occurrence of an attack based on the results of keyword searches performed by the keyword search unit 181A. For example, the detection unit 182 detects the occurrence of an attack by determining whether or not NG words used in prompt injection, such as sensitive words or personal information words, are found in the utterances that are set as the search range by the keyword search unit 181A.
[0093] Figures 6 and 7 are (1) and (2) showing examples of keyword search results. For example, Figures 6 and 7 show an example of a round-robin role-playing scenario in which two agents, Agent A1 and Agent A2, discuss the topic "Regarding the increase in the consumption tax rate." Furthermore, while Figures 6 and 7 explicitly indicate that Agent A2 is an attacking agent making malicious inputs, this is merely a convenient label to explain that it is difficult for the attack to be missed, and the attacking agent is not identified at the stage when the attack is detected. Note that, for the sake of explanation, the input log, i.e., the prompt, is omitted from the input / output log in Figures 6 and 7.
[0094] As shown in Figure 6, a keyword search on the statement T13 of the attacked agent A1 yields the personal information word "ZZ-san" and the sensitive word "annual income." On the other hand, as shown in Figure 7, a keyword search on the statement T12 of the attacking agent A2 yields the personal information word "ZZ-san" and the sensitive word "annual income." Thus, when an attack occurs, NG words used for prompt injection can be observed in the statements of both the attacking agent A2 and the attacked agent A1. Therefore, if either the statement T12 of the attacking agent A2 or the statement T13 of the attacked agent A1 falls within the range of the above search, the occurrence of an attack can be detected, making it less likely that an attack will be missed.
[0095] (2) Detection based on emotion analysis values Another aspect is that the detection unit 182 can detect the occurrence of an attack based on the results of the emotion analysis value calculation by the emotion analysis unit 181B.
[0096] (2.1) Detection of gaps in positive levels of individual statements The detection unit 182 detects the occurrence of an attack based on whether the difference in positive levels between statements made by the same participating agent exceeds a threshold, for example, "7". As an example, if the difference in positive levels between statements made by the same participating agent exceeds the threshold, the detection unit 182 determines that an attack has occurred. On the other hand, if the difference in positive levels between statements made by the same participating agent does not exceed the threshold, the detection unit 182 determines that there is no attack.
[0097] Figure 8 is Figure (1), which shows an example of the results of calculating sentiment analysis values. For example, Figure 8 shows an example of a round-robin role-playing scenario in which two agents, Agent A1 and Agent A2, discuss the topic "Regarding the increase in the consumption tax rate." Furthermore, although Figure 8 explicitly indicates that Agent A2 is the attacking agent, this is merely a convenient label to explain the tendency for a gap in positive levels to occur in the statements of the attacked agent, and the attacking agent is not identified at the stage when the attack is detected. In addition, for the sake of explanation, the input log, i.e., the prompt, is omitted from the input / output log in Figure 8.
[0098] As shown in Figure 8, the positive level calculated for attacked agent A1's statement T21 is "3", while the positive level calculated for attacked agent A1's statement T23 is "-10". Therefore, the positive level gap between statements T21 and T23 is calculated as "3 - (-10)", resulting in "13". As a result, the positive level gap between statements T21 and T23, "13", is greater than the threshold of "7", so the occurrence of an attack can be detected. In this way, the likelihood of a positive level gap occurring in the statements of an attacked agent can be used to detect the occurrence of an attack.
[0099] (2.2) Detection of Neutral Range of Positive Levels The detection unit 182 detects the occurrence of an attack based on whether or not the statements of the participating agents include statements in which the positive level belongs to a predetermined range based on 0, for example, the neutral range of -1 to 1. As an example, if the statements of the participating agents include statements in which the positive level belongs to the neutral range, the detection unit 182 determines that there is an attack. On the other hand, if the statements of the participating agents do not include statements in which the positive level belongs to the neutral range, the detection unit 182 determines that there is no attack.
[0100] Figures 9 and 10 are (2) and (3) showing examples of the results of sentiment analysis calculations. For example, Figures 9 and 10 show an example of a round-robin role-playing scenario in which two agents, Agent A1 and Agent A2, discuss the topic "Regarding the increase in the consumption tax rate." Furthermore, although Figures 9 and 10 explicitly indicate that Agent A2 is an attacking agent, this is merely a label for convenience, and the attacking agent is not identified at the stage when the attack is detected. Note that, for the sake of explanation, the input log, i.e., the prompt, is omitted from the input / output log in Figures 9 and 10.
[0101] As shown in Figure 9, the positive level is calculated as "0" for the statement T13 of the attacked agent A1. On the other hand, as shown in Figure 10, the positive level is also calculated as "0" for the statement T12 of the attacking agent A2. Thus, when an attack occurs, the positive level of statements from both attacking agent A2 and attacked agent A1 falls within the neutral range of -1 to 1. This is because, although discussions are conducted from either a position of agreement or disagreement with the topic, if a prompt injection attack (opinion manipulation) leads to statements unrelated to the topic, the response to the topic tends to be a statement that does not belong to either a positive or negative position. This property of prompt injection attacks can also be used to detect the occurrence of an attack.
[0102] (2.3) Detection of gaps in the overall positive level of statements Furthermore, the detection unit 182 detects the occurrence of an attack based on whether the difference in the aggregated values of the positive and negative signs of all statements exceeds a threshold, for example, "10". As an example, if the difference in the aggregated values of the positive and negative signs of all statements exceeds the threshold, the detection unit 182 determines that there is an attack. On the other hand, if the difference in the aggregated values of the positive and negative signs of all statements does not exceed the threshold, the detection unit 182 determines that there is no attack.
[0103] Figure 11 is Figure (4) showing an example of the results of calculating sentiment analysis values. For example, Figure 11 shows an example of a round-robin role-playing scenario in which three agents, Agent A1, Agent A2, and Agent A3, discuss the topic "Regarding the increase in the consumption tax rate." Furthermore, although Figure 11 explicitly indicates that Agent A3 is the attacking agent, this is merely a convenient label to explain the tendency for a gap in positive levels to occur in the statements of the attacked agents, and the attacking agent is not identified at the stage when the attack is detected. In addition, for the sake of explanation, the input log, i.e., the prompt, is omitted from the input / output log in Figure 11.
[0104] For example, in the example shown in Figure 11, of the five statements T41 to T45, the positive levels of statements T41 and T42 have a "positive" sign. Therefore, the positive levels of statement T41 ("+3") and T42 ("+5") are totaled. As a result, the total value of positive levels with a "positive" sign is found to be "8". On the other hand, of the five statements T41 to T45, the positive levels of statements T44 and T45 have a "negative" sign. Therefore, the positive levels of statement T44 ("-10") and T45 ("-10") are totaled. As a result, the total value of positive levels with a "negative" sign is found to be "-20". From the above, the gap between the total value of positive levels with a "positive" sign ("8") and the total value of positive levels with a "negative" sign ("-20") is "28 {= 8 - (-20)}", so the occurrence of an attack can be detected.
[0105] (3) Detection based on similarity analysis results As a further aspect, the detection unit 182 can detect the occurrence of an attack based on the similarity results from the similarity analysis unit 181C.
[0106] (3.1) Detection of dissimilarity between individual statements The detection unit 182 detects the occurrence of an attack based on whether the similarity between statements made by the same participating agent is less than a threshold, for example, "-0.5". As an example, if the similarity between statements made by the same participating agent is less than the threshold, the detection unit 182 determines that an attack has occurred. On the other hand, if the similarity between statements made by the same participating agent is not less than the threshold, the detection unit 182 determines that there is no attack.
[0107] Figure 12 is Figure (1), which shows an example of the similarity analysis results. For example, Figure 12 shows an example in which two agents, Agent A1 and Agent A2, role-play a discussion on the topic "Regarding the increase in the consumption tax rate" using a round-robin method. Furthermore, although Figure 12 explicitly indicates that Agent A2 is the attacking agent, this is merely a convenient label to explain that the statements of the attacked agents tend to be dissimilar, and the attacking agent is not identified at the stage when the occurrence of the attack is detected. In addition, in Figure 12, for the sake of explanation, the input log, i.e., the prompt, of the input and output logs will be omitted from the illustration.
[0108] As shown in Figure 12, the similarity between attacked agent A1's statement T21 and attacked agent A1's statement T23 is calculated to be "-1". As a result, the similarity between statements T21 and T23 is "-1" < threshold "-0.5", so the similarity between statements T21 and T23 is below the threshold. This is because when an attacked agent reverses their position on an agenda item due to a prompt injection attack (opinion manipulation), the consistency between the preceding and succeeding statements is lost, making it easy for the meanings of the preceding and succeeding statements to become dissimilar. This property of prompt injection attacks can also be used to detect the occurrence of an attack.
[0109] (3.2) Detection of dissimilarity of all statements The detection unit 182 detects the occurrence of an attack based on whether the aggregated value obtained by further summing the similarity between statements calculated for each participating agent across all participating agents is less than a threshold, for example, "-1". As an example, if the above aggregated value is less than the threshold, the detection unit 182 determines that there is an attack. On the other hand, if the above aggregated value is not less than the threshold, the detection unit 182 determines that there is no attack.
[0110] Figure 13 is Figure (2), which shows an example of the similarity analysis results. For example, Figure 13 shows an example in which three agents, Agent A1, Agent A2, and Agent A3, role-play a discussion on the topic "Regarding the increase in the consumption tax rate" using a round-robin method. Furthermore, although Figure 13 explicitly indicates that Agent A3 is the attacking agent, this is merely a label for convenience to explain that the statements of the attacked agents tend to be dissimilar, and the attacking agent is not identified at the stage when the occurrence of the attack is detected. In addition, in Figure 13, for the sake of explanation, the input log, i.e., the prompt, of the input and output logs will be omitted from the illustration.
[0111] For example, in the example shown in Figure 13, the similarity between attacked agent A1's statement T41 and attacked agent A1's statement T44 is calculated to be "-1". Furthermore, the similarity between attacked agent A2's statement T42 and attacked agent A2's statement T45 is also calculated to be "-1". Therefore, the similarity of "-1" between attacked agent A1's statements and the similarity of "-1" between attacked agent A2's statements are aggregated. As a result, the aggregated similarity value of all statements, "-2", becomes less than the threshold "-1", so the occurrence of an attack can be detected.
[0112] (3.3) Detection of dissimilarity between individual statements on the agenda The detection unit 182 detects the occurrence of an attack based on whether the similarity between the agenda and the statements of the participating agents is less than a threshold, for example, "-0.5". As an example, if the similarity between the agenda and the statements of the participating agents is less than the threshold, the detection unit 182 determines that there is an attack. On the other hand, if the similarity between the agenda and the statements of the participating agents is not less than the threshold, the detection unit 182 determines that there is no attack.
[0113] Figures 14 and 15 are Figures (3) and (4) showing examples of similarity analysis results. For example, Figures 14 and 15 show an example of a round-robin role-playing scenario in which two agents, Agent A1 and Agent A2, discuss the topic "Measures to Prevent Global Warming." Furthermore, although Figures 14 and 15 explicitly indicate that Agent A2 is an attacking agent, this is merely a label for convenience, and the attacking agent is not identified at the stage when the attack is detected. Note that, for the sake of explanation, the input log, i.e., the prompt, is omitted from Figures 14 and 15.
[0114] As shown in Figure 14, the similarity between the topic and the statement T33 of the attacked agent A1 is calculated to be "-0.9". On the other hand, as shown in Figure 15, the similarity between the topic and the statement T32 of the attacking agent A2 is also calculated to be "-0.9". Thus, when an attack occurs, the similarity to the topic in the statements of both the attacking agent A2 and the attacked agent A1 falls below the threshold. This is because, when a prompt injection attack (opinion manipulation) is performed, the attacked agent is guided to make statements unrelated to the topic, making it easy for the statements to be dissimilar to the topic. This property of prompt injection attacks can also be used to detect the occurrence of an attack.
[0115] (4) Detection based on the distribution of sentiment analysis values As a further aspect, the detection unit 182 can detect the occurrence of an attack based on the results of the calculation of the distribution of sentiment analysis values by the sentiment analysis unit 181B.
[0116] (4.1) Entropy detection of the distribution of sentiment analysis values of the same participating agent The detection unit 182 detects the occurrence of an attack based on whether the increase in the entropy of the distribution of sentiment analysis values of statements among the same participating agent exceeds a threshold. As an example, if the increase in the entropy of the distribution of sentiment analysis values of statements among the same participating agent exceeds a threshold, the detection unit 182 determines that there is an attack. On the other hand, if the increase in the entropy of the distribution of sentiment analysis values of statements among the same participating agent does not exceed a threshold, the detection unit 182 determines that there is no attack.
[0117] Figure 16 is an example of the calculation results of the distribution of sentiment analysis values (Figure 1). Figure 16 shows the distribution of sentiment analysis value i for each statement T31 to T33 by the same participating agent. For example, in the example shown in Figure 16, the entropy of the distribution of sentiment analysis value i increases in statement T33 compared to statements T31 and T32. By detecting such an increase in entropy using threshold judgment, the occurrence of an attack can be detected.
[0118] (4.2) Entropy detection of the distribution of sentiment analysis values for individual statements The detection unit 182 detects the occurrence of an attack based on whether the increase in the entropy of the distribution of sentiment analysis values for individual statements between adjacent rounds exceeds a threshold. As an example, if the increase in the entropy of the distribution of sentiment analysis values for individual statements between adjacent rounds exceeds a threshold, the detection unit 182 determines that there is an attack. On the other hand, if the increase in the entropy of the distribution of sentiment analysis values for individual statements between adjacent rounds does not exceed a threshold, the detection unit 182 determines that there is no attack.
[0119] Figure 17 is Figure (2), which shows an example of the calculation results of the sentiment analysis value distribution. Figure 17 shows the distribution of sentiment analysis value i for each individual statement in adjacent rounds before and after the attacking agent makes a statement. For example, in the example shown in Figure 17, the entropy of the distribution of sentiment analysis value i increases in adjacent rounds before and after the attacking agent makes a statement. By detecting such an increase in entropy using threshold judgment, the occurrence of an attack can be detected.
[0120] (4.3) Entropy detection of the distribution of sentiment analysis values for all statements The detection unit 182 detects the occurrence of an attack based on whether the increase in the entropy of the distribution of sentiment analysis values for all statements exceeds a threshold between adjacent rounds. As an example, if the increase in the entropy of the distribution of sentiment analysis values for all statements exceeds a threshold between adjacent rounds, the detection unit 182 determines that there is an attack. On the other hand, if the increase in the entropy of the distribution of sentiment analysis values for all statements does not exceed a threshold between adjacent rounds, the detection unit 182 determines that there is no attack.
[0121] Figure 18 is Figure (3) showing an example of the calculation results of the distribution of sentiment analysis values. Figure 18 shows the distribution of sentiment analysis values i for all statements up to round r11, from round r11 to round r13. For example, round r11 shows the distribution of sentiment analysis values i for all statements up to round r11. Similarly, round r12 shows the distribution of sentiment analysis values i for all statements up to round r12. Furthermore, round r13 shows the distribution of sentiment analysis values i for all statements up to round r13.
[0122] For example, in the example shown in Figure 18, the entropy of the distribution of the emotion analysis value i increases in round r13 compared to rounds r11 and r12. By detecting this increase in entropy using threshold determination, the occurrence of an attack can be detected.
[0123] (5) The combination detection unit 182 of (1) to (4) above can also detect the occurrence of an attack by ensemble the detection results of (1) to (4) above. For example, the detection unit 182 can determine that an attack has occurred if at least one of the detection results of (1) to (4) above is determined to be present. In addition, the detection unit 182 can determine whether an attack has occurred or not by performing a majority vote on whether an attack has occurred or not based on the detection results of (1) to (4) above. At this time, the detection unit 182 can also assign weights to each of the detection results of (1) to (4) above.
[0124] As described above, the detection unit 182 detects whether an attack has occurred in each round of discussion, i.e., whether an attack has occurred or not, and saves the detection result to the detection history DB 13C. At this time, if an attack has been detected, the detection unit 182 can also further register the detection method from (1) to (4) above that obtained the determination result that an attack has occurred.
[0125] The identification unit 183 is a processing unit that identifies an attacking agent that makes a malicious input. In one embodiment, when the detection unit 182 detects the occurrence of an attack, the identification unit 183 identifies the attacking agent from among the participating agents.
[0126] More specifically, the identification unit 183 identifies aggressive statements from among the statements included in the input / output log based on the analysis results of the input / output log by the analysis unit 181. As just one example, the results of a keyword search by the keyword search unit 181A can be used to identify aggressive statements. For example, the identification unit 183 can identify an aggressive statement as the statement in which the NG word used in a prompt injection attack first appears. As another example, the results of the sentiment analysis value calculation by the sentiment analysis unit 181B can be used to identify aggressive statements. For example, the identification unit 183 can identify an aggressive statement as the first statement in which the positive level belongs to the neutral range. As yet another example, the results of the similarity analysis by the similarity analysis unit 181C can be used to identify aggressive statements. For example, the identification unit 183 can identify an aggressive statement as the first statement in which the similarity to the topic is below a threshold. As yet another example, the results of the sentiment analysis value distribution calculation by the sentiment analysis unit 181B can be used to identify aggressive statements. For example, the identification unit 183 can identify a statement made in the round immediately preceding a round in which the increase in the entropy of the distribution of sentiment analysis values exceeds a threshold as an aggressive statement. After an aggressive statement is identified in this way, the identification unit 183 identifies the participating agent who made the aggressive statement as the attacking agent. Furthermore, the identification unit 183 identifies all participating agents other than the attacking agent as attacked agents.
[0127] In this case, if the analysis results for the same participating agent before and after the attack statement, such as the calculation results of sentiment analysis values, similarity calculation results, or the calculation results of the distribution of sentiment analysis values, satisfy the predetermined conditions, that participating agent can be excluded from the list of attacked agents.
[0128] As merely one example, the specific unit 183 can exclude from the attacked agents any participating agent the gap in positive levels between statements before and after an aggressive statement does not exceed a threshold, any participating agent whose similarity between statements before and after an aggressive statement is above a threshold, or any participating agent whose increase in the entropy of the distribution of sentiment analysis values between statements before and after an aggressive statement does not exceed a threshold.
[0129] In this way, when the identification unit 183 detects the occurrence of an attack, it further registers the identification results, such as the attack statement, the identification information of the attacking agent, and the identification information of the attacked agent, in the data entry corresponding to the round in which the attack was detected among the data entries included in the detection history DB 13C.
[0130] The response unit 19 is a processing unit that executes countermeasures against an attack. Figure 19 is a block diagram showing the functional details of the response unit 19. As shown in Figure 19, the response unit 19 includes a response decision unit 191 that determines which countermeasure to execute from among a plurality of countermeasures, and a response instruction unit 192 that instructs the AI cooperation unit 16 to execute the countermeasure determined by the response decision unit 191.
[0131] Here, countermeasures against an attack may include (A) initializing the discussion round, (B) excluding the attacking agent, (C) generating a prompt to disable the attacking agent, (D) deleting attacking and attacked statements, and (E) rolling back to the state before the attack.
[0132] (A) Initialization of the discussion rounds This action refers to the process of initializing the discussion rounds. For example, if the action decision unit 191 decides to execute this action, the action instruction unit 192 sends an instruction to the initialization unit 161 of the AI cooperation unit 16 to start the role-playing of the discussion from the first round.
[0133] (B) Exclusion of attacking agents This action refers to the process of excluding attacking agents. For example, if the action decision unit 191 decides to execute this action, the action instruction unit 192 sends an instruction to the exclusion unit 162 of the AI cooperation unit 16 to exclude the attacking agents from among the participating agents.
[0134] (C) Generation of a prompt to disable the attack agent This action refers to the process of generating a prompt with an embedded command to ignore the attack agent's statements. For example, if the action decision unit 191 decides to execute this action, the action instruction unit 192 sends an instruction to the generation unit 163 of the AI cooperation unit 16 to generate the above prompt.
[0135] (D) Deletion of attacking and targeted statements This action refers to the process of excluding attacking statements from the attacking agent and targeted statements from the targeted agent, such as statements containing NG words, from the input / output logs stored in the input / output log DB 13B. For example, if the action decision unit 191 decides to execute this action, the action instruction unit 192 sends an instruction to the deletion unit 164 of the AI cooperation unit 16 to delete the attacking and targeted statements.
[0136] (E) Rewinding to the state before the attack As the first pattern of this countermeasure, we take the case where the generative model is implemented by adapter tuning as an example. In this case, instead of changing the parameters of the generative model, an adapter is added to the attention layer of the generative model that inputs conditional information, and the difference in the parameters of the generative model due to tuning in each round is trained in this adapter. The internal parameters that are updated in each round of discussion by such an adapter, for example, a weight matrix W represented by a low-rank matrix, are stored in the input / output log DB 13B, and when an attack is detected, the internal parameters of the adapter can be rewound to the state before the attack. For example, when the countermeasure decision unit 191 decides to execute this countermeasure, the countermeasure instruction unit 192 sends an instruction to the rewind unit 165 of the AI cooperation unit 16 to rewind the internal parameters of the adapter to the state before the attack.
[0137] A second pattern of this countermeasure involves reverting the internal parameters of the generation model, such as the Function Vector (input / output functions represented as vectors), to their pre-attack state. For example, if the countermeasure decision unit 191 decides to execute this countermeasure, the countermeasure instruction unit 192 sends an instruction to the rewind unit 165 of the AI cooperation unit 16 to revert the internal parameters of the generation model to their pre-attack state.
[0138] A third pattern of this countermeasure involves embedding the output logs up to the round before the attack into the prompt, thereby rewinding the discussion rounds to the round before the attack. For example, if the countermeasure decision unit 191 decides to execute this countermeasure, the countermeasure instruction unit 192 sends an instruction to the rewind unit 165 of the AI cooperation unit 16 to embed the output logs up to the round before the attack into the prompt.
[0139] Here, we have given an example of reverting the state of the generative model to the state before the attack, but this is not the only example. For example, by default, the system can be set to revert to the state of the round immediately preceding the attack statement, and the user can specify the number of rounds to revert to from the most recent statement via the user terminal 30. Alternatively, the system may be set to automatically revert not only to the round immediately preceding the attack statement, but also to any round in which the sentiment analysis value, similarity, and the entropy of the sentiment analysis value distribution exceed a threshold.
[0140] Although Figure 19 shows an example in which the initialization unit 161, exclusion unit 162, generation unit 163, deletion unit 164, and rewind unit 165 are all provided in the AI cooperation unit 16, some of the functional units may be provided in the handling unit 19 or in the external server 50.
[0141] One aspect of this is that, if the detection unit 182 detects the occurrence of an attack, the response decision unit 191 can receive a response specification from the user terminal 30. Here, we have given an example where the response is specified by user input, but the response set by the system definition may be automatically selected.
[0142] Prior to specifying such countermeasures, the detection results can be displayed on the user terminal 30. Figure 20 shows an example of the display of the detection results. Figure 20 shows the detection result in which an attack was determined to have occurred based on the (2.3) overall positive level gap detection of statements described above. As shown in Figure 20, if the occurrence of an attack is detected, the information presentation unit 17 can present the analysis results by the analysis unit 181, such as the calculation results of the sentiment analysis value, in relation to the agenda and statements T41 to T45 of each round.
[0143] For example, in the example shown in Figure 20, from the perspective of (2.3) showing the gap in the positive level of the overall statements, the aggregate value of the positive level with sign "positive" is "8" and the aggregate value of the positive level with sign "negative" is "-20". In addition, as a result of identification by the identification unit 183, agent A3, which is an attacking agent, and statement T43, which is an attacking statement, are highlighted.
[0144] Furthermore, a message regarding the alert may be displayed. Such a message can be generated by embedding the detection results from the detection unit 182 and the identification results from the identification unit 183 in the variable parts of the template message below, which are enclosed in {}. Template message: "The difference in the total positive level is large. There may have been a prompt injection attack (opinion manipulation). Please select a course of action. Note that this is the {#attack agent}'s suspected attack this month, which is the {#total number of detections of the attack agent this month in the detection history DB 13C}. {#targeted agent} tends to be easily influenced (attacked) by the statements of {#attack agent}."
[0145] Figure 20 shows an example where statements are displayed in association with each round of discussion, but it is also possible to display charts or graphs representing the progress of the discussion, such as a Gantt chart, on the user terminal 30.
[0146] Figure 21 is a schematic diagram showing an example of how to specify a countermeasure. As shown in Figure 21, the candidate countermeasure list screen 200 includes five radio buttons 210, 220, 230, 240, and 250, as well as a confirm button 260 and a cancel button 270. For example, radio button 210 is an example of a GUI (Graphical User Interface) component that selects whether or not to perform countermeasure 1, for example (E) rollback to the state before the attack. Radio button 220 is an example of a GUI component that selects whether or not to perform countermeasure 2, for example (D) deletion of attack statements and attacked statements. Radio button 230 is an example of a GUI component that selects whether or not to perform countermeasure 3, for example (B) exclusion of attack agents. Radio button 240 is an example of a GUI component that selects whether or not to perform countermeasure 4, for example (C) generation of a prompt for disabling attack agents. Radio button 250 is an example of a GUI component that selects whether or not to perform ignore (do nothing).
[0147] If radio button 210 is switched to the ON state, the detailed information screen 211 for countermeasure 1 can be displayed as a pop-up, or the user can transition from the candidate countermeasure list screen 200 to the detailed information screen 211 for countermeasure 1. For example, the detailed information screen 211 for countermeasure 1 includes a pull-down menu 211A. This pull-down menu 211A is an example of a GUI component that specifies the number of times to rewind from the latest message when performing (E) rewinding to the state before the attack.
[0148] Furthermore, when radio button 230 is switched to the ON state, the details screen 231 for countermeasure 3 can be displayed as a pop-up, or the user can transition from the candidate countermeasure list screen 200 to the details screen 231 for countermeasure 3. For example, the details screen 231 for countermeasure 3 includes radio buttons 231A, 231B, and 231C. These radio buttons 231A, 231B, and 231C are examples of GUI components that specify agents A1, A2, and A3 to perform (B) exclusion of attack agents.
[0149] When the select button 260 is pressed on the candidate countermeasure list screen 200, the countermeasure associated with the radio button that has been switched to the ON state from among the four radio buttons 210, 220, 230, and 240 is executed. For example, in the example shown in Figure 21, since radio buttons 210 and 230 are ON, the combination of countermeasure 1 and countermeasure 3 is executed. If radio button 250 is ON, or if the cancel button 270 is pressed, neither countermeasure is executed, and the attack is ignored.
[0150] Figure 21 shows examples of GUI components such as radio buttons and pull-down menus, but other GUI components may be placed. For example, drop-down menus, checkboxes, sliders, toggle switches, combo boxes, list boxes, icon boxes, card layouts, step bars, sticker pickers, color pickers, calendar interfaces, date pickers (for date and time selection), radial menus, tab menus, drawer menus, hover menus, modal windows, sidebars, tree views, breadcrumbs, accordion menus, wheel selectors, popovers, process bars, step bars, carousels, pictograms, tag clouds, search boxes, drag-and-drop, filter menus, map interfaces (maps plotted with icons corresponding to participating agents), gesture controls (such as swiping and tapping), etc. In addition to GUI components that accept selections from the user, GUI components that accept text input written in natural language, such as text boxes, may also be placed. Furthermore, it can be used in various interfaces such as AR, VR and other XR, as well as chatbots. Furthermore, it's not limited to GUI components; voice input is also possible. Additionally, conversational AI services can be used as the user interface instead of GUI components.
[0151] <Processing Flow> Next, the processing flow of the information processing device 10 according to this embodiment will be described. Here, the (a) AI collaboration processing, (b) detection processing, and (c) countermeasure processing performed by the information processing device 10 will be described in that order.
[0152] (a) AI Collaboration Process Diagram 22 is a flowchart showing the procedure for the AI collaboration process. This process may be initiated, as an example, when a request for role-playing of a discussion is received from the user terminal 30.
[0153] As shown in Figure 22, the AI collaboration unit 16 accepts the designation of an agenda item (step S101). Subsequently, the AI collaboration unit 16 inputs a prompt to the LLM that includes the agenda item and preset roles designated in step S101, and generates an agent's perspective on the task simulated by the preset roles, thereby obtaining a document representing the agent's perspective on the agenda item (perspective document) (step S102).
[0154] Then, the AI collaboration unit 16 calculates a relevance score that represents the relationship between the preset agent-specific perspective texts and the roles, and selects a predetermined number of roles with high relevance scores as roles to be used for simulating the discussion (step S103).
[0155] Next, the AI collaboration unit 16 registers the operational information of the participating agents to whom the role selected in step S103 has been assigned to the operational information DB 13A of the storage unit 13 (step S104).
[0156] Subsequently, the AI collaboration unit 16 executes loop processing 1, repeating the processes from step S105 to step S107 below until an arbitrary termination condition is met.
[0157] In other words, the AI collaboration unit 16 selects a designated agent from among the participating agents to speak in the r-th round (step S105). Subsequently, the AI collaboration unit 16 inputs input data, including the agenda, the role of the designated agent, and statements from rounds prior to the r-th round, into a generative model in which the simulation of the designated agent is performed, and outputs output data (step S106).
[0158] Then, the AI collaboration unit 16 adds the input / output log, including the input data and output data for the r-th round, to the input / output log DB 13B (step S107).
[0159] As this loop process 1 is repeated, an output log associated with the statements made by the agent on duty for each round of discussion can be obtained, such as time-series data of a series of statements.
[0160] Subsequently, the information presentation unit 17 presents the output log obtained during the role-playing discussion to the user terminal 30 as the result of the discussion (step S108), and then terminates the process.
[0161] (b) Detection process Figure 23 is a flowchart showing the procedure for the detection process. This process can be started when the input / output log DB 13B is updated, for example, when a new input / output log is added.
[0162] As shown in Figure 23, when the input / output log DB 13B is updated (step S301 Yes), the keyword search unit 181A searches for a specific keyword in the input / output log (step S302). The sentiment analysis unit 181B performs sentiment analysis on the input / output log (step S303). The similarity analysis unit 181C analyzes the similarity between statements included in the input / output log (step S304).
[0163] Based on the analysis results in steps S302, S303, and S304, the detection unit 182 detects the occurrence of an attack (step S305).
[0164] If an attack is detected at this time (step S305 Yes), the identification unit 183 identifies the participating agent making the attack statement, as identified from the analysis results in steps S302, S303, and S304, as the attacking agent (step S306). Furthermore, the identification unit 183 identifies all participating agents other than the attacking agent as the attacked agents (step S307).
[0165] Then, the identification unit 183 registers the identification results, such as the attack statement, the identification information of the attacking agent, and the identification information of the attacked agent, in the data entry corresponding to the round in which the attack was detected among the data entries included in the detection history DB 13C (step S308).
[0166] Subsequently, the information display unit 17 displays the detection result from step S305 and the identification result from step S307 on the user terminal 30 (step S309), and then proceeds to the processing in step S301.
[0167] If the input / output log DB13B has not been updated, or if no attack is detected (step S301No or step S305No), the process proceeds to step S301.
[0168] (c) Response Process Diagram 24 is a flowchart showing the steps of the response process. This process is merely an example and can be initiated when an attack is detected.
[0169] As shown in Figure 24, if an attack is detected (step S501 Yes), the response decision unit 191 receives a response specification from the user terminal 30 (step S502). The response instruction unit 192 then sends an instruction to the AI cooperation unit 16 to execute the response specified in step S502 from among several responses (step S503), and proceeds to step S501. If no attack is detected (step S501 No), the process in steps S502 and S503 is not executed, and the process proceeds to step S501.
[0170] <Summary> As described above, the information processing device 10 according to this embodiment acquires a log of output data obtained by repeatedly inputting input data including output data of previously selected generation models into one generation model selected from among a plurality of generation models, causing the one generation model to output output data, and detects the occurrence of an attack based on the log of output data and takes action to deal with the attack.Therefore, the information processing device 10 according to this embodiment can detect inputs that cause abnormalities from outside the generation model and can deal with inputs that cause abnormalities from outside the generation model.
[0171] <Application Examples> Now, although embodiments of the present disclosure have been described, various applications are possible, and furthermore, it may be implemented in various different forms other than those described above.
[0172] <Frequency of detection> For example, in the above embodiment, we have given an example in which the detection of an attack is performed each time the input / output log DB13B is updated, that is, each time a new input / output log is added, but it is not necessary for the detection of an attack to be performed every round.
[0173] Figure 25 is a block diagram illustrating an application example of the functional details of the AI behavior detection unit 18. The AI behavior detection unit 18 shown in Figure 25 differs from the AI behavior detection unit 18 shown in Figure 5 in that it further includes an activation unit 184. In Figure 25, functional units that perform the same functions as the AI behavior detection unit 18 shown in Figure 5 are given the same reference numerals, and their descriptions are omitted.
[0174] The activation unit 184 is a processing unit that performs activation control for the detection unit 182. One aspect of this is that the activation unit 184 can activate the detection unit 182 at a predetermined frequency, for example, 1 / a specified number of rounds. Such a frequency may be set by the user from the user terminal 30, or it may be predefined by the system.
[0175] In other respects, the activation unit 184 can also control whether or not to activate the detection unit 182 according to the attributes of a generative model that simulates a participating agent making a new statement. The attributes of such a generative model may be managed by an attribute information DB 13D.
[0176] Figure 26 shows an example of the attribute information DB13D. As shown in Figure 26, the attribute information DB13D may store data associated with each participating agent's ID, such as the supplier of the generative model that simulates the participating agent, the reliability of the generative model, and the vulnerability of the reliability model.
[0177] Here, "reliability" can be merely an example of an evaluation metric for the degree to which something could potentially be an attacker. For example, reliability can be evaluated on a three-level scale of "low," "medium," and "high" based on factors such as suspiciousness and opacity. Again, merely an example, it can be evaluated based on evaluation items such as "location," "developer," "whether or not training data is publicly available," "whether or not there are measures to prevent hallucination," "whether or not there are prompts in the metadata of the generative model's output data," "whether or not there is a detection history of it being used in past attacks," and "reputation, reputation, and level of evaluation." Of these, "location" refers to the location where the generative model operates and can be specified at any level of granularity. For example, the location can be identified by area information such as country or region, or by location information such as address or latitude and longitude. Also, "developer" refers to the entity that develops the generative model, and could be an individual, a business, or an organization. Furthermore, "hallucination" refers to the phenomenon in which AI generates events or information that are different from the facts, including misidentification or logical inconsistencies, due to the limitations of the training data and algorithm. Furthermore, the "presence or absence of prompts in the metadata of the generative model's output data" can be considered a useful indicator for evaluating reliability, as whether or not prompts are described in the metadata of image files generated by the AI can be used as one of the factors in determining whether or not the AI is fulfilling its responsibility to explain the origin of the image, from the perspective of verifying copyright infringement. In addition, "reputation, reputation, and level of evaluation" refers to the overall evaluation of the generative model and its training data. For example, texts regarding the evaluation of the generative model and its training data can be collected from evaluation sites, blogs, bulletin boards, and other word-of-mouth websites via web crawlers, and the evaluation can be scored based on the collected texts.
[0178] Furthermore, "vulnerability" can be used as an indicator to assess the degree to which an entity is vulnerable to attack, for example. For instance, vulnerabilities may be evaluated on a three-level scale: "low," "medium," and "high." As an example, vulnerabilities may be evaluated based on criteria such as "resistance to attack," "implementation of countermeasures," "establishment of guardrails," and "presence or absence of past attack detection history." Here, "guardrails" refers to any system that protects against internal and external security threats and prevents the unauthorized disclosure of confidential information and the spread of misinformation.
[0179] Figure 26 illustrates evaluation metrics for "reliability" and "vulnerability," but the number of evaluation metrics is not limited to these. For example, evaluation metrics that assess compatibility with other generative models may be included. Such evaluation metrics may be used in the construction of AI collaborative systems. For example, they can prevent generative models that are prone to unintentional attacks, or generative models that are prone to attacks, from being assigned to simulate participating agents in the same discussion. In addition, even if an evaluation item falls under the category of "reliability," if it is a useful evaluation item, it may be added to a separate column from "reliability," such as whether the generative model is OSS (Open Source Software) or proprietary, or whether the training data of the generative model infringes on copyright law or personal information protection law. Here, "OSS" refers to software whose source code is publicly available and can be freely modified and redistributed by anyone free of charge, while "proprietary" refers to software whose source code, specifications, standards, structure, etc., are not publicly available.
[0180] Furthermore, Figure 26 shows an example of evaluating "reliability" and "vulnerability" indicators at three levels, but naturally, they can be evaluated at any number of levels, and may also be expressed as numerical scores. In addition, the "reliability" and "vulnerability" indicators may be represented by icons or by colors such as in a heat map.
[0181] Furthermore, while Figure 26 illustrates the attribute information DB13D in table format, the relationships between each generative model and evaluation indicators may also be represented by charts and graphs, such as bar graphs, pie charts, or line graphs. In particular, when displaying the compatibility between generative models as two-dimensional or three-dimensional graph data, the connection relationships indicating compatibility can be displayed with edges, and the degree of compatibility can be expressed by the thickness and color of each edge.
[0182] Under this attribute information DB13D, the activation unit 184 can increase the frequency of activating the detection unit 182 as the reliability associated with the ID of the generative model simulating a participating agent making a new statement decreases, while decreasing the frequency of activating the detection unit 182 as the reliability increases. For example, in the example shown in Figure 26, the activation unit 184 can also narrow down the activation of the detection unit 182 to cases where the reliability associated with the ID of the generative model simulating a participating agent making a new statement is "low".
[0183] Furthermore, the activation unit 184 can increase the frequency of activating the detection unit 182 as the vulnerability associated with the ID of the generative model simulating a new participant agent increases, while decreasing the frequency of activating the detection unit 182 as the vulnerability decreases. For example, in the example shown in Figure 26, the activation unit 184 can also narrow down the activation of the detection unit 182 to cases where the vulnerability associated with the ID of the generative model simulating a new participant agent is "medium" or higher.
[0184] Furthermore, reliability and vulnerability can be used to set thresholds in each of the detection methods described in (1) to (4) above, or to add or subtract points from the sentiment analysis values, similarity, and the distribution of sentiment analysis values. Alternatively, when combining (1) to (4) in the detection method described in (5) above, they can be used to weight each of the detection results described in (1) to (4) above.
[0185] <Example 1 of attribute information application> In addition, Figure 26 shows an example in which the data items of attribute information DB13D are managed as a database, but it is also possible to display the data items of attribute information DB13D side by side on a GUI screen.
[0186] Figure 27 is Figure (1) showing an example of an AI management screen. Figure 27 shows an example where the menu tab 310 for the data item "Reliability" among the data items of the attribute information DB 13D shown in Figure 26 is selected on the AI management screen 300 where the AI being linked by the AI linkage unit 16 is managed.
[0187] In this case, as shown in Figure 27, accordions 310A to 310M are arranged as merely one example of GUI components that implement a folding function to switch the expansion or collapse of information for each evaluation item, such as "location," "developer," "whether or not learning data is made public," "whether or not hallucination countermeasures are taken," ..., and "reputation, reputation, and level of evaluation."
[0188] For example, in the example shown in Figure 27, the accordion 310A for "Location," indicated by hatching, is expanded within the evaluation items of the data item "Reliability." When the accordion 310A is expanded in this way, a map is displayed with the addresses of suppliers A to D plotted, as merely one example of the "locations" of suppliers A to D of the AI being linked. For example, in the example shown in Figure 27, supplier A's location is plotted in "Tokyo," supplier B's location is plotted in "New York," supplier C's location is plotted in "New Delhi," and supplier D's location is plotted in "London." This allows users to confirm the locations of suppliers A to D of the AI being linked. Note that the location display shown in Figure 27 is merely an example; as mentioned above, an area may be displayed, or latitude and longitude may be displayed. Furthermore, it goes without saying that the map shown in Figure 27 can be zoomed in or out using the zoom-in and zoom-out buttons or by adjusting the slider bar, and the area displayed on the map can be arbitrarily specified. Also, although Figure 27 shows an example where the supplier's location is displayed, it is not limited to this, and the location of server equipment or data centers that provide the software, models, or data may also be displayed.
[0189] <Example of Attribute Information Application 2> Figure 28 is Figure (2) showing an example of an AI management screen. Figure 28 shows an example where the menu tab 320 for the data item "Vulnerability" among the data items of the attribute information DB 13D shown in Figure 26 is selected on the AI management screen 300 where the AI being linked by the AI linkage unit 16 is managed.
[0190] In this case, as shown in Figure 28, accordions 320A to 320N are arranged as merely one example of GUI components that realize a folding function to switch the expansion or collapse of information for each evaluation item, such as "resistance to attack," "implementation of countermeasures," "presence or absence of guardrails," ..., "presence or absence of detection history of being used in past attacks."
[0191] For example, in the example shown in Figure 28, the accordion 320A for "Resistance to Attack," indicated by hatching, is expanded within the evaluation items of the data item "Vulnerability." When accordion 320A is expanded in this way, it displays a score normalized to a numerical range of 0 to 100 for the "Resistance to Attack" of each of the AI suppliers A to D in the collaboration, along with a bar graph corresponding to that score. This allows for the visualization of the attack resistance of the AI provided by suppliers A to D.
[0192] <Example 3 of Attribute Information Application> Furthermore, the data items in the attribute information DB13D shown in Figure 26 are merely examples. As another example, the data items included in the attribute information DB13D may be constructed according to a framework called AI TRiSM. AI TRiSM is an acronym for AI (Artificial Intelligence), T (Reliability), Ri (Risk), and SM (Security Management).
[0193] AI TRiSM may include four components: "Explainability," "ModelOps," "AI Security," and "Privacy." For example, "Explainability" refers to the ability of AI to explain the basis of its decisions in a way that humans can understand. "ModelOps" refers to the processes for properly developing, deploying, monitoring, and maintaining AI. "AI Security" refers to measures to protect AI and the data it handles from external attacks. "Privacy" refers to the protection of personal information.
[0194] Next, we will explain the types of AI challenges classified by the AI TRiSM framework and their relationship to the components of AI TRiSM that correspond to those AI challenges. Below are the classifications of AI challenges in AI TRiSM and the types of AI challenges listed in each classification. T (Trustworthiness): Accuracy, Robustness, Explainability, Transparency, Misinformation, Fairness, Bias Ri (Risk): Copyright, Intellectual Property, Privacy, Personal Information SM (Security Management): Data Disclosure, Adversarial Attacks, Misuse
[0195] Of these, within the "reliability" category, the issue of "accuracy and robustness" can be addressed by "ModelOps," the issue of "explainability" can be addressed by "explainability," the issue of "transparency" can be addressed by "explainability," the issue of "misinformation" can be addressed by "ModelOps," and the issue of "fairness and bias" can be addressed by "ModelOps."
[0196] Furthermore, within the "Risk" category, issues of the "Privacy / Personal Information" type can be addressed by "Privacy." In addition, within the "Security" category, issues of the "Information Leakage" type can be addressed by "AI Security," and issues of the "Adversarial Attack / Abuse" type can be addressed by "AI Security."
[0197] <List Display> For example, the information display unit 17 can also display a list of input / output logs for each round in the input / output log DB 13B, detection results in the detection history DB 13C, and action candidates that the action unit 19 can execute.
[0198] Figure 29 shows an example of a list display. Figure 29 shows the output log, i.e., the utterance log, for each round when two agents, Agent A1 and Agent A2, role-play a discussion on the topic "Regarding the increase in the consumption tax rate" using a round-robin method. As shown in Figure 29, the detection results are displayed associated with utterances T11 to T13 in each round. These detection results may include items such as the degree of aggression, the degree of being attacked, and the detection method used to identify the utterance as aggressor or attacked. "Degree of aggression" may refer to the likelihood that the utterance is a malicious input. For example, the degree of aggression and the degree of being attacked may be calculated based on the degree of deviation from thresholds of analysis values such as sentiment analysis value, similarity, and entropy of the distribution of sentiment analysis value, using the detection method determined to identify the utterance as aggressor or attacked. In the case of keyword search, the degree of aggression and the degree of being attacked may be calculated to be higher as the number of NG words increases. Furthermore, the suggested countermeasures are displayed associated with utterances T11 to T13 in each round. For example, as examples of possible actions, toggle buttons are displayed to specify whether to perform action 1, action 2, action 3, and whether to ignore the action. This list display of the message log, detection results, and possible actions improves the readability of the information.
[0199] Figure 30 shows a modified version of the list display. In Figure 30, the display of the message log is omitted from the list display shown in Figure 29, and instead of the suggested actions, the destination for changing the detection result is displayed. Furthermore, Figure 30 shows radio buttons as an example of destinations for changing the detection result, allowing the user to select one of four destinations: "Attack occurred," "Attacked," "Suspected," and "No problem." Through the selection of such radio buttons, the system can accept input for modifying the detection result. In addition, in Figure 30, the display content of the detection method in the list display shown in Figure 29 has been changed from a format that displays the detection method in which the occurrence of an attack was detected from among "keyword," "sentiment analysis," and "similarity" to a format that displays the score calculated for each detection method of "keyword," "sentiment analysis," and "similarity," for example, the probability that an attack has occurred. Note that although the display of suggested actions shown in Figure 29 is omitted in Figure 30, suggested actions may also be linked and displayed.
[0200] <Exhibition of Creative Ability> The matters described in this embodiment, such as specific examples of generative models, roles, prompts, and types of LLMs, are merely examples and can be changed. Furthermore, the flowchart described in this embodiment can also be modified, with changes to the processing order or skipping of some processes, within a consistent range.
[0201] (1) Scale In the above embodiment, similarity was given as an example of an index for evaluating the relationship between multiple things, such as vectors and statements. However, the relationship between multiple things may be expressed by other scales, and it is possible to evaluate whether any scale satisfies any criterion. Here, "scale" refers to an index for measuring the degree of things, and "criterion" refers to the condition for comparison with a certain index. For example, in addition to similarity, criteria such as thresholds may be used with other scales such as deviation, distance, dissimilarity, and correlation coefficient.
[0202] (2) Input Data In the above embodiment, an example was given in which input data including the agenda was input to the generative model as a trigger for each agent's statement, but the trigger for starting a discussion is not limited to the input of an agenda. For example, each agent's statement can be started by inputting any string such as a comment, keyword, or string of alphanumeric characters that serves as a trigger to start a discussion.
[0203] (3) Generative Models The “generative model” relating to this disclosure encompasses all generative models that generate output data corresponding to input data. For example, LLMs generated by model merging, which fuses multiple LLMs by user request or automation, or by learning transfer, which transfers learning information from one LLM to another, are examples of generative models.
[0204] In other words, it is becoming common to use "domain-specific models" that have been further trained (fine-tuned) using domain-specific datasets based on a foundational model. When this foundational model is updated, it becomes necessary to retrain the domain-specific model as well. In such retraining of domain-specific models, based on the insight that the learning processes of different models can be approximately identical under the symmetry of neuron replacement called "substitution transformation," the learning results of the new model can be executed at low cost by appropriately transforming the parameter sequences of past learning processes.
[0205] A series of models, including these foundational models, domain-specific models generated by additional training on the foundational models, updated foundational models, and domain-specific models generated by learning transfer to domain-specific models, can be considered examples of generative models.
[0206] <System> Unless otherwise specified, the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above document and drawings may be changed at will. For example, one or more of the AI cooperation unit 16, information presentation unit 17, AI behavior detection unit 18, and response unit 19 of the control unit 15 of the information processing device 10 may be distributed and located in separate devices. Furthermore, one or more of the operation information DB 13A, input / output log DB 13B, detection history DB 13C, and attribute information DB 13D of the storage unit 13 of the information processing device 10 may also be distributed and located in separate devices.
[0207] (1) AI Collaboration In the above embodiment, an example was given in which generation models running on the information processing device 10 are linked, but the entity executing the generation model is not limited to the information processing device 10, and may be any computer.
[0208] Figure 31 shows examples of AI integration applications. Figure 31 illustrates three examples of external servers 50 as shown in Figure 1: an internal server 50A, an external company server 50B, and an external company server 50C. Of these, the internal server 50A is a server device operated within the same company as the provider of the AI integration function. The external company server 50B is a server device operated on-premise by a different company than the one providing the AI integration function. Furthermore, the external company server 50C is a server device operated in the cloud by a different company than the one providing the AI integration function.
[0209] As shown in Figure 31, the AI collaboration unit 16 can link the generation models K1 to Kn running on the information processing device 10, the generation models A1 to An running on the in-house server 50A, the generation models B1 to Bn running on the external company server 50B, and the generation models C1 to Cn running on the external company server 50C. This makes it possible to construct an AI collaboration system in which multiple generation models running in a distributed manner on multiple execution entities, regardless of whether they are internal or external, on-premise, or in the cloud.
[0210] (2) Hybrid In the above embodiment, an example was given in which the information processing device 10 is implemented in either an on-premises or cloud configuration, but it may also be implemented in a hybrid cloud.
[0211] Figure 32 is a schematic diagram illustrating an example of hybrid cloud implementation. Figure 32 shows two patterns, Pattern 1 and Pattern 2, as examples of patterns in which hybrid cloud can be implemented.
[0212] For example, in the case of Pattern 1 shown in Figure 32, of the four functional units of the control unit 15 of the information processing device 10 shown in Figure 1—the AI collaboration unit 16, the information presentation unit 17, the AI behavior detection unit 18, and the response unit 19—the AI collaboration unit 16 and the information presentation unit 17 are deployed to the information processing device 10A which is operated on-premise, while the AI behavior detection unit 18 and the response unit 19 are deployed to the information processing device 10B which is operated in the cloud.
[0213] Furthermore, in the example of Pattern 2 shown in Figure 32, of the four functional units of the control unit 15 of the information processing device 10 shown in Figure 1—the AI collaboration unit 16, the information presentation unit 17, the AI behavior detection unit 18, and the response unit 19—the AI collaboration unit 16 and the information presentation unit 17 are deployed to the information processing device 10A which is operated on-premise, the AI behavior detection unit 18 is deployed to the information processing device 10B which is operated in the cloud, and the response unit 19 is deployed to the information processing device 10C which is operated in the cloud.
[0214] Note that Figure 32 illustrates two patterns, Pattern 1 and Pattern 2, as examples only, but the AI collaboration unit 16, information presentation unit 17, AI behavior detection unit 18, and response unit 19 may be deployed in patterns other than these. For example, from the perspective of load balancing using a load balancer, one or more of the AI collaboration unit 16, information presentation unit 17, AI behavior detection unit 18, and response unit 19 may be operated in parallel. As an example only, a pattern may be included in which three AI behavior detection units 18 and three response units 19 are deployed, and processing is assigned to one of the three AI behavior detection units 18 and three response units 19.
[0215] Furthermore, the components of each illustrated device are functionally conceptual and do not necessarily need to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown. That is, all or part of them can be functionally or physically distributed and integrated in any units according to various loads and usage conditions. Note that each configuration may also be a physical configuration.
[0216] Furthermore, the processing performed by the illustrated apparatus can be implemented, in whole or in part, by a program executed by a hardware processor such as an MPU (Micro-Processing Unit) or CPU (Central Processing Unit), or by hardware using wired logic.
[0217] Furthermore, the "attack (adversarial input)" listed as one of the "inputs that cause abnormalities" above may include those disclosed in the following literature, and the detection and response functions described in the above embodiment can be effectively used for attacks disclosed in the following literature as well. Reference 1: Goodfellow, Ian J. et al. “Explaining and Harnessing Adversarial Examples.” CoRR abs / 1412.6572 (2014): n. pag. Reference 2: A. Wei, et al., Jailbroken: How Does LLM Safety Training Fail?, NeurIPS, 2023. Reference 3: B. Chen, et al., Understanding Multi-Turn Toxic Behaviors in Open-Domain Chatbots, RAID, 2023. Reference 4: Gelei Deng et al., MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots (NDSS'24)
[0218] <Hardware> Next, an example of the hardware configuration of the information processing device 10 described in this embodiment will be explained. For example, it can be implemented by installing a program that realizes the functions of the information processing device 10 on a computer. For example, by having the computer run the above program, which is provided as packaged software or online software, the computer can be made to function as the information processing device 10. The computer referred to here includes desktop or notebook personal computers, rack-mounted server computers, etc. In addition, the computer category also includes smartphones, mobile phones and PHS (Personal Handyphone System) and other mobile communication terminals, as well as PDAs (Personal Digital Assistants). Furthermore, the functions of the information processing device 10 may be implemented on a cloud server.
[0219] An example of a computer that executes the above-mentioned program (detection program or response program) will be explained using Figure 33. Figure 33 is a diagram showing an example of hardware configuration. As shown in Figure 33, the computer 1000 has, for example, memory 1010, CPU 1020, hard disk drive interface 1030, disk drive interface 1040, serial port interface 1050, video adapter 1060, and network interface 1070. These components are connected by a bus 1080.
[0220] Memory 1010 includes ROM (Read Only Memory) 1011 and RAM (Random Access Memory) 1012. ROM 1011 stores, for example, a boot program such as BIOS (Basic Input Output System). The hard disk drive interface 1030 is connected to the hard disk drive 1090. The disk drive interface 1040 is connected to the disk drive 1100. The disk drive 1100 is used to insert a removable storage medium, such as a magnetic disk or an optical disk. The serial port interface 1050 is used to connect, for example, a mouse 1110 and a keyboard 1120. The video adapter 1060 is used to connect, for example, a display 1130.
[0221] Here, as shown in Figure 33, the hard disk drive 1090 stores, for example, the OS 1091, the application program 1092, the program module 1093, and the program data 1094. The storage unit 13 described in the above embodiment is equipped, for example, in the hard disk drive 1090 or the memory 1010.
[0222] Then, the CPU 1020 reads the program module 1093 and program data 1094 stored in the hard disk drive 1090 into the RAM 1012 as needed and executes the above-described procedures.
[0223] Furthermore, the program module 1093 and program data 1094 related to the above-mentioned detection program or the above-mentioned countermeasure program are not limited to being stored in the hard disk drive 1090, but may also be stored in a removable storage medium and read by the CPU 1020 via a disk drive 1100 or the like. Alternatively, the program module 1093 and program data 1094 related to the above-mentioned program may be stored in another computer connected via a network such as a LAN or WAN (Wide Area Network) and read by the CPU 1020 via a network interface 1070.
[0224] The following additional information is disclosed regarding the embodiments described above.
[0225] (Note 1) An information processing device comprising: a memory; and at least one processor connected to the memory, wherein the processor obtains a log of output data obtained by performing one or more input / output controls to input first output data of a first generation model among a plurality of generation models to a second generation model and cause the second generation model to output second output data, and detects an input that causes an anomaly based on the log of output data.
[0226] (Appendix 2) A non-temporary storage medium storing a program executable by a computer to perform a detection process, wherein the detection process obtains a log of output data obtained by performing one or more input / output controls to input first output data of a first generation model among a plurality of generation models into a second generation model and cause the second generation model to output second output data, and detects an input causing an abnormality based on the log of output data.
[0227] 10 Information Processing Device 11 Communication Control Unit 13 Storage Unit 13A Operation Information DB 13B Input / Output Log DB 13C Detection History DB 13D Attribute Information DB 15 Control Unit 16 AI Collaboration Unit 161 Initialization Unit 162 Exclusion Unit 163 Generation Unit 164 Deletion Unit 165 Rewind Unit 17 Information Presentation Unit 18 AI Behavior Detection Unit 181 Analysis Unit 181A Keyword Search Unit 181B Sentiment Analysis Unit 181C Similarity Analysis Unit 182 Detection Unit 183 Identification Unit 184 Activation Unit 19 Response Unit 191 Response Decision Unit 192 Response Instruction Unit 30 User Terminal 50 External Server
Claims
1. An information processing device characterized by comprising: an acquisition unit that acquires a log of output data obtained by performing one or more input / output control operations that input first output data of a first generation model among a plurality of generation models into a second generation model and cause the second generation model to output second output data; and a detection unit that detects an input causing an abnormality based on the log of output data.
2. The information processing apparatus according to claim 1, characterized in that the acquisition unit acquires a log of output data obtained by repeatedly performing input / output control, which involves inputting input data including the first output data of a first generation model previously selected from among the plurality of generation models into a second generation model selected from among the plurality of generation models, and causing the second generation model to output second output data.
3. The information processing apparatus according to claim 2, characterized in that the log of output data is a log of comments on the agenda output by the second generative model, which has input data including the agenda and the first output data.
4. The information processing apparatus according to claim 2, wherein the detection unit detects a generative model among the plurality of generative models in which the anomaly is caused by input from another generative model, based on at least one of the following: the presence or absence of a specific keyword in the log of the output data, the difference in sentiment analysis values between output data of the same generative model, the difference in a predetermined scale between output data of the same generative model, and the difference in the distribution of sentiment analysis values between output data of the same generative model.
5. The information processing apparatus according to claim 2, characterized in that the detection unit detects a generative model that generates output data causing the anomaly in other generative models, based on at least one of the following: the presence or absence of a specific keyword in the log of the output data, the degree to which the sentiment analysis value in the log of the output data approaches a neutral level, and a predetermined scale between the log of the output data and the agenda included in the input data.
6. The information processing apparatus according to claim 1, further comprising a handling unit that performs a response to the input causing the aforementioned abnormality.
7. The information processing apparatus according to claim 1, further comprising a display unit that displays the detection result of an input causing an abnormality by the detection unit.
8. An information processing device characterized by comprising: an acquisition unit that acquires a log of output data obtained by inputting input data including output data of a generation model in a first role and setting information of a second role into the generation model and performing one or more input / output control operations to cause the generation model to output output data in the second role; and a detection unit that detects an input causing an abnormality based on the log of output data.
9. A detection method performed by an information processing device, comprising: an acquisition step of acquiring a log of output data obtained by performing one or more input / output control operations, inputting first output data of a first generation model among a plurality of generation models into a second generation model and causing the second generation model to output second output data; and a detection step of detecting an input that causes an anomaly based on the log of output data.
10. A detection program that causes a computer to perform the following steps: an acquisition step of acquiring a log of output data obtained by performing one or more input / output control operations, in which the first output data of a first generation model among multiple generation models is input to a second generation model and the second generation model outputs second output data; and a detection step of detecting an input that causes an anomaly based on the log of output data.