Information processing device, information processing method, and information processing program

The information processing device optimizes AI model usage by caching API requests, addressing inefficiencies in conventional automated response services, and enhancing operational efficiency and user data protection.

JP7734787B2Active Publication Date: 2025-09-05PAYPAY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024075359
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-05-07
Publication Date
2025-09-05
Estimated Expiration
2043-10-16

AI Technical Summary

Technical Problem

Conventional automated response services do not efficiently utilize models that generate sentences (automatic generation AI), leading to inefficient operation due to charging based on the amount and number of prompts.

Method used

An information processing device that includes a receiving unit, determining unit, generating unit, and executing unit to manage and cache API requests generated by a trained model, optimizing the use of AI models to reduce fees and enhance efficiency.

Benefits of technology

Enables efficient operation of AI services by caching API requests, reducing model usage frequency, and providing appropriate content to users while protecting user information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007734787000001
    Figure 0007734787000001
  • Figure 0007734787000002
    Figure 0007734787000002
  • Figure 0007734787000003
    Figure 0007734787000003
Patent Text Reader

Abstract

To provide an information processing device for efficiently operating services using automatically generated AI.SOLUTION: An information processing device includes a receiving unit, a determination unit, a generation unit, and an execution unit. The receiving unit receives a sentence in natural language representing an execution instruction input by a user. The determination unit determines whether or not execution information for executing a process corresponding to the execution instruction of the sentence received by the receiving unit is cached. When the determination unit determines that the execution information is not cached, the generation unit generates the execution information using a model that has been trained to generate an answer to a question. The execution unit executes the process using the cached execution information or the execution information generated by the generation unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device, an information processing method, and an information processing program. [Background technology]

[0002] Conventionally, automated response services such as chatbots have become popular. For example, a technology has been proposed for such services that extracts questions and their answers similar to a question entered by a user from a database and provides the extracted answers to the user (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2019 / 185578 Summary of the Invention [Problem to be solved by the invention]

[0004] However, conventional technologies have not taken into consideration the use of models that generate sentences (so-called automatic generation AI). For example, when using such models, fees are charged depending on the amount and number of prompts, so efficient operation is required.

[0005] The present invention has been made in consideration of the above, and aims to provide an information processing device, an information processing method, and an information processing program that can efficiently operate services using automatically generated AI. [Means for solving the problem]

[0006] In order to solve the above-mentioned problems and achieve the object, an information processing device according to the present invention includes a receiving unit, a determining unit, a generating unit, and an executing unit. The receiving unit receives a natural language sentence indicating an execution instruction input by a user. The determining unit determines whether execution information for executing a process corresponding to the execution instruction of the sentence received by the receiving unit is cached. If the determining unit determines that the execution information is not cached, the generating unit generates the execution information using a model trained to generate an answer to a question. The executing unit executes the process using the cached execution information or the execution information generated by the generating unit. [Effects of the Invention]

[0007] According to the present invention, it is possible to provide appropriate content to a user. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram illustrating an example of information processing according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of a user interface according to the embodiment. [Figure 3] FIG. 3 is a block diagram illustrating an example of the configuration of the information processing device according to the embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of information stored in a prompt dictionary storage unit according to the embodiment. [Figure 5] FIG. 5 is a diagram illustrating an example of information stored in an API request storage unit according to the embodiment. [Figure 6] FIG. 6 is an explanatory diagram of the process of generating an extended prompt according to the embodiment. [Figure 7] FIG. 7 is an explanatory diagram of the process of generating an extended prompt according to the embodiment. [Figure 8] FIG. 8 is an explanatory diagram of the process of generating an extended prompt according to the embodiment. [Figure 9] FIG. 9 is an explanatory diagram of the process of generating an extended prompt according to the embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of an extended prompt according to the embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of a response to an extended prompt according to the embodiment. [Figure 12] FIG. 12 is a diagram illustrating an example of the answer screen according to the embodiment. [Figure 13] FIG. 13 is a flowchart illustrating an example of the providing process according to the embodiment. [Figure 14] FIG. 14 is a hardware configuration diagram showing an example of a computer that realizes the functions of the information processing device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, an information processing device, an information processing method, and an information processing program according to the present application will be described in detail with reference to the drawings. Note that the information processing device, the information processing method, and the information processing program according to the present application are not limited to the embodiments.

[0010] [Embodiment] [1. Information Processing] First, an example of information processing according to the embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of information processing according to the embodiment. Note that the information processing according to the embodiment is realized by an information processing device 1 shown in Fig. 1.

[0011] 1 is a server device operated by an electronic payment service for electronic money payments for each user. In this embodiment, the information processing device 1 provides an automatic response service that generates answers to questions posed to the user.

[0012] For example, the information processing device 1 can accept questions about the amount of electronic money used, questions about coupons and stores, questions about person-to-person electronic money transfers, etc. For example, these questions require executing one of a plurality of APIs (Application Programming Interfaces) to acquire data. In this embodiment, as will be described later, an API request (a query for executing an API) is generated using a second model M2. Note that the second model M2 is a GPT (Generative Pre-trained Transformer) model.

[0013] The user terminal 100 shown in Fig. 1 is an information processing device used by a user. The user terminal 100 is realized, for example, by a smartphone, a tablet terminal, a notebook PC (Personal Computer), a desktop PC, a mobile phone, a PDA (Personal Digital Assistant), or the like. The user terminal 100 displays information distributed by the information processing device 1 and the server device 200 using a web browser or an application. In the example shown in Fig. 1, the user terminal 100 is a smartphone.

[0014] 1 is an information processing device that provides electronic payment services related to electronic payments using the user terminal 100 and performs various payments. The server device 200 manages the accounts of providers (businesses) of transaction objects and users to whom transaction objects are provided, and realizes various payments by transferring electronic money between accounts in accordance with payment requests from users.

[0015] Here, prior to the provision process executed by the information processing device 1, an example of payment (electronic payment) using the user terminal 100 will be described. Note that the following description will be given of an example in which a user makes a payment using the user terminal 100 using a two-dimensional code (QR code (registered trademark)) placed in store A, which indicates store identification information that identifies store A, but the embodiment is not limited to this. The example of payment described below can also be applied to a case in which an arbitrary user makes a payment at an arbitrary store using an arbitrary user terminal 100. Furthermore, the store identification information may be not only a QR code, but also a barcode, a predetermined mark, a number, or the like.

[0016] For example, when a user makes a payment for using or purchasing a payment object (transaction object) such as various products or services at store A, the user launches a payment app pre-installed on user terminal 100. Then, the user photographs store identification information installed in store A via the payment app. In such a case, user terminal 100 displays a screen for inputting the price of the payment object and accepts input of the payment amount from the user or a store clerk at store A. Then, user terminal 100 transmits to server device 200 user identification information that identifies the user, store identification information (or information indicated by the store identification information, i.e., information indicating store A (or the operator of store A) (e.g., store ID)), and payment information indicating the payment amount.

[0017] In such a case, the server device 200 transfers electronic money in the amount indicated by the payment amount from the user's account indicated by the user identification information to the account of store A indicated by the store identification information. Then, the server device 200 transmits a notification that the payment has been completed to the user terminal 100. In such a case, the user terminal 100 notifies the user that the payment has been made using electronic money by outputting a screen or a predetermined sound indicating that the payment has been completed.

[0018] A more detailed example will be described. For example, the store identification information installed in store A is a URL set for each store, and is linked to group identification information indicating a group to which store A belongs and group store identification information identifying store A within the group, and is managed so that the server device 200 can refer to the URL. The URL serving as the store identification information is a URL for accessing the server device 200. When the user terminal 100 photographs the store identification information, it accesses the URL indicated by the photographed store identification information and transmits the user identification information. In such a case, the server device 200 identifies the group identification information corresponding to the accessed URL and identifies the electronic money account (sometimes referred to as a "wallet") associated with the identified group identification information. Next, the server device 200 displays an amount input screen on the user terminal 100 and prompts the user to input an amount. The server device 200 then identifies the group identification information from the wallet associated with the user identification information received from the user terminal 100, and transfers the input amount of electronic money to the wallet associated with the identified group identification information. The server device 200 may transfer the electronic money to a wallet linked to the group identification information and the group store identification information.

[0019] It should be noted that payment using the user terminal 100 is not limited to the above-described process. For example, payment using the user terminal 100 may be made using a store terminal installed in store A. For example, the user terminal 100 displays user identification information for identifying the user on a screen. In such a case, the store terminal installed in store A reads the user identification information displayed on the user terminal 100 and transmits payment information indicating the user identification information (or information indicated by the user identification information, i.e., information indicating the user (e.g., user ID)), the payment amount, and information identifying store A) to the server device 200. In such a case, the server device 200 may transfer electronic money in the amount indicated by the payment amount from the user's account indicated by the user identification information to store A's account, and notify the store terminal of store A or the user terminal 100 that the payment has been completed by outputting a screen or a predetermined sound indicating that the payment has been completed.

[0020] More specifically, the user terminal 100 sends a payment request to the server device 200 along with user identification information. In such a case, the server device 200 generates a one-time code, associates the generated one-time code with the user identification information, and transmits the one-time code to the user terminal 100. The user terminal 100 then displays the one-time code (i.e., information that identifies the user) on its screen. In such a case, the store terminal reads the one-time code displayed on the user terminal 100 and transmits the read one-time code, group identification information, group store identification information, and payment amount to the server device 200. The server device 200 then transfers electronic money equivalent to the payment amount from the wallet associated with the user identification information associated with the one-time code to the wallet associated with the group identification information and group store identification information.

[0021] Furthermore, payment using the user terminal 100 may not only be a process of transferring electronic money from an account to which the user has previously charged electronic money to an account at store A, but may also be a payment using a credit card that the user has previously registered. In such a case, for example, the user terminal 100 may transfer the electronic money of the payment amount to the account at store A, and may also bill the operating company (card company) of the user's credit card for the payment amount.

[0022] Recently, a model trained to generate answers to questions (so-called auto-generating AI) has been attracting attention. In this embodiment, such a model is utilized to provide an automatic response service.

[0023] In such an automatic response service, when it is necessary to execute a process to respond to a question from a user, it is necessary to convert the text entered by the user into a format compatible with the corresponding API.

[0024] In this embodiment, the above format conversion is performed using an automatically generated AI (first model and second model, which will be described later).

[0025] 1, the information processing device 1 receives a question from a user via the user terminal 100 (step S1). The question is a sentence in a natural language, and the information processing device 1 can receive the question via a UI, which will be described later with reference to FIG.

[0026] The information processing device 1 generates a prompt to be input to the first model (step S2). The first model is a model such as GPT or BERT (Bidirectional Encoder Representation from Transformers). The first model is an example of an internal model.

[0027] In this embodiment, the information processing device 1 performs prompt engineering using the first model to generate an extended prompt from a question from a user. Note that prompt engineering is a process for optimizing the content of a prompt.

[0028] As will be described later, the information processing device 1 generates an extended prompt by inputting a prompt into a first model. That is, in this embodiment, a desired answer can be obtained from a second model M2, which will be described later, by performing prompt engineering using the first model. The generation of an extended prompt will be described later with reference to FIGS. 6 to 10.

[0029] Next, the information processing device 1 inputs the generated extension prompt to a second model M2 external to the information processing device 1 (step S3), and acquires an answer corresponding to the extension prompt from the second model M2 (step S4). The above-mentioned extension prompt includes an instruction sentence that causes the second model M2 to generate an answer in an API format, and the answer acquired from the second model M2 is in, for example, a JSON format.

[0030] In this way, the information processing device 1 can convert the natural language sentence by the user into the JSON format, which is the API format, using the second model M2. That is, the information processing device 1 can obtain the answer of the second model M2 to the extended prompt as an API request.

[0031] Then, the information processing device 1 registers the answer (API request) acquired from the second model M2 in a cache (step S5). For example, the information processing device 1 tags the answer acquired from the second model M2 with the question from the user and registers the answer in the cache. For example, the information processing device 1 extracts words from the question entered by the user by morphological analysis or the like, and registers these words as tags in the cache, linking them to the answer acquired from the second model M2. Note that the information processing device 1 may omit the process of registering the answer in the cache if an answer corresponding to a past tag that corresponds to the current tag has already been registered in the cache.

[0032] Thereafter, the information processing device 1 generates an API request based on the response acquired from the second model M2 (step S6). The API request is an example of execution information, and includes various parameters for executing the API.

[0033] Next, the information processing device 1 executes the API by providing an API request to the server device 200 (step S7). As a result, the server device 200 outputs content according to the API request to the information processing device 1 (step S8).

[0034] At this time, the information processing device 1 assigns an identifier (e.g., a user ID) of the user who is making the inquiry to the API request and provides the request to the server device 200. In other words, the information processing device 1 generates an API request without passing on the user's identification information to the external second model M2. This allows the information processing device 1 to protect the user information.

[0035] Then, the information processing device 1 provides a response to the user terminal 100 based on the content acquired from the server device 200 (step S9). A specific example of the response that the information processing device 1 provides to the user will be described later with reference to FIG.

[0036] In this way, the information processing device 1 according to the embodiment uses the first model to generate an extended prompt from a question from a user, and inputs the generated extended prompt into the second model M2 to obtain an answer to the API request from the second model M2.

[0037] That is, by using a model such as GPT or BERT, the information processing device 1 can automatically convert a natural language sentence input by a user into an API request. Then, by executing a process corresponding to the API request, the information processing device 1 can appropriately execute a process corresponding to the natural language sentence input by the user.

[0038] Furthermore, when the information processing device 1 receives a new question, it generates a tag indicating the content of the question and searches the cache for a corresponding API request based on the tag. If the API request corresponding to the tag is cached, the information processing device 1 executes processing using the API request, and if the API request corresponding to the tag is not cached, the information processing device 1 generates an API request using the second model M2 based on the first prompt and the second prompt as described above, and executes processing.

[0039] In other words, the information processing device 1 can reduce the frequency of use of the second model M2 by linking the API request with a tag and registering it in the cache, thereby reducing the usage fee for the second model M2 and enabling efficient operation.

[0040] Next, a user interface for accepting a question from a user will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of a user interface according to an embodiment. The information processing device 1 accepts a question from a user through a payment application provided by the server device 200.

[0041] As shown in Figure 2, a QR code and selection buttons for selecting each application are displayed on the home screen HG of the payment app. For example, when the user selects selection button B, the screen transitions from the home screen HG to the chat screen CG.

[0042] As shown in Figure 2, the talk screen CG displays a tutorial area C that displays a tutorial on the function, a possible question Q where the user can select a question, and an input area A where the user can enter a question in text.

[0043] The information processing device 1 accepts a question in natural language from the user by the user selecting an expected question Q or inputting a sentence into the input area A. Then, the information processing device 1 provides an answer to the question in a chat format through a talk screen CG.

[0044] [2. Information Processing Device] Next, a configuration example of the information processing device 1 according to the embodiment will be described with reference to Fig. 3. Fig. 3 is a block diagram showing the configuration example of the information processing device 1 according to the embodiment. As shown in Fig. 3, the information processing device 1 includes a communication unit 2, a storage unit 3, and a control unit 4. Note that the information processing device 1 may also include an input unit (e.g., a keyboard or a mouse) that accepts various operations from an administrator who uses the information processing device 1, and a display unit (e.g., a liquid crystal display) that displays various information.

[0045] The communication unit 2 is realized by, for example, a network interface card (NIC), etc. The communication unit 2 is connected to a communication network such as 4G (4th Generation) or 5G (5th Generation) by wire or wirelessly, and transmits and receives information to and from each of the user terminal 100, the server device 200, etc. via the communication network.

[0046] The storage unit 3 is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 3 includes a prompt dictionary storage unit 31, an API request storage unit 32, and a first model storage unit 33.

[0047] The prompt dictionary storage unit 31 stores a prompt dictionary. The prompt dictionary is a dictionary related to the first prompt and the second prompt described above. Fig. 4 is a diagram showing an example of information stored in the prompt dictionary storage unit 31 according to the embodiment.

[0048] As shown in Fig. 4, the prompt dictionary storage unit 31 stores information on items such as "category question," "category," and "API question" in association with each other. A "category question" is a question used to ask the first model about a category corresponding to a user's question. Specific examples of category questions will be described later with reference to Fig. 6.

[0049] In the example shown in FIG. 4, all API questions are of the same category, "category question #1." However, the prompt dictionary storage unit 31 may store multiple patterns of category questions.

[0050] "Category" indicates the category of the sentence entered by the user. Specific examples of categories will be described later with reference to FIG. 6. "API question" is a sentence used to ask the first model the type of API to execute the process corresponding to the user's question or the type of value to be obtained. The API question also includes an answer format corresponding to each API. In the example of FIG. 4, the prompt dictionary storage unit 31 stores one API question in association with one category, but it may also store multiple API questions in association with one category.

[0051] 3, the API request storage unit 32 will be described. The API request storage unit 32 stores the API request generated by the second model M2. The API request storage unit 32 corresponds to an example of a cache.

[0052] 5 is a diagram illustrating an example of information stored in the API request storage unit 32 according to the embodiment. As illustrated in FIG. 5, the API request storage unit 32 stores information items such as "tag" and "API request" in association with each other.

[0053] "Tag" is a tag indicating the content of the question by the user. "API request" is an API request generated by the second model M2. In other words, every time a new API request is generated by the second model M2, the new API request is accumulated in the API request storage unit 32.

[0054] Returning to the description of FIG. 3 , the first model storage unit 33 will be described. The first model storage unit 33 stores the first model. As described above, the first model is a model trained to generate answers for natural language sentences, such as GPT or BERT. In this embodiment, the first model is a model used to perform the prompt engineering described above.

[0055] Next, we will explain the control unit 4. The control unit 4 is a controller, and is realized by, for example, a CPU (Central Processing Unit) or an MPU (Micro Processing Unit) executing various programs (corresponding to examples of information processing programs) stored in a storage device inside the information processing device 1 using RAM as a work area. The control unit 4 is also, for example, a controller, and is realized by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0056] As shown in Fig. 3, the control unit 4 includes a reception unit 41, an estimation unit 42, a determination unit 43, a generation unit 44, an execution unit 45, and a provision unit 46, and realizes or executes the functions and actions of the information processing described below. Note that the internal configuration of the control unit 4 is not limited to the configuration shown in Fig. 3, and may be other configurations as long as they perform the information processing described below. Furthermore, the connection relationship between the processing units included in the control unit 4 is not limited to the connection relationship shown in Fig. 3, and may be other connection relationships.

[0057] The receiving unit 41 receives a sentence in a natural language indicating an execution instruction input by a user. As described above, the receiving unit 41 receives a question document, which is a sentence related to a question in a natural language, from a user through a payment application provided by the server device 200.

[0058] The receiving unit 41 may receive instructions from the user to execute various services provided by the server device 200, such as instructions to remit electronic money or instructions to make a payment.

[0059] The estimation unit 42 estimates the content of the sentence received by the reception unit 41. For example, the estimation unit 42 extracts predetermined attribute words from the question sentence input by the user by natural language processing, and estimates the content of the sentence input by the user based on the extracted attribute words.

[0060] For example, suppose the question entered by the user is "Please tell me how much electronic money I have spent this month." In this case, the estimation unit 42 extracts "electronic money" and "amount spent" as attribute words.

[0061] The estimation unit 42 then associates the extracted attribute words with the question sentence as tags. Note that the meaning of words indicating a date, time, or period, such as "this month," "this week," or "today," changes depending on the date and time when the user inputs the sentence, whereas words such as "2023," "July 2023," or "July 1, 2023" can uniquely identify a date, time, or period.

[0062] Therefore, the estimation unit 42 may associate "2023," "July 2023," "July 1, 2023," etc. as tags with the question sentence, but may not associate "this month," "this week," "today," etc. as tags with the question sentence. In other words, the estimation unit 42 may not generate tags for words that include specific variables.

[0063] The estimation unit 42 may also estimate the content of the question sentence and generate tags using a model that generates tags from the question sentence. For example, such a model can be realized by BERT (Bidirectional Encoder Representation from Transformers) or the like.

[0064] The determination unit 43 determines whether or not execution information (corresponding to an API request) for executing a process corresponding to the execution instruction of the sentence accepted by the acceptance unit 41 is cached. Based on the content of the question sentence estimated by the estimation unit 42, the determination unit 43 determines whether or not an API request corresponding to the content of the question sentence is stored in the API request storage unit 32, which serves as a cache.

[0065] Specifically, the determination unit 43 searches the API request storage unit 32 based on the tag generated by the estimation unit 42, and determines whether the API request has been cached. For example, the determination unit 43 searches for an API request associated with the same tag as the tag generated by the estimation unit 42, and determines whether the API request has been cached.

[0066] If the determination unit 43 determines that the API request is not cached, it outputs an instruction to generate a prompt to the generation unit 44, and if it determines that the API request is cached, it passes the cached API request to the execution unit 45.

[0067] If the determining unit 43 determines that the API request is not cached, the generating unit 44 generates an API request using the second model M2 that has been trained to generate an answer to the question.

[0068] First, prior to generating an API request, the generation unit 44 performs prompt engineering to generate an extended prompt using the first model stored in the first model storage unit 33. Here, the process of generating an extended prompt will be described with reference to FIGS.

[0069] 6 to 9 are explanatory diagrams of the process of generating an extended prompt according to the embodiment. In the following, a case will be described in which the user's question sentence is "Did I spend more this month?"

[0070] For example, the prompt to be input to the first model includes the category of the user's question sentence described above, a question sentence inquiring about the type of value (Value as the Key) that will be the key for performing processing for the question, and an instruction sentence to output the answer in JSON format.

[0071] Then, the generation unit 44 inputs such a prompt into the first model and obtains the answer. For example, as shown in Fig. 7, it can be seen from the answer from the first model that the category ID is "0" and the value type is "Expenses (Spend data)".

[0072] Next, the generation unit 44 determines a prompt to be input to the first model based on the category ID and the type of value included in these answers. For example, the generation unit 44 refers to the prompt dictionary storage unit 31 and selects an API question sentence based on the category ID and the type of value.

[0073] For example, the generation unit 44 selects an API question sentence with a matching category ID from the prompt dictionary storage unit 31. The generation unit 44 generates a prompt to be further input to the first model from the selected API question sentence.

[0074] For example, the generation unit 44 generates the prompt shown in Fig. 8. In the prompt shown in Fig. 8, the question asked by the user (Did you spend more this month?) is replaced with a question from each user.

[0075] As shown in Figure 8, the prompt here includes a statement inquiring about the type of API to be executed ("normalized metric") and the values ​​to be obtained ("current period start datetime", "current period end datetime", "previous period start datetime", "previous period end datetime"), as well as an instruction statement to output the answer in JSON format.

[0076] The generation unit 44 inputs the prompt shown in Fig. 8 to the first model and obtains an answer from the first model. The first model outputs the answer shown in Fig. 9, for example, in response to such a prompt.

[0077] As shown in Figure 9, the answers from the first model include answers to the items specified by the prompt: “normalized metric,” “current period stat datetime,” “current period end datetime,” “previous period start datetime,” and “previous period end datetime.”

[0078] Then, the generation unit 44 generates an extended prompt based on these answers. Fig. 10 is a diagram showing an example of an extended prompt according to the embodiment. The generation unit 44 generates, for example, an extended prompt as shown in Fig. 10. The extended prompt is a prompt that is ultimately input to the second model M2 and is a prompt for generating an API request in the second model M2.

[0079] In the example shown in FIG. 10, the extended prompt includes a request to input items described in the request body and a request to describe a user response.

[0080] The generation unit 44 then inputs the extension prompt to the second model M2 and obtains a response to the extension prompt from the second model M2. The output result from the second model M2 will now be described with reference to Fig. 11. Fig. 11 is a diagram illustrating an example of a response to the extension prompt according to the embodiment.

[0081] 11, the second model M2 generates a response in JSON format that describes the API type, authentication, request body, and user response. In other words, the second model M2 outputs the API request as a response.

[0082] In this way, the generation unit 44 can generate an extended prompt in which the answer of the second model M2 becomes an API request by executing an engineering prompt using the first model based on the question sentence of the user. Therefore, in this embodiment, an appropriate API request can be generated from the question of the user.

[0083] Returning to the explanation of Fig. 3, the execution unit 45 will now be described. The execution unit 45 executes processing using cached execution information (corresponding to the API request) or execution information generated by the generation unit 44. If an API request corresponding to the execution instruction has been cached, the execution unit 45 extracts the corresponding API request from the API request storage unit 32, associates the API request with the user ID, and transmits the API request to the server device 200.

[0084] Furthermore, if the API request corresponding to the execution instruction has not been cached, the execution unit 45 associates the user ID with the API request generated by the generation unit 44 and transmits the request to the corresponding server device 200.

[0085] This allows the execution unit 45 to obtain the content corresponding to the API request from the server device 200.

[0086] The execution unit 45 then associates the tag generated by the estimation unit 42 with the API request newly generated by the generation unit 44 and registers the request in the API request storage unit 32 (corresponding to the cache).

[0087] As a result, when the information processing device 1 receives a new question text, if the content of the text is similar, it can acquire data by utilizing past API requests registered in the cache.

[0088] Although the case where all API requests are generated by the second model M2 has been described here, the present invention is not limited to this. For example, API requests generated in advance by an administrator or the like may be cached in advance.

[0089] Furthermore, although the case where the information processing device 1 acquires data through an API has been described, the present invention is not limited to this. For example, the information processing device 1 may cache a sentence or a URL that serves as a corresponding answer.

[0090] For example, in response to a question about "resetting a password," the information processing device 1 may return a URL of a page about "resetting a password."

[0091] The providing unit 46 generates an answer to the user's question and provides it to the user who has asked the question. For example, the providing unit 46 generates an answer to the question based on data acquired by the executing unit 45 from the API.

[0092] The providing unit 46 then provides the generated answer to the user through the talk screen CG (see FIG. 2). Note that the providing unit 46 may generate and provide an answer by graphing the data acquired from the API.

[0093] Here, an example of the answer screen provided by the providing unit 46 will be described with reference to Fig. 12. Fig. 12 is a diagram showing an example of the answer screen according to the embodiment. In Fig. 12, as described above, a case where the user's question is "Did I spend more this month" will be described.

[0094] In response to such an input question, the execution unit 45 transmits an API request to the server device 200 (see FIG. 1) to obtain content from the server device 200. The providing unit 46 generates, for example, text and graphs shown in FIG. 12 based on the content obtained from the server device 200 and provides them to the user. Note that the graph shown in FIG. 12 is generated by the server device 200 by transmitting a predetermined API request to the server device 200, for example.

[0095] Through this series of processes, the information processing device 1 provides an answer to the user's question. As a result, for example, the information processing device 1 can generate and provide an appropriate answer to the user's question.

[0096] [3. Processing flow] Next, a processing procedure executed by the information processing device 1 according to the embodiment will be described with reference to Fig. 13. Fig. 13 is a flowchart showing an example of a providing process according to the embodiment. Note that the processing procedure shown below is repeatedly executed by the information processing device 1 using the reception of a question sentence from a user as a key.

[0097] 13, the information processing device 1 receives a question written in a natural language from a user (step S101), and then generates a tag indicating the content of the question (step S102).

[0098] Next, the information processing device 1 searches the cache based on the tag (step S103) and determines whether the API request is in the cache (step S104). If the information processing device 1 determines that the API request is not in the cache (step S104; No), it generates a prompt (step S105). Note that the prompt here includes a first prompt and a second prompt.

[0099] Next, the information processing device 1 acquires a response from the model (step S106). That is, the information processing device 1 acquires an API request from the second model M2. Next, the information processing device 1 transmits the API request to the API and acquires data (step S107).

[0100] Then, the information processing device 1 provides an answer to the user's question (step S108), associates the tag with the API request, and registers it in the cache (step S109), and ends the process. If the information processing device 1 determines in step S104 that the cache exists, it extracts the corresponding API request from the cache and then proceeds to the process of step S107.

[0101] [4. Modifications] In the above-described embodiment, the case where the user inputs a sentence as text has been described, but the present invention is not limited to this. The information processing device 1 may also be configured to accept a sentence in a natural language from the user by voice or the like.

[0102] In addition, although the above-described embodiment has been described as a case where a question about electronic money is received, the present invention is not limited to this. For example, the present invention may be applied to other services, such as shopping sites, that acquire user account information through APIs.

[0103] In the above embodiment, the case where a question is received from a user is described, but the present invention is not limited to this. For example, an instruction to execute a remittance or the like may be received from a user.

[0104] For example, if the user writes "Please send 3,000 yen to A" but "A" cannot be identified, the system may ask the user to confirm A's account before sending the money.

[0105] [5. Effects] The information processing device 1 according to the embodiment includes a reception unit 41 that receives a natural language sentence indicating an execution instruction input by a user, a determination unit 43 that determines whether execution information for executing a process corresponding to the execution instruction of the sentence received by the reception unit 41 is cached, a generation unit 44 that generates execution information using a model trained to generate an answer to a question when the determination unit 43 determines that the execution information is not cached, and an execution unit 45 that executes a process using the cached execution information or the execution information generated by the generation unit 44.

[0106] The information processing device 1 also includes an estimation unit 42 that estimates the content of the sentence using the reception unit 41, and a determination unit 43 determines whether the execution information is cached based on the content of the sentence estimated by the estimation unit 42.

[0107] Furthermore, when the execution unit 45 executes processing using the execution information generated by the generation unit 44, the execution unit 45 adds the content of the sentence estimated by the estimation unit 42 to the execution information and registers it in the cache.

[0108] The generation unit 44 generates execution information by inquiring of the model about the type of application interface for executing the process and the content of the execution information to be executed by the application interface. The generation unit 44 also causes the model to generate the content of the execution information in the format of the application interface.

[0109] By performing any one or a combination of the above-described processes, the information processing device according to the present application can efficiently operate services using automatically generated AI.

[0110] [6. Hardware Configuration] The information processing device 1 according to the embodiment described above is realized by, for example, a computer 1000 configured as shown in Fig. 14. Fig. 14 is a hardware configuration diagram showing an example of a computer that realizes the functions of the information processing device according to the embodiment. The computer 1000 has a CPU 1100, a RAM 1200, a ROM 1300, an HDD 1400, a communication interface (I / F) 1500, an input / output interface (I / F) 1600, and a media interface (I / F) 1700.

[0111] The CPU 1100 operates and controls each unit based on programs stored in the ROM 1300 or the HDD 1400. The ROM 1300 stores a boot program executed by the CPU 1100 when the computer 1000 starts up, programs that depend on the hardware of the computer 1000, and the like.

[0112] The HDD 1400 stores programs executed by the CPU 1100, data used by such programs, etc. The communication interface 1500 receives data from other devices via a network (communication network) N and sends the data to the CPU 1100, and transmits data generated by the CPU 1100 to other devices via the network N.

[0113] The CPU 1100 controls output devices such as a display and a printer, and input devices such as a keyboard and a mouse (in FIG. 14, the output devices and input devices are collectively referred to as "input / output devices") via the input / output interface 1600. The CPU 1100 acquires data from the input devices via the input / output interface 1600. The CPU 1100 also outputs generated data to the output devices via the input / output interface 1600.

[0114] Media interface 1700 reads a program or data stored in recording medium 1800 and provides it to CPU 1100 via RAM 1200. CPU 1100 loads the program or data from recording medium 1800 onto RAM 1200 via media interface 1700 and executes the loaded program. Recording medium 1800 is, for example, an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.

[0115] For example, when the computer 1000 functions as the information processing device according to the embodiment, the CPU 1100 of the computer 1000 executes programs loaded onto the RAM 1200 to realize the functions of the control unit 4. The CPU 1100 of the computer 1000 reads and executes these programs from the recording medium 1800, but as another example, the CPU 1100 may obtain these programs from another device via the network N.

[0116] [7. Other] Although the embodiments of the present application have been described above, the present invention is not limited to the contents of these embodiments. Furthermore, the above-described components include those that can be easily imagined by a person skilled in the art, those that are substantially the same, and those that are within the scope of so-called equivalents. Furthermore, the above-described components can be combined as appropriate. Furthermore, various omissions, substitutions, or modifications of the components can be made without departing from the spirit of the above-described embodiments.

[0117] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.

[0118] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.

[0119] For example, the above-mentioned information processing device may be realized using multiple server computers, and depending on the function, the configuration can be flexibly changed, such as by calling an external platform using an API (Application Programming Interface) or network computing.

[0120] Furthermore, the above-described embodiments and modifications can be combined as appropriate within the scope of not causing any contradiction in the processing content.

[0121] Furthermore, the above-mentioned "section, module, unit" can be read as "means" or "circuit," etc. For example, an acquisition unit can be read as an acquisition means or an acquisition circuit. [Explanation of symbols]

[0122] 1. Information processing equipment 2. Communications Department 3 Storage section 4. Control Unit 31 Prompt Dictionary Storage 32 API request storage 41 Reception 42 Estimation part 43 Judgment section 44 Generation part 45 Executive Department 46 Providing Department 100 user terminals 200 Server device

Claims

1. a receiving unit that receives a natural language sentence indicating an execution instruction input by a user; a determination unit that determines whether execution information for executing a process corresponding to the execution instruction of the sentence accepted by the acceptance unit is cached based on whether the execution information corresponding to a tag formed by a word extracted from the sentence is cached; a generation unit that generates the execution information using a model that has been trained to generate an answer to a question when the determination unit determines that the execution information is not cached; an execution unit that executes the processing using the cached execution information or the execution information generated by the generation unit, and assigns an identifier of the user who input the sentence to the execution information used for the processing before the processing; Equipped with The generation unit Prompt engineering is performed using an internal model that is different from the model and has been trained to generate answers to questions, the execution information is generated by inputting the extended prompt generated by the prompt engineering into the model, the model is made to select an application interface that executes processing corresponding to the sentence from among a plurality of application interfaces, and the execution information is generated in a format that can be executed by the application interface.

1. An information processing device comprising:

2. The execution unit: When the process is executed using the execution information generated by the generation unit, the content of the sentence is added to the execution information and registered in the cache.

2. The information processing device according to claim 1,

3. 1. A computer-implemented information processing method, comprising: a receiving step of receiving a natural language sentence indicating an execution instruction input by a user; a determining step of determining whether execution information for executing a process corresponding to the execution instruction of the sentence accepted by the accepting step is cached based on whether the execution information corresponding to a tag formed by words extracted from the sentence is cached; a generating step of generating the execution information using a model trained to generate an answer to a question when it is determined by the determining step that the execution information is not cached; an execution step of executing the processing using the cached execution information or the execution information generated by the generation step, and assigning an identifier of the user who input the sentence to the execution information used in the processing before the processing; Including, The generating step includes: Prompt engineering is performed using an internal model that is different from the model and has been trained to generate answers to questions, the execution information is generated by inputting the extended prompt generated by the prompt engineering into the model, the model is made to select an application interface that executes processing corresponding to the sentence from among a plurality of application interfaces, and the execution information is generated in a format that can be executed by the application interface.

1. An information processing method comprising:

4. an acceptance step for accepting a natural language sentence indicating an execution instruction input by a user; a determination step of determining whether execution information for executing a process corresponding to the execution instruction of the sentence accepted by the acceptance step is cached based on whether the execution information corresponding to a tag formed by words extracted from the sentence is cached; a generating step of generating the execution information using a model trained to generate an answer to a question when it is determined by the determining step that the execution information is not cached; an execution procedure of executing the processing using the cached execution information or the execution information generated by the generation procedure, and before the processing, assigning an identifier of the user who input the sentence to the execution information used in the processing; on the computer, The generating procedure includes: Prompt engineering is performed using an internal model that is different from the model and has been trained to generate answers to questions, the execution information is generated by inputting the extended prompt generated by the prompt engineering into the model, the model is made to select an application interface that executes processing corresponding to the sentence from among a plurality of application interfaces, and the execution information is generated in a format that can be executed by the application interface. An information processing program characterized by:

Citation Information

Patent Citations

  • SQL (Structured Query Language) statement acquisition method and device, report generation method and device, computer equipment and storage medium device

    CN116541411A

  • Human-computer interaction method

    CN116701601A

  • Method, computer device, and computer program for providing dialogue dedicated to domain by using language model

    JP2023076413A

  • Organic semiconducting compounds

    WO2019185578A1