How to tune AI models, etc.

The system addresses the lack of clarity in identifying good responses in conversational AI by incorporating user feedback and rewards to refine the AI model, ensuring accurate and continuous learning.

JP7777298B1Active Publication Date: 2025-11-28ENGINEERING SAMURAI CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025067780
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-11-28
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

Conversational AI systems struggle to identify and learn from good responses, as they can only recognize bad answers but lack clarity on what constitutes good answers, and their memory capacity limits long-term improvement.

Method used

A system that includes a server, questioner, collaborator, and approver terminals, and a reward unit, where users rate answers as good or bad, and approvers validate and reward corrected responses to refine the AI model through relearning.

Benefits of technology

This system enables the AI model to learn from good examples, improving its responses by incorporating validated corrections, ensuring accurate and continuous updates, and motivating users with rewards, thus enhancing the model's practicality and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007777298000001_ABST
    Figure 0007777298000001_ABST
Patent Text Reader

Abstract

The goal is to provide a tuning method for an AI model that can teach a good answer as soon as a bad answer is given. [Solution] A tuning method 40 for an AI model executed by a system 10 including a server 1 that stores an AI model 19, an approver terminal 13, and a reward unit 14, the tuning method including the steps of the server 1 obtaining a question 2a from a questioner X, generating an answer 3a to the question, obtaining a revised answer 4a to the answer by a collaborator if the answer is inappropriate, relearning the AI ​​model using a pair of the question and the approved revised answer 4b, the approver terminal sending approval of the revised answer by approver Z to the server, and the reward unit paying a reward to the collaborator.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a tuning method for an AI model, a server used in the tuning method for an AI model, a server program, and a system for the tuning method for an AI model. [Background technology]

[0002] Patent Document 1 discloses an information processing system including a processor, which is capable of reducing the sense of discomfort felt during a conversation. Specifically, in this information processing system, the processor acquires user utterance data indicating user utterances in a dialogue in an acquisition step. In a determination step, the processor determines the matters requiring confirmation in a dialogue in which the utterance indicated by the acquired user utterance data was made based on relationship information indicating the relationship between the characteristics of the general utterance data indicating utterances in a general dialogue and the matters requiring confirmation that should be confirmed in the dialogue. In a presentation step, the processor presents the determined matters requiring confirmation to the user. In a reply step, the processor outputs a reply to the user based on the content of the utterance indicated by the acquired user utterance data. If there is an answer to the presented matters requiring confirmation, the processor outputs a reply reflecting the content of the answer. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 7462995 Summary of the Invention [Problem to be solved by the invention]

[0004] In conversational AI, if the user evaluates the AI's answer as BAD (off-topic), the AI ​​understands that the answer is BAD. However, it remains unclear what would be evaluated as GOOD. In other words, it is in a state where it is only being criticized unilaterally.

[0005] Therefore, the present invention aims to provide a tuning method for an AI model that can teach a good answer when a bad answer is given, a server used in the tuning method for an AI model, a server program, and a system for the tuning method for an AI model. [Means for solving the problem]

[0006] (1) The tuning method of the present invention is a tuning method for an AI model executed by a system including a server that stores the AI ​​model, an approver terminal, and a reward unit, and is characterized in that it includes the steps of the server acquiring a question from a questioner, generating an answer to the question, acquiring a revised answer to the answer from a collaborator if the answer is inappropriate, relearning the AI ​​model using a pair of the question and the approved revised answer, the approver terminal sending approval of the revised answer from the approver to the server, and the reward unit paying a reward to the collaborator.

[0007] (2) It is preferable that the same person be in charge of both the questioner and the collaborator in such a tuning method.

[0008] (3) It is also preferable that the method further comprises a step in which the server acquires from the questioner the part of the answer where the problem exists.

[0009] (4) Another aspect of the server of the present invention is a server used in a tuning method for an AI model that constitutes a system together with an approver terminal and a reward unit, and is characterized by comprising: a question acquisition unit that acquires a question from a questioner; an answer generation unit that generates an answer to the question; a revised answer acquisition unit that acquires a revised answer to the answer from a collaborator if the answer is inappropriate and acquires approval of the revised answer from the approver terminal based on input from the approver; and a learning unit that re-trains the AI ​​model using a pair of the question and the approved revised answer.

[0010] (5) Another aspect of the program of the present invention is a server program used in a tuning method for an AI model that includes an approver terminal and a reward unit and constitutes a system, and is characterized in that it causes a computer to acquire a question from a questioner, generate an answer to the question, and if the answer is inappropriate, acquire a revised answer to the answer from a collaborator, and obtain approval of the revised answer from the approver terminal through input by the approver, and re-train the AI ​​model using a pair of the question and the approved revised answer.

[0011] (6) Another aspect of the system of the present invention is a system for tuning an AI model, comprising a server that stores the AI ​​model, an approver terminal, and a reward unit, wherein the server comprises a question acquisition unit that acquires a question from a questioner, an answer generation unit that generates an answer to the question, a step of acquiring a revised answer to the answer from a collaborator if the answer is inappropriate, and a learning unit that retrains the AI ​​model using a pair of the question and the approved revised answer, wherein the approver terminal transmits approval of the revised answer by the approver to the server, and the reward unit pays a reward to the collaborator based on the approval. [Effects of the Invention]

[0012] The AI ​​model tuning method, server used in the AI ​​model tuning method, server program, and system for the AI ​​model tuning method of the present invention are capable of fine-tuning an AI model. [Brief explanation of the drawings]

[0013] [Figure 1] FIG. 1 is a schematic diagram illustrating an embodiment of a system for a method for tuning an AI model. [Figure 2] FIG. 1 is a functional block diagram illustrating one embodiment of a system for a method for tuning an AI model. [Figure 3]FIG. 2 is a schematic diagram illustrating an example of a hardware configuration of a server. [Figure 4] FIG. 2 is a schematic diagram illustrating an example of a hardware configuration of a user terminal. [Figure 5] 10 is a flow chart illustrating an embodiment of a process flow of an apparatus for indicating learning progress. [Figure 6] FIG. 2 is a schematic diagram showing an example of the data structure of a question and answer database. [Figure 7] FIG. 10 is a schematic diagram showing an example of the data structure of a collaborator database. [Figure 8] FIG. 2 is a schematic diagram illustrating an example of a data structure of an approver database. DETAILED DESCRIPTION OF THE INVENTION

[0014] [1. Brief description] (Introduction) The system for tuning AI models (hereinafter simply referred to as the system) relates to an interactive learning method for AI models. In particular, in semi-supervised learning, in addition to pre-learning, which is unsupervised learning, the system improves the model by fine-tuning (fine adjustments through supervised learning) using human training data. Existing conversational AI has acquired basic language skills through pre-training using large amounts of data, but it may produce inappropriate (BAD) responses depending on individual needs or specific situations. Currently, while AI can understand user ratings of "BAD," specific examples of appropriate (GOOd) responses remain unclear. Furthermore, there is a capacity limit to the AI's memory function for storing "GOOd" responses, limiting long-term improvement.

[0015] Therefore, this system proposes a system that allows the AI ​​model to learn good answer examples through fine tuning immediately when a bad answer is generated. This makes it possible to efficiently learn many more model answers despite the limitations of memory function. Specifically, we will create a system that rewards users who cooperate in fine-tuning the AI ​​model as an incentive for learning. At the same time, we will introduce an approval process by human approvers into the system to prevent inappropriate learning by malicious users.

[0016] (System 10) First, the system 10 will be described using Figure 1. Figure 1 shows the system 10. In this embodiment, the system 10 includes, for example, a server 1 that stores an AI model 19, a questioner terminal 11 held by a questioner X who uses the system 10, a collaborator terminal 12 held by a collaborator Y, an approver terminal 13 held by an approver Z, and a reward server 14 equipped with a reward unit 14a. The server 1, the questioner terminal 11, the collaborator terminal 12, the approver terminal 13, and the reward server 14 are each communicatively connected via a communication network 15. Note that it is sufficient that there are one or more of each of the terminals 11, 12, and 13. In the following description, the terminals 11, 12, and 13 may be collectively referred to simply as terminals. Furthermore, in the following description, the inquirer X, the collaborator Y, and the approver Z may be collectively referred to as users.

[0017] (Questioner X) The questioner X is a person who inputs a question to the AI ​​model 19 via the questioner terminal 11. The questioner X also receives an answer output from the AI ​​model 19, selects GOOD or BAD for the quality of the answer, and transmits the result to the system 10 via the questioner terminal 11.

[0018] (Collaborator Y) The helper Y checks the question and the answer to that question and corrects part or all of the answer. The correction consists of pointing out the part to be corrected and the correction content. If the entire answer is to be corrected, only the correction content is required. In this embodiment, when the questioner X gives a BAD rating, the server 1 transmits the question and the answer to the helper terminal 12.

[0019] (Approver Z) Collaborator Y obtains the revised response from system 10, checks whether the revised response is appropriate, and if it is appropriate, sends a notice of approval to system 10 via approver terminal 13. If it is not appropriate, he / she inputs a notice of denial into approver terminal 13 and notifies the collaborator via system 10. In this embodiment, if denial is made, the notice is sent together with the reason, but it is also possible to send either the notice of denial or the reason for denial.

[0020] The inquirer X, the collaborator Y, and the approver Z are users of the system 10. Note that operations including input to the terminals 11, 12, and 13 may be performed by representatives of the inquirer, the collaborator, and the approver, respectively.

[0021] [2. Details of each component] (Server 1) Fig. 2 shows an embodiment of a functional block diagram of system 10. Server 1 shown in Fig. 2 mainly includes a question acquisition unit 2 that acquires a question 2a from a questioner X, an answer generation unit 3 that generates an answer 3a to the question, a revised answer acquisition unit 4 that acquires a revised answer 4a, and a learning unit 5 that retrains AI model 19 using a pair of question 2a and approved revised answer 4b. Server 1 also includes a GOOD / BAD acquisition unit 6 and an approval acquisition unit 7.

[0022] (Question acquisition part 2, question 2a) The question acquisition unit 2 acquires a question from a questioner X. The question 2a may be text information, audio information, or image information, or a combination of two or more of these pieces of information, and may be data that can be converted into these pieces of information.

[0023] (Answer generator 3, answer 3a) The answer generation unit 3 generates an answer 3a to the question 2a via an AI model 19. The AI ​​model 19 is used to generate the answer 3a. The answer 3a may be text information, audio information, or image information, or a combination of two or more of these pieces of information, and may be data that can be converted into these pieces of information.

[0024] (AI model 19) The AI ​​model 19 is an interactive AI model that generates an answer 3a corresponding to a question 2a. The interactive AI model has acquired basic language capabilities through prior learning using large amounts of data. The AI ​​model 19 uses publicly known technology.

[0025] An AI model19, for example, has the following capabilities: natural language understanding, which is the ability to understand the meaning and intent of words used by humans (natural language); dialogue management, which is the ability to grasp the flow of a conversation and control it to generate appropriate responses; natural language generation, which is the ability to generate responses in natural language that is easy for humans to understand based on the understood content and purpose; and context understanding, which is the ability to remember the content of past conversations and understand the context of current utterances while advancing the dialogue.

[0026] (Corrected answer acquisition part 4, modified answer 4a) The corrected response acquisition unit 4 acquires the corrected response 4a from the collaborator terminal 12. The corrected response 4a may be text information, audio information, or image information, or a combination of a plurality of these pieces of information, and may be data that can be converted into these pieces of information.

[0027] (Study Section 5) The learning unit 5 retrains the AI ​​model 19 using a pair of the question 2a and the revised answer 4b approved by the approver Z.

[0028] (GOOD / BAD acquisition part 6, GOOD / BAD information 6a) The GOOD / BAD acquisition unit 6 acquires GOOD / BAD information 6a from the questioner terminal 11. The GOOD / BAD information 6a is input to the questioner terminal 11 by the questioner X. The GOOD / BAD information 6a is transmitted to the server 1 by selecting / touching a GOOD / BAD button displayed on the display unit 16a of the questioner terminal 11. The GOOD / BAD information 6a may be text information, audio information, or image information, or a combination of two or more of these types of information, and may be any data that can be converted into these types of information.

[0029] (Approval acquisition unit 7, approval / denial information 7a) The approval acquisition unit 7 acquires approval / denial information 7a from the approver terminal 13. The approval / denial information 7a is input into the approver terminal 13 by the approver Z. The approval information is sent to the correction response acquisition unit and reward unit 14a. On the other hand, the denial information is sent to the collaborator terminal 12. The reason for denial may be sent together with the denial information. Alternatively, only the reason for denial may be sent. Note that the denial information does not have to be sent to the collaborator terminal 12. The approval / denial information 7a may be text information, audio information, or image information, or a combination of two or more of these types of information, and may be any data that can be converted into these types of information.

[0030] (Questioner Terminal 11) Questioner terminal 11 is a device capable of communicating with server 1. A smartphone, a tablet PC, a personal computer, or the like may be used as questioner terminal 11. Questioner terminal 11 mainly includes display processing unit 16 and input processing unit 17.

[0031] (Display processing unit 16) The display processing unit 16 displays information acquired by the questioner terminal 11 on the display unit 16a. As the display unit 16a, for example, a device for displaying images such as a liquid crystal display (LCD), a plasma display panel (PDP), or an organic electroluminescence (EL) display is used. It is preferable that the display unit 16a has a touch panel function.

[0032] (Input processing unit 17) The input processing unit 17 receives operation input from the user 8 via the display unit 16a and sends a control signal corresponding to the operation content to the CPU 30 (see FIG. 3). In this embodiment, for example, a part of the input processing unit 17 and the display unit 16a may be integrated into a touch panel. Alternatively, the input processing unit 17 may receive operation input via a keyboard external to the questioner terminal 11.

[0033] (Transmitter / Receiver 18) The transmitter / receiver 18 acquires information transmitted from another terminal or the server 1, and transmits information to be sent to another terminal or the server 1. For example, the questioner terminal 11 transmits a question 2a and GOOD / BAD information 6a to the server 1, and receives an answer 3a.

[0034] (Collaborator terminal 12, approver terminal 13) The helper terminal 12 and the approver terminal 13 are almost the same as the questioner terminal 11 described above, so detailed explanations of the helper terminal 12 and the approver terminal 13 will be omitted.

[0035] (Reward Server 14, Reward Unit 14a, Reward Information 14b) The reward server 14 includes a reward unit 14a. The reward unit 14a acquires reward information 14b from the server 1. The reward information 14b is information including an identification number (ID) of the collaborator Y who will receive the reward, and may also include the amount of reward. The reward unit 14a pays reward to the collaborator Y based on the reward information 14b. The reward unit 14a may calculate the occurrence of reward and the unit price of reward based on predetermined reward conditions. The reward is paid to the collaborator Y as cash, electronic money, points that can be converted into cash or electronic money, or used to purchase goods or services, or the like.

[0036] (Communication Network 15) The communication network 15 is, for example, the Internet. The communication network 15 may also be an intranet or a communication network using wired or wireless communication. Examples of wireless communication include Bluetooth (registered trademark), Wi-Fi (registered trademark), and LPWA (Low Power Wide Area). LPWA may use LTE (registered trademark) or a frequency band that does not require a license.

[0037] [3. Hardware configuration] (Server 1 hardware configuration) Next, the hardware configuration of the server 1 will be described with reference to Fig. 3. As shown in Fig. 3, the server 1 of this embodiment uses, for example, a computer. The server 1 is equipped with a CPU (or GPU) 30. To the CPU 30, for example, a memory (hereinafter referred to as a storage unit) 31, a connection port 33 for connecting / reading a storage device 32, etc., and a communication circuit 34 for communicating with the outside via a network are connected via a bus line 35.

[0038] (Storage unit 31) The storage unit 31 mainly stores a program 36 for processing the operation of the system 10, and a question and answer database (hereinafter simply referred to as the question and answer DB) 20 that stores a question 2a, an answer 3a to the question 2a, a revised answer 4a, or an approved revised answer. The question and answer DB 20 only needs to store at least a valid answer to the question 2a. The storage unit 31 also stores the question and answer DB 20, a questioner database 21, a helper database 22, and an approver database 23. The storage unit 31 may further store a browser program 37 and even an OS 38 (operating system). The program 36 is installed on the server 1 by the storage device 32.

[0039] In this embodiment, the program 36 may operate in cooperation with the OS 38 and the browser program 37. The program 36 may also operate independently without using the browser program 37 and the OS 38.

[0040] In the hardware configuration of the program 36 described above, the functions shown in the functional block diagram of FIG. 2 are realized, for example, using a CPU 30 and the program 36, but some or all of them may be sequence-controlled using a logic circuit such as a microcomputer or a PLC (programmable logic controller).

[0041] (Hardware configuration of terminals 11, 12, and 13) Next, the hardware configuration of terminals 11, 12, and 13 will be described with reference to Figure 4. The hardware configuration of terminals 11, 12, and 13 is almost the same as that of server 1, so the same parts are given the same reference numerals and their description will be omitted. A program 39 for processing the operations of terminals 11, 12, and 13 is stored in storage unit 31 of terminals 11, 12, and 13. A browser program 37 and an OS 38 (operating system) may also be stored.

[0042] [4. Program] (Flowchart showing the processing of the system 10) 5 is a flowchart showing one embodiment of the processing flow of the system 10. The flowchart shows a program 36 of the server 1, programs 39, 39a, 39b of the terminals 11, 12, 13, and a tuning method 40 for an AI model.

[0043] (S01: User registration) In this embodiment, user registration (not shown) is required when using the system 10. However, the system 10 may be used without user registration.

[0044] (T1) An unregistered user who wishes to use the system 10 performs user registration via the server 1 or an external pre-registered login server (hereinafter referred to as the server, etc.). The user requests the server, etc. to send a registration form for user registration via their respective terminal. For example, in this embodiment, the function of a web browser is used.

[0045] (T2) The server or the like transmits a registration form, in which the user name and password used for logging in to the system 10, the handle name to be used in the system 10, and the like are entered.

[0046] The collaborator Y and / or approver Z may be required to have a predetermined qualification, experience, or a person who is presumed to have an equivalent in the field of question 2a so that they can appropriately answer or consider question 2a from questioner X. The operator of system 10 may register in advance persons who meet the requirements in storage unit 31. Also, for example, information certifying that the person has a qualification, experience, or a person who is presumed to have an equivalent may be registered in collaborator DB 22 and approver DB 23.

[0047] (S02:Login) A user 8 logs in to the system 10 (not shown). Note that the user may use the system 10 without logging in.

[0048] (S1: Enter and submit question 2a) The questioner X inputs a question 2a via the input processing unit 17 of the questioner terminal 11. The input question 2a is transmitted to the server 1.

[0049] (S2: Acquisition of Question 2a) Server 1 gets question 2a.

[0050] (S3: Generation of answer 3a) The server 1 generates an answer 3a to the question 2a using the AI ​​model 19. The server 1 transmits the generated answer 3a to the questioner terminal 11.

[0051] (S4: Acquisition and display of answer 3a) The questioner terminal 11 acquires the answer 3a from the server 1. The questioner terminal 11 displays the answer 3a on the display unit 16a.

[0052] (S5: Enter and submit your evaluation) The questioner X inputs GOOD / BAD information 6a as an evaluation of the answer 3a via the input processing unit 17 of the questioner terminal 11. The input GOOD / BAD information 6a is transmitted to the server 1.

[0053] (S6: Send question 2a if BAD) The server 1 acquires the GOOD / BAD information 6a from the questioner terminal 11. If the GOOD / BAD information 6a is BAD, the server 1 publishes the question 2a on a website or SNS site on the communication network 15. The helper Y uses the helper terminal 12 to request the website or SNS site to download the question 2a. The server 1 transmits the question 2a to the helper terminal 12. The website or SNS site may be viewable by anyone or may be restricted to only the specific helper Y. Furthermore, the server 1 may directly transmit the question 2a to the helper terminal 12 of one or more helpers, without publishing the question 2a on a website or an SNS site. The server 1 may publish the answer 3a together with the question 2a on a website or an SNS site, or may transmit the answer 3a to the helper terminal 12.

[0054] (S7: Acquire and send revised answer 4a) The helper terminal 12 acquires the question 2a from the server 1. The helper Y inputs the corrected answer 4a via the helper terminal 12. The helper terminal 12 transmits the corrected answer 4a to the server 1. The helper terminal 12 may acquire the answer 3a from the server 1 along with the question 2a. When there are a plurality of helpers Y who input corrected answers 4a to the question 2a, the plurality of corrected answers 4a are input via the helper terminals 12 of the helpers, and are transmitted to the server 1. In preparation for the case where a large number of revised answers 4a are sent to the server 1, the server 1 may set the number of accepted revised answers 4a or a predetermined acceptance time. The acceptance time may be, for example, the time from when the server 1 obtains a BAD answer from the GOOD / BAD information 6a or when the question 2a is published until a predetermined time has elapsed.

[0055] (S8: Show revised answer 4a) The server 1 obtains the revised answer 4a from the collaborator terminal 12. The server 1 publishes the revised answer 4a via a website or SNS site. The approver Z uses the approver terminal 13 to request the website or SNS site to download the revised answer 4a. The website or SNS site may not only make the revised answer 4a available to anyone, but may also impose restrictions such as only allowing the specific approver Z to view it. Furthermore, the server 1 may transmit the corrected response 4a to the approver terminal 13 of one or more approvers without publishing it on a website or an SNS site. The server 1 may publish the question 2a and / or the answer 3a together with the corrected answer 4a on a website or an SNS site, or may transmit the same to the approver terminal 13.

[0056] (S9: Obtaining and determining corrected answer 4a) The approver terminal 13 obtains the revised answer 4a from the server 1. The approver Z checks the revised answer 4a via the approver terminal 13. The approver Z decides to approve or reject the revised answer 4a. If there are multiple revised answers 4a for one question 2a, the approver Z selects and approves the most appropriate revised answer 4a and rejects the other revised answers 4a. The approver terminal 13 may obtain the question 2a and / or answer 3a from the server 1 along with the revised answer 4a.

[0057] (S10: Enter and send approval / denial information 7a) The approver Z inputs approval / denial information 7a via the approver terminal 13. The approval / denial information 7a is information indicating approval or denial. The approver terminal 13 transmits the approval / denial information 7a to the server 1. In this embodiment, when denying, the approver Z may input the reason for denial via the approver terminal 13. The approver terminal 13 transmits the reason for denial to the server 1 together with the approval / denial information 7a.

[0058] (If there are multiple approvers Z) If multiple revised answers 4a are approved by multiple approvers Z, the server 1 selects the revised answer 4a that has received the most approvals and rejects the others. Alternatively, the server 1 may select the approved revised answer 4a that the server 1 obtained first and reject the others.

[0059] (If there is no revised answer 4a to approve) If there is no revised answer 4a to be approved, approver Z rejects all revised answers 4a. Server 1 transmits to questioner terminal 11 that it is unable to provide an appropriate answer.

[0060] (others) Instead of rejecting, the approver terminal 13 may transmit the reason for rejection to the server 1. Alternatively, the approver Z may input a comment to be sent to the collaborator Y who input the revised answer 4a, and transmit it to the server 1. For example, if the revised answer 4a to the question 2a is appropriate but not the optimal revised answer 4a, the approver Z may comment that it was an appropriate revised answer 4a.

[0061] (S11: If approved) When the server 1 acquires the approved information, it transmits the revised answer 4a to the questioner terminal 11. The server 1 also transmits the reward information 14b to the reward server 14.

[0062] (S12: Acquire and display revised answers) The questioner terminal 11 acquires the approved revised answer 4b from the server 1, and the questioner terminal 11 displays the approved revised answer 4b on the display unit 16a.

[0063] (S13: Payment of Remuneration) The reward server 14 acquires the reward information 14b from the server 1. The reward server 14 pays the reward to the collaborator Y as cash, electronic money, points that can be converted into cash or electronic money, or used to purchase goods or services.

[0064] (S14: Re-learning) The server 1 stores the question 2a input from the questioner terminal 11 and the corresponding revised answer 4b provided from the helper terminal 12 and finally approved as new learning data in the question and answer DB 20. In re-learning (S14), the server 1 uses the stored pair of question 2a and approved revised answer 4b as training data to tune (update) its own AI model 19.

[0065] (Question / Answer DB20) 6 is a schematic diagram showing an example of the data structure of the question and answer DB 20. The question and answer DB 20 shown in the figure includes a plurality of questions 2a and answers 3a or approved revised answers 4b corresponding to those questions 2a. Furthermore, if the answer 3a to the question 2a is GOOD, the question and answer DB 20 does not necessarily include the approved revised answers 4b. Furthermore, the question and answer DB 20 may include an ID 22a, which is the identification number of the collaborator Y. (Questioner DB21) The following describes an example of the data structure of the questioner database (hereinafter referred to as questioner DB) 21. The questioner DB 21 includes an ID, which is an identification number of the questioner X, and the user name of the questioner (not shown).

[0066] (Contributor DB22) FIG. 7 is a schematic diagram showing an example of the data structure of a collaborator database (hereinafter referred to as collaborator DB) 22. The collaborator DB 22 shown in the figure is a compilation of personal information on multiple collaborators Y. The collaborator DB 22 mainly includes, for example, an ID 22a which is the identification number of the collaborator Y, a collaborator's username 22b, and a reward recipient 22c. The collaborator DB 22 may also include a specialty field 22d. There may be multiple specialty fields.

[0067] (Approver DB23) FIG. 8 is a schematic diagram showing an example of the data structure of the approver database (hereinafter referred to as the approver DB) 23. The approver DB 23 shown in the figure is a compilation of personal information on multiple approvers Z. The approver DB 23 mainly includes, for example, an ID 23a which is the identification number of the approver Z, and a user name 23b of the approver. The approver DB 23 may also include a field of expertise 23d. There may be multiple fields of expertise.

[0068] (Learning method) As an example, the following learning method can be considered. (1) Updating the natural language processing model: The model learns the context and keywords of the question 2a, as well as the expression patterns of appropriate corrective answers 4a, to improve the accuracy of answers to new questions. For example, the model analyzes the frequency of occurrence and co-occurrence of words and phrases included in the question, and adjusts parameters to generate more appropriate answers. (2) Strengthening the similar question classification model: Based on past questions 2a and their answers, we retrain the model to more accurately determine the similarity between new questions and existing questions. This allows us to more appropriately utilize past answers to similar questions. (3) Optimization of the answer ranking model: Retrain the model to rank the revised answers 4a that are likely to be accepted higher among the candidates for revised answers 4a provided by multiple contributors. Analyze the relationship between the content of the question and the characteristics of the revised answers 4a provided by contributors (keywords, expressions, revision history, etc.) to improve the accuracy of the ranking. (4) Expansion of the knowledge base: Pairs of new questions 2a and approved revised answers 4b stored in the question-answer DB 20 are integrated into the knowledge held by the server 1. This enables the server 1 to respond to a wider range of questions and to provide a wider variety of answers. Here, a knowledge base is, for example, a question and answer database 20, a dictionary that registers important words (keywords) contained in questions and answers and the concepts and meanings they represent, a database that classifies and organizes the categories to which past questions belong (e.g., technical questions, fact-checking questions, opinion-seeking questions, etc.), and the intentions of questioners, or rules and expression patterns for generating answers learned from past question and answer pairs.

[0069] Re-learning can be performed at an appropriate frequency depending on the system operation status, such as at regular intervals or when a certain amount of new data has accumulated in the question and answer DB 20. When re-learning, it is also effective to consider the balance with past learning data and apply a method to prevent over-learning (such as regularization). Through this re-learning process, Server 1 constantly learns the most recent questions and their optimal answers, improving the answer accuracy and user experience of the entire system.

[0070] (S15: If denied) When the server 1 acquires the rejected information, the server 1 transmits the reason to the helper terminal 12 of the helper Y who input the corrected answer 4a.

[0071] 5. Other Embodiments Next, we will explain modifications and other embodiments of the system 10. The modifications and other embodiments explained below are almost the same as the system 10 described above, so the same parts are given the same reference numerals and their explanations will be omitted.

[0072] (Variation) We will now explain modified examples of the system 10. In the system according to Modification 1, when the revised answer 4a is approved (step S11), the subsequent steps of step S12 (obtaining and displaying the revised answer 4a), step S13 (paying the reward), and step S14 (relearning) may be performed in any order.

[0073] (Second embodiment) A description will be given of another embodiment of the system 10. In the system according to the second embodiment, the questioner X answers a corrected answer 4a as the helper Y. Therefore, the questioner terminal 11 and the helper terminal 12 are the same terminal.

[0074] (Third embodiment) Next, a further embodiment of the system 10 will be described. In the system according to the third embodiment, when an answer 3a is BAD, the GOOD / BAD acquisition unit 6 of the server 1 acquires a defective portion 2b (see the dashed line in FIG. 2) of the answer 3a. The defective portion 2b is stored in the question and answer DB 20.

[0075] (Fourth embodiment) Next, another embodiment of the system 10 will be described. The system according to the fourth embodiment includes a category analysis unit 24 and a selection unit 25 (see the two-dot chain line in FIG. 2). When the answer 3a receives a BAD rating, the category analysis unit 24 determines to which category the question 2a belongs by linguistic analysis of the content of the question 2a. Next, the selection unit 25 selects the ID 21a of the helper Y and the ID 22a of the approver Z, which belong to the same or similar category as the question 2a, from the helper DB 22 and the approver DB 23, respectively. The server 1 then transmits the question 2a to the helper terminal 12 corresponding to the acquired ID 21a. The server 1 also transmits the revised answer 4a to the approver terminal 13 corresponding to the acquired ID 22a.

[0076] The language analysis may be performed using a conventionally known method, such as morphological analysis that breaks down a sentence into its smallest units and identifies parts of speech, syntactic analysis that analyzes the structure and relationships of a sentence and identifies the type of question, semantic analysis that understands the meaning of words or entire sentences and recognizes similar questions, keyword extraction that extracts important words and phrases that represent the content, or named entity extraction that identifies proper nouns such as people's names and place names.

[0077] A conventionally known method may be used to identify the field, such as keyword matching, which collates the question sentence with a field-specific keyword list to identify fields with many matches, rule-based classification, which determines the field using rules based on keywords and syntactic patterns, machine learning classification, which inputs the characteristics of the question sentence into a trained model to predict the field, or semantic similarity classification, which compares the meaning of the question sentence with representative sentences in the field to identify fields with high similarity.

[0078] The server 1 may use a conventionally known method to search for appropriate collaborators Y and approvers Z from the collaborator DB 22 and approver DB 23. As a search method, for example, a field matching method may be used, which compares the fields of expertise of the question field and the collaborators and approvers and prioritizes matching personnel, or, if there is no exact match, a similar field search method may be used, which searches for personnel with partial matches or related keywords.

[0079] (Fifth embodiment) Next, another embodiment of the system 10 will be described. In the system according to the fifth embodiment, the question acquisition unit 2 of the server 1 acquires a question 2a and a field 2c corresponding to the question 2a from the asker X via the asker terminal 11. The input may be performed using a function of a web browser. Next, the selection unit 9 selects the ID 21a of the helper Y and the ID 22a of the approver Z, which are in the field 2c or a similar field. The server 1 then transmits the question 2a to the helper terminal 12 corresponding to the acquired ID 21a. The server 1 also transmits the revised answer 4a to the approver terminal 13 corresponding to the acquired ID 22a.

[0080] [6. Other] (1) The items described as "other" in the above-described embodiments can be used in appropriate combinations. (2) Any or all of the question and answer DB 20, the questioner DB 21, the helper DB 22, and the approver DB 23 may be stored in an external server (not shown). (3) All or part of the functions of the learning unit 5 may be executed by an external server (not shown). (4) In the above-described modified examples and systems according to the first to fifth embodiments, the approver terminal 13 may not be provided. In this case, the corrected answer 4a by the helper Y is transmitted to the questioner terminal 11. (5) All or part of the functions of the server 1 may be executed by an external server (not shown).

[0081] [7. Summary] (1), (4-6) A tuning method 40 for an AI model 19 is executed by a system 10 including a server 1 that stores the AI ​​model 19, an approver terminal 13, and a reward unit 14a, and is characterized in that it includes the steps of the server 1 acquiring a question 2a from a questioner X, generating an answer 3a to the question 2a, acquiring a revised answer 4a to the answer 3a from a collaborator Y if the answer 3a is inappropriate, relearning the AI ​​model 19 using a pair of the question 2a and the approved revised answer 4b, the approver terminal 13 sending approval of the revised answer 4a from the approver Z to the server 1, and the reward unit 14a paying a reward to the collaborator Y.

[0082] By obtaining question 2a, it is possible to develop an AI model 19 that is in line with the questions and issues of actual users, and to reflect the needs of questioner X. Furthermore, if answer 3a is inappropriate, incorporating corrected answer 4a from collaborator Y improves the accuracy of the initial answer of AI model 19. Furthermore, the quality of the answer is guaranteed by going through the approval step of the revised answer 4a by Approver Z. Furthermore, by retraining using a combination of Question 2a and the approved revised answer 4b, the AI ​​model 19 can continuously update its knowledge and evolve. Furthermore, the payment of rewards by the reward unit 14a encourages active participation by the collaborator Y, and it can be expected that the collaborator Y will provide a high-quality revised answer 4a. As a result, the tuning method 40 contributes to the construction and operation of more practical and reliable AI models 19.

[0083] (2) In this tuning method 40, the same person is in charge of both the questioner X and the collaborator Y, so the following effects can be expected. Since questioner X, who best understands the intent of question 2a, makes the corrections himself, the corrections are more accurate, there is less rework, and the improvement cycle for the AI ​​model 19 can be significantly shortened. In addition, communication between questioner X and collaborator Y is no longer necessary, which saves time and effort. If questioner X has expertise in the field, he or she can make more specialized and higher-quality corrections directly to the AI ​​model19. Furthermore, since questioner X’s own question 2a and corrected answer 4a directly lead to improvements in AI model 19, it is expected that questioner X’s motivation will increase. Furthermore, because high-quality corrected answers 4a are directly entered, approver Z can concentrate more on the confirmation work, reducing the burden of approval. In this way, by having the same person act as both questioner X and collaborator Y, more direct and efficient tuning of the AI ​​model 19 is achieved, thereby improving the accuracy and practicality of the AI ​​model 19.

[0084] (3) Furthermore, if there is a problem with the answer 3a, the server 1 has a mechanism for directly obtaining information about the specific part 2b from the questioner X. This allows the AI ​​model 19 to re-train, not just by knowing that the answer 3a was wrong, but also by knowing in detail "which part" and "how" it was wrong, allowing the AI ​​model 19 to more accurately correct errors and acquire improved knowledge. As a result, re-training that further improves the performance of the AI ​​model 19 becomes possible. [Explanation of symbols]

[0085] 1 server 2 Question acquisition part 2a question 2b Location of defect 2c Field 3 Answer generation part 3a answer 4 Modified answer acquisition section 4a Revised answer 4b Approved Revised Answer 5. Learning Department 6 GOOD / BAD acquisition part 6a GOOD / BAD information 7 Approval Department 7a Approval / Disapproval Information 10 Systems 10a Variations 11 Questioner terminal 12 Collaborator terminal 13 Approver terminal 14 Reward Server 14a Compensation Department 14b Remuneration Information 15. Communication Networks 16 Display processing section 16a Display section 17 Input processing section 18 Transmitter / Receiver 19 AI models 20 Question and Answer Database (Question and Answer DB) 21 Questioner Database 22 Collaborator Database 23 Approver Database 24 Field Analysis Department 25 Selection section 30 CPU 31 Memory 32 Storage Devices 33 connection ports 34 Communication Circuit 35 Bus Line 36 Programs 36a Program 37 Browser Program 38 OS 39, 39a, 39b Terminal Programs 40 Tuning Method X Questioner Y Collaborator Z approver

Claims

1. A method for tuning an AI model, which is executed by a system including a server that stores an AI model, an approver terminal, and a reward unit, The server: obtaining a question from a questioner; generating an answer to the question; When information indicating that the answer by the questioner is inappropriate is obtained, a step of obtaining a revised answer by a collaborator to the answer is obtained; retraining the AI ​​model using the pair of questions and approved revised answers; The approver terminal, sending an approval of the revised response by the approver to the server; The reward department, and paying a reward to said collaborators.

2. 2. The tuning method according to claim 1, wherein the questioner and the collaborator are the same person.

3. 2. The tuning method according to claim 1, further comprising a step of the server acquiring a defective portion of the answer from the questioner.

4. A server used in a tuning method for an AI model that configures a system together with an approver terminal and a reward unit, a question acquisition unit that acquires a question from a questioner; an answer generation unit that generates an answer to the question; a corrected answer acquisition unit that, when acquiring information that the answer by the questioner is inappropriate, acquires a corrected answer by a helper to the answer and acquires approval of the corrected answer from the approver terminal according to an input by the approver; A server comprising a training unit that retrains the AI ​​model using the set of questions and the approved revised answers.

5. A server program used in a tuning method for an AI model that includes an approver terminal and a reward unit and configures a system, On the computer, Obtaining questions from questioners, generating an answer to said question; When information indicating that the answer by the questioner is inappropriate is acquired, a corrected answer to the answer by a helper is acquired, and approval of the corrected answer is acquired from the approver terminal by input by an approver; retraining the AI ​​model using the pair of questions and the approved revised answers.

6. A system for tuning an AI model, comprising: The system includes a server that stores the AI ​​model, an approver terminal, and a reward unit, The server: a question acquisition unit that acquires a question from a questioner; an answer generation unit that generates an answer to the question; a corrected answer acquisition unit that acquires a corrected answer to the answer by a helper when information indicating that the answer by the questioner is inappropriate is acquired; a training unit that retrains the AI ​​model using the set of questions and approved revised answers; the approver terminal transmits approval by the approver for the revised response to the server, The system, wherein the reward unit pays a reward to the collaborator based on the approval.

Citation Information

Patent Citations

  • Enterprise management large model fine tuning method, device and equipment and storage medium

    CN118428490A

  • Accounting processing system and accounting processing method

    JP2018205851A

  • Semi-automatic labeling of datasets

    JP2018537798A

  • Domain-Based Machine-Learned Classifiers

    US20250086500A1

  • Information processing system, information processing method, and program

    JP7462995B1