Information processing system and information processing method
The information processing system addresses the challenges of operating large language models by using a search engine to review responses for reliability, reducing the likelihood of hallucinations and enhancing the convenience and reliability of information processing.
Patent Information
- Application Number
- JP2024214383
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-14
- Filing Date
- 2024-12-09
- Publication Date
- 2025-06-26
Smart Images

Figure 2025096203000001_ABST
Abstract
Description
Technical Field
[0001] One aspect of the present invention relates to an information processing system and an information processing method.
[0002] Note that one aspect of the present invention is not limited to the above technical field. The technical field of one aspect of the invention disclosed in this specification and the like relates to an article, a method, or a manufacturing method. Alternatively, one aspect of the present invention relates to a process, a machine, a manufacture, or a composition of matter. Therefore, more specifically, as the technical field of one aspect of the present invention disclosed in this specification, semiconductor devices, display devices, light-emitting devices, power storage devices, storage devices, their driving methods, or their manufacturing methods can be cited as an example.
Background Art
[0003] In recent years, the development of language models using neural networks has been actively carried out, and in particular, large language models (LLMs) have attracted attention. A large language model is a natural language processing model trained using a large amount of data. With a large language model, for example, a dialogue model that answers user instructions can be realized. In Non-Patent Document 1, GPT-4 (Generative Pre-trained Transformer 4) (registered trademark) is disclosed as a large language model, and ChatGPT is disclosed as a dialogue model.
[0004] By using a large language model, the capabilities of natural language processing models have been significantly improved. On the other hand, due to the enlargement of language models, it is difficult to incorporate and operate a language model on one's own in terms of facilities and costs. Therefore, using an external service that provides a language model has become one form of using a language model.
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] One aspect of the present invention aims to provide a novel information processing system excellent in convenience, usefulness, or reliability. Or, one aspect of the present invention aims to provide a novel information processing method excellent in convenience, usefulness, or reliability. Or, one aspect of the present invention aims to provide a novel information processing system, a novel information processing method, or a novel semiconductor device.
[0007] Note that the description of these problems does not prevent the existence of other problems. Note that one aspect of the present invention does not necessarily need to solve all of these problems. Note that other problems will become apparent from the descriptions in the specification, drawings, claims, etc., and it is possible to extract these other problems from the descriptions in the specification, drawings, claims, etc.
Means for Solving the Problems
[0008] (1) Also, one aspect of the present invention is an information processing system having a first component, a second component, and a third component.
[0009] The first component has a function of receiving a questionnaire and passing it to the third component. Note that the questionnaire is described in natural language. Also, the first component has a function of receiving and providing a first response.
[0010] The second component has a function of receiving the first instruction and passing the answer to the third component. Note that the second component has a function of performing processing using a large language model. Also, the large language model has learned a dataset, and the large language model has a function of generating an answer according to the first instruction.
[0011] The third component has a function of creating the first instruction and passing it to the second component. Note that the first instruction includes a questionnaire, and the third component has a function of performing processing using a search engine. Also, the search engine has a function of using the questionnaire as a query to obtain search results from a database. Also, the database stores at least a part of the information not adopted in the dataset.
[0012] The third component has a function of reviewing the answer using the search results and generating a review result. When the review result is true, the third component has a function of creating the first answer sheet using the answer and passing it to the first component.
[0013] When the review result is false, the third component has a function of creating the first answer sheet using the search results and passing it to the first component.
[0014] As a result, the information processing system according to one aspect of the present invention can evaluate the likelihood that information not based on facts generated by the large language model (also referred to as hallucinations) is included in the description of the first response. In addition, the information processing system according to one aspect of the present invention can determine the appropriateness of the description in the first response. In addition, the information processing system according to one aspect of the present invention can evaluate the reliability of the description in the first response. In addition, the information processing system according to one aspect of the present invention can review the response using information not adopted in the dataset used for learning the large language model. In addition, the information processing system according to one aspect of the present invention can review the response using information collected from the database using the questionnaire as a query. In addition, the information processing system according to one aspect of the present invention can reflect the review result in the first response. In addition, the information processing system according to one aspect of the present invention can obtain an answer without determining whether the dataset is appropriate or inappropriate for the content of the questionnaire. In addition, the information processing system according to one aspect of the present invention can utilize the large language model without unnecessary fine-tuning. In addition, the information processing system according to one aspect of the present invention can utilize the large language model without unnecessarily updating the dataset and re-learning the large language model. In addition, the information processing system according to one aspect of the present invention can utilize the large language model without using the Retrieval Augmented Generation (RAG) method. In addition, since the information processing system according to one aspect of the present invention does not need to include the search result in the first instruction, the degree of freedom of the questionnaire is high. As a result, a novel information processing system excellent in convenience, usefulness, or reliability can be provided.
[0015] (2) Further, one aspect of the present invention is the above information processing system, wherein the third component includes a morphological analyzer.
[0016] The morphological analyzer extracts morphemes from the response to create a first array. In addition, the morphological analyzer extracts morphemes from the search result to create a second array.
[0017] The third component has a function of calculating the content rate of the morphemes included in the second array in the first array. Note that the examination result includes true or false determined based on the content rate.
[0018] (3) Also, one aspect of the present invention is the above information processing system in which the third component has a function of converting the answer and the search result into a distributed expression and calculating the similarity. Note that the examination result includes true or false determined based on the similarity.
[0019] (4) Also, one aspect of the present invention is the above information processing system in which the third component includes a text implication relation recognizer.
[0020] The text implication relation recognizer has a function of determining whether the answer and the search result are in a text implication relation. Note that the examination result includes true or false determined based on the determination of the text implication relation recognizer.
[0021] (5) Also, one aspect of the present invention is the above information processing system in which the first component has a function of receiving and providing the second answer sheet.
[0022] The second component has a function of receiving the second instruction sentence and passing the second answer sheet to the third component. Note that the large language model has a function of generating the second answer sheet according to the second instruction sentence.
[0023] When the examination result is false, the third component has a function of creating the second instruction sentence and passing it to the second component. The second instruction sentence includes the questionnaire and the search result. Also, the third component has a function of receiving the second answer sheet and passing it to the first component.
[0024] As a result, the information processing system according to one aspect of the present invention can generate a second response using the search expansion generation method when the large language model generates a hallucination. The information processing system according to one aspect of the present invention can prevent the generation of hallucinations using the search expansion generation method. Further, when the search expansion generation method is not used, the information processing system according to one aspect of the present invention does not need to create a second instruction including a query and search results. Further, when the search expansion generation method is not used, the information processing system according to one aspect of the present invention can input a query with a high degree of freedom. As a result, a novel information processing system excellent in convenience, usefulness, or reliability can be provided.
[0025] (6) Further, one aspect of the present invention is an information processing system having a first component, a second component, and a third component.
[0026] The first component has a function of receiving a query and passing it to the third component. Note that the query is described in natural language. Further, the first component has a function of receiving and providing a response.
[0027] The second component has a function of receiving an instruction and passing the response to the third component. Note that the second component has a function of performing processing using a large language model. Further, the large language model has learned a dataset, and the large language model has a function of generating a response according to the instruction.
[0028] The third component has a function of creating an instruction and passing it to the second component. Note that the instruction includes a query, and the third component has a function of performing processing using a search engine. Further, the search engine has a function of using the response as a query to obtain search results from a database. Further, the database stores at least a part of information not adopted in the dataset.
[0029] The third component has a function of reviewing an answer using search results and generating a review result. When the review result is true, the third component has a function of creating an answer sheet using the answer and delivering it to the first component.
[0030] When the review result is false, the third component has a function of creating an answer sheet using the search results and delivering it to the first component.
[0031] Accordingly, the information processing system according to one aspect of the present invention can review an answer using information not adopted in the dataset used for learning the large language model. Also, the information processing system according to one aspect of the present invention can review an answer using information collected from a database using a questionnaire as a query. Also, the information processing system according to one aspect of the present invention can review an answer using information collected from a database using the answer as a query. Also, the information processing system according to one aspect of the present invention can reflect the review result in the answer sheet. As a result, a novel information processing system excellent in convenience, usefulness, or reliability can be provided.
[0032] (7) Also, one aspect of the present invention is an information processing method having the first step to the ninth step.
[0033] In the first step, the first component receives a questionnaire and delivers the questionnaire to the second component.
[0034] In the second step, the second component creates an instruction statement and delivers it to the third component. Note that the instruction statement includes the questionnaire.
[0035] In the third step, the third component receives the instruction text and passes the response text to the second component. Note that the third component has a function of performing processing using a large language model, the large language model has learned a dataset, and the large language model has a function of generating a response text according to the instruction text.
[0036] In the fourth step, the second component uses the questionnaire as a query to obtain search results from the database. Note that the database stores at least a part of the information not adopted in the dataset.
[0037] In the fifth step, the second component reviews the response text using the search results and generates a review result.
[0038] In the sixth step, when the review result is true, the process proceeds to the seventh step. Also, when the review result is false, the process proceeds to the eighth step.
[0039] In the seventh step, the second component creates a response document using the response text, passes it to the first component, and then the process proceeds to the ninth step.
[0040] In the eighth step, the second component creates a response document using the search results, passes it to the first component, and then the process proceeds to the ninth step.
[0041] In the ninth step, the first component provides the response document.
[0042] As a result, the information processing system according to one aspect of the present invention can evaluate the likelihood that the hallucination generated by the large language model is included in the description of the answer sheet. In addition, the information processing system according to one aspect of the present invention can determine the appropriateness of the description in the answer sheet. In addition, the information processing system according to one aspect of the present invention can evaluate the reliability of the description in the answer sheet. In addition, the information processing system according to one aspect of the present invention can review the answer using information not adopted in the dataset used for learning the large language model. In addition, the information processing system according to one aspect of the present invention can review the answer using the information collected from the database using the questionnaire as a query. In addition, the information processing system according to one aspect of the present invention can reflect the review result in the answer sheet. In addition, the information processing system according to one aspect of the present invention can obtain an answer without determining whether the dataset is appropriate or inappropriate for the content of the questionnaire. In addition, the information processing system according to one aspect of the present invention can utilize the large language model without unnecessary fine-tuning. In addition, the information processing system according to one aspect of the present invention can utilize the large language model without unnecessarily updating the dataset and re-learning the large language model. In addition, the information processing system according to one aspect of the present invention can utilize the large language model without using the search expansion generation method. In addition, since the information processing system according to one aspect of the present invention does not need to include the search result in the instruction sentence, the degree of freedom of the questionnaire is high. As a result, it is possible to provide a novel information processing method excellent in convenience, usefulness, or reliability.
[0043] (8) Further, one aspect of the present invention is an information processing method having the first step to the ninth step.
[0044] In the first step, the first component receives the questionnaire and passes the questionnaire to the second component.
[0045] In the second step, the second component creates an instruction sentence and passes the instruction sentence to the third component. Note that the instruction sentence includes the questionnaire.
[0046] In the third step, the third component receives the instruction text and passes the response answer to the second component. Note that the third component has a function of performing processing using a large language model, the large language model has learned a dataset, and the large language model has a function of generating a response answer according to the instruction text.
[0047] In the fourth step, the second component uses the response answer as a query to obtain search results from the database. Note that the database stores at least a part of the information not adopted in the dataset.
[0048] In the fifth step, the second component reviews the response answer using the search results and generates a review result.
[0049] In the sixth step, when the review result is true, the process proceeds to the seventh step. Also, when the review result is false, the process proceeds to the eighth step.
[0050] In the seventh step, the second component creates a response document using the response answer, passes it to the first component, and then the process proceeds to the ninth step.
[0051] In the eighth step, the second component creates a response document using the search results, passes it to the first component, and then the process proceeds to the ninth step.
[0052] In the ninth step, the first component provides the response document.
[0053] As a result, it is possible to review the answer using information not adopted in the dataset used for learning the large language model. Also, it is possible to review the answer using the questionnaire as a query and the information collected from the database, and using the answer as a query and the information collected from the database. Further, the review result can be expanded and reflected in the answer sheet. As a result, it is possible to provide a new information processing method excellent in convenience, usefulness, or reliability.
[0054] (9) Also, one aspect of the present invention is an information processing method having a first step to an eleventh step.
[0055] In the first step, the first component receives the questionnaire and passes the questionnaire to the second component.
[0056] In the second step, the second component creates a first instruction and passes the first instruction to the third component. Note that the first instruction includes the questionnaire.
[0057] In the third step, the third component receives the first instruction and passes the answer to the second component. Note that the third component has a function of performing processing using a large language model, the large language model has learned a dataset, and the large language model has a function of generating an answer according to the first instruction.
[0058] In the fourth step, the second component uses the questionnaire as a query to obtain a search result from the database. Note that the database stores at least a part of the information not adopted in the dataset.
[0059] In the fifth step, the second component reviews the answer using the search result and generates a review result.
[0060] In the 6th step, when the review result is true, proceed to the 7th step; when the review result is false, proceed to the 8th step.
[0061] In the 7th step, the second component creates the first response using the response answer, transfers it to the first component, and then proceeds to the 11th step.
[0062] In the 8th step, the second component creates the first response using the search result, transfers it to the first component, then creates the second instruction, and transfers it to the third component. Note that the second instruction includes the questionnaire and the search result.
[0063] In the 9th step, the third component receives the second instruction and transfers the second response to the second component. Note that the large language model has the function of generating the second response according to the second instruction.
[0064] In the 10th step, the second component receives the second response, transfers it to the first component, and then proceeds to the 11th step.
[0065] In the 11th step, when the review result is true, the first component provides the first response; when the review result is false, the first component provides the first response and the second response.
[0066] As a result, the information processing system according to one aspect of the present invention can generate a second response using the search expansion generation method when the large language model generates hallucinations. Further, the information processing system according to one aspect of the present invention can prevent the generation of hallucinations using the search expansion generation method. Further, when the search expansion generation method is not used, the information processing system according to one aspect of the present invention does not need to create a second instruction including a query and search results. Further, when the search expansion generation method is not used, the information processing system according to one aspect of the present invention can input a query with a high degree of freedom. As a result, a novel information processing method excellent in convenience, usefulness, or reliability can be provided.
[0067] In the drawings attached to this specification, the components are classified by function and shown as independent blocks in a block diagram. However, it is difficult to completely separate the actual components by function, and one component may be related to multiple functions.
Advantages of the Invention
[0068] According to one aspect of the present invention, a novel information processing system excellent in convenience, usefulness, or reliability can be provided. Further, one aspect of the present invention can provide a novel information processing method excellent in convenience, usefulness, or reliability. Further, a novel information processing system can be provided. Further, a novel information processing method can be provided.
[0069] Note that the description of these effects does not prevent the existence of other effects. Note that one aspect of the present invention does not necessarily have all of these effects. Note that other effects will be obvious from the description in the specification, drawings, claims, etc., and it is possible to extract these other effects from the description in the specification, drawings, claims, etc.
Brief Description of the Drawings
[0070]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
MODE FOR CARRYING OUT THE INVENTION
[0071] An information processing system according to one aspect of the present invention includes a first component, a second component, and a third component. The first component has a function of receiving a questionnaire and passing it to the third component. Note that the questionnaire is described in natural language. The first component has a function of receiving and providing a first response. The second component has a function of receiving a first instruction and passing a response to the third component. Note that the second component has a function of performing processing using a large language model, and the large language model has learned a dataset. The large language model also has a function of generating a response according to the first instruction. The third component has a function of creating a first instruction and passing it to the second component, and the first instruction includes the questionnaire. The third component also has a function of performing processing using a search engine. Note that the search engine has a function of using the questionnaire as a query to obtain search results from a database. The database stores at least a part of information not adopted in the dataset. The third component has a function of reviewing a response using the search results and generating a review result. When the review result is true, the third component has a function of creating a first response using the response and passing it to the first component. When the review result is false, the third component has a function of creating a first response using the search results and passing it to the first component.
[0072] As a result, a novel information processing system excellent in convenience, usefulness, or reliability can be provided.
[0073] The embodiments will be described in detail with reference to the drawings. However, the present invention is not limited to the following description, and it will be easily understood by those skilled in the art that the form and details thereof can be variously changed without departing from the spirit and scope of the present invention. Therefore, the present invention is not to be construed as limited to the description of the embodiments shown below. In the configuration of the invention described below, the same reference numerals are commonly used among different drawings for the same part or parts having the same or similar functions, and the repeated description thereof will be omitted.
[0074] (Embodiment 1) In the present embodiment, an information processing system according to an aspect of the present invention will be described with reference to FIGS. 1 to 5.
[0075] FIG. 1 is a diagram for explaining the configuration of an information processing system according to an aspect of the present invention.
[0076] FIG. 2 is a schematic diagram for explaining the configuration of an instruction sentence used by an information processing system according to an aspect of the present invention.
[0077] FIG. 3 is a diagram for explaining the configuration of an information processing system according to an aspect of the present invention.
[0078] FIG. 4 is a schematic diagram for explaining the configuration of an instruction sentence used by an information processing system according to an aspect of the present invention.
[0079] FIG. 5 is a block diagram for explaining the configuration of an information processing apparatus that can be used in an information processing system according to an aspect of the present invention.
[0080] <Example Configuration 1 of Information Processing System> The information processing system described in the present embodiment includes a component 30, a component 20, and a component 21 (see FIG. 1).
[0081] 《Example Configuration 1 of Component 30》 Component 30 has a function of receiving a questionnaire QRE and delivering the questionnaire QRE to component 21. Note that the questionnaire QRE is described in natural language.
[0082] Component 30 is provided with a function of receiving the response form ANS1 created by component 21 and providing it to the user of the information processing system.
[0083] 《Configuration Example 1 of Component 20》 Component 20 is provided with a function of receiving the instruction text PT1 created by component 21 and passing the response answer Drf to component 21.
[0084] Component 20 is provided with a function of performing processing using the large language model LLM. Note that the large language model LLM has learned the dataset DS.
[0085] In addition, the large language model LLM is provided with a function of generating the response answer Drf according to the instruction text PT1.
[0086] For example, BERT (Bidirectional Encoder Representations from Transformers), GPT-3, GPT-3.5, GPT-4 (registered trademark), LaMDA (Language Model for Dialogue Applications), PaLM (Pathways Language Model), Llama2, ALBERT, XLNet, etc. can be used for the large language model LLM.
[0087] 《Configuration Example 1 of Component 21》 Component 21 is provided with a function of creating the instruction text PT1 and passing it to component 20. Note that the instruction text PT1 includes the questionnaire QRE (see Figure 2).
[0088] In addition, information specifying the range of topics to which the questionnaire QRE belongs can be included in the instruction PT1. Specifically, terms such as "Politics", "Economics", "Culture", "Science", "Society", "Events", etc. can be used as information for specifying the range of topics. Also, terms such as "Natural Science", "Technology & Engineering", "Biology & Agriculture", "Medicine & Pharmacy & Psychology", etc. can be used as information for specifying the range of topics. Further, terms such as "Materials Engineering", "Architecture & Architectural Engineering", "Transportation Engineering", "Energy Engineering", "Information & Communication & Electronics Engineering", "Food Engineering", "Military Engineering", "Safety Engineering & Disasters", etc. can be used as information for specifying the range of topics. Thereby, a response Drf containing specialized information can be obtained.
[0089] Moreover, information specifying the format of the response Drf can be included in the instruction PT1. Specifically, statements such as "Answer as an expert would", "Answer as a researcher would", "Answer as an engineer would", etc. can be used as information for specifying the format of the response Drf. Thereby, a response Drf with a specialized structure can be obtained.
[0090] Component 21 has a function of performing processing using the search engine SE. The search engine SE has a function of using the questionnaire QRE as a query qu to obtain a search result SR from the database DB. Note that the database DB stores at least a part of the information not adopted in the dataset DS.
[0091] For example, an external search service providing information on websites on the Internet can be used for the search engine SE and the database DB. Also, an archive of official documents or private documents can be used for the database DB. Further, a database for managing confidential information within the organization to which the user of the information processing system belongs can be used for the database DB.
[0092] Component 21 has a function of examining the response Drf using the search result SR and generating an examination result ER.
[0093] When the review result ER is true, the component 21 has a function of creating the response document ANS1 using the response answer Drf and delivering it to the component 30. Also, when the review result ER is false, the component 21 has a function of creating the response document ANS1 using the search result SR and delivering it to the component 30. When the response document ANS1 is created using the response answer Drf, for example, it can be displayed on the component 30 as "This response document is generated by AI. This response document has been reviewed using search results." Also, when the response document ANS1 is created using the search result SR, for example, it can be displayed on the component 30 as "This response document is created based on the search results."
[0094] As a result, the information processing system according to one aspect of the present invention can evaluate the likelihood that the hallucination generated by the large language model LLM is included in the description of the answer sheet ANS1. Further, the information processing system according to one aspect of the present invention can determine the appropriateness of the description in the answer sheet ANS1. Further, the information processing system according to one aspect of the present invention can evaluate the reliability of the description in the answer sheet ANS1. Further, the information processing system according to one aspect of the present invention can review the answer Drf using information not adopted in the dataset DS used for the learning of the large language model LLM. Further, the information processing system according to one aspect of the present invention can review the answer Drf using the information collected from the database DB using the questionnaire QRE as the query qu. Further, the information processing system according to one aspect of the present invention can reflect the review result ER in the answer sheet ANS1. Further, the information processing system according to one aspect of the present invention can obtain an answer without determining whether the dataset DS is appropriate or inappropriate for the content of the questionnaire QRE. Further, the information processing system according to one aspect of the present invention can utilize the large language model LLM without unnecessary fine-tuning. Further, the information processing system according to one aspect of the present invention can utilize the large language model LLM without unnecessarily updating the dataset DS and re-learning the large language model LLM. Further, the information processing system according to one aspect of the present invention can utilize the large language model LLM without using the search expansion generation method. Further, since the information processing system according to one aspect of the present invention does not need to include the search result SR in the instruction text PT1, the degree of freedom of the questionnaire QRE is high. As a result, a novel information processing system excellent in convenience, usefulness, or reliability can be provided.
[0095] <<Configuration Example 2 of Component 21>> Component 21 is equipped with a morphological analyzer MP_A. For example, "MeCab", "Sudachi", etc. can be used as the morphological analyzer MP_A. Specifically, even for languages without a clear delimiter between consecutive words, such as Japanese, words can be delimited. Also, the part-of-speech of words can be obtained. Further, appropriate parts-of-speech, such as nouns or adjectives, can be selected. Additionally, consecutive words can be combined to select appropriate phrases.
[0096] Note that the morphological analyzer MP_A extracts morphemes from the answer Drf and creates an array AL1. Also, the morphological analyzer MP_A extracts morphemes from the search result SR and creates an array AL2.
[0097] Component 21 has a function, for example, of calculating the content rate CntR of the morphemes included in array AL2 in array AL1. The review result ER includes true or false determined based on the content rate CntR.
[0098] For example, when the content rate CntR is equal to or greater than a predetermined value, the review result ER can be set to true, and when it is less than the predetermined value, the review result ER can be set to false. Specifically, the content rate CntR is preferably 0.9 or more, and more preferably 1. Note that when the content rate CntR is 1, all of the morphemes extracted from the answer Drf are included in the morphemes extracted from the search result SR. Also, when the phrases extracted from the answer Drf are included in the similar phrase group of the phrases extracted from the search result SR, it also contributes to the content rate CntR.
[0099] 《Configuration Example 3 of Component 21》 Component 21 has a function of converting the answer Drf and the search result SR into a distributed representation and calculating the similarity SIM. For example, using a large language model, the answer Drf and the search result SR1 can be converted into a distributed representation. Specifically, the answer Drf and the search result SR1 can be converted into a distributed representation using BERT, GPT-3, GPT-3.5, GPT-4 (registered trademark), LaMDA, PaLM, Llama2, ALBERT, XLNet, etc. Note that by converting a sentence into a distributed representation, the semantic similarity can be known. Since the proximity of the meaning of the sentence included in the answer Drf and the meaning of the sentence included in the search result SR can be evaluated, fluctuations in expressions such as synonyms can be absorbed. In particular, fluctuations in verb expressions can be absorbed.
[0100] The review result ER includes true or false determined based on the similarity SIM.
[0101] For example, when the cosine similarity is equal to or greater than a predetermined value, the review result ER can be set to true, and when it is less than the predetermined value, the review result ER can be set to false. Specifically, it is preferable that the cosine similarity is 0.8 or more.
[0102] 《Configuration Example 4 of Component 21》 Component 21 includes a text entailment relation recognizer RTE_A. For example, a large language model can be used for the text entailment relation recognizer RTE_A. Specifically, BERT, GPT-3, GPT-3.5, GPT-4 (registered trademark), LaMDA, PaLM, Llama2, ALBERT, XLNet, etc. can be used for the text entailment relation recognizer RTE_A. Note that the evaluation of the text entailment relation recognizer is the most direct evaluation for evaluating whether the meanings of documents are the same.
[0103] The text entailment relation recognizer RTE_A has a function of determining whether the answer Drf and the search result SR are in a text entailment relation.
[0104] The examination result ER includes true or false determined based on the determination of the text implicature relation recognizer RTE_A.
[0105] <Configuration Example 2 of the Information Processing System> The information processing system described in this embodiment has a component 30, a component 20, and a component 21 (see Fig. 3). Note that Configuration Example 2 of the information processing system described in this embodiment is different from Configuration Example 1 of the information processing system in that when the result of examining the answer Drf is false, the component 20 is caused to generate an answer ANS2 using an instruction text PT2 including a questionnaire QRE and a search result SR. Here, the different parts will be described in detail, and for the same configurations, the above description will be incorporated by reference.
[0106] 《Configuration Example 2 of Component 30》 Component 30 has a function of receiving the answer ANS2 created by component 21 and providing it to the user of the information processing system.
[0107] 《Configuration Example 2 of Component 20》 Component 20 has a function of receiving the instruction text PT2 created by component 21 and passing the answer ANS2 to component 21. Note that the large language model LLM has a function of generating the answer ANS2 according to the instruction text PT2.
[0108] 《Configuration Example 5 of Component 21》 Component 21 has a function of creating an instruction text PT2 and passing it to component 20 when the examination result ER is false. Note that the instruction text PT2 includes a questionnaire QRE and a search result SR (see Fig. 4). Also, the instruction text PT2 includes a requirement to answer the questionnaire QRE with reference to the search result SR. Note that the method of using the instruction text PT2 including the requirement to answer the questionnaire QRE with reference to the search result SR, the questionnaire QRE, and the search result SR when causing the large language model LLM to generate the answer ANS2 can be said to be an aspect of the search expansion generation method.
[0109] Component 21 has a function of receiving the response document ANS2 generated by Component 20 and passing it to Component 30. When the response document ANS2 is a response to the questionnaire QRE with reference to the search result SR, for example, it can be displayed to Component 30 as "This response document is generated by AI with reference to the search result."
[0110] Thereby, the information processing system according to one aspect of the present invention can generate the response document ANS2 using the search expansion generation method when the large language model LLM generates hallucinations. Also, the information processing system according to one aspect of the present invention can prevent the generation of hallucinations using the search expansion generation method. Also, when the information processing system according to one aspect of the present invention does not use the search expansion generation method, it is not necessary to create the instruction text PT2 including the questionnaire QRE and the search result SR. Also, when the information processing system according to one aspect of the present invention does not use the search expansion generation method, a questionnaire QRE with a high degree of freedom can be input. As a result, a novel information processing system excellent in convenience, usefulness, or reliability can be provided.
[0111] For example, information related to the details of a company's research and development operations is not included in the dataset DS used for the training of the large language model LLM. Therefore, for example, the large language model LLM may not be able to answer accurately to questions regarding the details of integrated circuit design operations. The search expansion generation method that provides information for reference in answering, together with questions regarding the details of integrated circuit design operations, is useful.
[0112] Also, the questions asked by the users of the information processing system are diverse, and there may be cases where the questions can be answered without using the search expansion generation method or are not suitable for search. Therefore, by evaluating the answer generated by the large language model using the search result, a reliable answer can be obtained and a quick answer can be obtained. Also, the efficiency of integrated circuit design operations can be improved.
[0113] <Configuration Example 3 of Information Processing System> Configuration example 3 of the information processing system described in this embodiment is different from configuration example 1 and configuration example 2 of the information processing system in that the search engine SE of component 21 has a function of obtaining a search result SR from the database DB using the answer Drf as a query qu. Here, the different parts will be described in detail, and for the same configurations, the above description will be incorporated by reference.
[0114] 《Configuration example 6 of component 21》 Component 21 has a function of performing processing using the search engine SE. The search engine SE has a function of obtaining a search result SR from the database DB using the answer Drf as a query qu. Note that the database DB stores at least a part of the information not adopted in the dataset DS.
[0115] Thereby, the information processing system according to one aspect of the present invention can review the answer Drf using the information not adopted in the dataset DS used for the learning of the large language model LLM. Also, the information processing system according to one aspect of the present invention can review the answer Drf using the information collected from the database DB using the questionnaire QRE as a query qu. Also, the information processing system according to one aspect of the present invention can review the answer Drf using the information collected from the database DB using the answer Drf as a query qu. Also, the information processing system according to one aspect of the present invention can reflect the review result ER in the answer sheet ANS1. As a result, a new information processing system excellent in convenience, usefulness, or reliability can be provided.
[0116] <Configuration example 4 of the information processing system> The information processing system described in this embodiment has component 30, component 21, and component 20 (see FIG. 1 or FIG. 3).
[0117] For example, an information processing apparatus that performs the functions of component 30, an information processing apparatus that performs the functions of component 21, and an information processing apparatus that performs the functions of component 20 can constitute an information processing system according to an aspect of the present invention. Note that the number of information processing apparatuses constituting the information processing system according to an aspect of the present invention is one or more. Also, for example, a plurality of information processing apparatuses can be connected using network 51 to constitute an information processing system according to an aspect of the present invention.
[0118] When an information processing system according to an aspect of the present invention is configured using a plurality of information processing apparatuses, the load related to information processing can be dispersed.
[0119] <<Configuration Example 1 of Information Processing Apparatus>> Configuration Example 1 of the information processing apparatus described in this embodiment can be used for component 30. Configuration Example 1 of the information processing apparatus can also be called a client computer or the like. For example, a desktop computer can be used for component 30.
[0120] Configuration Example 1 of the information processing apparatus can receive data input by a user of the information processing system according to an aspect of the present invention. Also, Configuration Example 1 of the information processing apparatus can provide data output by the information processing system according to an aspect of the present invention to the user.
[0121] For example, dedicated application software or a web browser or the like operates. A user of the information processing system according to an aspect of the present invention can access the information processing system via any of them. Thereby, a service using the information processing system according to an aspect of the present invention can be enjoyed.
[0122] <<Configuration Example 2 of Information Processing Apparatus>> Configuration Example 2 of the information processing apparatus described in this embodiment can be used for component 21. For example, a workstation, a server computer, a supercomputer, or the like can be used for component 21.
[0123] In addition, Configuration Example 2 of the information processing apparatus preferably has a function as a parallel computer. By using it as a parallel computer, for example, large-scale calculations required for learning and inference of artificial intelligence (AI) can be performed.
[0124] Also, Configuration Example 2 of the information processing apparatus can perform processing using a natural language processing model using AI.
[0125] For example, processing (Natural Language Processing) using natural language models such as BERT, T5 (Text-to-Text Transfer Transformer), GPT-3, GPT-3.5, GPT-4 (registered trademark), LaMDA, PaLM, Llama2, etc. can be executed.
[0126] 《Configuration Example 3 of the Information Processing Apparatus》 For example, Configuration Example 3 of the information processing apparatus described in this embodiment can be used for Component 20. Note that Component 20 is larger in scale and higher in computing power than Component 21. For example, a large computer such as a server computer or a supercomputer can be used for Component 20.
[0127] In addition, Configuration Example 3 of the information processing apparatus preferably has a function as a parallel computer. By using it as a parallel computer, for example, large-scale calculations required for learning and inference of AI can be performed.
[0128] Also, Configuration Example 3 of the information processing apparatus can perform processing using a natural language processing model using AI. In particular, processing using a general-purpose language processing model capable of performing various natural language processing tasks can be executed.
[0129] For example, processing using natural language models such as BERT, T5, GPT-3, GPT-3.5, GPT-4 (registered trademark), LaMDA, PaLM, Llama2, etc. can be executed. In particular, it is preferable to be able to execute processing using GPT-4 (registered trademark). For example, compared with conventional natural language models, if processing using a large-scale language model can be performed, more natural text generation or conversation, etc. can be realized.
[0130] Note that a person who provides a service using the information processing system according to an aspect of the present invention does not necessarily need to own Configuration Example 3 of the information processing apparatus by themselves. For example, a service provider can use a part of the services provided by other operators or the like using Configuration Example 3 of the information processing apparatus.
[0131] 《Configuration Example of Network 51》 The network 51 that can be used in the information processing system according to an aspect of the present invention can connect a plurality of information processing apparatuses. Thereby, the plurality of connected information processing apparatuses can transmit and receive data to and from each other. Also, the load related to information processing can be dispersed.
[0132] When performing wireless communication, as a communication protocol or communication technology, communication standards such as the 4th generation mobile communication system (4G), the 5th generation mobile communication system (5G), the 6th generation mobile communication system (6G), or specifications standardized by the IEEE such as Wi-Fi (registered trademark), Bluetooth (registered trademark), etc. can be used.
[0133] For example, a local network can be used as the network 51. Also, an intranet or an extranet can be used as the network 51. Further, a PAN (Personal Area Network), a LAN (Local Area Network), a CAN (Campus Area Network), a MAN (Metropolitan Area Network), a WAN (Wide Area Network), a GAN (Global Area Network), etc. can be used as the network 51.
[0134] Also, for example, a global network can be used as the network 51. Specifically, the Internet, which is the basis of the World Wide Web (WWW), can be used.
[0135] Also, a person who provides a service using the information processing system according to an aspect of the present invention can provide a service using the information processing method according to an aspect of the present invention via, for example, the network 51.
[0136] Note that when the information processing system according to an aspect of the present invention is constructed within a local network, the possibility of, for example, confidential information leakage can be reduced as compared with the case of using the Internet.
[0137] 《Configuration Example 4 of Information Processing Apparatus》 An information processing apparatus that can be used in the information processing system according to an aspect of the present invention has, for example, an input unit 110, a storage unit 120, a processing unit 130, an output unit 140, and a transmission path 150 (see FIG. 5).
[0138] Note that in the drawings attached to this specification, the components are classified by function and shown as block diagrams with independent blocks for each other. However, in actuality, it is difficult to completely separate the components by function, and one component may be related to multiple functions. For example, a part of the processing unit 130 may function as the input unit 110. Also, one function may be related to multiple components. For example, the processing performed by the processing unit 130 may be executed on different servers depending on the processing.
[0139] [Input unit 110] The input unit 110 can receive data from outside the information processing apparatus. For example, the input unit 110 receives data via the network 51. Specifically, a device such as a personal computer equipped with a communication port or communication function can be used.
[0140] The input unit 110 supplies the received data to one or both of the storage unit 120 and the processing unit 130 via the transmission path 150.
[0141] [Storage unit 120] The storage unit 120 has a function of storing the programs executed by the processing unit 130. Also, the storage unit 120 can have a function of storing the data generated by the processing unit 130 (for example, calculation results, analysis results, inference results), and the data received by the input unit 110, etc.
[0142] The storage unit 120 can have a database. Also, the information processing apparatus can have a database separately from the storage unit 120. The information processing apparatus can have a function of retrieving data from a database existing outside the storage unit 120, outside the information processing apparatus, or outside the information processing system. Also, the information processing apparatus can have a function of retrieving data from both its own database and an external database.
[0143] One or both of the storage and the file server can be used as the storage unit 120. Also, a database recording the paths of the files stored in the file server can be used as the storage unit 120.
[0144] The storage unit 120 has at least one of a volatile memory and a non-volatile memory. Examples of the volatile memory include DRAM (Dynamic Random Access Memory) and SRAM (Static Random Access Memory). Examples of the non-volatile memory include ReRAM (Resistive Random Access Memory, also referred to as a resistive change type memory), PRAM (Phase change Random Access Memory), FeRAM (Ferroelectric Random Access Memory), MRAM (Magnetoresistive Random Access Memory, also referred to as a magnetic resistance type memory), and flash memory. Also, the storage unit 120 can have at least one of NOSRAM (registered trademark) and DOSRAM (registered trademark). Further, the storage unit 120 can have a recording media drive. Examples of the recording media drive include a hard disk drive (HDD) and a solid state drive (SSD).
[0145] NOSRAM is an abbreviation for "Nonvolatile Oxide Semiconductor Random Access Memory (RAM)". NOSRAM refers to a memory in which the memory cell is a 2-transistor type (2T) or 3-transistor type (3T) gain cell, and the transistor is a transistor that uses a metal oxide for the channel formation region (also referred to as an OS transistor). The OS transistor has an extremely small current flowing between the source and the drain in the off state, that is, a leakage current. NOSRAM can be used as a non-volatile memory by holding charges corresponding to data in the memory cell using the characteristic of an extremely small leakage current. In particular, since NOSRAM can read the stored data without destroying it (non-destructive readout), it is suitable for arithmetic processing that repeatedly performs a large number of only data readout operations. Since the data capacity of NOSRAM can be increased by stacking it, the high performance of a semiconductor device can be achieved by using it as a large-scale cache memory, main memory, or storage memory.
[0146] DOSRAM is an abbreviation for "Dynamic Oxide Semiconductor RAM" and refers to a RAM having a 1T (transistor) 1C (capacitance) type memory cell. DOSRAM is a DRAM formed using an OS transistor, and DOSRAM is a memory that temporarily stores information sent from the outside. DOSRAM is a memory that utilizes the small off-current of the OS transistor.
[0147] In this specification and the like, a metal oxide is an oxide of a metal in a broad sense. Metal oxides are classified into oxide insulators, oxide conductors (including transparent oxide conductors), oxide semiconductors (also referred to as Oxide Semiconductor or simply OS), and the like. For example, when a metal oxide is used for the semiconductor layer of a transistor, the metal oxide may be referred to as an oxide semiconductor.
[0148] The metal oxide included in the channel formation region preferably contains indium (In). When the metal oxide included in the channel formation region is a metal oxide containing indium, the carrier mobility (electron mobility) of the OS transistor increases. Further, the metal oxide included in the channel formation region is preferably an oxide semiconductor containing element M. Element M is preferably at least one of aluminum (Al), gallium (Ga), and tin (Sn). Other elements applicable to element M include boron (B), silicon (Si), titanium (Ti), iron (Fe), nickel (Ni), germanium (Ge), yttrium (Y), zirconium (Zr), molybdenum (Mo), lanthanum (La), cerium (Ce), neodymium (Nd), hafnium (Hf), tantalum (Ta), and tungsten (W). However, there may be cases where a plurality of the aforementioned elements are combined as element M. Element M is, for example, an element having a high binding energy with oxygen. For example, it is an element having a higher binding energy with oxygen than indium. Further, the metal oxide included in the channel formation region is preferably a metal oxide containing zinc (Zn). A metal oxide containing zinc may be likely to crystallize.
[0149] The metal oxide included in the channel formation region is not limited to a metal oxide containing indium. The metal oxide included in the channel formation region may be, for example, a metal oxide containing zinc but not indium, such as zinc tin oxide or gallium tin oxide, a metal oxide containing gallium, or a metal oxide containing tin.
[0150] [Processing unit 130] The processing unit 130 has a function of performing processes such as calculation, analysis, and inference using data supplied from one or both of the input unit 110 and the storage unit 120. The processing unit 130 can supply the generated data (for example, calculation result, analysis result, inference result) to one or both of the storage unit 120 and the output unit 140.
[0151] The processing unit 130 has a function of acquiring data from the storage unit 120. Further, the processing unit 130 can have a function of recording or registering data in the storage unit 120.
[0152] The processing unit 130 can have, for example, an arithmetic circuit. The processing unit 130 can have, for example, a central processing unit (CPU). Further, the processing unit 130 can have a graphics processing unit (GPU).
[0153] The processing unit 130 can have a microprocessor such as a digital signal processor (DSP). The microprocessor can be realized by a programmable logic device (PLD) such as a field programmable gate array (FPGA) or a field programmable analog array (FPAA). Further, the processing unit 130 can have a quantum processor. The processing unit 130 can perform various data processes and program controls by interpreting and executing instructions from various programs by the processor. Programs that can be executed by the processor are stored in at least one of the memory area of the processor and the storage unit 120.
[0154] The processing unit 130 can have a main memory. The main memory has at least one of a volatile memory such as a RAM and a non-volatile memory such as a read only memory (ROM). Further, the main memory can have at least one of the above-described NOSRAM and DOSRAM.
[0155] As the RAM, for example, DRAM, SRAM, etc. are used, and a memory space is virtually allocated and used as the working space of the processing unit 130. The operating system, application programs, program modules, program data, look-up tables, etc. stored in the storage unit 120 are loaded into the RAM for execution. These data, programs, and program modules loaded into the RAM are directly accessed and operated on by the processing unit 130 respectively.
[0156] The ROM can store the BIOS (Basic Input / Output System), firmware, etc. that do not require rewriting. Examples of the ROM include mask ROM, OTPROM (One Time Programmable Read Only Memory), EPROM (Erasable Programmable Read Only Memory), etc. Examples of the EPROM include UV-EPROM (Ultra-Violet Erasable Programmable Read Only Memory) that enables erasure of stored data by ultraviolet irradiation, EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory, etc.
[0157] The processing unit 130 can have one or both of an OS transistor and a transistor having silicon in the channel formation region (Si transistor).
[0158] The processing unit 130 preferably has an OS transistor. Since the OS transistor has an extremely small off-current, by using it as a switch for holding the charge (data) flowing into the capacitive element that functions as a memory element, the data retention period can be ensured over a long term. By using this characteristic for at least one of the register and the cache memory that the processing unit 130 has, the processing unit can be operated only when necessary, and in other cases, the processing unit 130 can be turned off by saving the information of the previous processing in the storage element. That is, normal-off computing becomes possible, and power consumption of the information processing system can be reduced.
[0159] The information processing apparatus preferably uses AI for at least some of the processing.
[0160] In particular, the information processing apparatus preferably uses an artificial neural network (ANN: Artificial Neural Network, hereinafter also simply referred to as a neural network). The neural network is realized by a circuit (hardware) or a program (software).
[0161] In this specification and the like, the neural network refers to a model in general that mimics the neural circuit network of a living organism, determines the connection strength between neurons by learning, and has problem-solving ability. The neural network has an input layer, an intermediate layer (hidden layer), and an output layer.
[0162] In this specification and the like, when describing the neural network, determining the connection strength (also referred to as a weight coefficient) between neurons from existing information may be referred to as "learning".
[0163] In this specification and the like, constructing a neural network using the connection strength obtained by learning and deriving a new conclusion therefrom may be referred to as "inference".
[0164] [Output unit 140] The output unit 140 can output at least one of the calculation result, analysis result, and inference result in the processing unit 130 to the outside of the information processing apparatus. For example, the output unit 140 can transmit data via the network 51. Specifically, a device such as a personal computer equipped with a communication port or communication function can be used. Also, a device equipped with a communication function may be used for the input unit 110 and the output unit 140.
[0165] [Transmission path 150] The transmission path 150 has a function of transmitting data. The data transmission and reception between the input unit 110, the storage unit 120, the processing unit 130, and the output unit 140 can be performed via the transmission path 150. Specifically, a LAN or the Internet can be used.
[0166] Note that this embodiment can be appropriately combined with other embodiments shown in this specification.
[0167] (Embodiment 2) In this embodiment, an information processing method according to an aspect of the present invention will be described with reference to FIGS. 6 and 7.
[0168] FIG. 6 is a diagram for explaining an information processing method according to an aspect of the present invention.
[0169] FIG. 7 is a diagram for explaining an information processing method according to an aspect of the present invention.
[0170] <Example 1 of information processing method> An information processing method according to an aspect of the present invention includes steps S1 to S9 (see FIG. 6).
[0171] [Step S1] In step S1, the component 30 receives the questionnaire QRE from the user of the information processing system and passes the questionnaire QRE to the component 21.
[0172] [Step S2] In step S2, component 21 creates instruction PT1 and transfers it to component 20. Note that instruction PT1 includes questionnaire QRE.
[0173] [Step S3] In step S3, component 20 receives instruction PT1 and transfers response Drf to component 21. Note that component 20 has a function to perform processing using large language model LLM. Also, the large language model LLM has learned dataset DS, and the large language model LLM has a function to generate response Drf according to instruction PT1.
[0174] [Step S4] In step S4, component 21 uses questionnaire QRE as query qu to obtain search result SR from database DB. Note that database DB stores at least a part of the information not adopted in dataset DS.
[0175] [Step S5] In step S5, component 21 reviews response Drf using search result SR and generates review result ER.
[0176] [Step S6] In step S6, when review result ER is true, the process proceeds to step S7, and when review result ER is false, the process proceeds to step S8.
[0177] [Step S7] In step S7, component 21 creates answer sheet ANS1 using response Drf. Also, after transferring answer sheet ANS1 to component 30, the process proceeds to step S9.
[0178] [Step S8] In step S8, component 21 creates response document ANS1 using search result SR. After passing response document ANS1 to component 30, the process proceeds to step S9.
[0179] [Step S9] In step S9, component 30 provides response document ANS1 to the user of the information processing system.
[0180] Thereby, the information processing system according to one aspect of the present invention can evaluate the likelihood that hallucinations generated by the large language model LLM are included in the description of response document ANS1. Also, the information processing system according to one aspect of the present invention can determine the appropriateness of the description in response document ANS1. Also, the information processing system according to one aspect of the present invention can evaluate the reliability of the description in response document ANS1. Also, the information processing system according to one aspect of the present invention can review response Drf using information not adopted in dataset DS used for learning the large language model LLM. Also, the information processing system according to one aspect of the present invention can review response Drf using information collected from database DB using questionnaire QRE as query qu. Also, the information processing system according to one aspect of the present invention can reflect review result ER in response document ANS1. Also, the information processing system according to one aspect of the present invention can obtain an answer without determining whether dataset DS is appropriate or inappropriate for the content of questionnaire QRE. Also, the information processing system according to one aspect of the present invention can utilize the large language model LLM without unnecessary fine-tuning. Also, the information processing system according to one aspect of the present invention can utilize the large language model LLM without unnecessarily updating dataset DS and re-learning the large language model LLM. Also, the information processing system according to one aspect of the present invention can utilize the large language model LLM without using the search expansion generation method. Also, since it is not necessary to include search result SR in instruction text PT1, the degree of freedom of questionnaire QRE is high. As a result, it is possible to provide a novel information processing method excellent in convenience, usefulness, or reliability.
[0181] <Example 2 of Information Processing Method> The information processing method according to one aspect of the present invention includes steps S1 to S9 (see FIG. 6). Note that in Example 2 of the information processing method described in this embodiment, in step S4, when component 21 obtains the search result SR from the database DB, the reply answer Drf is used as the query qu instead of the questionnaire QRE, which is different from Example 1 of the information processing method. Here, the different parts will be described in detail, and for the same configurations, the above description will be incorporated by reference.
[0182] [Step S4] In step S4, component 21 uses the reply answer Drf as the query qu to obtain the search result SR from the database DB. Note that the database DB stores at least a part of the information not adopted in the data set DS.
[0183] Thereby, the information processing system according to one aspect of the present invention can review the reply answer Drf using the information not adopted in the data set DS used for the learning of the large language model LLM. Also, the information processing system according to one aspect of the present invention can review the reply answer Drf using the information collected from the database DB using the questionnaire QRE as the query qu and the information collected from the database DB using the reply answer Drf as the query qu. Also, the information processing system according to one aspect of the present invention can expand the review result ER and reflect it in the answer sheet ANS1. As a result, a novel information processing method excellent in convenience, usefulness, or reliability can be provided.
[0184] <Example 3 of Information Processing Method> The information processing method according to one aspect of the present invention includes steps S1 to S11 (see FIG. 7). Note that in Example 3 of the information processing method described in this embodiment, in step S6, the information processing process when it is determined that the review result ER is false is different from Example 1 or Example 2 of the information processing method. Here, the different parts will be described in detail, and for the same configurations, the above description will be incorporated by reference.
[0185] [Step S1] In step S1, component 30 receives the questionnaire QRE from the user of the information processing system and passes the questionnaire QRE to component 21.
[0186] [Step S2] In step S2, component 21 creates the instruction PT1 and passes the instruction PT1 to component 20. Note that the instruction PT1 includes the questionnaire QRE.
[0187] [Step S3] In step S3, component 20 receives the instruction PT1 and passes the response Drf to component 21. Note that component 20 has a function to perform processing using the large language model LLM. Also, the large language model LLM has learned the dataset DS, and the large language model LLM has a function to generate the response Drf according to the instruction PT1.
[0188] [Step S4] In step S4, component 21 uses the questionnaire QRE as a query qu to obtain the search result SR from the database DB. Note that the database DB stores at least a part of the information not adopted in the dataset DS.
[0189] [Step S5] In step S5, component 21 reviews the response Drf using the search result SR and generates the review result ER.
[0190] [Step S6] In step S6, when the review result ER is true, the process proceeds to step S7, and when the review result ER is false, the process proceeds to step S8.
[0191] [Step S7] In step S7, component 21 creates response document ANS1 using the response answer Drf. After passing response document ANS1 to component 30, the process proceeds to step S11.
[0192] [Step S8] In step S8, component 21 creates response document ANS1 using the search result SR. Also, response document ANS1 is passed to component 30. Further, component 21 creates instruction PT2 and passes it to component 20. Note that instruction PT2 includes the questionnaire QRE and the search result SR.
[0193] [Step S9] In step S9, component 20 receives instruction PT2 and passes response document ANS2 to component 21. Note that component 20 has a function to perform processing using the large language model LLM, and the large language model LLM has a function to generate response document ANS2 according to instruction PT2.
[0194] [Step S10] In step S10, component 21 receives response document ANS2, passes it to component 30, and then proceeds to step S11.
[0195] [Step S11] In step S11, when the review result ER is true, component 30 provides response document ANS1 to the user of the information processing system. Also, when the review result ER is false, component 30 provides response document ANS1 and response document ANS2 to the user of the information processing system.
[0196] As a result, the information processing system according to one aspect of the present invention can generate the answer document ANS2 using the search extension generation method when the large language model LLM generates hallucinations. Further, the information processing system according to one aspect of the present invention can prevent the generation of hallucinations using the search extension generation method. Further, when the search extension generation method is not used, the information processing system according to one aspect of the present invention does not need to create the instruction document PT2 including the question document QRE and the search result SR. Further, when the search extension generation method is not used, the information processing system according to one aspect of the present invention can input the question document QRE with a high degree of freedom. As a result, a novel information processing method excellent in convenience, usefulness, or reliability can be provided.
[0197] Note that this embodiment can be appropriately combined with other embodiments shown in this specification.
Explanation of Signs
[0198] DB Database Drf Answer QRE Question Document qu Query SR Search Result 20 Component 21 Component 30 Component 51 Network 110 Input Unit 120 Storage Unit 130 Processing Unit 140 Output Unit 150 Transmission Path
Claims
1. A first component; A second component; and and a third component, The first component has a function of receiving a questionnaire and transferring it to the third component; The questionnaire is written in a natural language; The first component has a function of accepting and providing a first response document; the second component has a function of receiving a first instruction sentence and passing a proposed answer to the third component; the second component has a function of performing processing using a large-scale language model; the large-scale language model has been trained on a dataset; The large-scale language model has a function of generating the answer plan according to the first instruction sentence, the third component has a function of creating the first instruction statement and transferring it to the second component; the first set of instructions includes the questionnaire; the third component has a function of performing processing using a search engine; the search engine has a function of retrieving search results from a database using the questionnaire as a query; The database stores at least a portion of the information not included in the data set; the third component has a function of evaluating the proposed answers using the search results to generate an evaluation result; the third component has a function of creating the first response document using the proposed response when the examination result is true, and transferring the first response document to the first component; The third component has a function of creating the first response document using the search result when the examination result is false, and transferring the first response document to the first component.
2. the third component comprises a morphological analyzer; The morphological analyzer extracts morphemes from the answer suggestions to create a first array; The morphological analyzer extracts morphemes from the search results to create a second array; the third component has a function of calculating a content rate of the morphemes included in the second sequence in the first sequence; The information processing system according to claim 1 , wherein the review result includes true or false determined based on the content rate.
3. the third component has a function of converting the answer proposal and the search result into a distributed representation and calculating a similarity; The information processing system according to claim 1 , wherein the examination result includes true or false determined based on the degree of similarity.
4. the third component comprises a textual entailment recognizer; The textual entailment recognizer has a function of determining whether or not the answer plan and the search result have a textual entailment relationship; The information processing system of claim 1 , wherein the examination result includes true or false determined based on a determination of the textual entailment recognizer.
5. The first component has a function of accepting and providing a second response document; the second component has a function of receiving a second instruction sentence and transferring the second response document to the third component; the large-scale language model has a function of generating the second response document in accordance with the second instruction sentence; the third component has a function of creating the second instruction sentence and transferring it to the second component when the checking result is false; the second instruction includes the questionnaire and the search results; 5. The information processing system according to claim 1, wherein the third component has a function of receiving the second response document and transferring it to the first component.
6. A first component; A second component; and and a third component, The first component has a function of receiving a questionnaire and transferring it to the third component; The questionnaire is written in a natural language; The first component has a function of accepting and providing a response document; The second component has a function of receiving an instruction sentence and passing a proposed answer to the third component; the second component has a function of performing processing using a large-scale language model; the large-scale language model has been trained on a dataset; The large-scale language model has a function of generating the answer proposal according to the instruction sentence, the third component has a function of creating the instruction statement and transferring it to the second component; The instructions include the questionnaire, the third component has a function of performing processing using a search engine; the search engine has a function of retrieving search results from a database using the answer proposal as a query; The database stores at least a portion of the information not included in the data set; the third component has a function of evaluating the proposed answers using the search results to generate an evaluation result; the third component has a function of creating the response document using the proposed response when the examination result is true, and transferring the response document to the first component; The third component has a function of creating the response document using the search result when the examination result is false, and transferring the response document to the first component.
7. An information processing method having first to ninth steps, In the first step, the first component receives a questionnaire and passes the questionnaire to the second component; In the second step, the second component creates an instruction statement and passes it to a third component; The instructions include the questionnaire, In the third step, the third component receives the instruction sentence and passes a proposed answer to the second component; the third component has a function of performing processing using a large-scale language model; the large-scale language model has been trained on a dataset; The large-scale language model has a function of generating the answer proposal according to the instruction sentence, In the fourth step, the second component retrieves search results from a database using the questionnaire as a query; The database stores at least a portion of the information not included in the data set; In the fifth step, the second component evaluates the proposed answers using the search results to generate an evaluation result; In the sixth step, when the checking result is true, the process proceeds to the seventh step, and when the checking result is false, the process proceeds to the eighth step; In the seventh step, the second component creates a response document using the proposed response and passes it to the first component, and then the process proceeds to the ninth step; In the eighth step, the second component creates the answer sheet using the search result and passes it to the first component, and then the process proceeds to the ninth step; In the ninth step, the first component provides the response document.
8. An information processing method having first to ninth steps, In the first step, the first component receives a questionnaire and passes the questionnaire to the second component; In the second step, the second component creates an instruction statement and passes it to a third component; The instructions include the questionnaire, In the third step, the third component receives the instruction sentence and passes a proposed answer to the second component; the third component has a function of performing processing using a large-scale language model; the large-scale language model has been trained on a dataset; The large-scale language model has a function of generating the answer proposal according to the instruction sentence, In the fourth step, the second component retrieves search results from a database using the proposed answers as queries; The database stores at least a portion of the information not included in the data set; In the fifth step, the second component evaluates the proposed answers using the search results to generate an evaluation result; In the sixth step, when the checking result is true, the process proceeds to the seventh step, and when the checking result is false, the process proceeds to the eighth step; In the seventh step, the second component creates a response document using the proposed response and passes it to the first component, and then the process proceeds to the ninth step; In the eighth step, the second component creates the answer sheet using the search result and passes it to the first component, and then the process proceeds to the ninth step; In the ninth step, the first component provides the response document.
9. An information processing method having first to eleventh steps, In the first step, the first component receives a questionnaire and passes the questionnaire to the second component; In the second step, the second component creates a first instruction statement and passes it to a third component; the first set of instructions includes the questionnaire; In the third step, the third component receives the first instruction sentence and passes a proposed answer to the second component; the third component has a function of performing processing using a large-scale language model; the large-scale language model has been trained on a dataset; The large-scale language model has a function of generating the answer plan according to the first instruction sentence, In the fourth step, the second component retrieves search results from a database using the questionnaire as a query; The database stores at least a portion of the information not included in the data set; In the fifth step, the second component evaluates the proposed answers using the search results to generate an evaluation result; In the sixth step, when the checking result is true, the process proceeds to the seventh step, and when the checking result is false, the process proceeds to the eighth step; In the seventh step, the second component creates a first answer document using the proposed answer and passes it to the first component, and then the process proceeds to the eleventh step; In the eighth step, the second component creates the first response document using the search result and passes it to the first component, and then creates a second instruction sentence and passes it to the third component; the second instruction includes the questionnaire and the search results; In the ninth step, the third component receives the second instruction sentence and passes a second response document to the second component; the large-scale language model has a function of generating the second response document in accordance with the second instruction sentence; In the tenth step, the second component receives the second response document and passes it to the first component, and then the process proceeds to the eleventh step; An information processing method in which, in the eleventh step, the first component provides the first response document when the examination result is true, and provides the first response document and the second response document when the examination result is false.