Information processing apparatus and information processing method
Patent Information
- Application Number
- JP2026032233
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-07
- Filing Date
- 2026-03-02
- Publication Date
- 2026-09-17
AI Technical Summary
【0021】 本発明の一態様により、知的財産に係る文書作成を補助する情報処理装置、情報処理方法、及び情報処理システムを提供できる。又は、本発明の一態様により、短時間で知的財産に係る文書作成を行うことができる情報処理装置、情報処理方法、及び情報処理システムを提供できる。又は、本発明の一態様により、ユーザが少ない労力で知的財産に係る文書作成を行うことができる情報処理装置、情報処理方法、及び情報処理システムを提供できる。又は、本発明の一態様により、知的財産に係る文書を完成度高く作成できる情報処理装置、情報処理方法、及び情報処理システムを提供できる。
Smart Images

Figure 2026148504000001_ABST
Abstract
Description
[Technical Field]
[0001] One aspect of the present invention relates to an information processing apparatus. One aspect of the present invention also relates to an information processing method using the information processing apparatus. Furthermore, one aspect of the present invention relates to an information processing system including the information processing apparatus.
[0002] Note that one aspect of the present invention is not limited to the above technical field. The technical field of one aspect of the invention disclosed in this specification and the like relates to an article, a method, or a manufacturing method. Alternatively, one aspect of the present invention relates to a process, a machine, a manufacture, or a composition of matter. Therefore, as more specific examples of the technical field of one aspect of the present invention disclosed in this specification, an information processing apparatus, a semiconductor device, a memory device, a driving method thereof, or a manufacturing method thereof can be given. [Background Art]
[0003] Documents related to patent applications, such as patent specifications, need to describe technical ideas while satisfying description requirements. This requires advanced knowledge and skills related to intellectual property, and is difficult for inexperienced persons. Patent Document 1 discloses a method for automatically complementing descriptions in patent specifications.
[0004] In recent years, there has been a surge in the development of language models using artificial neural networks (ANNs, also simply referred to as neural networks), with large-scale language models (LLMs) attracting particular attention. Large-scale language models are natural language processing models trained using large amounts of data. Large-scale language models can be used to realize, for example, dialogue models that respond to user instructions. Non-patent document 1 discloses GPT-4 (registered trademark) (Generative Pre-trained Transformer 4) as a large-scale language model, and ChatGPT as a dialogue model. In addition to the above, other large-scale language models include LaMDA (Language Model for Dialogue Applications), PaLM (Pathways Language Model), Llama2, and Llama3. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] Japanese Patent Publication No. 2022-24112 [Non-patent literature]
[0006] [Non-Patent Document 1] Summary of ChatGPT / GPT-4 Research and Perspective Towards the Future of Large Language Models, Yiheng Liu et al. (Submitted on 4 Apr 2023, [online], Internet)<URL:https: / / arxiv.org / abs / 2304.01852> [Overview of the Initiative] [Problems that the invention aims to solve]
[0007] For example, a patent specification must meet the requirements for description while differentiating it from the invention disclosed in prior art documents. This requires a thorough understanding of the invention in question and the invention disclosed in prior art documents, as well as a deep knowledge of patent law. Furthermore, even slight differences in the description of a patent specification can significantly affect the scope of patent rights. Therefore, utmost care must be taken in writing the patent specification. As a result, drafting a patent specification requires a great deal of time and effort, and it is difficult for someone with little experience to do so in a short amount of time. For example, it may be difficult for the inventor themselves to write a patent specification in between research and development work.
[0008] Therefore, one aspect of the present invention aims to provide an information processing device, an information processing method, and an information processing system that assist in the creation of documents relating to intellectual property. Alternatively, one aspect of the present invention aims to provide an information processing device, an information processing method, and an information processing system that can create documents relating to intellectual property in a short amount of time. Alternatively, one aspect of the present invention aims to provide an information processing device, an information processing method, and an information processing system that enable users to create documents relating to intellectual property with minimal effort. Alternatively, one aspect of the present invention aims to provide an information processing device, an information processing method, and an information processing system that can create documents relating to intellectual property with a high degree of completion.
[0009] Alternatively, one aspect of the present invention aims to provide an information processing device, an information processing method, and an information processing system that are excellent in convenience and usefulness. Alternatively, one aspect of the present invention aims to provide a novel information processing device, an information processing method, and an information processing system.
[0010] Furthermore, the description of these problems does not preclude the existence of other problems. Moreover, one aspect of the present invention does not need to solve all of these problems. Furthermore, those skilled in the art can naturally identify other problems from the description in the specification, drawings, claims, etc., and it is possible to extract other problems from the description in the specification, drawings, claims, etc. [Means for solving the problem]
[0011] One aspect of the present invention is an information processing method that receives a first document and a second document, obtains a summary of the second document using a first language model, obtains differences in the first document from the summary using a second language model, the second language model is trained using a first data point and a second data point, the first data point comprises a first command, a first training document, a second training document, and a first training response, the second data point comprises a first command, a third training document, a fourth training document, and a second training response, the first training response indicates the differences in the first training document from the second training document, and the second training response indicates that there are no differences in the third training document from the fourth training document.
[0012] Alternatively, in the above embodiment, differences from the summary are presented, a judgment is received on whether the differences from the summary are valid, if the differences from the summary are not valid, after receiving an instruction document, the differences from the summary in the first document are re-acquired using the second language model based on the instruction document, and then a judgment is received again on whether the differences from the summary are valid, if the differences from the summary are valid, the second language model is used to confirm that all of the differences from the summary are not described in the second document, and the second language model learns using the third and fourth data points. The third data point includes a second command, a fifth learning sentence, a sixth learning sentence, a first learning instruction document, and a third learning response. The fourth data point includes a second command, a seventh learning sentence, an eighth learning sentence, a second learning instruction document, and a fourth learning response. The third learning response indicates the differences between the fifth learning sentence and the sixth learning sentence, based on the first learning instruction document. The fourth learning response may indicate that the second learning instruction document is not appropriate as a basis for judging the differences between the seventh learning sentence and the eighth learning sentence.
[0013] Alternatively, one aspect of the present invention comprises steps 1 to 12, in which a first document and a second document are received; in the second step, a first prompt for generating a summary of the second document is created and sent to a first language model; in the third step, a second prompt for presenting differences in the first document to the summary is created and sent to a second language model; in the fourth step, a determination is received as to whether the differences are valid; if the differences are not valid, in the fifth step, a first instruction document is received. Then, in step 6, based on the first instruction document, create a third prompt to re-present the differences in the first document and send it to the second language model, then repeat step 4, and if the differences are valid, in step 7, create a fourth prompt to confirm that not all of the differences are described in the second document and send it to the second language model, if at least some of the differences are described in the second document, repeat step 5, and if not all of the differences are described in the second document, proceed to step 8. In the second step, a fifth prompt is created and sent to the first language model to generate a second document and a third document based on the differences. In the ninth step, a decision is received as to whether the third document needs to be modified. If it does, in the tenth step, a second instruction document is received. In the eleventh step, a sixth prompt is created and sent to the first language model to modify the third document based on the second instruction document, and then the ninth step is repeated. If it does not need to be modified, in the twelfth step, the third document is output, and the second language... The word model is trained using a first data point and a second data point, the first data point having a first command, a first training sentence, a second training sentence, and a first training response sentence, the second data point having a first command, a third training sentence, a fourth training sentence, and a second training response sentence, the first training response sentence indicating the differences between the first training sentence and the second training sentence, and the second training response sentence indicating that there are no differences between the third training sentence and the fourth training sentence, this is an information processing method.
[0014] Alternatively, in the above embodiment, steps 13 through 15 are performed, and if not all of the differences with respect to the summary are described in the second document, in step 13, before step 8, a seventh prompt is created to determine whether there is insufficient information to generate the third document and sent to the third language model; if there is insufficient information, in step 14, after receiving additional information, in step 15, the seventh prompt is created again, including the additional information, and sent to the first language model; if there is no insufficient information, step 8 is repeated and the third language model is generated The system is trained using a fifth data point and a sixth data point, the fifth data point having a third command, a ninth training sentence, a tenth training sentence, and a fifth training response, the sixth data point having a third command, an eleventh training sentence, a twelfth training sentence, and a sixth training response, the fifth training response may indicate missing information when generating a document based on the ninth training sentence and the tenth training sentence, and the sixth training response may indicate that there is no missing information when generating a document based on the eleventh training sentence and the twelfth training sentence.
[0015] Alternatively, one aspect of the present invention includes a receiving unit, an output unit, and a processing unit, wherein the receiving unit has the function of receiving a first document, a second document, a first instruction document, and a second instruction document; the output unit has the function of supplying a first prompt, a fifth prompt, and a sixth prompt to a first language model, and a second prompt, a third prompt, and a fourth prompt to a second language model; the output unit has the function of outputting a third document; and the processing unit has the function of supplying a first prompt for generating a summary of the second document. The second language model has the following functions: creating a data point; creating a second prompt to present the differences in the summary in the first document; having the user determine whether the differences are valid; creating a third prompt to re-present the differences in the first document based on the first instruction document if the differences are not valid; creating a fourth prompt to confirm that not all of the differences are described in the second document if the differences are valid; creating a third prompt based on the first instruction document if at least some of the differences are described in the second document; creating a fifth prompt to generate the second document and the third document based on the differences if not all of the differences are described in the second document; having the user determine whether the third document needs to be modified; and creating a sixth prompt to modify the third document based on the second instruction document if it needs to be modified. The second language model has the following functions: creating a data point, and the second An information processing device that learns using data points, wherein the first data point has a first command statement, a first learning sentence, a second learning sentence, and a first learning response statement, the second data point has a first command statement, a third learning sentence, a fourth learning sentence, and a second learning response statement, the first learning response statement indicates the differences between the first learning sentence and the second learning sentence, and the second learning response statement indicates that there are no differences between the third learning sentence and the fourth learning sentence.
[0016] Alternatively, in the above embodiment, the second language model is trained using a third data point and a fourth data point, wherein the third data point comprises a second command, a fifth learning sentence, a sixth learning sentence, a first learning instruction document, and a third learning response sentence, and the fourth data point comprises a second command, a seventh learning sentence, an eighth learning sentence, a second learning instruction document, and a fourth learning response sentence, wherein the third learning response sentence indicates the differences between the fifth learning sentence and the sixth learning sentence, based on the first learning instruction document, and the fourth learning response sentence indicates that the second learning instruction document is not appropriate as a basis for judging the differences between the seventh learning sentence and the eighth learning sentence.
[0017] Alternatively, in the above embodiment, the first instruction document may include at least one of instructions to amend the first document and comments on differences with respect to the summary, and the second instruction document may include at least one of instructions to amend the third document and comments on the third document.
[0018] Alternatively, in the above embodiment, the receiving unit has a function to receive additional information, the output unit has a function to supply a seventh prompt to the third language model, the processing unit has a function to create a seventh prompt for determining whether there is insufficient information to generate the third document when not all of the differences with respect to the summary sentence are described in the second document, to recreate the seventh prompt including additional information if there is insufficient information, and to create a fifth prompt if there is no insufficient information, the third language model has a fifth data point, and The system is trained using the sixth data point, the fifth data point having the third command, the ninth training text, the tenth training text, and the fifth training response, the sixth data point having the third command, the eleventh training text, the twelfth training text, and the sixth training response, the fifth training response may indicate missing information when generating a document based on the ninth training text and the tenth training text, and the sixth training response may indicate that there is no missing information when generating a document based on the eleventh training text and the twelfth training text.
[0019] Alternatively, in the above embodiment, the first document and the second document may each be documents indicating a creation.
[0020] Alternatively, in the above embodiment, the second document may be a patent document or a utility model document, and the third document may be a specification belonging to a patent application or a specification belonging to a utility model registration application. [Effects of the Invention]
[0021] According to one aspect of the present invention, there can be provided an information processing apparatus, an information processing method, and an information processing system that assist document creation relating to intellectual property. Alternatively, according to one aspect of the present invention, there can be provided an information processing apparatus, an information processing method, and an information processing system that can create a document relating to intellectual property in a short time. Alternatively, according to one aspect of the present invention, there can be provided an information processing apparatus, an information processing method, and an information processing system that allow a user to create a document relating to intellectual property with less effort. Alternatively, according to one aspect of the present invention, there can be provided an information processing apparatus, an information processing method, and an information processing system that can create a document relating to intellectual property with high completeness.
[0022] Alternatively, according to one aspect of the present invention, there can be provided an information processing apparatus, an information processing method, and an information processing system excellent in convenience and usefulness. Alternatively, according to one aspect of the present invention, there can be provided a novel information processing apparatus, a novel information processing method, and a novel information processing system.
[0023] It should be noted that the description of these effects does not preclude the existence of other effects. One aspect of the present invention does not necessarily need to have all of these effects. Effects other than these can be naturally found by those skilled in the art from the description of the specification, drawings, claims, etc., and it is possible to extract effects other than these from the description of the specification, drawings, claims, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] [Figure 1] FIG. 1 is a schematic diagram showing an example of an information processing system. [Figure 2] FIG. 2 is a block diagram showing an example of an information processing system. [Figure 3] FIG. 3 is a flow diagram showing an example of an information processing method. [Figure 4] FIG. 4 is a flow diagram showing an example of an information processing method. [Figure 5] FIG. 5 is a schematic diagram showing an example of an information processing method. [Figure 6] FIGS. 6(A) and 6(B) are schematic diagrams showing an example of an information processing method. [Figure 7] Figures 7(A) and 7(B) are schematic diagrams illustrating an example of an information processing method. [Figure 8] Figures 8(A) and 8(B) are schematic diagrams illustrating an example of an information processing method. [Figure 9] Figures 9(A) and 9(B) are schematic diagrams illustrating an example of an information processing method. [Figure 10] Figure 10 is a flowchart illustrating an example of an information processing method. [Figure 11] Figures 11(A) and 11(B) are schematic diagrams illustrating an example of an information processing method. [Figure 12] Figure 12 is a flowchart illustrating an example of an information processing method. [Figure 13] Figure 13 is a schematic diagram illustrating an example of an information processing method. [Modes for carrying out the invention]
[0025] Embodiments will be described in detail with reference to the drawings. However, it will be readily apparent to those skilled in the art that the present invention is not limited to the following description, and that its form and details can be modified in various ways without departing from the spirit and scope of the present invention. Accordingly, the present invention is not to be interpreted as being limited to the descriptions of the embodiments shown below. In the configuration of the invention described below, the same reference numerals are used in common across different drawings for the same parts or parts having similar functions, and repeated descriptions are omitted.
[0026] In the drawings attached to this specification, components are classified by function and shown as independent blocks in block diagrams. However, in reality, it is difficult to completely separate components by function, and one component may be involved in multiple functions.
[0027] In this specification, the terms "first" and "second" may be used for convenience in understanding the technical content or to identify each component. Therefore, the terms "first" and "second" do not limit the number of components. Nor do the terms "first" and "second" limit the order of the components. Furthermore, the terms "first" and "second" or the identification codes used in this specification may not correspond to the terms or identification codes in the claims of this patent.
[0028] (Embodiment) This embodiment describes an information processing system according to one aspect of the present invention. Furthermore, this embodiment describes an information processing device included in the information processing system and an information processing method using the information processing system.
[0029] One aspect of the present invention relates to an information processing system that generates a third document using a language model based on a first document and a second document. The first document and the second document are documents that demonstrate a creation, for example, a document that demonstrates an invention or a utility model. Specifically, the first document may be a document that demonstrates a creation for which an application is filed. The second document may be a prior art document, for example, a patent document or a utility model document. The third document may be a patent application document or a utility model application document, for example, a specification belonging to a patent application or a specification belonging to a utility model registration application. Here, the specification belonging to a patent application and the specification belonging to a utility model registration application are collectively referred to simply as a specification. The second document and the third document may be, for example, a research paper. In this case, the first document may be a document that demonstrates research results.
[0030] Both inventions and utility models are creations of technical ideas that utilize the laws of nature. Therefore, the term "creation" includes both inventions and utility models. Here, inventions are protected by patent law, while utility models are protected by utility model law.
[0031] In one aspect of the present invention, an information processing method first generates a summary of a second document using a text generation language model. Then, a difference-displaying language model presents the differences between the first document and the summary of the second document. Specifically, these differences can be features described in the first document but not in the summary of the second document, such as technical features.
[0032] The difference-displaying language model is a finely tuned model using multiple data points, each containing an instruction, two training texts, and a training response. The instruction includes an instruction that presents the differences between the two training texts. One of the two training texts corresponds to the first document mentioned above. The other of the two training texts corresponds to a summary of the second document mentioned above. Of the training response sentences included in each of the multiple data points, some training response sentences indicate that there are no differences between the two training texts. The remaining training response sentences indicate differences between the two training texts.
[0033] By fine-tuning the language model used for highlighting differences, the occurrence of hallucination can be suppressed compared to cases where fine-tuning is not performed. In particular, fine-tuning can prevent the language model from highlighting differences even when there are no differences between the first document and the summary of the second document. Therefore, even someone with limited experience in intellectual property work can create a high-quality specification.
[0034] In this specification, hallucination refers to the phenomenon in which artificial intelligence (AI) generates information that is not based on facts.
[0035] The user of the information processing system determines whether the above-mentioned differences are valid. If the differences are not valid, the user of the information processing system creates a first instruction document. Based on the first instruction document, the language model for presenting differences presents the differences in the first document to the summary of the second document again. The first instruction document may include, for example, instructions to revise the first document and at least one of the above-mentioned differences. Here, the above-mentioned comments to the differences may be, for example, opinions on the above-mentioned differences. For example, if the user of the information processing system believes that there are differences in the first document to the second document other than the differences presented, a statement to that effect may be included in the first instruction document. Also, if the user of the information processing system believes that at least a part of the presented differences are described in the second document, a statement to that effect may be included in the first instruction document.
[0036] If the aforementioned differences are valid, the language model for presenting differences will verify that not all of those differences are described in the second document. The verification results will be presented to the user of the information processing system.
[0037] If at least some of the differences are described in the second document, the user of the information processing system creates the first instruction document, as in the case where the differences described above are not valid. If not all of the differences are described in the second document, the language model for document generation generates the second document and a third document based on those differences.
[0038] The user of the information processing system reviews the third document. For example, if the third document needs to be modified, the user of the information processing system creates a second instruction document. The text generation language model modifies the third document based on the second instruction document. Here, the second instruction document may include comments regarding the contents of the third document. Such comments may be, for example, questions regarding the contents of the third document. In this case, the text generation language model can generate answers to such comments. Note that the text generation language model may generate only answers to comments and not modify the third document. When the third document is completed and no further modifications are needed, the third document is output. For example, the third document is supplied to an information terminal owned by the user of the information processing system.
[0039] As described above, by using the information processing system according to one aspect of the present invention, a third document can be generated, taking into account the differences between the first document and the second document. For example, the specification relating to the creation shown in the first document can be written in a way that emphasizes technical features not disclosed in the prior art documents. As described above, the information processing system according to one aspect of the present invention can assist in the creation of documents related to intellectual property, such as the creation of specifications. Therefore, even a person with little experience in intellectual property work can create documents related to intellectual property, such as specifications, in a short time by using the information processing system according to one aspect of the present invention. Furthermore, users can create documents related to intellectual property, such as specifications, with less effort.
[0040] Furthermore, in one aspect of the information processing system of the present invention, when a user creates a first instruction document, the difference-displaying language model can present the differences between the first document and the second document multiple times. Specifically, the difference-displaying language model can present the above-mentioned differences multiple times, taking into account the content of the first instruction document. Moreover, in one aspect of the information processing system of the present invention, when a user creates a second instruction document, the document generation language model can modify the third document one or more times. As a result, a user of one aspect of the information processing system of the present invention can create the third document in an interactive format. This allows, for example, a person with little experience in intellectual property work to create a highly complete specification.
[0041] <Example of an information processing system configuration 1> Figure 1 is a schematic diagram showing an example configuration of an information processing system according to one aspect of the present invention. The information processing system according to one aspect of the present invention comprises an information processing device 10 and an information processing device 40. The information processing device 10 and the information processing device 40 are connected via a network 30 and can transmit and receive data.
[0042] The information processing device 40 has a function to perform processing using a language model. The language model has a function to generate a response sentence based on a prompt. A prompt can be described as an input sentence that causes the language model to perform a desired action. The language model processes the prompt by separating it into tokens (tokenization). For tokenization, word tokenization, character tokenization, or subword tokenization can be used.
[0043] A prompt is supplied, for example, from the information processing device 10 to the information processing device 40. The response sentence generated by the language model is supplied, for example, from the information processing device 40 to the information processing device 10. The information processing device 10 may also have a function to perform processing using the language model.
[0044] Furthermore, in one embodiment of the present invention, the information processing system may be configured so that a user can input documents, etc., by directly operating the information processing device 10, or, as shown in Figure 1, it may be configured so that a user can input documents, etc., using an information terminal 20 connected to the information processing device 10 via a network 30.
[0045] The following describes an example configuration of the information processing device 10, the information processing device 40, the information terminal 20, and the network 30.
[0046] 《Example of the configuration of the information processing device 10》 Figure 2 is a block diagram showing an example configuration of the information processing device 10. The information processing device 10 includes a receiving unit 110, a storage unit 120, a processing unit 130, an output unit 140, and a transmission line 150. Figure 2 also shows an information terminal 20 and an information processing device 40.
[0047] [Reception Desk 110] The reception unit 110 has the function of receiving data from outside the information processing device 10. The reception unit 110 has the function of receiving data from the information terminal 20, for example, data indicating documents, etc. The reception unit 110 has the function of receiving data from the information processing device 40, for example, data indicating response statements, etc.
[0048] The receiving unit 110 has the function of supplying the received data to one or both of the storage unit 120 and the processing unit 130 via the transmission line 150. For example, a wired communication port, wireless communication port, or optical communication port can be used as the receiving unit 110.
[0049] [Storage section 120] The storage unit 120 has the function of storing the program executed by the processing unit 130. The storage unit 120 may also have the function of storing data created by the processing unit 130 (for example, calculation results, analysis results, and inference results), and data received by the receiving unit 110.
[0050] The storage unit 120 may have a database. The information processing device 10 may also have a database separate from the storage unit 120. The information processing device 10 may have a function to retrieve data from a database located outside the storage unit 120, outside the information processing device 10, or outside the information processing system. Furthermore, the information processing device 10 may have a function to retrieve data from both its own database and an external database.
[0051] Either or both of the storage and / or file server can be used in the storage unit 120. Furthermore, a database recording the paths of files stored on the file server can also be used in the storage unit 120.
[0052] [Processing step 130] The processing unit 130 has the function of performing calculations, analyses, inferences, and other processing using data supplied from either or both of the receiving unit 110 and the storage unit 120. The processing unit 130 can supply the created data (e.g., calculation results, analysis results, and inference results) to either or both of the storage unit 120 and the output unit 140. The processing unit 130 also has the function of creating prompts. The processing unit 130 may also have the function of performing processing using a language model.
[0053] The processing unit 130 has the function of acquiring data from the storage unit 120. The processing unit 130 may also have the function of recording or registering data in the storage unit 120.
[0054] [Output section 140] The output unit 140 has the function of outputting at least one of the calculation results, analysis results, and inference results from the processing unit 130 to the outside of the information processing device 10. As the output unit 140, devices such as a wired communication port, wireless communication port, or optical communication port can be used.
[0055] For example, the output unit 140 has the function of supplying data indicating prompts, etc., to the information processing device 40. The output unit 140 also has the function of supplying data indicating documents, etc., to the information terminal 20.
[0056] [Transmission path 150] The transmission line 150 has the function of transmitting data. Data can be transmitted and received between the receiving unit 110, the storage unit 120, the processing unit 130, and the output unit 140 via the transmission line 150. As the transmission line 150, for example, a bus line on a motherboard, a wired communication cable, or an optical communication cable can be used.
[0057] 《Example of the configuration of the information processing device 40》 The information processing device 40 can process the received data and transmit the processing results. For example, it can perform calculations and other processing using the data supplied from the information processing device 10. Furthermore, the information processing device 40 can supply the processing results to the information processing device 10. This reduces the computational burden on the information processing device 10.
[0058] As described above, the information processing device 40 can perform processing using language models. For example, it can perform processing using language models such as BERT (Bidirectional Encoder Representations from Transformers) and T5 (Text-to-Text Transfer Transformer). In addition, the information processing device 40 can perform processing using models that utilize language models (text generation models, dialogue models, etc.).
[0059] As described above, the language model generates a response sentence based on the prompt. For example, a prompt created by the processing unit 130 is supplied to the language model of the information processing device 40 via the output unit 140. The response sentence generated by the language model is supplied via the receiving unit 110 to, for example, one or both of the storage unit 120 and the output unit 140.
[0060] Furthermore, the information processing device 40 can perform processing using a general-purpose language processing model that can handle various natural language processing tasks.
[0061] The information processing device 40 is a large computer such as a server computer or a supercomputer. Preferably, the information processing device 40 also has the functionality of a parallel computer. By using the information processing device 40 as a parallel computer, for example, large-scale calculations necessary for artificial intelligence learning and inference can be performed.
[0062] Furthermore, the information processing device 40 is a computer with higher processing power compared to the information processing device 10. For example, if both the information processing device 10 and the information processing device 40 have the functionality of parallel computers, the information processing device 40 will have higher processing power than the information processing device 10 and will be able to perform large-scale calculations. Also, for example, if both the information processing device 10 and the information processing device 40 can perform processing using models that utilize language models, the information processing device 40 will be able to perform processing using larger-scale models compared to the information processing device 10.
[0063] Furthermore, service providers are not necessarily required to own the information processing device 40 themselves. For example, a service provider can utilize some of the services provided by other businesses using the information processing device 40.
[0064] 《Example configuration of information terminal 20》 The information terminal 20 can receive data input by a user of an information processing system according to one embodiment of the present invention. Furthermore, the information terminal 20 can present data output by the information processing system according to one embodiment of the present invention to the user by displaying it on the display unit of the information terminal 20. Alternatively, the information terminal 20 can present data output by the information processing system according to one embodiment of the present invention to the user by printing it on the printing unit of the information terminal 20.
[0065] Furthermore, the information terminal 20 can supply data received from the user to the information processing device 10. The information terminal 20 can also present data supplied from the information processing device 10 to the user.
[0066] Furthermore, the information terminal 20 can supply data created based on data received from the user to the information processing device 10. The information terminal 20 can also present data created based on data supplied by the information processing device 10 to the user.
[0067] The information terminal 20 has, for example, dedicated application software, a web browser, etc. installed on it. The user can access the information processing device 10 through either of these. As a result, the user can enjoy services using an information processing system according to one embodiment of the present invention, using, for example, a computer with lower processing power than the information processing device 10 as the information terminal 20.
[0068] The information terminal 20 can also be referred to as a client computer or the like. In any case, the information terminal 20 is an information terminal device used by a user of an information processing system according to one embodiment of the present invention.
[0069] For example, a desktop computer 20a, a notebook computer 20b, a smartphone 20c, or a tablet computer 20d can be used as the information terminal 20. The tablet computer 20d can also be used as a notebook computer by connecting it to a casing 21 with a keyboard.
[0070] Network 30 Network 30 connects information processing device 10 and information processing device 40. Network 30 also connects multiple information terminals 20 to information processing device 10. This enables the transmission and reception of input and processed data between the two. Furthermore, it allows for the distribution of the information processing load.
[0071] The following describes an information processing method using the information processing system shown in Figures 1 and 2. Specifically, the operation method of the information processing device 10 will be described. In one embodiment of the present invention, a third document is generated using a language model based on a first document and a second document.
[0072] <Information Processing Method 1> Figures 3 and 4 are flowcharts illustrating an example of an information processing method according to one aspect of the present invention, and more specifically, flowcharts illustrating an example of the operation method of the information processing device 10.
[0073] An information processing method according to one embodiment of the present invention is "started," and in step S101 of Figure 3, the receiving unit 110 receives a first document and a second document. The first document can be a document showing the creation for which an application is to be filed, specifically an invention, a design, etc. The second document can be a prior art document, as described above, and can be, for example, a patent document or a utility model document. Examples of patent documents include published patent gazettes and patent publications. Examples of utility model documents include utility model gazettes. However, the second document is not limited to patent documents and utility model documents. The second document may be a design document, or a paper, book, magazine, etc. Furthermore, the second document is not limited to published documents, but may be unpublished documents.
[0074] The second document may consist of one or more references. The second document may, for example, be at least a portion of references stored in a database. The second document may, for example, be references retrieved from the database based on the first document.
[0075] In this specification, "search" refers to finding documents that are highly relevant to the search query document from among multiple searchable documents. In the example above, the search query document is the first document.
[0076] When performing a search, the first document and the references may be converted into vector data, and the similarity between the vector data may be calculated. Vector data represents multidimensional numerical data composed of integers from 0 to 9, as opposed to text data consisting of strings of characters (natural language) such as sentences. Vector data can also be described as data in a format that can be processed using arithmetic. The conversion of text data to vector data can be performed using AI, and it is preferable to do so using a neural network, for example. A neural network can be implemented by a circuit (hardware) or a program (software). Specific methods for converting text data to vector data include Bag of Words, distributed representation, and embedding representation. Furthermore, cosine similarity can be used as one of the indicators to represent the similarity between vector data. In addition, query rewriting, reranking, and agentization may be used for the search.
[0077] In this specification, the term "neural network" refers to any model that mimics the neural network of living organisms, determines the strength of connections between neurons through learning, and possesses problem-solving capabilities. A neural network has an input layer, an intermediate layer (hidden layer), and an output layer.
[0078] In this specification and other documents, when discussing neural networks, the process of determining the connection strength (also called weight coefficient) between neurons from existing information is sometimes referred to as "learning."
[0079] In this specification and other documents, the process of constructing a neural network using connection strengths obtained through learning and deriving new conclusions from it may be referred to as "inference."
[0080] Furthermore, a search may be performed using a search expression. A "search expression" is a string that contains at least one search term. A search expression may contain multiple search terms. A search expression may also contain search conditions, classification codes, etc. Here, for example, a search expression may include at least a portion of the string contained in the first document. Also, the search term may include, for example, the creative field indicated by the first document.
[0081] Figure 5 shows an example of search results. In the example shown in Figure 5, the search results are shown as Table 200. Table 200 can be displayed on the display unit of the information terminal 20. Figure 5 can also be considered an example of a graphical user interface (GUI) related to an information processing system according to one embodiment of the present invention. The forms, icons, tables, etc. in the drawings relating to the GUI exemplified in this embodiment, such as Figure 5, are examples and are not particularly limited. The GUI can be configured as a web page accessed by a user of the information processing system via the network 30. Alternatively, the GUI can be configured as the screen of a program application executed on the information terminal 20.
[0082] Table 200 has columns for "Reference Number" and "Content." The "Reference Number" column shows a number that identifies the searched document. In Figure 5, "xxx," "yyy," and "zzz" are shown as reference numbers.
[0083] The "Content" column displays at least a portion of the content of the searched document. Examples of document numbers include application numbers, publication numbers, and registration numbers. Alternatively, a company's internal identification number may be used. The content can include, for example, one or both of the text and / or drawings contained in the document. For example, in the case of patent documents or utility model documents, the abstract may be shown in the "Content" column.
[0084] Furthermore, the content of a document may be displayed by selecting a document number, for example. For example, the content of a document may be displayed by clicking on a document number. For example, the specification, drawings, claims, etc. of a document may be displayed by clicking on a document number. Note that the document number may be selected using a keyboard, for example. Also, if the information terminal 20 has a touch panel, the document number may be selected by touching it. The selection operations shown below can also be performed in the same manner. Furthermore, in the GUI shown below, when information about a document, such as a document number, is displayed, the content of the document may be displayed by selecting the displayed information.
[0085] Table 200 shows a checkbox 201 for each document. Documents to be designated as the second document can be selected by checking the checkbox 201. Figure 5 shows a button 203 labeled "Select". By selecting button 203, the documents with checkboxes 201 checked can be accepted as the second document by, for example, the reception unit 110. Note that instead of checkboxes 201, radio buttons or text boxes may be displayed.
[0086] In the example shown in Figure 5, checkboxes 201 for the document with document number "xxx" and the document with document number "zzz" are checked, i.e., selected. When button 203 is selected in this state, the document with document number "xxx" and the document with document number "zzz" are supplied as second documents to the reception unit 110, for example, from the database. For example, button 203 can be selected by clicking it. Alternatively, button 203 can be selected using the keyboard. Furthermore, if the information terminal 20 has a touch panel, button 203 can be selected by touching it. The buttons shown below can be selected in the same manner.
[0087] Here, Table 200 can include not only documents retrieved from the database based on the first document, but also documents specified by, for example, the user of the information processing system. The user can specify the documents to be included in Table 200, for example, by specifying the document number. Furthermore, documents retrieved using, for example, a search expression specified by the user can also be included in Table 200.
[0088] Next, in step S102 of Figure 3, the processing unit 130 creates a first prompt and transmits it to the information processing device 40 via the output unit 140. Specifically, the first prompt is transmitted to the language model of the information processing device 40. The first prompt is a prompt for generating a summary of the second document. The first prompt contains the second document. The first prompt also contains an instruction such as "Please summarize the following document." Here, the language model to which the first prompt is transmitted is called the document generation language model.
[0089] In this specification, the information contained in a prompt is referred to as context. A prompt has an instruction and a context. The context may be information related to the instruction. In the first prompt, the second document may be the context.
[0090] The information processing device 40 generates a first response sentence based on a first prompt using a text generation language model and supplies it to the receiving unit 110. The information processing device 10 then obtains the first response sentence. The first response sentence includes a summary of the second document. If the first prompt includes multiple second documents, the text generation language model can generate a summary for each of the multiple second documents and supply each to the receiving unit 110.
[0091] Furthermore, when the second document is a patent document or a utility model document, it is preferable to generate the abstract using only the specification, for example. That is, it is preferable to generate the abstract without using the claims and abstract included in the patent document or utility model document. In this case, the command in the first prompt can be, for example, "Summarize the specification included in the following document." Furthermore, the claims, abstract, etc. included in the patent document or utility model document do not need to be included in the first prompt. Also, when the second document is a design publication, the language model for text generation can generate the abstract using, for example, the description of the article to which the design relates and / or the description of the design.
[0092] Next, in step S103 of Figure 3, the processing unit 130 creates a second prompt and transmits it to the information processing device 40 via the output unit 140. Specifically, the second prompt is transmitted to the language model of the information processing device 40. The second prompt is a prompt for presenting the differences between the first document and the summary of the second document. Specifically, the differences can be features described in the first document that are not described in the summary, such as technical features, as described above. Here, the language model to which the second prompt is transmitted is called the difference-presentation language model. The difference-presentation language model may be a different language model from the document generation language model, or it may be the same language model.
[0093] The second prompt has the first document and a summary of the second document as context. That is, the second prompt has the first document and a summary of the document selected in the GUI shown in Figure 5, for example. Here, the second prompt can be used to present differences in features not mentioned in any of the summaries when there are multiple summaries of the second document, i.e., when multiple documents are selected in the GUI shown in Figure 5, for example. In this case, the second prompt has a command statement such as, "Please present the differences in the documents among the features described in the input document that are not mentioned in any of the document summaries. Please present the differences in a bulleted list." Here, the input document refers to the first document. The second prompt may also be used to present differences in each of the multiple summaries compared to the first document.
[0094] The information processing device 40 generates a second response sentence based on the second prompt using a language model for presenting differences and supplies it to the receiving unit 110. As a result, the information processing device 10 obtains the second response sentence. The second response sentence contains the differences between the first document and the summary of the second document.
[0095] The language model for highlighting differences is a finely tuned model, specifically a model trained using multiple first data points and multiple second data points. The first data point consists of a first command, a first training sentence, a second training sentence, and a first training response. The second data point consists of a first command, a third training sentence, a fourth training sentence, and a second training response.
[0096] The first instruction can be the same as the instruction given by the second prompt described above. The first data point and the second data point can each have the same instruction.
[0097] The first and third learning materials correspond to the first document mentioned above. The second and fourth learning materials correspond to summaries of the second document mentioned above.
[0098] The first learning response is a response that shows the differences between the first learning text and the second learning text. The second learning response is a response that shows that there are no differences between the third learning text and the fourth learning text.
[0099] In the fine-tuning described above, the language model for highlighting differences is trained so that when a prompt containing the first command, the first training text, and the second training text is input, the first training response is output. Furthermore, the language model for highlighting differences is trained so that when a prompt containing the first command, the third training text, and the fourth training text is input, the second training response is output. Fine-tuning can be performed using supervised learning with the first and second training response sentences as labels.
[0100] As described above, by fine-tuning the language model for presenting differences, the occurrence of hallucination can be suppressed compared to the case where no fine-tuning is performed. In particular, by performing fine-tuning using a second data point that includes a second training response sentence, it is possible to suppress the language model for presenting differences from presenting differences even when there are no differences between the first document and the summary of the second document. It is preferable that the number of first data points used to train the language model for presenting differences be greater than or equal to the number of second data points, more preferably twice or more the number of second data points, and even more preferably three times or more. By increasing the number of first data points, the language model for presenting differences can more easily and accurately present the differences between the first document and the summary of the second document.
[0101] Next, in step S104 of Figure 3, the processing unit 130 presents the differences presented by the difference-presentation language model to the user of the information processing system. This allows the processing unit 130 to allow the user to determine whether the differences are valid or not. For example, the processing unit 130 creates data containing the content indicated by the second response statement and supplies it to the information terminal 20 via the output unit 140. As a result, the differences are displayed on the display unit of the information terminal 20.
[0102] Figure 6(A) is a diagram showing an example of information displayed on the display unit of the information terminal 20 in step S104. Figure 6(A) can also be considered an example of a GUI related to an information processing system according to one embodiment of the present invention. The following diagrams showing examples of information displayed on the display unit can also be considered examples of a GUI related to an information processing system according to one embodiment of the present invention.
[0103] In the example shown in Figure 6(A), the differences between the first document and the summary of the second document are shown as "Differences to the References." The differences shown as "Differences to the References" are denoted as difference 211. In the example shown in Figure 6(A), difference 211 includes "aaa.", "bbb.", and "ccc.". Here, "aaa." is designated as difference 211[1], "bbb." as difference 211[2], and "ccc." as difference 211[3]. Note that similarities may also be shown in addition to differences to the references. In this case, the second prompt is a prompt that presents not only the differences between the first document and the summary of the second document, but also the similarities. Specifically, the command in the second prompt includes a sentence indicating that the similarities will be presented.
[0104] Furthermore, the example shown in Figure 6(A) shows Table 212. Table 212 has a "Document Number" column and a "Summary" column. The "Document Number" column shows a number that identifies the second document, for example, the document number selected in the GUI shown in Figure 5. In Figure 6(A), "xxx" and "zzz" are shown as document numbers. The "Summary" column shows the summary generated in step S102 of Figure 3. Although not shown in Figure 6(A), the first document may also be displayed on the display unit of the information terminal 20.
[0105] In the example shown in Figure 6(A), the question "Is the difference valid?" is presented. A button 213 labeled "Yes" and a button 214 labeled "No" are also shown. The user of the information processing system selects button 213 if they determine that the difference 211 is valid. The user selects button 214 if they determine that the difference 211 is not valid. If button 213 is selected, data indicating that the difference 211 is valid is transmitted from the information terminal 20 to the reception unit 110. In this case, at the branching point in step S105 in Figure 3, the process proceeds to step S108. If button 214 is selected, data indicating that the difference 211 is not valid is transmitted from the information terminal 20 to the reception unit 110. In this case, at the branching point in step S105 in Figure 3, the process proceeds to step S106. Therefore, step S105 can be described as the step in which the reception unit 110 receives a judgment on whether the difference 211 is valid or not, specifically the data indicating that judgment.
[0106] In step S106, the reception unit 110 receives the first instruction document. Figure 6(B) shows an example of the information displayed on the display unit of the information terminal 20 when the button 214 shown in Figure 6(A) is selected.
[0107] In the example shown in Figure 6(B), Form 221 is shown below the text, "Please write any proposed revisions to the input document, comments on any differences, etc." Here, the input document refers to the first document. Below Form 221, a button 223 labeled "Submit" is shown.
[0108] A user of the information processing system enters text into form 221, for example. Then, by selecting button 223, the contents of form 221 are supplied from the information terminal 20 to the reception unit 110 as a first instruction document.
[0109] The first instruction document may include, for example, instructions to modify the first document and at least one of the following: comments on difference 211. Here, comments on difference 211 may be, for example, opinions on difference 211. For example, if a user of the information processing system believes that there are differences in the first document from the second document other than difference 211, a statement to that effect may be included in the first instruction document. Also, if a user of the information processing system believes that at least a part of difference 211 is described in the second document, a statement to that effect may be included in the first instruction document.
[0110] Users of the information processing system can enter sentences into form 221 such as, for example, "Please modify the input document as follows," "I believe ddd is not disclosed in any document," and "I believe aaa is disclosed in paragraph [a1b1c1d1] (a1, b1, c1, and d1 are integers between 0 and 9) of document number xxx." Users of the information processing system may also enter information into form 221 in bullet points. If the language model for presenting differences determines in step S103 of Figure 3 that there are no differences 211, the processing unit 130 can display, for example, a sentence indicating that there are no differences, table 212, form 221, and button 223 on the display unit of the information terminal 20.
[0111] Form 221 and Button 223 may be displayed in the GUI shown in Figure 6(A). In this case, Buttons 213 and 214 do not need to be displayed. In the above case, for example, if Button 223 is selected while Form 221 is blank, the process can proceed to step S108 at the branching point in step S105 in Figure 3. Also, for example, if Button 223 is selected while text is entered in Form 221, the process can proceed to step S106 at the branching point in step S105 in Figure 3. Note that the process can proceed to step S108 not only when Button 223 is selected while Form 221 is blank, but also when Button 223 is selected while only predetermined characters are entered in Form 221. For example, if Button 223 is selected while only blank characters, symbols, and special characters are entered in Form 221, the process can proceed to step S108.
[0112] Figure 7(A) shows an example of the information displayed on the display unit of the information terminal 20 when no differences between the first document and the summary of the second document are presented in step S103. In the example shown in Figure 7(A), the message "There are no differences in the documents." is displayed. Also, in the example shown in Figure 7(A), the form 221 and button 223 shown in Figure 6(B) are shown. The form 221 shown in Figure 7(A) can include, for example, instructions to modify the first document.
[0113] By fine-tuning the language model for displaying differences using the first and second data points described above, if there are no differences between the first document and the summary of the second document, it is possible to display that there are no differences, as shown in Figure 7(A). This prevents the display of difference 211, as shown in Figure 6(A), even when there are no differences between the first document and the summary of the second document. In other words, it is possible to suppress the occurrence of hallucination. As a result, even someone with little experience in intellectual property work can create a high-quality specification.
[0114] Following step S106 in Figure 3, in step S107, the processing unit 130 creates a third prompt based on the first instruction document and transmits it to the information processing device 40 via the output unit 140. Specifically, the third prompt is transmitted to the difference-presentation language model of the information processing device 40. The third prompt is a prompt for presenting the differences between the first document and the summary of the second document, taking into account the content of the first instruction document.
[0115] The third prompt has the first document, the summary of the second document, and the first instruction document as context. The third prompt has command statements such as, "Please revise the entered document and answer the questions based on the information below. Also, please identify any features described in the entered document that are not mentioned in the summaries of any of the documents as differences from the documents. Please list the differences in bullet points."
[0116] The processing unit 130 may send a third prompt to both the language model for text generation and the language model for presenting differences. In this case, the instruction for the third prompt sent to the language model for text generation may be, for example, "Please modify the input document and answer the questions based on the information shown below." The instruction for the third prompt sent to the language model for presenting differences may be, for example, "Based on the information shown below, please identify the features described in the input document that are not described in the summaries of any of the documents, and present them as differences from the documents. Please present the differences in a bulleted list." The first document, the summary of the second document, and the first instruction document can be included in both the third prompt sent to the language model for text generation and the third prompt sent to the language model for presenting differences. Furthermore, when sending a third prompt to both the text generation language model and the difference presentation language model, the processing unit 130 may send the third prompt to the text generation language model to obtain a response sentence, and then include the response sentence in the third prompt before sending it to the difference presentation language model.
[0117] The information processing device 40 generates a third response sentence based on the third prompt using a language model for presenting differences and supplies it to the receiving unit 110. As a result, the information processing device 10 obtains the third response sentence. The third response sentence, like the second response sentence, includes differences between the first document and the summary of the second document. These differences may be based on the content of the first instruction document. Furthermore, the third response sentence may include responses to comments included in the first instruction document.
[0118] The language model for highlighting differences can be fine-tuned using multiple third and fourth data points, in addition to the first and second data points described above. The third data point includes the second command, the fifth training text, the sixth training text, the first training instruction document, and the third training response text. The fourth data point includes the second command, the seventh training text, the eighth training text, the second training instruction document, and the fourth training response text.
[0119] The second instruction can be the same as the instruction given by the third prompt described above. The third data point and the fourth data point can each have the same instruction.
[0120] The fifth and seventh learning texts correspond to the first document mentioned above. The sixth and eighth learning texts correspond to the summaries of the second document mentioned above. The first and second learning instruction documents correspond to the first instruction document mentioned above.
[0121] The third learning response is a response that points out the differences between the fifth learning text and the sixth learning text. The differences included in the third learning response are based on the first learning instruction document. The fourth learning response is a response that indicates that the second learning instruction document is not appropriate as a basis for judging the differences between the seventh learning text and the eighth learning text.
[0122] In the fine-tuning described above, the language model for highlighting differences is trained so that when a prompt containing the second command, the fifth training text, the sixth training text, and the first training instruction document is input, the third training response is output. Furthermore, the language model for highlighting differences is trained so that when a prompt containing the second command, the seventh training text, the eighth training text, and the second training instruction document is input, the fourth training response is output. Fine-tuning can be performed using supervised learning with the third and fourth training response sentences as labels.
[0123] Figure 7(B) shows an example of the information displayed on the display unit of the information terminal 20 when the language model for presenting differences determines in step S107 that the content of the first instruction document is not valid as a basis for determining the differences between the first document and the summary of the second document. In the example shown in Figure 7(B), the message "The entered content is not appropriate. Please rewrite the proposed revisions to the input document, comments on the differences, etc." is displayed. Also, in the example shown in Figure 7(B), the form 221 and button 223 shown in Figure 6(B), etc., are shown.
[0124] As described above, by performing fine-tuning on the language model for presenting differences using the third and fourth data points, it is possible to prompt the user of the information processing system to re-enter the first instruction document if the content of the first instruction document is inappropriate. This makes it easier for the user of the information processing system according to one embodiment of the present invention to create a highly complete specification. It is preferable that the number of third data points used to train the language model for presenting differences be at least the number of fourth data points, more preferably twice the number of fourth data points, and even more preferably three times the number of third data points. By increasing the number of third data points, the language model for presenting differences can more easily and accurately present the differences between the first document and the summary of the second document.
[0125] After step S107 in Figure 3, steps S104 and S105 are repeated. Here, the display unit of the information terminal 20 can display, for example, the differences 211, table 212, button 213, and button 214 shown in Figure 6(A), as well as the response to the comments included in the first instruction document. Also, if the first document is modified based on the first instruction document, the modified first document can be displayed on the display unit of the information terminal 20.
[0126] In step S108 of Figure 3, the processing unit 130 creates a fourth prompt and transmits it to the information processing device 40 via the output unit 140. Specifically, the fourth prompt is transmitted to the difference-presentation language model of the information processing device 40. The fourth prompt is a prompt to confirm that all of the differences 211 shown in Figure 6(A) are not described in the second document. In other words, the fourth prompt is a prompt to detect differences 211 that are not described in the summary of the second document but are described in the second document itself.
[0127] The fourth prompt has the context of Difference 211 and the second document. The fourth prompt also has a command statement such as, "Confirm that the differences to the document are not disclosed in either document. If there are any disclosed differences, please provide them."
[0128] The information processing device 40 generates a fourth response statement based on the fourth prompt using a language model for presenting differences and supplies it to the receiving unit 110. As a result, the information processing device 10 obtains the fourth response statement. The fourth response statement includes the confirmation results. It also includes whether or not each difference 211 is described in the second document.
[0129] Next, in step S109 of Figure 3, the processing unit 130 presents the above-mentioned verification result to the user of the information processing system. For example, the processing unit 130 creates data containing the content indicated by the fourth response statement and supplies it to the information terminal 20 via the output unit 140. As a result, the above-mentioned verification result is displayed on the display unit of the information terminal 20.
[0130] Figure 8(A) shows an example of the information displayed on the display unit of the information terminal 20 when it is determined that difference 211[2] is described in the second document. In the example shown in Figure 8(A), difference 211[1] and difference 211[3] are shown as differences with respect to the document. On the other hand, unlike the example shown in Figure 6(A), difference 211[2] is not shown. In addition, the confirmation result shows the sentence, "The following differences are not included in the abstract, but are disclosed in the document," and difference 211[2] is shown as such a difference. Furthermore, the form 221 and button 223 shown in Figure 6(B) are shown. Note that in the example shown in Figure 8(A), the document number of the document in which difference 211[2] is described is not shown, but the document number may be shown.
[0131] Figure 8(B) shows an example of the information displayed on the display unit of the information terminal 20 when it is determined that all of the differences 211[1], differences 211[2], and differences 211[3] shown in Figure 6(A) are not described in the second document. In the example shown in Figure 8(A), in addition to the differences 211 and table 212 shown in Figure 6(A), the confirmation result shows the statement, "None of the differences were disclosed in the document." Furthermore, in the example shown in Figure 8(B), a button 217 labeled "Accept" is shown.
[0132] If it is determined that at least some of the differences 211 are described in the second document, that is, if the confirmation result shown in Figure 8(A) is displayed on the display unit of the information terminal 20, the branching in step S110 of Figure 3 proceeds to step S106. Here, the first instruction document may include comments on the confirmation result shown in Figure 8(A). Comments on the confirmation result may be, for example, opinions on the confirmation result. For example, if a user of the information processing system believes that the differences 211[2] are not described in the second document, a statement to that effect may be included in the first instruction document.
[0133] If it is determined that all of the differences 211 are not described in the second document, that is, if the confirmation result shown in Figure 8(B) is displayed on the display unit of the information terminal 20, then the process proceeds to connector A in the branching step S110 of Figure 3. Specifically, the process proceeds to connector A when the user of the information processing system selects button 217. Note that although Figure 8(B) does not show, for example, the button for performing the process in step S106, such a button may be shown. For example, by selecting this button, form 221 and button 223 may be displayed on the display unit of the information terminal 20.
[0134] In the process shown in Figure 3, steps S102 and steps S108 to S110 may be omitted. In this case, in step S103, the processing unit 130 creates a second prompt to present the differences between the first document and the second document itself. If the user of the information processing system determines that the difference 211 is valid, the process proceeds to connector A at the branch in step S105.
[0135] By performing steps S102 and S108 to S110, the processing unit 130 can more easily present the differences 211. For example, it becomes easier to compare the technical features described in the first document with the technical features described in the second document, and to present the differences 211 based on the comparison results. On the other hand, by not performing steps S102 and S108 to S110, the number of processes performed by the information processing system can be reduced.
[0136] Connector A is connected to step S111 shown in Figure 4. In step S111, the processing unit 130 creates a fifth prompt and transmits it to the information processing device 40 via the output unit 140. Specifically, the fifth prompt is transmitted to the document generation language model of the information processing device 40. The fifth prompt is a prompt for generating a third document based on the second document and the differences 211.
[0137] The fifth prompt has the second document and the differences 211 as context. The fifth prompt also has a command statement such as, "Generate the document based on the references and the differences." The fifth prompt may also have the first document. In this case, it may be possible to prevent the structure shown in the first document from not being shown in the third document. On the other hand, if the fifth prompt does not have the first document, it may be possible to prevent the third document from being generated in a format different from, for example, the format of the second document.
[0138] The information processing device 40 generates a fifth response sentence based on the fifth prompt using a language model for text generation and supplies it to the receiving unit 110. As a result, the information processing device 10 obtains the fifth response sentence. The fifth response sentence includes the third document.
[0139] Since the third document is generated based on the second document, it can be of the same type as the second document. For example, if the second document is a patent application or a utility model application, the third document can also be a patent application or a utility model application. Specifically, if the second document is a specification, the third document can also be a specification.
[0140] Next, in step S112 of Figure 4, the processing unit 130 presents the third document generated by the text generation language model to the user of the information processing system. This allows the processing unit 130 to confirm the third document with the user. For example, the processing unit 130 creates data containing the content indicated by the fifth response statement and supplies it to the information terminal 20 via the output unit 140. As a result, the third document is displayed on the display unit of the information terminal 20.
[0141] Figure 9(A) shows an example of the information displayed on the display unit of the information terminal 20 in step S112, following step S111. In Figure 9(A), the third document generated by the language model for document generation is referred to as document 231. In subsequent figures, the third document may also be referred to as document 231. Document 231 is also called the generated document.
[0142] The example shown in Figure 9(A) includes document 231, a button 233 labeled "Accept," form 235, and a button 237 labeled "Submit." Figure 9(A) shows an example where form 235 and button 237 are displayed below the text, "Please enter any suggested revisions, comments, etc."
[0143] For example, if the second document is a specification, document 231 can be a document that conforms to the format of the specification. Figure 9(A) shows an example in document 231 where each paragraph is numbered. In the example shown in Figure 9(A), the paragraph numbers are
[0001] ,
[0002] , and [n1n2n3n4] (where n1, n2, n3, and n4 are integers between 0 and 9). Paragraphs containing differences 211, for example, may be highlighted. Methods of highlighting include changing the background color, changing the text color, increasing the font size, changing the font type, making it bold, and underlining.
[0144] For example, if a user of the information processing system determines that document 231 is complete and does not need to be modified, they select button 233. Alternatively, if they determine that document 231 needs to be modified, the user enters instructions for the modification into form 235 and then selects button 237. If button 233 is selected, data indicating that document 231 does not need to be modified is sent from the information terminal 20 to the reception unit 110. In this case, at the branching point in step S113 in Figure 4, the process proceeds to step S116. If button 237 is selected, data indicating that document 231 needs to be modified is sent from the information terminal 20 to the reception unit 110. In this case, at the branching point in step S113 in Figure 4, the process proceeds to step S114. Thus, step S113 can be described as the step in which the reception unit 110 receives a determination on whether or not document 231 needs to be modified, specifically the data indicating that determination.
[0145] Instructions for revisions include, for example, "Add a description of the ddd configuration between paragraphs [p1p2p31] and [p1p2p32] (where p1, p2, and p3 are integers between 0 and 9, respectively)," "Revise the eee configuration described in paragraph [p4p5q1q2] (where p4, p5, q1, and q2 are integers between 0 and 9, respectively) to the fff configuration," "Explain the fff configuration, then explain the ddd configuration," "Add a more detailed explanation of the term eee based on reference xxx," "Add the effect of the fff configuration," and "Delete paragraph [q3p6q4p7] (where p6, p7, q3, and q4 are integers between 0 and 9, respectively)." Users of the information processing system may also enter information into Form 235 in bullet points.
[0146] Users of the information processing system can enter information necessary for modifying document 231 into form 235. For example, they can enter documents not included in the second document into form 235. Specifically, they can enter patent application documents, papers, and other documents that were not selected as the second document into form 235. They can also enter addresses of relevant web pages, explanatory texts defining terms, and other information into form 235.
[0147] Form 235 may include comments regarding document 231. These comments may, for example, be questions regarding document 231. In this case, form 235 does not need to include instructions for modification. That is, even if the user of the information processing system determines that there is no need to modify document 231, if there are matters they would like to comment on regarding document 231, they can proceed to step S114 at the branching point in step S113 in Figure 4. Examples of questions regarding document 231 include, "What does ggg mean in paragraph [r1r2r3r4] (where r1, r2, r3, and r4 are integers between 0 and 9)?" and "What effect does the configuration shown in paragraph [s1s2s3s4] (where s1, s2, s3, and s4 are integers between 0 and 9) produce?". If the user of the information processing system does not want to modify document 231, they may write, "Please do not modify the document," etc., on form 235.
[0148] Note that button 233 does not need to be displayed in the GUI shown in Figure 9(A). In this case, for example, if button 237 is selected while form 235 is blank, the process can proceed to step S116 at the branching point in step S113 in Figure 4. Similarly, if button 223 is selected while only predetermined characters are entered in form 221, the process may also proceed to step S116 if button 237 is selected while only predetermined characters are entered in form 235.
[0149] In step S114 of Figure 4, the reception unit 110 receives the second instruction document. Specifically, the text entered in form 235 can be used as the second instruction document.
[0150] Following step S114 in Figure 4, in step S115, the processing unit 130 creates a sixth prompt based on the second instruction document and transmits it to the information processing device 40 via the output unit 140. Specifically, the sixth prompt is transmitted to the document generation language model of the information processing device 40. The sixth prompt is a prompt for modifying the third document based on the second instruction document.
[0151] The sixth prompt has the third document and the second instruction document as context. The sixth prompt also has an instruction such as, "Please provide the answer and revise the generated document regarding the following. Please provide the answer in bullet points." The sixth prompt may also have the second document and difference 211, similar to the fifth prompt. Furthermore, the sixth prompt may have the first document. For example, the sixth prompt may have the first document if the fifth prompt has the first document.
[0152] The information processing device 40 generates a sixth response sentence based on the sixth prompt using a language model for text generation and supplies it to the receiving unit 110. The information processing device 10 then obtains the sixth response sentence. The sixth response sentence may include, for example, a modified third document. It may also include a response to a comment contained in the second instruction document.
[0153] After step S115 in Figure 4, step S112 is performed again. Figure 9(B) shows an example of the information displayed on the display unit of the information terminal 20 in step S112, following step S115. The following will mainly explain the differences from Figure 9(A).
[0154] In the example shown in Figure 9(B), the modified parts in document 231 are indicated with a hatching pattern. Specifically, Figure 9(B) shows an example where "fff" is added and a hatching pattern is applied. Highlighting the modified parts in document 231 in this way is preferable because it allows users of the information processing system to easily recognize the modified parts. Methods of highlighting include changing the background color, changing the font color, increasing the font size, changing the font type, making it bold, and underlining, as mentioned above. Highlighting may be applied to each sentence or to each paragraph. That is, even if only a part of a sentence is modified, the entire sentence may be highlighted. Similarly, even if only a part of a paragraph is modified, the entire paragraph may be highlighted.
[0155] Furthermore, in the example shown in Figure 9(B), the answer to the comment included in the second instruction document is shown in area 239. In Figure 9(B), an example of an answer to a comment is shown as, "ggg means kkk." For example, if the second instruction document includes the comment, "What does ggg mean?", the answer shown in Figure 9(B) can be displayed in area 239. Area 239 may also show either or both the modified parts of document 231 and the content of the modifications. Alternatively, either or both the modified parts of document 231 and the content of the modifications may be shown in an area different from area 239.
[0156] In step S116 of Figure 4, the processing unit 130 outputs a third document. The processing unit 130 supplies the third document to, for example, a database via the output unit 140. The processing unit 130 can also supply the third document to, for example, an information terminal 20 via the output unit 140. The third document supplied to the information terminal 20 can be displayed on, for example, a display unit. The third document supplied to the information terminal 20 can also be stored in, for example, the information terminal 20. With this, the information processing method according to one embodiment of the present invention is "completed".
[0157] As described above, by using the information processing system according to one aspect of the present invention, a third document can be generated, taking into account the differences between the first document and the second document. For example, the specification relating to the creation shown in the first document can be written in a way that emphasizes technical features not disclosed in the prior art documents. As described above, the information processing system according to one aspect of the present invention can assist in the creation of documents related to intellectual property, such as the creation of specifications. Therefore, even a person with little experience in intellectual property work can create documents related to intellectual property, such as specifications, in a short time by using the information processing system according to one aspect of the present invention. Furthermore, users can create documents related to intellectual property, such as specifications, with less effort.
[0158] Furthermore, in one aspect of the information processing system of the present invention, when a user creates a first instruction document, the difference-displaying language model can present the differences between the first document and the second document multiple times. Specifically, the difference-displaying language model can present the above-mentioned differences multiple times, taking into account the content of the first instruction document. Moreover, in one aspect of the information processing system of the present invention, when a user creates a second instruction document, the document generation language model can modify the third document one or more times. As a result, a user of one aspect of the information processing system of the present invention can create the third document in an interactive format. This allows, for example, a person with little experience in intellectual property work to create a highly complete specification.
[0159] Based on the above, one aspect of the present invention provides an information processing device, an information processing method, and an information processing system that are excellent in convenience and usefulness.
[0160] <Information Processing Method 2> Figure 10 is a flowchart showing a different example from Figure 4 for the processing after connector A. In the example shown in Figure 10, steps S121, S122, S123, S124, and S125 are shown before step S111.
[0161] In step S121, the processing unit 130 creates a seventh prompt and transmits it to the information processing device 40 via the output unit 140. Specifically, the seventh prompt is transmitted to the language model of the information processing device 40. The seventh prompt is a prompt for determining whether or not there is insufficient information to generate the third document. The seventh prompt is also a prompt for indicating the type of information that is missing if such information is found to be missing. Here, the language model to which the seventh prompt is transmitted is called the language model for indicating missing information. The language model for indicating missing information may be a different language model from the language model for indicating differences, or it may be the same language model. Furthermore, the language model for document generation, the language model for indicating differences, and the language model for indicating missing information may be different language models from each other, or they may all be the same language model.
[0162] The seventh prompt has the second document and the differences 211 as context. The seventh prompt also has a command statement such as, "Determine whether there is any missing information when generating a document that satisfies the enablement requirements based on the document and the differences. If there is missing information, indicate what information is missing." The seventh prompt may also have the first document. For example, if the processing unit 130 creates the fifth prompt to include the first document in the later step S111, it is preferable that the seventh prompt also includes the first document.
[0163] The information processing device 40 generates a seventh response statement based on the seventh prompt using a language model for presenting missing information and supplies it to the receiving unit 110. As a result, the information processing device 10 obtains the seventh response statement. The seventh response statement includes the result of determining whether or not there is missing information for generating the third document. If there is missing information, the seventh response statement also includes the type of missing information.
[0164] For example, if the third document includes examples, and the experimental conditions are not sufficiently described in the first document, the experimental conditions can be considered missing information. Similarly, if the first document describes the transistor configuration, and the second document describes both the transistor configuration and the manufacturing method, but the first document does not describe the manufacturing method, the manufacturing method can be considered missing information. Furthermore, if the first document describes the physical properties of a material, such as electrical resistivity, but does not specify the name of the material, the name of the material can be considered missing information. In addition, any matter not described in either the second document or Difference 211 can be considered not described in the first document either.
[0165] The language model for presenting missing information can be a finely tuned model. Fine tuning can be performed using multiple fifth and sixth data points. The fifth data point includes the third command, the ninth training sentence, the tenth training sentence, and the fifth training response sentence. The sixth data point includes the third command, the eleventh training sentence, the twelfth training sentence, and the sixth training response sentence.
[0166] The third instruction can be the same as the instruction given by the seventh prompt described above. The fifth data point and the sixth data point can each have the same instruction.
[0167] The ninth and eleventh learning texts correspond to the second document mentioned above. The tenth and twelfth learning texts correspond to the difference 211 mentioned above.
[0168] The fifth training response is a response that indicates missing information when generating a document based on the ninth training text and the tenth training text. The sixth training response is a response that indicates that no information is missing when generating a document based on the eleventh training text and the twelfth training text. These documents correspond to the third document mentioned above.
[0169] In fine-tuning the language model for presenting missing information, the model is trained so that when a prompt containing the third command, the ninth training sentence, and the tenth training sentence is input, the fifth training response sentence is output. Furthermore, the model is trained so that when a prompt containing the third command, the eleventh training sentence, and the twelfth training sentence is input, the sixth training response sentence is output. Fine-tuning can be performed using supervised learning with the fifth and sixth training response sentences as labels.
[0170] As described above, by fine-tuning the language model for presenting missing information, the occurrence of hallucination can be suppressed compared to the case where no fine-tuning is performed. In particular, by fine-tuning using the sixth data point, which includes the sixth training response sentence, it is possible to suppress the language model for presenting missing information from presenting missing information even though there is sufficient information to generate the third document. It is preferable that the number of fifth data points used for training the language model for presenting missing information be greater than or equal to the number of sixth data points, more preferably twice or more the number of sixth data points, and even more preferably three times or more. By increasing the number of fifth data points, the language model for presenting missing information can more easily and accurately present the types of information necessary to generate the third document.
[0171] If there is insufficient information to generate the third document, the branch in step S122 proceeds to step S123. If there is sufficient information to generate the third document, the branch in step S122 proceeds to step S111.
[0172] In step S123, the processing unit 130 presents the user of the information processing system with the information missing to generate the third document. For example, the processing unit 130 creates data containing the content indicated by the seventh response statement and supplies it to the information terminal 20 via the output unit 140. As a result, the information missing to generate the third document is displayed on the display unit of the information terminal 20.
[0173] Figure 11(A) is a diagram showing an example of the information displayed on the display unit of the information terminal 20 in step S123. In the example shown in Figure 11(A), the information that is missing to generate the third document is shown as "missing information". The information shown as "missing information" is referred to as information 241. In the example shown in Figure 11(A), information 241 is shown as "hhh", "iii", and "jjj". Here, "hhh" is referred to as information 241[1], "iii" as information 241[2], and "jjj" as information 241[3].
[0174] In the example shown in Figure 11(A), Form 243 is shown below the text, "Please enter the missing information and select the 'Submit' button." Below Form 243, a button 245 labeled "Submit" is shown.
[0175] The user of the information processing system enters the missing information into form 243, based on information 241, to generate the third document. Then, by selecting button 245, the contents of form 243 are supplied as additional information from the information terminal 20 to the reception unit 110.
[0176] For example, if information 241[1] is experimental conditions, the user of the information processing system enters the experimental conditions into form 243. For example, if information 241[2] is a method for fabricating a transistor, the user of the information processing system enters the method for fabricating a transistor into form 243. For example, if information 241[3] is a conductive material with an electrical resistivity of xΩ·cm or less, the user of the information processing system enters the specific name of the conductive material into form 243. In this way, the user of the information processing system can enter additional information corresponding to information 241 into form 243. Note that form 243 may allow input of not only text but also, for example, drawings. Furthermore, form 243 may allow input of additional information in a table format for each piece of information 241.
[0177] In step S124 of Figure 10, the reception unit 110 receives the additional information mentioned above. Specifically, information such as text entered in form 243 can be used as additional information.
[0178] In step S125, the processing unit 130 creates a seventh prompt again, including the additional information, and sends it to the information processing device 40 via the output unit 140. If, after step S125, there is still insufficient information to generate the third document, the process proceeds to step S123 again at the branch in step S122. If the lack of information to generate the third document is resolved, the process proceeds to step S111 at the branch in step S122. In step S111, the processing unit 130 creates a fifth prompt that includes the additional information. In this case, the fifth prompt includes the second document, the differences 211, and additional information. The instruction statement for the fifth prompt can also be, "Generate a document based on the document, the differences, and the additional information," etc.
[0179] As described above, the information processing methods shown in Figures 10 and 11(A) allow for the addition of missing information before the language model for document generation generates the third document. This improves the completeness of the third document.
[0180] In step S121, if it is determined that there is no missing information to generate the third document, this fact may be presented to the user of the information processing system. After the user of the information processing system confirms that there is no missing information to generate the third document, the processing unit 130 may perform the processing in step S111. Figure 11(B) is a diagram showing an example of the information displayed on the display unit of the information terminal 20 when it is determined in step S121 that there is no missing information to generate the third document. In the example shown in Figure 11(B), the determination result is shown as, "No missing information was found." Also in the example shown in Figure 11(B), a button 242 labeled "Accept" is shown. By the user of the information processing system selecting button 242, the processing unit 130 can perform the processing in step S111.
[0181] In the example shown in Figure 11(B), below the text, "If you have any additional information, please enter it in the form below and select the 'Submit' button," are shown a form 243 and a button 246 labeled "Submit." In this case, even if the user of the information processing system determines in step S121 that there is no missing information to generate the third document, the user can still enter additional information in form 243. By selecting button 246, the processing unit 130 can perform step S111 after performing step S124. Alternatively, after button 246 is selected, the processing unit 130 may perform step S125 instead of step S111. Furthermore, form 243 and button 246 do not necessarily have to be displayed in the GUI shown in Figure 11(B).
[0182] By fine-tuning the language model for presenting missing information using the fifth and sixth data points described above, it is possible to suppress the display of information 241 shown in Figure 11(A) instead of the screen shown in Figure 11(B), even though there is sufficient information to generate the third document. In other words, it is possible to suppress the occurrence of hallucination.
[0183] Based on the above, we can provide an information processing device, an information processing method, and an information processing system that are excellent in terms of convenience and usefulness.
[0184] <Information Processing Method 3> Figure 12 is a flowchart showing an example of an information processing method according to one aspect of the present invention, different from that shown in Figures 3 and 4. In Figure 12, steps that are not performed in the information processing methods shown in Figures 3 and 4 are indicated by thick borders.
[0185] An information processing method according to one embodiment of the present invention is "started," and in step S131, the receiving unit 110 receives a base document in addition to the first document and the second document. The base document is the document that serves as the basis when generating the third document.
[0186] The base document shall be a document that describes a creation related to the first document. Preferably, the base document shall be a pre-publication document. For example, the base document may be a specification belonging to a pre-publication patent application or a specification belonging to a utility model registration application. Alternatively, the base document may be an application form belonging to a pre-publication design registration application. Furthermore, the base document may be a paper, book, journal, etc. For example, the base document may be a document stored in a database.
[0187] The receiving unit 110 may accept multiple base documents. In this case, in a later step, a third document can be generated based on parts common to multiple base documents. This makes it possible to prevent the inclusion of a description of the main feature of the invention related to the patent application in the third document, for example, when the base document is a specification belonging to a patent application.
[0188] Next, step S102 is performed. Then, in step S132, the processing unit 130 creates a second prompt and transmits it to the information processing device 40 via the output unit 140. Specifically, the second prompt is transmitted to the difference presentation language model of the information processing device 40, similar to step S103 in Figure 3. In step S132, the second prompt is a prompt to present the differences between the first document and the summary of the second document, as well as the differences between the first document and the base document.
[0189] The second prompt in step S132 includes the first document, a summary of the second document, and the base document. The second prompt in step S132 also includes instructions such as, "Please list the features described in the input document that are not described in the base document as differences from the base document. Also, please list the features that are not described in the summary of any of the documents as differences from the documents. Please list the differences in bullet points."
[0190] The information processing device 40 generates a second response statement based on the second prompt using a language model for presenting differences and supplies it to the receiving unit 110. The information processing device 10 then obtains the second response statement. This second response statement includes differences between the first document and the summary of the second document, as well as differences between the first document and the base document. Here, if there are multiple base documents, the second prompt in step S132 can be a prompt to present features not described in any of the base documents as differences. In one embodiment of the present invention, the information processing system may perform step S103 instead of step S132. That is, even in the information processing method shown in Figure 12, the information processing system does not need to present differences between the first document and the base document.
[0191] Next, in step S133, the processing unit 130 presents the differences presented by the difference-presentation language model to the user of the information processing system. Specifically, the processing unit 130 presents to the user of the information processing system the differences 215 between the first document and the base document, in addition to the differences 211 between the first document and the summary of the second document. This allows the user to determine whether these differences are valid or not. For example, the processing unit 130 creates data containing the content indicated by the second response statement, similar to step S104 in Figure 3, and supplies it to the information terminal 20 via the output unit 140. As a result, the differences are displayed on the display unit of the information terminal 20.
[0192] Figure 13 shows an example of the information displayed on the display unit of the information terminal 20 in step S133. In Figure 13, in addition to the differences 211 with respect to the document (second document) shown in Figure 6(A), a table 212 which is a list of the second documents, a button 213 which is selected when the differences are valid, and a button 214 which is selected when the differences are not valid, an example is shown in which differences 215 with respect to the base document are displayed. In the example shown in Figure 13, the differences 215 are shown as "kkk." and "lll." Here, "kkk." is referred to as difference 215[1] and "lll." is referred to as difference 215[2]. Note that not only differences with respect to the base document but also similarities may be shown. In this case, the second prompt is a prompt that presents not only the differences in the first document with respect to the base document but also the similarities. Specifically, the command statement of the second prompt includes a sentence indicating that the similarities will be presented.
[0193] Next, steps S105 to S110 are performed. Here, the first instruction document may include comments regarding difference 215. The third prompt may be a prompt to present the differences in the first document from the base document, in addition to the differences in the summary of the second document, based on the content of the first instruction document.
[0194] In step S108, if it is determined that at least some of the differences 211 are described in the second document, the branch in step S110 proceeds to step S106, similar to the example shown in Figure 3. In step S108, if it is determined that all of the differences 211 are not described in the second document, the branch in step S110 proceeds to step S134. In step S134, the processing unit 130 creates a fifth prompt and transmits it to the information processing device 40 via the output unit 140. Specifically, the fifth prompt is transmitted to the document generation language model of the information processing device 40, similar to step S111 in Figure 4. The fifth prompt in step S134 is a prompt for generating a third document, which includes the differences 211 and 215, based on the base document.
[0195] Here, difference 211 is a difference in the first document with respect to, for example, publicly available literature. On the other hand, difference 215 is a difference in the first document with respect to, for example, unpublished literature. Therefore, it is preferable that the fifth prompt in step S134 be a prompt to generate the third document such that difference 211 is emphasized more than difference 215.
[0196] The fifth prompt in step S134 includes the second document and difference 211, as well as the base document and difference 215. The fifth prompt in step S134 also includes instructions such as, "Generate a new document based on the base document, including the differences in the input document to the references and the differences to the base document. The differences to the references should be expressed more emphatically than the differences to the base document." Note that the fifth prompt does not necessarily have to include difference 215. For example, if step S103 is performed instead of step S132 as described above, the fifth prompt will not include difference 215.
[0197] The information processing device 40 generates a fifth response sentence based on the fifth prompt using a language model for text generation and supplies it to the receiving unit 110. As a result, the information processing device 10 obtains the fifth response sentence. The fifth response sentence includes the third document as described above.
[0198] Here, if the fifth prompt has multiple base documents, the language model for document generation can generate a third document based on, for example, parts common to the multiple base documents, as described above. This makes it possible to suppress the inclusion of a description of the essential structure of the invention related to the patent application in the third document, for example, if the base document is a specification belonging to a patent application.
[0199] Since the third document in step S134 is generated based on the base document, it can be of the same type as the base document. For example, if the base document is a patent application or a utility model application, the third document can also be a patent application or a utility model application. Specifically, if the base document is a specification, the third document can also be a specification.
[0200] Steps S121 to S125 shown in Figure 10 may be performed before step S134. That is, the language model for document generation may add the missing information necessary to generate the third document before generating the third document in step S134.
[0201] Next, steps S112 to S116 in Figure 4 are performed. With this, the information processing method according to one embodiment of the present invention is "completed".
[0202] As described above, the information processing method shown in Figures 12 and 13 allows for the generation of a third document based on a base document. This makes it possible to generate a third document that includes, for example, information commonly found in a company's patent applications. Furthermore, it is possible to suppress the phrasing and style of the third document from becoming too similar to that of other companies' documents. As a result, the completeness of the third document can be improved.
[0203] Furthermore, a third document of a different type from the second document can be generated. For example, by using a patent document or utility model document as the base document, even if at least a part of the second document consists of a paper, book, journal, etc., a specification belonging to a patent application or utility model registration application can be generated as the third document. Moreover, by generating the third document to include the differences between the first document and the base document, it becomes possible to cover, for example, more configurations in one of the company's applications. Thus, a comprehensive network of rights, such as a patent network, can be easily created.
[0204] Based on the above, we can provide an information processing device, an information processing method, and an information processing system that are excellent in terms of convenience and usefulness.
[0205] <Example of an information processing system configuration 2> The following describes a detailed configuration example of an information processing system according to one embodiment of the present invention. Specifically, a configuration example of the storage unit 120, the processing unit 130, and the network 30 that is not shown in <Configuration Example 1 of Information Processing System> will be described.
[0206] The storage unit 120 includes at least one of volatile memory and non-volatile memory. Examples of volatile memory include DRAM (Dynamic Random Access Memory) and SRAM (Static Random Access Memory). Examples of non-volatile memory include ReRAM (Resistive Random Access Memory), PRAM (Phase Change Random Access Memory), FeRAM (Ferroelectric Random Access Memory), MRAM (Magnetoresistive Random Access Memory), and flash memory. The storage unit 120 may also include at least one of NOSRAM (registered trademark) and DOSRAM (registered trademark). The storage unit 120 may also include a recording media drive. Examples of recording media drives include hard disk drives (HDD) and solid state drives (SSD).
[0207] NOSRAM is an abbreviation for "Nonvolatile Oxide Semiconductor Random Access Memory." NOSRAM is a type of memory where the memory cell is a 2-transistor (2T) or 3-transistor (3T) gain cell, and the transistors are transistors that use metal oxide in the channel formation region (also called OS transistors). OS transistors have an extremely small current flowing between the source and drain when off, i.e., a leakage current. By utilizing this extremely low leakage current characteristic, NOSRAM can be used as a non-volatile memory by holding a charge corresponding to the data within the memory cell. In particular, NOSRAM can read the stored data without destroying it (non-destructive read), making it suitable for computational processing that repeatedly performs a large number of data read operations. Because the data capacity of NOSRAM can be increased by stacking it, it can be used as a large-scale cache memory, main memory, storage memory, etc., to improve the performance of semiconductor devices.
[0208] DOSRAM is an abbreviation for "Dynamic Oxide Semiconductor RAM," and refers to RAM with a 1T (transistor) 1C (capacitance) type memory cell. DOSRAM is a DRAM formed using OS transistors and is a memory that temporarily stores information sent from the outside. DOSRAM is a memory that takes advantage of the low off-current of OS transistors.
[0209] In this specification, "metal oxide" refers to an oxide of a metal in a broad sense. Metal oxides are classified into oxide insulators, oxide conductors (including transparent oxide conductors), oxide semiconductors (also called oxide semiconductors or simply OS), etc. For example, when a metal oxide is used in the semiconductor layer of a transistor, that metal oxide may be called an oxide semiconductor.
[0210] The metal oxide in the channel-forming region preferably contains indium (In), and for example, indium oxide is preferred. When the metal oxide in the channel-forming region is an indium-containing metal oxide, the carrier mobility (electron mobility) of the OS transistor is increased. Furthermore, the metal oxide in the channel-forming region is preferably an oxide semiconductor containing element M. Element M is preferably at least one of aluminum (Al), gallium (Ga), and tin (Sn). Other elements applicable to element M include boron (B), silicon (Si), titanium (Ti), iron (Fe), nickel (Ni), germanium (Ge), yttrium (Y), zirconium (Zr), molybdenum (Mo), lanthanum (La), cerium (Ce), neodymium (Nd), hafnium (Hf), tantalum (Ta), and tungsten (W). However, multiple elements mentioned above may be combined as element M. Element M is, for example, an element with a high bond energy with oxygen. For example, it is an element whose bonding energy with oxygen is higher than that of indium. Furthermore, the metal oxide containing the channel-forming region is preferably a metal oxide containing zinc (Zn). Zinc-containing metal oxides may be more prone to crystallization.
[0211] The metal oxides present in the channel-forming regions are not limited to indium-containing metal oxides. For example, the metal oxides present in the channel-forming regions may be zinc-tin oxides, gallium-tin oxides, or other metal oxides that do not contain indium but contain zinc, gallium, or tin.
[0212] The processing unit 130 may, for example, have an arithmetic circuit. The processing unit 130 may, for example, have a central processing unit (CPU). Furthermore, the processing unit 130 may have a graphics processing unit (GPU).
[0213] The processing unit 130 may have a microprocessor such as a DSP (Digital Signal Processor). The microprocessor may be implemented using a PLD (Programmable Logic Device) such as an FPGA (Field Programmable Gate Array) or FPAA (Field Programmable Analog Array). The processing unit 130 may also have a quantum processor. The processing unit 130 can perform various data processing and program control by interpreting and executing instructions from various programs via the processor. Programs that can be executed by the processor are stored in at least one of the processor's memory area and the storage unit 120.
[0214] The processing unit 130 may have main memory. The main memory includes at least one of volatile memory such as RAM (Random Access Memory) and non-volatile memory such as ROM (Read Only Memory). The main memory may also include at least one of the above-mentioned NOSRAM and DOSRAM.
[0215] For RAM, for example, DRAM or SRAM is used, and a virtual memory space is allocated and used as the workspace for the processing unit 130. The operating system, application programs, program modules, program data, lookup tables, etc., stored in the storage unit 120 are loaded into RAM for execution. These data, programs, and program modules loaded into RAM are directly accessed and manipulated by the processing unit 130.
[0216] ROM can store data that does not require rewriting, such as BIOS (Basic Input / Output System) and firmware. Examples of ROM include mask ROM, OTPROM (One Time Programmable Read Only Memory), and EPROM (Erasable Programmable Read Only Memory). Examples of EPROM include UV-EPROM (Ultra-Violet Erasable Programmable Read Only Memory), which allows data to be erased by ultraviolet irradiation, EEPROM (Electrically Erasable Programmable Read Only Memory), and flash memory.
[0217] The processing unit 130 may have either or both an OS transistor and a transistor having silicon in its channel formation region (Si transistor).
[0218] The processing unit 130 preferably has an OS transistor. Because the OS transistor has an extremely small off-current, using the OS transistor as a switch to hold the charge (data) that has flowed into a capacitive element that functions as a memory element ensures that the data retention period can be ensured over a long period. If at least one of the registers and cache memory of the processing unit 130 has this characteristic, the processing unit 130 can be operated only when necessary, and in other cases the information of the previous processing is saved to the memory element and the processing unit 130 is turned off. In other words, normally-off computing becomes possible, and the power consumption of the information processing system can be reduced.
[0219] Network 30 can, for example, use the Internet, which is the foundation of the World Wide Web (WWW), as a global network. Network 30 can also use a local network. Furthermore, an intranet or extranet can be used as Network 30. In addition, PAN (Personal Area Network), LAN (Local Area Network), CAN (Campus Area Network), MAN (Metropolitan Area Network), WAN (Wide Area Network), GAN (Global Area Network), etc., can be used as Network 30.
[0220] When a service provider using an information processing method according to one aspect of the present invention and a user enjoying such a service belong to the same organization, such as a company, it is preferable that data transmission and reception between the information terminal 20 and the information processing device 10 be performed, for example, using a network established within that organization. This allows for more secure data transmission and reception between the information terminal 20 and the information processing device 10 compared to transmission via the Internet. Furthermore, it prevents the leakage of confidential information within the organization to external parties.
[0221] When performing wireless communication, communication protocols or technologies such as 4G, 5G, 6G, or other communication standards, or specifications standardized by the IEEE (Institute of Electrical and Electronics Engineers) such as Wi-Fi® and Bluetooth®, may be used. [Examples]
[0222] In this example, we will describe the results of fine-tuning and evaluating the language model.
[0223] In this example, we prepared language models A, B, C, D, and E. All four models were based on Llama 3.1 8B. Fine-tuning was not performed on Model A. Fine-tuning was performed on Models B, C, D, and E. For fine-tuning, we used the rag-dataset-12000 dataset. The rag-dataset-12000 contains data points including commands, contexts, and responses. The commands represent questions. Fine-tuning was performed using PEFT (Parameter-Efficient Fine Tuning). In this example, LoRA (Low-Rank Adaptation) was used as the PEFT. LoRA is a method for updating language model parameters using low-rank matrices. Using LoRA reduces the number of parameters to update compared to not using LoRA.
[0224] In this example, a first dataset, a second dataset, a third dataset, and a fourth dataset were prepared based on rag-dataset-12000. The data points for the first dataset were the data points themselves from rag-dataset-12000.
[0225] The data points in the second dataset were obtained by omitting the context from the data points in rag-dataset-12000. In other words, the data points in the second dataset contain commands and responses, but do not have context.
[0226] The data points in the third dataset are the data points of rag-dataset-12000, but with noise contexts added. The noise contexts were randomly extracted from each context included in rag-dataset-12000. Here, for each data point, the context that becomes the noise context was extracted so that the noise context is not equal to the context included in that data point. The noise context of a data point is a context that is unrelated to the command statement that the data point possesses. Thus, the third dataset contains command statements, contexts, noise contexts, and response statements. Furthermore, each data point included in the third dataset is designed to have a different noise context from one another.
[0227] The fourth dataset has a first data point and a second data point. In this embodiment, the first data point is the same as the data point in the first dataset, using the data point itself from rag-dataset-12000. In this embodiment, the second data point replaces the context of rag-dataset-12000 with the noise context described above, and sets the response to the phrase "I cannot answer this question" (also called a refusal response or non-answer). Of the data points included in the fourth dataset, 80% are the first data points, and the remaining 20% are the second data points. Thus, the fourth dataset is a dataset in which 20% of the data points included in the first dataset have their context replaced with a noise context, and the response is set to a refusal response.
[0228] Hyperparameter optimization (HPO) was performed during the fine-tuning of Models B through E. The hyperparameters selected for optimization were the optimization algorithm (optimizer), learning rate (LR), number of epochs, batch size, and the rank of the low-rank matrix in LoRA (Low-Rank Adaptation), also known as the LoRA rank. For hyperparameter optimization, 9598 data points out of the 12000 data points included in rag-dataset-12000 were used.
[0229] Table 1 shows the hyperparameter search space. Table 2 shows the optimized hyperparameter values for models B through E. In hyperparameter optimization, each hyperparameter was optimized within the search space shown in Table 1.
[0230] [Table 1]
[0231] [Table 2]
[0232] Note that the number of warm-up steps and the per-device training batch size are hyperparameters that were not optimized. In models B through E, the number of warm-up steps was set to 10, and the per-device training batch size was set to 1.
[0233] In models B through E, after hyperparameter optimization, training and evaluation were performed using 10-fold cross-validation (CV). In 10-fold cross-validation, 11,997 data points out of the 12,000 data points in rag-dataset-12000 were divided into 1st through 10th data point groups. Then, the 1st through 9th data point groups were used as training data for training, and the 10th data point group was used as evaluation data for evaluation. Subsequently, the 1st through 8th data point groups and the 10th data point group were used as training data for training, and the 9th data point group was used as evaluation data for evaluation. Training and evaluation were performed sequentially in this manner, and finally, the 2nd through 10th data point groups were used as training data, and the 1st data point group was used as evaluation data for evaluation. Here, in model A, where fine tuning was not performed as described above, the evaluation was performed using the 10th data point group as evaluation data, then the 9th data point group as evaluation data, and so on, up to evaluation using the 1st data point group as evaluation data. Based on the above, evaluations were performed 10 times for each of Models A through E.
[0234] The data points used to train Model B were designated as the first dataset. The data points used to train Model C were designated as the second dataset; that is, the context was omitted. The data points used to train Model D were designated as the third dataset; that is, noise context was added. The data points used to train Model E were designated as the fourth dataset; that is, for 20% of the data points, the context was replaced with noise context, and the response sentences were set as rejection responses. Models B through E were trained using supervised learning with the response sentences as labels.
[0235] The evaluation criteria were Context Recall (CR), Precision (P), and Negative Rejection (NR).
[0236] Context recall is the value obtained by dividing the number of tokens in both the response sentences (also called golden answers) included in the evaluation data mentioned above and the response sentences generated when that evaluation data is input into each model (also called generated responses) by the total number of tokens in the generated responses. Precision is the value obtained by dividing the number of tokens in both the context included in the evaluation data mentioned above and the generated responses by the total number of tokens in the generated responses. Here, the data points used for evaluating context recall and the data points used for evaluating precision were the data points of rag-dataset-12000 themselves. That is, in the evaluation of any of Models A to E, no context was omitted, no noise context was added, etc., to the data point set used as evaluation data.
[0237] Negative rejection was evaluated using data in which the context of each data point included in the evaluation data described above was randomly swapped. This data is called the NR evaluation data. In the NR evaluation data, the golden answer was defined as the rejection response. Negative rejection is the percentage of responses generated when the NR evaluation data is input into each model that are rejection responses, which are the golden answers.
[0238] Based on the above, it can be said that models with higher context recall, accuracy, and negative rejection are less likely to generate hallucination.
[0239] Table 3 shows the context recall (CR) for Models A through E. Table 4 shows the precision (P) for Models A through E. Table 5 shows the negative rejection (NR) for Models A through E. In this example, as described above, context recall, precision, and negative rejection were each evaluated 10 times for each language model. Tables 3 through 5 show the macro mean values of the 10 evaluations and the 95% confidence intervals calculated based on the 10 evaluations.
[0240] [Table 3]
[0241] [Table 4]
[0242] [Table 5]
[0243] Tables 3 and 4 show that models B through E, which underwent fine-tuning, exhibited higher context recall and accuracy than model A, which did not undergo fine-tuning. On the other hand, model C, which was fine-tuned without using context, had lower context recall and accuracy than model B, which was fine-tuned using context. This suggests that the language model Llama3.1 8B is capable of learning while taking context into account.
[0244] Furthermore, Tables 3 and 4 confirm that Model D, which includes the noise context, exhibited similar context recall and accuracy to Model B, which does not include the noise context. This is likely due to hyperparameter optimization performed during fine-tuning.
[0245] Table 5 shows that Models B, C, and D exhibited significantly lower negative rejection than Model A. This is likely due to overfitting in Models B, C, and D. On the other hand, Tables 3 to 5 show that Model E, which replaced some contexts with noise contexts and used rejection responses for some responses, demonstrated context recall and accuracy comparable to Model B, while maintaining negative rejection comparable to Model A. Thus, it was confirmed that fine-tuning using rejection responses for some labels can improve context recall and accuracy while preventing a decrease in negative rejection compared to not performing fine-tuning.
[0246] In the above-described embodiment, the second training response sentence, the fourth training response sentence, and the sixth training response sentence correspond to the rejection responses in this embodiment. Therefore, it is suggested that by performing fine tuning using the second and fourth data points in the above-described embodiment, a language model for presenting differences with suppressed hallucination can be provided. Furthermore, it is suggested that by performing fine tuning using the sixth data point, a language model for presenting missing information with suppressed hallucination can be provided. [Explanation of Symbols]
[0247] 10: Information processing device, 20: Information terminal, 20a: Desktop computer, 20b: Notebook computer, 20c: Smartphone, 20d: Tablet computer, 21: Enclosure, 30: Network, 40: Information processing device, 110: Reception unit, 120: Storage unit, 130: Processing unit, 140: Output unit, 150: Transmission line, 201: Check box, 203: Button, 211[1]: Differences, 211[2]: Differences, 211[ 3]: Differences, 211: Differences, 213: Button, 214: Button, 215[1]: Differences, 215[2]: Differences, 215: Differences, 217: Button, 221: Form, 223: Button, 231: Document, 233: Button, 235: Form, 237: Button, 239: Area, 241[1]: Information, 241[2]: Information, 241[3]: Information, 241: Information, 242: Button, 243: Form, 245: Button, 246: Button
Claims
1. We have received the first document and the second document. The summary of the second document is obtained using the first language model, The differences between the first document and the summary are obtained using the second language model. The second language model is trained using the first and second data points. The first data point comprises a first command statement, a first learning sentence, a second learning sentence, and a first learning response statement. The second data point comprises the first command statement, the third learning sentence, the fourth learning sentence, and the second learning response statement. The first training response sentence illustrates the differences between the first training text and the second training text. An information processing method that indicates that the second training response sentence is identical to the third training text in that there are no differences between it and the fourth training text.
2. In claim 1, The differences from the aforementioned summary will be presented, and a judgment will be accepted on whether the differences from the aforementioned summary are valid or not. If the differences with respect to the summary are not valid, after receiving the instruction document, the differences with respect to the summary in the first document are re-acquired using the second language model based on the instruction document, and then the determination of whether the differences with respect to the summary are valid is received again. If the differences from the summary are valid, then the second language model is used to confirm that not all of the differences from the summary are described in the second document. The second language model is trained using the third and fourth data points. The third data point comprises a second command statement, a fifth learning sentence, a sixth learning sentence, a first learning instruction document, and a third learning response statement. The fourth data point comprises the second command statement, the seventh learning sentence, the eighth learning sentence, the second learning instruction document, and the fourth learning response statement. The third learning response sentence, based on the first learning instruction document, shows the differences between the fifth learning text and the sixth learning text. The fourth learning response sentence is an information processing method that indicates the second learning instruction document is not valid as a basis for judging the differences between the seventh learning document and the eighth learning document.
3. The process comprises a first to a twelfth step, In the first step described above, the first document and the second document are received, In the second step, a first prompt for generating a summary of the second document is created and sent to the first language model. In the third step, a second prompt is created to present the differences in the first document from the summary and sent to the second language model. In the fourth step described above, a judgment is made as to whether the differences are valid or not. If the aforementioned differences are not valid, in the fifth step, the first instruction document is received, and in the sixth step, a third prompt is created based on the first instruction document to re-present the aforementioned differences in the first document and sent to the second language model, and then the fourth step is repeated. If the differences are valid, in the seventh step, a fourth prompt is created and sent to the second language model to confirm that all of the differences are not described in the second document. If at least some of the differences are described in the second document, repeat the fifth step: If all of the aforementioned differences are not described in the second document, in the eighth step, create the second document and a fifth prompt for generating a third document based on the aforementioned differences and send it to the first language model. In step 9 above, the third document accepts a determination as to whether or not it needs to be modified. If the aforementioned modification is necessary, in step 10, a second instruction document is received, and in step 11, a sixth prompt for modifying the third document based on the second instruction document is created and sent to the first language model, and then step 9 is repeated. If the above modification is not necessary, in step 12, the third document is output. The second language model is trained using the first and second data points. The first data point comprises a first command statement, a first learning sentence, a second learning sentence, and a first learning response statement. The second data point comprises the first command statement, the third learning sentence, the fourth learning sentence, and the second learning response statement. The first training response sentence illustrates the differences between the first training text and the second training text. An information processing method that indicates that the second training response sentence is identical to the third training text in that there are no differences between it and the fourth training text.
4. In claim 3, The second language model is trained using the third and fourth data points. The third data point comprises a second command statement, a fifth learning sentence, a sixth learning sentence, a first learning instruction document, and a third learning response statement. The fourth data point comprises the second command statement, the seventh learning sentence, the eighth learning sentence, the second learning instruction document, and the fourth learning response statement. The third learning response sentence, based on the first learning instruction document, shows the differences between the fifth learning text and the sixth learning text. The fourth learning response sentence is an information processing method that indicates the second learning instruction document is not valid as a basis for judging the differences between the seventh learning document and the eighth learning document.
5. In claim 3, The first instruction document includes at least one of instructions for modifying the first document and comments regarding differences with respect to the summary document. The second instruction document is an information processing method comprising at least one of instructions for modifying the third document and comments relating to the third document.
6. In claim 3, The process comprises steps 13 through 15, If not all of the differences with respect to the summary are described in the second document, then in the 13th step, before the 8th step, a 7th prompt is created to determine whether there is insufficient information to generate the 3rd document, and is sent to the 3rd language model. If the aforementioned information is insufficient, in step 14, after receiving additional information, in step 15, the prompt 7 is recreated including the additional information and sent to the first language model. If the aforementioned information is not missing, repeat step 8 above. The third language model is trained using the fifth and sixth data points. The fifth data point comprises a third command statement, a ninth learning sentence, a tenth learning sentence, and a fifth learning response statement. The sixth data point comprises the third command statement, the eleventh learning sentence, the twelfth learning sentence, and the sixth learning response statement. The fifth training response sentence indicates the missing information when generating a document based on the ninth training sentence and the tenth training sentence. The sixth learning response sentence is an information processing method that indicates that no information is missing when generating a document based on the eleventh learning document and the twelfth learning document.
7. In any one of claims 3 to 6, The first document and the second document are information processing methods, each being a document demonstrating creation.
8. In claim 7, The second document mentioned above is a patent document or a utility model document. The third document is an information processing method which is a specification belonging to a patent application or a specification belonging to a utility model registration application.
9. It has a reception unit, an output unit, and a processing unit. The reception unit has the function of receiving the first document, the second document, the first instruction document, and the second instruction document. The output unit has the function of supplying a first prompt, a fifth prompt, and a sixth prompt to a first language model, and the function of supplying a second prompt, a third prompt, and a fourth prompt to a second language model. The output unit has the function of outputting a third document, The aforementioned processing unit, A process for creating the first prompt for generating a summary of the second document, A process for creating a second prompt to present the differences in the first document compared to the summary, A process to allow the user to determine whether the aforementioned differences are valid or not, If the aforementioned differences are not valid, the process of creating a third prompt to re-present the aforementioned differences in the first document based on the first instruction document, If the aforementioned differences are valid, the process of creating a fourth prompt to confirm that not all of the aforementioned differences are described in the second document, If at least some of the differences are described in the second document, the process of creating the third prompt based on the first instruction document, If all of the aforementioned differences are not described in the second document, the process of creating the second document and the fifth prompt for generating the third document based on the aforementioned differences, The process involves allowing the user to determine whether the third document needs to be modified, The system has a function to perform the following: if the aforementioned modification is necessary, create a sixth prompt for modifying the third document based on the second instruction document; The second language model is trained using the first and second data points. The first data point comprises a first command statement, a first learning sentence, a second learning sentence, and a first learning response statement. The second data point comprises the first command statement, the third learning sentence, the fourth learning sentence, and the second learning response statement. The first training response sentence illustrates the differences between the first training text and the second training text. The second training response statement indicates that there are no differences between the third training text and the fourth training text.
10. In claim 9, The second language model is trained using the third and fourth data points. The third data point comprises a second command statement, a fifth learning sentence, a sixth learning sentence, a first learning instruction document, and a third learning response statement. The fourth data point comprises the second command statement, the seventh learning sentence, the eighth learning sentence, the second learning instruction document, and the fourth learning response statement. The third learning response sentence, based on the first learning instruction document, shows the differences between the fifth learning text and the sixth learning text. The fourth learning response statement indicates that the second learning instruction document is not valid as a basis for determining the differences between the seventh learning document and the eighth learning document.
11. In claim 9, The first instruction document includes at least one of instructions for modifying the first document and comments regarding differences with respect to the summary document. The second instruction document includes at least one of instructions for modifying the third document and comments relating to the third document.
12. In claim 9, The aforementioned reception unit has a function to receive additional information, The output unit has the function of supplying a seventh prompt to the third language model. The processing unit includes a process for creating a seventh prompt to determine whether there is insufficient information to generate the third document if not all of the differences with respect to the summary are described in the second document, If the aforementioned information is insufficient, the process involves creating the seventh prompt again, including the aforementioned additional information. The system has a function to perform the process of creating the fifth prompt if the aforementioned information is not insufficient, The third language model is trained using the fifth and sixth data points. The fifth data point comprises a third command statement, a ninth learning sentence, a tenth learning sentence, and a fifth learning response statement. The sixth data point comprises the third command statement, the eleventh learning sentence, the twelfth learning sentence, and the sixth learning response statement. The fifth training response sentence indicates the missing information when generating a document based on the ninth training sentence and the tenth training sentence. The sixth learning response statement is an information processing device that indicates that no information is missing when generating a document based on the eleventh learning document and the twelfth learning document.
13. In any one of claims 9 to 12, The first document and the second document are information processing devices that are documents indicating creation, respectively.
14. In claim 13, The second document mentioned above is a patent document or a utility model document. The third document is an information processing device, which is a specification belonging to a patent application or a specification belonging to a utility model registration application.
Citation Information
Patent Citations
Patent document drafting device, method, computer program, computer-readable recording medium, server, and system
JP2022024112A