Verification apparatus, verification method, and verification program
By designing a verification device and method, using the relationship between supporting and refuting text for text editing, the problem of difficult to identify and correct false content in natural language sentences generated by large-scale language models is solved, and the effective recognition and correction of content is achieved.
Patent Information
- Application Number
- JP2023188309
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-02
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art is difficult to effectively identify and correct false content in natural language sentences generated by large-scale language models.
A verification device and method is designed to identify and correct false content by extracting the text parts that need to be verified, collect support and refutation texts, and edit text based on the relationship between the two.
It realizes effective identification and correction of false content in natural language sentences generated by large-scale language models, and improves the reliability and accuracy of the content.
Smart Images

Figure 2025076620000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to natural language processing technology, and more particularly to technology for verifying sentences including parts whose truth or falsity is unknown, such as natural language sentences generated by a large-scale language model. [Background technology]
[0002] Large-scale language models are currently attracting a great deal of attention. In particular, some large-scale language models are available online, and when a character string called a prompt is input, they output a natural language sentence following the prompt. Such large-scale language models can output a variety of sentences for a variety of topics. Therefore, large-scale language models can be effectively used when people obtain information, get new ideas, and create sentences.
[0003] However, although the output of a large-scale language model is a natural sentence, there is a problem that the content may be erroneous, because the large-scale language model basically only determines the output words by calculating the occurrence probability of each word according to parameters statistically acquired by prior learning.
[0004] A technique for detecting errors and proposing correction candidates in text generated by an information processing device or in text created by a human being is disclosed in Patent Document 1, which will be described later. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] JP 2023-83926 A Summary of the Invention [Problem to be solved by the invention]
[0006] The technology disclosed in Patent Document 1 is for correcting an expression that does not conform to the sentence expression rules in a sentence so that the expression conforms to the sentence expression rules. When the output is a natural expression, such as the output of a large-scale language model, the technology disclosed in Patent Document 1 cannot be applied.
[0007] It is widely known that the output of large-scale language models can contain falsehoods, but no technology has yet been proposed to correct or point out those falsehoods.
[0008] Therefore, an object of the present invention is to provide a verification device, a verification method, and a verification program that can verify the content of sentences that appear natural on the surface but contain falsehoods, such as the output of a large-scale language model, and perform appropriate processing. [Means for solving the problem]
[0009] A verification device according to a first aspect of the present invention includes a target portion extraction means for extracting a portion to be verified from an input sentence, a text collection means for collecting supporting text that supports the content of the portion to be verified and contradicting text that denies it from a collection of existing texts, and a selective editing means for executing a process of editing the portion to be verified in different ways depending on whether a predetermined relationship exists between the set of supporting text and the set of contradicting text collected by the text collection means.
[0010] Preferably, the text collection means includes an answer collection means which generates supporting specific questions for obtaining answers that support the expressions in the portion to be verified and contradiction specific questions for obtaining answers that contradict the contents of the portion to be verified, and collects supporting text and contradictory text by recursively executing a process of obtaining answers from a collection of existing texts for each of the questions.
[0011] More preferably, the rewriting means includes a contradictory text selection means for selecting one of the contradictory texts in accordance with a predetermined criterion, an insertion point determination means for determining an insertion point in the input sentence where the corrected new text is to be inserted, and an insertion means for generating new text based on the contradictory text selected by the contradictory text selection means and inserting the new text at the insertion point.
[0012] More preferably, the rewriting means further includes a text adding means for inputting new text into the large-scale language model and adding, consecutively to the new text, text output by the large-scale language model.
[0013] Preferably, the rewriting means further includes a deleting means for deleting at least a portion of the contradictory text.
[0014] More preferably, the process of editing the portion to be verified includes a text selection process of selecting either the contradictory text or the supporting text in accordance with predetermined criteria, and an editing process of editing at least a portion of the portion to be verified based on the contradictory text or the supporting text selected in the text selection process.
[0015] More preferably, the process of editing the portion to be verified further includes a text adding process of inputting new text into the large-scale language model and adding text output by the large-scale language model subsequent to the new text.
[0016] Preferably, the process of editing the portion to be verified further includes a deletion means for deleting at least a portion of the contradictory text.
[0017] A verification method according to a second aspect of the present invention includes the steps of: a computer extracting a portion to be verified from an input sentence; a computer collecting supporting text that supports the content of the portion to be verified and contradicting text that denies it from a collection of existing text; and a computer selectively executing a process of editing the portion to be verified according to different methods depending on whether a predetermined relationship exists between the set of supporting text and the set of contradicting text collected in the collecting step.
[0018] A verification program according to a third aspect of the present invention causes a computer to function as a target portion extraction means for extracting a portion to be verified from an input sentence, a text collection means for collecting supporting text that supports the content of the portion to be verified and contradicting text that denies it from a collection of existing texts, and a selective editing means for executing a process of editing the portion to be verified according to different methods depending on whether a predetermined relationship exists between the supporting text and the contradicting text collected by the text collection means.
[0019] The above and other objects, features, aspects and advantages of the present invention will become apparent from the following detailed description of the invention taken in conjunction with the accompanying drawings. [Brief description of the drawings]
[0020] [Figure 1] FIG. 1 is a block diagram showing the functional configuration of a text dialogue system according to a first embodiment of the present invention. [Diagram 2] FIG. 2 is a block diagram showing a functional configuration of the text selection unit shown in FIG. [Diagram 3] FIG. 3 is a diagram showing a schematic shape of a semantic network generated by the contradiction network storage unit and the support network storage unit shown in FIG. [Figure 4] FIG. 4 is a block diagram showing a functional configuration of the contradiction network storage unit shown in FIG. [Diagram 5]FIG. 5 is a flowchart representing a control structure of a program implementing the functions of the text editing device shown in FIG. [Figure 6] FIG. 6 is a block diagram showing an example of the configuration of a text generator that realizes the process of generating alternative text. [Figure 7] FIG. 7 is an external view of a computer that realizes a dialogue system according to each embodiment of the present invention. [Figure 8] FIG. 8 is a hardware block diagram of the computer shown in FIG. [Figure 9] FIG. 9 is a block diagram showing a functional configuration of a dialogue system according to the second embodiment of the present invention. [Figure 10] FIG. 10 is a functional block diagram of the network creation unit shown in FIG. [Figure 11] FIG. 11 is a flowchart showing a control structure of a program for implementing the network creation unit shown in FIG. [Figure 12] FIG. 12 is a flowchart showing a control structure of a recursive program for constructing a semantic network in the modification of the second embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0021] In the following description and drawings, the same parts are given the same reference numbers. Therefore, detailed description thereof will not be repeated. In the embodiment described below, each text is actually converted into a token string and then processed. However, in order to make the description easier to understand, the conversion into a token string and the reverse conversion from the token string to text will not be specifically shown in the following description.
[0022] First embodiment 1 Configuration 1.1 Dialogue System 50 Referring to FIG. 1, a dialogue system 50 according to a first embodiment of the present invention includes a large-scale language model 62 that generates and outputs a response sentence in a natural language in response to a text input by a user using a text input device 60, a verification device 64 that verifies the content of the response sentence that is the output of the large-scale language model 62 and edits and outputs the response sentence as necessary, a web-based question answering system 66 that is capable of communicating with a number of computers on the Internet 68 and that searches the web for one or more appropriate answers to a received question, and that is used by the verification device 64 to verify the response sentence, and an output device 70 that outputs the edited response sentence output by the verification device 64 as a response to the user's input.
[0023] In this embodiment, the large-scale language model 62 is a system separate from the verification device 64, and is an independent system capable of communicating with the verification device 64 via the Internet. It is assumed that the text input device 60 can communicate with the large-scale language model 62, and the output device 70 is located at the same position as the text input device 60. In this embodiment, the web-based question answering system 66 is also a system independent from the verification device 64, and has a function of processing inputs from various users, similar to the large-scale language model 62. It is assumed that the web-based question answering system 66 is capable of receiving a designation of the number of answers to one question, and outputting up to the designated number of answers to one question. In this embodiment, each answer output by the web-based question answering system 66 is assigned a score indicating how suitable the answer is as an answer to the question. An example of a service equivalent to the web-based question answering system 66 is "WISDOM X" provided by the applicant of the present application.
[0024] In this embodiment, the large-scale language model 62 and the web-based question-answering system 66 are separate services from the verification device 64. However, they may be provided inside the verification device 64. Furthermore, the web-based question-answering system 66 does not need to search the Internet 68 every time a question is received. For example, the web-based question-answering system 66 may collect and store a large amount of text from the Internet 68 in advance, and search for answer candidates to a question within that range. Furthermore, as described later, a pre-trained large-scale language model similar to the large-scale language model 62 may be used as part of the web-based question-answering system 66.
[0025] 1.2 Verification device 64 In this embodiment, the verification device 64 selects processing targets in order from the first sentence of the output of the large-scale language model 62, and verifies the output of the large-scale language model 62 by determining whether the contents of each processing target are appropriate. The processing target may be a sentence, or in the case of a compound sentence in which multiple sentences are connected to form one sentence, each sentence may be treated as a unit. In the case of a compound sentence in which another sentence is embedded, the embedded sentence may be processed first, and then the entire sentence may be processed. In reality, these processes are performed in a tangled manner. The method of determining these processing targets may be a rule-based method, or a machine learning model that has been trained in advance may be used to receive input in sentence units, break it down into text to be processed, and output it.
[0026] The verification device 64 includes a semantic network creation device 102 for sequentially selecting texts to be processed in a sentence output from the large-scale language model 62, collecting texts that support the content of the text to be processed (these are referred to as "supporting texts") and texts that contradict the content of the text to be processed (these are referred to as "contradictory texts") using a web-based question-answering system 66, and creating a support network, which is a semantic network made up of the supporting texts, and a contradiction network, which is a semantic network made up of the contradictory texts. The support network and the contradiction network will be described later with reference to FIG. 3.
[0027] The verification device 64 further includes a contradiction network storage unit 104 and a support network storage unit 106 for respectively storing the contradiction network and the supporting network created by the semantic network creation device 102, a support / conflict determination unit 107 for determining whether the text to be processed is supported by the text on the web or contradicts the text on the web by calculating a score by a predetermined calculation method for each of the contradiction network stored in the contradiction network storage unit 104 and the supporting network stored in the supporting network storage unit 106, and a text editing device 108 for selectively executing and outputting a process of maintaining the text to be processed as it is and a process of making necessary edits to the text to be processed according to the determination result by the support / conflict determination unit 107. The contradiction network storage unit 104 and the support network storage unit 106 may be provided in the same storage device or in different storage devices. Moreover, these storage devices may be provided in a device different from the semantic network creation device 102 or in the same device. In this embodiment, when it is determined that the text to be processed is supported by the text on the web, the text to be processed is maintained as it is. However, the present invention is not limited to such an embodiment. Such text may also be edited. For example, the text to be processed may be underlined in blue or other colors, and a link to a web document that has the highest score as text supporting the text to be processed from a collection of collected supporting texts may be added. Note that the case where nothing is done, as in this embodiment, can also be considered as one type of "editing process."
[0028] The semantic network creation device 102 includes a text storage unit 130 for storing natural language text output by the large-scale language model 62, a text selection unit 132 for selecting text to be processed from the text stored in the text storage unit 130 and outputting it in sequence, and a support network creation unit 134 and a contradiction network creation unit 136 for creating a support network and a contradiction network, respectively, using the text to be processed selected by the text selection unit 132.
[0029] In this embodiment, the text selection unit 132 has a function of decomposing the text stored in the text storage unit 130 into sentence units, and further decomposing the sentence into smaller texts, if necessary, by applying a predetermined rule to each text, and inputting them in sequence to each of the support network creation unit 134 and the contradiction network creation unit 136. Of course, the text selection unit 132 may not only decompose the text based on a rule, but also decompose the input text into multiple texts using a trained machine learning model.
[0030] 1.2.1 Support Network Creation Unit 134 2, support network creation unit 134 includes a recursive support text collection unit 180 for recursively executing a process of receiving a text to be processed from text selection unit 132, collecting descriptions supporting the content of the text from the web, and outputting the descriptions as candidate support texts. The meaning of "recursively executing a process of collecting descriptions supporting the content of the text from the web" will be described later.
[0031] The support network creation unit 134 further includes a support text verification unit 182 for verifying whether each candidate support text collected by the recursive support text collection unit 180 appropriately supports the text to be processed, a support text verification model 184 which is a pre-trained machine learning model used by the support text verification unit 182 when verifying the candidate support text, and a semantic network addition unit 186 for adding the candidate support text determined to be appropriate by the support text verification unit 182 as a new node to a semantic network consisting of the support text.
[0032] With reference to Fig. 3, in this embodiment, the support network 300 refers to a network (graph) in which the text to be processed is the root node 310, the support texts obtained during the process are nodes, and when a next support text is obtained from a certain support text, the lines connecting the nodes of these support texts are edges. In this case, each edge can be considered to correspond to an edge from a parent node (a node close to the root) to a child node (a node far from the root). When there are multiple answers to the same question, they can be treated as separate nodes and the edges can also be separate. The process of obtaining the next support text from a certain support text is performed recursively as already explained, and the details will be described later.
[0033] The recursive supporting text collection unit 180 includes a supporting question generation unit 210 for generating questions (called "supporting specific questions") from an input text that are expected to produce answers supporting the text, a question issuing unit 212 for inputting each of the questions generated by the supporting question generation unit 210 to the Web-based question answering system 66 (FIG. 1) to output one or more answers to each question from the Web-based question answering system 66, and an answer receiving unit 214 for receiving one or more answers output by the Web-based question answering system 66 and outputting them to the supporting text verification unit 182 together with information identifying the original question. Note that the text verified as an appropriate supporting text by the supporting text verification unit 182 is also provided to the supporting question generation unit 210.
[0034] The supporting question generator 210 not only generates questions based on the text provided by the target text selector 132, but also receives input of the text obtained by the recursive supporting text collector 180 and selected by the supporting text verifyer 182, and generates questions from the text. The supporting question generator 210 also generates further questions for the text obtained based on the answer to the question obtained from the target text. In this way, the recursive supporting text collector 180 operates in such a way that, not only for the target text, but also for text supporting the target text, the recursive supporting text collector 180 collects text supporting the supporting text, and further collects text supporting the supporting text. In other words, the operation of the recursive supporting text collector 180 is recursive.
[0035] The supporting question generation unit 210 generates questions as follows. A text to be processed is first input to the supporting question generation unit 210. The supporting question generation unit 210 applies a predetermined rule to the text to be processed to generate one or more questions. In this embodiment, the predetermined rule imposes a constraint that if the text to be processed is a positive sentence, then the question should be in the positive form, and if the text to be processed is a negative sentence, then the question should be in the negative form.
[0036] The questions that are generated here include, for example, what type questions, yes / no type questions, why type questions, how type questions, and the like.
[0037] For example, suppose the text to be processed is, "If we want to change our cars to combat global warming, we should switch to electric cars."
[0038] A "what" type question is a question that asks "what," such as "What type of car should we buy to combat global warming?" This question is used to check whether information about the content of the original text or alternative information is available on the Web.
[0039] A YES / NO question is one that can be answered with a YES or NO, such as, "If we want to change our cars to combat global warming, should we switch to electric cars?" This question is asked to check whether there is information on the Web that matches the content of the text to be processed, that is, information that serves as the basis for the content of the text to be processed. Looking at it from another perspective, this question can also be said to be a question that checks whether there is information on the Web that contradicts the content of the text to be processed.
[0040] Why-type questions are questions that ask for reasons, such as, "If you were to change your car to combat global warming, why would you choose an electric car?" This question is used to check whether there is evidence of the text to be processed on the Web.
[0041] Examples of how-type questions are questions such as "How should we change our cars to electric cars to combat global warming?" or "What will happen if we change our cars to electric cars to combat global warming?" These types of questions are questions that search the web for information about the history or subsequent developments of the event represented by the text to be processed, and confirm that content. If such information is available on the web, it is possible to determine whether the content is desirable or not. A model for natural language processing can be used to make this determination. Furthermore, a model can also be used to determine who the content is desirable for.
[0042] There are various other possible types of questions. However, in this embodiment, the rule for generating questions by the supporting question generator 210 is that the questions must not contradict the contents of the text to be processed.
[0043] In the case of a "what" type question, information equivalent to other options may be obtained along with information that matches the text to be processed. For example, information such as "hybrid car" may be obtained along with "electric car." In such a case, it is possible to add text such as "(Some people think that hybrid cars are good.)" to the end of the text to be processed.
[0044] 1.2.2 Contradiction Network Creation Unit 136 4, the contradiction network creation unit 136 includes a negated form generation unit 350 that converts the target text received from the text selection unit 132 into a negative form, and a recursive contradictory text collection unit 352 that receives the target text converted into a negative form from the negated form generation unit 350, collects statements supporting the content of the text, i.e., statements contradicting the original target content, from the web, and outputs them as candidates for contradictory text. Here, the meaning of "recursively collecting statements contradictory to the content of the text from the web" is the same as that explained for the supporting question generation unit 210. However, it differs from the supporting question generation unit 210 in that since the original target text is converted into a negative form, the text collected as a response to the question obtained from the text becomes a text contradictory to the original target text (contradiction text). In other words, the negative form question generated by the negative form generation unit 350 may be called a question for identifying a contradictory text that contradicts the target text (contradiction identification question). Of course, questions other than the above can be used as questions for collecting contradictory text. For example, a contradiction-identifying question may be generated by changing a specific noun, noun phrase, or verb that appears in the text to be processed to a noun, noun phrase, or verb that has an opposite meaning to the noun, noun phrase, or verb. For example, a combination such as "~ is useless" and "~ is useful" or expressions that can be paraphrased as the opposite meaning to each other, such as "decreases" and "increases," which are not necessarily antonyms, may be used.
[0045] The contradiction network creation unit 136 further includes a contradiction text verification unit 354 for verifying whether each contradiction text candidate collected by the recursive contradiction text collection unit 352 is appropriate as a candidate that contradicts the text to be processed, a contradiction text verification model 356 which is a pre-trained machine learning model used when the contradiction text verification unit 354 verifies the contradiction text candidate, and a semantic network addition unit 358 for adding the contradiction text candidate determined to be appropriate by the contradiction text verification unit 354 as a new node to the semantic network consisting of the contradiction text.
[0046] The recursive contradictory text collection unit 352 includes a contradictory question generation unit 370 for generating questions (contradiction-identifying questions) from the input text, which are expected to have answers that contradict the text to be processed, a question issuing unit 372 for inputting each of the questions generated by the contradictory question generation unit 370 to the web-based question answering system 66 (FIG. 1) to output one or more answers to each question from the web-based question answering system 66, and an answer receiving unit 374 for receiving one or more answers output by the web-based question answering system 66 and outputting them to the contradictory text verification unit 354 together with information identifying the original question. Note that the text verified by the contradictory text verification unit 354 as being an appropriate contradictory text is also provided to the contradictory question generation unit 370. The contradiction-identifying questions are obtained, for example, by negating the input text and then transforming it into a question.
[0047] The contradictory question generator 370 generates questions not only based on the text provided by the negated form generator 350, but also from the questions obtained by the recursive contradictory text collector 352 and selected by the contradictory text verifyer 354. The contradictory question generator 370 further generates questions for the text obtained by the recursive text collector 800 based on answers to questions obtained from text obtained by converting the target text into a negated form. The recursive contradictory text collector 352 repeats this process. That is, the operation of the recursive contradictory text collector 352 is also recursive.
[0048] In addition, the text to be processed is converted to a negative form in the contradiction network creation unit 136. The processing in the recursive contradiction text collection unit 352 and the processing by the contradiction text verification unit 354, the model for verifying contradiction text 356, and the semantic network addition unit 358 are substantially the same as those in the recursive supporting text collection unit 180, the supporting text verification unit 182, and the semantic network addition unit 186 shown in Fig. 2, respectively, but since the initial input is converted to a negative form, all of the texts obtained by the contradiction network creation unit 136 are contradictory texts.
[0049] The generation of questions by the contradiction question generator 370 is substantially the same as that by the supporting question generator 210 shown in FIG.
[0050] For example, if the text to be processed is "If you are asking how we should change our cars to combat global warming, it would be good to switch to electric cars," the negative form generation unit 350 will change this question to "If you are asking how we should change our cars to combat global warming, it would not be good to switch to electric cars."
[0051] In contrast, a "what" type question is a question that asks "what," such as "What kind of car is not good for combating global warming?" This question is used to check whether there is information on the Web that contradicts the content of the original text or alternative information.
[0052] An example of a YES / NO question about a contradictory text is, "Isn't it a good idea to switch to electric cars as a measure against global warming?" This is a question to check whether there is information on the Web that matches the content that negates the text being processed.
[0053] An example of a why-type question is, "If you were to change your car to combat global warming, why would it be bad to switch to an electric car?" This question is used to check whether there is evidence on the Web that would disprove the text being processed.
[0054] Examples of how-type questions are questions such as "How can we avoid switching cars to electric vehicles as a measure against global warming?" or "What would happen if we didn't switch cars to electric vehicles as a measure against global warming?" These questions are questions that check whether there is information on the Web about the history or subsequent developments of events that negate the text to be processed. If such information is available on the Web, there is a high possibility that the content that negates the text to be processed is a fact.
[0055] There are various other possible types of questions. However, unlike the supporting question generator 210 in Fig. 2, the rules for generating questions by the contradictory question generator 370 must be such that they generate questions to obtain answers that contradict the contents of the text to be processed. In other words, the rules for generating questions by the contradictory question generator 370 must be such that they generate questions that are consistent with the contents that contradict the text to be processed.
[0056] The recursive supporting text collection process by the support network creator 134 and the recursive contradictory text collection process by the recursive contradictory text collector 352 must be terminated when an appropriate termination condition is met. For example, the termination condition can be when the number of generated texts or their candidates since the start of text generation reaches an upper limit. Alternatively, the termination condition can be when the number of generated questions since the start of text generation reaches an upper limit. Another possible condition is when the sum of the number of generated texts and the number of questions reaches an upper limit.
[0057] In addition, when generating supporting questions and negative questions, an upper limit may be set on the number of questions generated for one input text, or an upper limit may be set on the number of answers (supporting text candidates or contradictory text candidates) obtained for one question. Furthermore, answers obtained for one question may be limited to those with a certain score (confidence level) or higher.
[0058] The generation of questions is performed while generating a semantic network such as that shown by the support network 300. In this case, the order of generating the network can be either depth-first or breadth-first. In the case of depth-first, it is desirable to stop searching when a certain depth (layer) is reached and backtrack. In the case of breadth-first, the generation of the network itself may be terminated when a certain depth (layer) is reached or when a score above a certain level is not obtained.
[0059] 1.2.3 Support / Conflict Judgment Unit 107 and Text Editing Device 108 Fig. 5 shows a control structure of a program for implementing the support / contradiction determination unit 107 and the text editing device 108 shown in Fig. 1 by a computer. Referring to Fig. 5, this program is executed after the generation of both the support network and the contradiction network is completed. This program includes step 400 of calculating a score Sp to be assigned to the support network and a score Sn to be assigned to the contradiction network, and step 402 of calculating the reliability of the content of the text to be processed based on the scores Sp and Sn calculated in step 400.
[0060] There are various methods for calculating the scores Sp and Sn. For example, the following methods are considered:
[0061] A) The total number of nodes in the resulting text (supportive or contradictory text) B) The sum of the number of questions (i.e. edges) generated during the generation of the semantic network. C) The sum of the number of nodes and the number of edges D) The sum of the scores assigned to each supporting text and the sum of the scores assigned to each contradictory text (The score for each text can be the score output by the model for verifying supporting text 184 for the supporting text. The same applies to the contradictory text. A constant can also be assigned as the score for each text. If this constant is set to "1", the result will be the same as C) above.) E) In the calculation of the various sums described above, a weight that decreases as the distance between each node and the root node increases (for example, the number of edges between each node and the root node, or the number of nodes between each node and the root node, etc.) is multiplied by the score of each node.
[0062] In the above embodiment, the support network creation unit 134 collects only the support texts, and the contradiction network creation unit 136 collects only the contradiction texts. However, the present invention is not limited to such an embodiment. When calculating the score of each text belonging to each semantic network, for example, when constructing a support network, a text that "contradicts a part of the support text" may be found. In such a case, the effectiveness of the support network will decrease. Regarding the "contradiction text", if a contradiction-specific question about the contradiction text finds a contradiction text, the effectiveness of the contradiction network will decrease. Such a contradiction relationship can be continued recursively. In the process, the contradiction of the contradiction will ultimately work to increase the effectiveness of the network. Also, the "contradiction of the contradiction of the contradiction" will work to decrease the effectiveness of the network. It is desirable to take these circumstances into consideration when calculating the scores Sp and Sn.
[0063] For this reason, it is preferable to perform two processes for calculating the scores of the support network and the contradiction network.
[0064] First, we consider that the contradictory texts for each text in each of these semantic networks are also included as members of that semantic network, in which case it is reasonable to reduce the score of the semantic network according to the score of the text.
[0065] The second is the idea of moving the above-mentioned text into the other semantic network. In this case, it is reasonable to operate it so that the score of the semantic network is the sum of the scores of the texts belonging to each semantic network multiplied by an appropriate weight. Incident texts found during the construction of a support network can be added as child nodes of any node in the contradiction network. However, it is desirable to move the nodes according to a certain policy, such as adding it directly under the root node, or adding it as a child node of any node at the same level (or one level above) where the text was found.
[0066] In both the first and second cases, the idea is the same: the score indicating the effectiveness of each semantic network is determined largely by its size, but if there is a text in the network whose meaning contradicts the meaning of the semantic network (supporting / contradicting), the score is reduced by an amount corresponding to that text.
[0067] In addition, it is believed that when such operations are repeated recursively, the relationship between the obtained text and the truth of the original sentence becomes weaker. Therefore, as described above, it is desirable to reduce the impact on the effectiveness of the network according to the number of recursion stages. Taking this into consideration, it is desirable to calculate the score of each semantic network based on its size, while also performing additions and subtractions that take into account the presence of contradictory text identified recursively and the number of recursion stages leading up to it.
[0068] In the determination by the support / contradiction determination unit 107, the simplest way is to follow the rule of adopting the one with the larger score between the support network and the contradiction network. For example, if Sp > Sn, it can be determined that the content of the text to be processed is reliable, and if Sp < Sn, it can be determined that the content of the text to be processed is unreliable. That is, the ratio of Sp to the sum of Sp and Sn is used as the reliability. Then, if Sp > 1 / 2, the target text is determined to be reliable, and if Sp < 1 / 2, it is determined to be unreliable.
[0069] In this embodiment, the reliability calculation uses the reliability as described above. However, in this embodiment, as the object to be compared with this reliability, for example, a first threshold larger than 1 / 2 and a second threshold smaller than 1 / 2 are provided. That is, referring to FIG. 5, this program further includes a step 404 of branching the control flow according to whether the reliability is greater than the first threshold, and when the determination in step 404 is affirmative, considering the text to be processed as reliable, and the text corresponding to the node with the highest score among the nodes of the support network (and the URL (Uniform Resource Locator) of the text or the message including the text) is embedded in the text part to be processed in the document in the form of an annotation or a link or the like as indicating the basis of the text to be processed. Note that in step 406, the document may not be particularly changed (not changing it is also a kind of "editing"). Or, the text part to be processed may be shown in bold or underlined in green to indicate that it is reliable.
[0070] The determination in step 404 is not limited to the above. For example, when the value of the score Sp exceeds a predetermined threshold, the determination in step 404 may be positive regardless of the value of the score Sn. Conversely, when the value of the score Sn exceeds a predetermined threshold, the determination in step 404 may be negative regardless of the value of the score Sp. Furthermore, when both values exceed the threshold, the larger value may be adopted. There are various other possible methods for making a determination based on the scores Sp and Sn.
[0071] This program further includes step 408 of branching the control flow according to whether the reliability is less than a second threshold value when the determination in step 404 is negative; step 410 of modifying the text to be processed by deleting at least a part of the text to be processed using the highest-scoring text among the contradictory texts and adding alternative text according to the content of the contradictory text to the deleted part or other appropriate part determined from the sentence pattern of the original sentence when the determination in step 408 is negative; and step 412 of editing the part of the text to be processed to indicate that the part cannot be judged to be reliable or unreliable when the determination in step 408 is negative. As an example of editing in step 412, it is possible to leave the relevant character string unchanged and to underline the part in red. In some cases, it is also possible to leave the relevant character string unchanged and quote the text to be processed that has the highest score among the contradictory texts in parentheses and add a character string such as "(Regarding this part, there are also opinions such as ... (quoted part) ...)".
[0072] This program further includes a step 414 which is executed after the determination in step 404 is positive and the processing of step 406 is completed, or after the determination in step 404 is negative and, as a result of the determination in step 408, the processing of step 410 or step 412 is completed, for updating the text stored in text storage unit 130 with the edited text, and a step 416 which selects the text next to the edited portion from the text stored in text storage unit 130 as the text to be processed, instructs text selection unit 132 (see Figure 1) to start the above-mentioned processing, and then terminates the processing of the current text to be processed.
[0073] Fig. 6 shows an example of the configuration of a text generator 420 for realizing the process of generating alternative text used in step 410 shown in Fig. 5, for example, in the text editing device 108 shown in Fig. 1. Referring to Fig. 6, the text generator 420 includes a text replacement unit 432 for performing editing in response to an input 430 including the original text stored in the text storage unit 130, the contradictory text, etc., by deleting a part of the original text and inserting the contradictory text at a predetermined position, and outputting the edited text.
[0074] The input 430 to the text exchanger 432 includes the original text stored in the text store 130. In the original text, the tokens at the start and end of the text to be processed are labeled. The input 430 also includes the contradiction text used in the editing process. The contradiction text is the highest scoring contradiction text in the first layer of the contradiction network. The input 430 also includes information indicating the type of question for which the contradiction text was obtained.
[0075] In this embodiment, the text replacement unit 432 performs this process using a text replacement model 436. Inputs 434 to the text replacement model 436 are the target text of the original text, a predetermined length of text before and after the target text (for example, one sentence before and after each), and the contradictory text that is the source of the text to be replaced. The text replacement model 436 has a function of replacing a part of the target text of the input text with the contradictory text or a part of the contradictory text, and outputting it as a replaced text 438. The text replacement model 436 is trained in advance to determine which part of the target text should be replaced with which part of the contradictory text, and how the part before and after the part to be replaced should be modified, depending on the type of question. The manner of replacement by the text replacement model 436 differs depending on the type of question.
[0076] The text generator 420 further includes a prompt generator 444 for generating a prompt to be input to the large-scale language model 446 based on the large-scale language model 446, the edited text 440 output by the text exchanger 432, and information 442 indicating the type of question on which the text is based, and inputting the generated prompt to the large-scale language model 446, and a text integration unit 450 for performing a process of integrating the text by adding the text 448 output by the large-scale language model 446 to the end of the edited text 440 output by the text exchanger 432, replacing the original text stored in the text storage unit 130 with the new text, and outputting the result as edited text 452. As shown in FIG. 1, the text integration unit 450 also instructs the text selection unit 132 to continue processing in the new text from immediately after the edited portion. When the text selected by the text selection unit 132 is finished, the verification process for the text output by the large-scale language model 62 is finished. In addition, if the text to be processed approaches the end of the text stored in the text storage unit 130, the end time may be brought forward by limiting the length of the text output from the large-scale language model 446.
[0077] 2 operations The dialogue system 50 according to the first embodiment operates as follows: A user inputs a prompt to the large-scale language model 62 using the text input device 60, and the large-scale language model 62 outputs text. This text is stored in the text storage unit 130 as original text.
[0078] The text selection unit 132 selects the first sentence in the text storage unit 130, and if necessary further divides this sentence into portions comprising the text to be processed, and inputs the first portion as the text to be processed to the support network creation unit 134 and the contradiction network creation unit 136.
[0079] The support question generator 210 (FIG. 2) of the support network creator 134 generates one or more questions for obtaining supporting text based on the input text to be processed. The question generator 212 provides these questions to the web-based question answering system 66. The web-based question answering system 66 searches the Internet 68 for one or more answers to each given question, and provides them to the answer receiver 214 of the support network creator 134. The answer receiver 214 inputs the answers received from the web-based question answering system 66 and the text (the text to be processed or the supporting text) on which the question generator 212 generated the questions to the support text verification model 184, and determines whether the answers obtained from the web-based question answering system 66 support the content of the text on which the questions were based. If the obtained answers support (serve as evidence for) the text on which the questions were based, the answer receiver 214 adopts the answers and provides them to the semantic network adder 186, and if not, discards the answers.
[0080] The supporting question generator 210 of the support network creator 134 further repeats the above-mentioned method for each answer adopted by the supporting text verification unit 182 among the answers obtained from the web-based question answering system 66, obtains each answer, and adopts text supporting the supporting text on which the answer is based. In this way, the support network creator 134 performs recursive processing such as obtaining supporting text for the first input, obtaining supporting text for that supporting text, and obtaining supporting text that supports that supporting text. The semantic network adder 186 of the support network creator 134 creates a support network using the obtained supporting text, and stores it in the support network storage unit 106. The support network creator 134 ends this processing when the total number of obtained supporting texts reaches an upper limit.
[0081] The contradiction network creation unit 136 transforms the input target text into a negative form, and then, similarly to the support network creation unit 134, obtains one or more answers to the question from the Web-based question-answering system 66 to collect contradictory text that supports (contradicts) the input target text transformed into a negative form. The contradiction network creation unit 136 further recursively executes a process of obtaining contradictory text from the Web-based question-answering system 66 using the collected contradictory text. The contradiction network creation unit 136 ends the collection of contradictory text when the total number of obtained contradictory texts reaches an upper limit. The contradiction network creation unit 136 generates a contradiction network using the collected contradictory text, and stores it in the contradiction network storage unit 104.
[0082] Referring to FIG. 5, the text editing device 108 shown in FIG. 1 calculates the obtained support network score Sp and contradiction network score Sn (step 400). The text editing device 108 calculates the reliability of the text to be processed from Sp and Sn (step 402). If the reliability is greater than a first threshold, the text editing device 108 basically maintains the input text to be processed and outputs it with an embedded reference (link) to the answer (or a passage including the answer) that provided the support text with the highest score as a basis. This embedding of the reference is not necessarily required. Instead of or in addition to this embedding, a description of a text relating to an alternative thing to the thing described in the text to be processed, or a reference to that text, may be embedded in parentheses after the text to be processed. Note that maintaining the description without changing it is also a type of "editing".
[0083] The text editing device 108 then replaces the original text stored in the text storage unit 130 with the changed text (step 414). Furthermore, the text editing device 108 instructs the text selection unit 132 to start the next process on the sentence or part of the sentence next to the changed portion of the original text as the text to be processed (step 416).
[0084] On the other hand, if the determination in step 404 is negative, the text editing device 108 determines whether the confidence level is less than a second threshold value (step 408). If the determination in step 408 is positive, the text editing device 108 executes a process of modifying the source text including the target text by using the contradictory text with the highest score among the contradictory texts (processing by the text generating device 420 in FIG. 6). Thereafter, control proceeds to step 414. The processes from step 414 onwards have been described above.
[0085] If the determination in step 408 is negative, the text editing device 108 modifies the original text so as to add an alert indication to the part of the text to be processed, indicating that the evidence for the text is weak, as an annotation. After this, the control proceeds to step 414. The processing from step 414 onwards has already been described. Note that the annotation may further include a reference to a contradictory text that contradicts the text.
[0086] In this way, the editing process for the original text held in the text storage unit 130 is performed by progressing through the text to be processed one location at a time. When there are no more locations in the final original text that can be further edited (or when some other predetermined end condition is satisfied, for example, when the unprocessed portion of the original text is equal to or less than a predetermined number of sentences), the process ends, and the text finally stored in the text storage unit 130 is output as the edited text.
[0087] In the above embodiment, it is assumed that the large-scale language model 62 (FIG. 1) and the large-scale language model 446 (FIG. 6) are separate. However, the present invention is not limited to such an embodiment. The two may be the same. Furthermore, each of the large-scale language model 62 and the large-scale language model 446 may be a part of the dialogue system 50, or may be included in a service provided by another entity outside the dialogue system 50.
[0088] Furthermore, in the above embodiment, the text exchange unit 432 shown in FIG. 6 edits the original text using a machine learning model. However, the present invention is not limited to such an embodiment. The original text may be edited using a rule base. Alternatively, a machine learning model that identifies a deleted portion of the original text, a machine learning model that identifies a portion of the original text where the changed text is to be inserted, a machine learning model that identifies a change in the text before and after the text is inserted, etc. may be used. Also, each model or rule may be changed depending on the type of question that is the basis of the supporting text or the contradictory text.
[0089] The function of the prompt creation unit 444 differs depending on the type of large-scale language model 446 shown in Fig. 6. For example, the text itself changed by the text exchange unit 432 may be input to the large-scale language model 446, and the large-scale language model 446 may be trained so that the large-scale language model 446 outputs the subsequent text. Alternatively, in order to control the output from the large-scale language model 446, some keyword may be input to the large-scale language model 446 in the form of a prompt.
[0090] In the above embodiment, the output of the large-scale language model is maintained as is or modified according to whether or not there is a statement that serves as evidence in the existing text. It is believed that most of the existing text has been edited or revised by humans. Therefore, even if the output of the large-scale language model appears natural at first glance, if no evidence can be found in the existing text as in the above embodiment or if a contradictory statement is found, the output can be determined to be unreliable. Furthermore, by editing such text based on an existing statement that has evidence as in the above embodiment, the reliability of the output of the large-scale language model can be increased.
[0091] In the above embodiment, the reliability of the text is calculated based on the scores of the support network and the contradiction network, and the scores are based on the size of each network. The support network and the contradiction network each contain a considerable number of support texts or contradiction texts. Moreover, most of the support texts and contradiction texts further have support texts and contradiction texts that are the basis for the texts. Therefore, the number of texts included in the support network and the contradiction network is large and the range is wide. When trying to determine the reliability of the output of a large-scale language model based on existing texts, there is a method of embedding erroneous information and false information in the web in order to make the determination incorrect. However, even if the information embedded in the web by such an attempt is based on many texts and even the basis for each text is collected as a network, as in the support network and the contradiction network, the amount of erroneous information and false information is relatively small compared to other reliable information. As a result, the above embodiment can reduce the possibility of making an incorrect judgment about whether the text is reliable.
[0092] In the above embodiment, even when a text is determined to be reliable, a reference to an existing text or a passage containing the text that is the basis for the determination can be embedded in the text. This allows the user to easily check the basis for the edited text. As a result, the user can use the output of the large-scale language model with confidence, and the usefulness of the large-scale language model can be improved.
[0093] In the above embodiment, the web-based question answering system 66 is used to obtain an answer to a question. However, the present invention is not limited to such an embodiment. For example, in the case of a language model such as the large-scale language model 62, it is possible to provide an answer to a question. The possibility that such an answer is incorrect is high compared to a system such as the web-based question answering system 66. However, as in this embodiment, if the web-based question answering system 66 is used in combination with the web-based question answering system 66, which allows the description on the basis of the obtained answer to be confirmed, the large-scale language model can be used in addition to a system such as the web-based question answering system 66. For example, an embodiment is conceivable in which a question is given to the large-scale language model in parallel with the web-based question answering system 66.
[0094] Most of the machine learning models used in the above embodiments are for processing natural language. For learning each model, learning data according to the function (combination of input and output) of each model in the above embodiments may be used.
[0095] 3. Hardware Configuration Fig. 7 is an external view of a computer system 600 that realizes the verification device 64 in the dialogue system 50 according to the first embodiment of the present invention shown in Fig. 1. Fig. 8 is a hardware block diagram of the computer system 600. The hardware configuration of the computer system 600 will be described below.
[0096] 7, this computer system 600 includes a computer 650 having a DVD (Digital Versatile Disc) drive 662, and a keyboard 654, a mouse 656, and a monitor 652 for interacting with a user, all of which are connected to the computer 650. Of course, these are just one example of a configuration for when interaction with an operator becomes necessary, and any general hardware and software (e.g., a touch panel, voice input, pointing device in general) that can be used for interacting with an operator can be used.
[0097] 7 and 8, the computer 650 includes a central processing unit (CPU) 710, a graphics processing unit (GPU) 712, and a bus 720 connected to the CPU 710, the GPU 712, and the DVD drive 662, in addition to a DVD drive 662. The computer 650 further includes a read-only memory (ROM) 714 connected to the bus 720 and storing a boot-up program of the computer 650, a random access memory (RAM) 716 connected to the bus 720 and storing instructions constituting a program, a system program, working data, and the like, and a solid state drive (SSD) 718, which is a non-volatile memory, connected to the bus 720. The SSD 718 is for storing programs executed by the CPU 710 and the GPU 712, and data used by the programs executed by the CPU 710 and the GPU 712, and the like. The computer 650 further includes a network I / F (Interface) 726 that provides a connection to a network enabling communication with other terminals, and a USB port 664 to which a USB (Universal Serial Bus) memory 702 can be attached / detached and that provides communication between the USB memory 702 and each part within the computer 650.
[0098] The computer 650 further includes an audio I / F 722 that is connected to the microphone 660 and the speaker 658 and the bus 720, and has the function of reading out audio signals, video signals, and text data generated by the CPU 710 and stored in the RAM 716 or the SSD 718 in accordance with instructions from the CPU 710, performing analog conversion and amplification processing to drive the speaker 658, and digitizing the analog audio signal from the microphone 660 and storing it at any address in the RAM 716 or the SSD 718 specified by the CPU 710.
[0099] In the above embodiment, the programs and the like for realizing each function of the verification device 64 shown in Fig. 1 are stored in, for example, the ROM 714, SSD 718, DVD 700, or USB memory 702 shown in Fig. 8, or a storage medium of an external device (not shown) connected via the network I / F 726 and the network 704. Typically, these data and parameters are written, for example, from outside to the SSD 718, and loaded into the RAM 716 when the computer 650 is executed.
[0100] 1 is stored in a DVD 700 mounted in a DVD drive 662, and transferred from the DVD drive 662 to the SSD 718. Alternatively, these programs are stored in a USB memory 702, which is mounted in a USB port 664 and the programs are transferred to the SSD 718. Alternatively, the programs may be transmitted to the computer 650 via the network 704 and stored in the SSD 718. Of course, a source program may be input using the keyboard 654, the monitor 652, and the mouse 656, and the compiled object program may be stored in the SSD 718.
[0101] The program is loaded into the RAM 716 when executed. If the program is written in a script language, the script input by the operator using the keyboard 654 or the like may be stored in the SSD 718. In the case of a program that runs on a virtual machine, a program that functions as a virtual machine must be installed in the computer 650 in advance. A machine learning model such as a deep neural network is used for the supporting text verification model 184 shown in FIG. 2, the contradictory text verification model 356 shown in FIG. 4, the text exchange model 436 and the large-scale language model 446 shown in FIG. 6, and the like. In the computer system 600, a machine learning model that has been trained in another device may be used, or the computer system 600 may be used as a learning device to train the machine learning model.
[0102] The CPU 710 reads a program from the RAM 716 according to an address indicated by a register (not shown) called a program counter inside the CPU 710 and interprets the instruction. The CPU 710 reads data required for executing the instruction from the RAM 716, the SSD 718, or other devices according to an address specified by the instruction, and executes the process specified by the instruction. The CPU 710 stores the execution result data in an address specified by the program, such as the RAM 716, the SSD 718, or a register in the CPU 710. Depending on the address, the execution result data is output from the computer to the outside via, for example, the network I / F 726. The output destination is, for example, the web-based question answering system 66 shown in FIG. 1. At this time, the value of the program counter is also updated by the program. The computer program may be directly loaded into the RAM 716 from the DVD 700, the USB memory 702, or via the network 704. Note that some tasks (mainly numerical calculations) among the programs executed by the CPU 710 are issued to the GPU 712 according to instructions included in the program or according to the analysis results when the CPU 710 executes the instructions.
[0103] The program for implementing the functions of each part of the verification device 64 (FIG. 1) according to the embodiment described above by the computer 650 includes a plurality of instructions written and arranged to operate the computer 650 to implement those functions. Some of the basic functions required to execute the instructions may be provided by an OS (Operating System) or a third party program running on the computer 650, various toolkit modules installed on the computer 650, or an execution environment of the program. Therefore, the program does not necessarily include all of the functions required to implement the system and method of this embodiment. The program may include only instructions to execute the operations of each of the above-mentioned devices and their components by statically linking appropriate functions or modules at compile time or dynamically calling them at run time in a controlled manner so as to obtain the desired results. The operation method of the computer 650 for this purpose is well known. Therefore, the description of the operation method of the computer 650 will not be repeated in this section.
[0104] The GPU 712 is capable of parallel processing, and can simultaneously execute a large amount of calculations associated with machine learning and inference in a parallel or pipelined manner. For example, parallel calculation elements found in a program when the program is compiled, or parallel calculation elements found when the program is executed, are dispatched from the CPU 710 to the GPU 712 as needed, and executed. The results are returned to the CPU 710 directly or via a predetermined address in the RAM 716, and assigned to a predetermined variable in the program.
[0105] Second embodiment 1 Configuration 9, a dialogue system 750 according to the second embodiment includes a text input device 60, a large-scale language model 62 which receives an output from the text input device 60 and outputs a text to be verified, a verification device 760 which verifies the content of a response sentence which is an output from the large-scale language model 62 and edits and outputs the response sentence as necessary, a web-based question answering system 66 which the verification device 760 uses to verify the response sentence, and an output device 70 which outputs the edited response sentence output by the verification device 760 as a response to a user's input.
[0106] The verification device 760 includes a semantic network creation device 770 for sequentially selecting text within a sentence output by the large-scale language model 62 and creating a semantic network including both supporting text and contradicting text related to the text to be processed based on the text to be processed, a network memory unit 772 for storing the semantic network created by the semantic network creation device 770, a support / contradiction determination unit 774 for determining whether the text to be processed is reliable or not based on the supporting text and contradicting text stored in the network memory unit 772, and a text editing device 108 for editing and outputting the text to be processed if necessary in accordance with the determination result by the support / contradiction determination unit 774.
[0107] The semantic network creation device 770 includes the same text storage unit 130 and text selection unit 132 as shown in FIG. 1, and a network creation unit 780 for creating one or more questions based on the text to be processed selected by the text selection unit 132 and providing the questions to the Web-based question-answering system 66, thereby collecting texts of documents on the Web that are output as answers from the Web-based question-answering system 66, and creating a semantic network including both supporting text and contradictory text based on the texts, and storing the network in the network storage unit 772.
[0108] 10, the network creation unit 780 shown in FIG. 9 includes a recursive text collection unit 800 for collecting many supporting texts and contradictory texts by performing recursive processing, in which a question is generated starting from the target text selected by the text selection unit 132, an answer is obtained using the web-based question answering system 66, and a new question is generated based on the answer to obtain the next answer; a text verification model 802 that is pre-trained to receive two texts and output a score indicating whether one of the texts supports the other text; and a text verification unit 804 for providing the target text and the individual texts collected by the recursive text collection unit 800 to the text verification model 802 to determine whether the text collected by the recursive text collection unit 800 supports the target text, and classifying the texts collected by the recursive text collection unit 800 into supporting texts and contradictory texts according to the determination result and outputting the classified texts.
[0109] The text verification model 802 is pre-trained to receive as input a string formed by concatenating the first text and the second text with a separation token between them, and to output the probability that the second text is a text that supports the first text and the probability that the second text is a text that contradicts the first text as a score for the second text.
[0110] The recursive text collection unit 800 includes a question generation unit 810 that generates and outputs a plurality of questions based on the input text using the method described in the first embodiment. However, unlike the supporting question generation unit 210 and the contradiction question generation unit 370 in the first embodiment, the question generation unit 810 has no particular restrictions on the questions to be generated, and generates both supporting and contradiction specific questions for the input text.
[0111] The recursive text collection unit 800 further includes a question issuing unit 212 for inputting the questions generated by the question generating unit 810 to the Web-based question answering system 66, and an answer receiving unit 214 for receiving one or more times output by the Web-based question answering system 66 for each question and providing the same to the text verification unit 804.
[0112] 2 Programmatic implementation Fig. 11 is a flowchart showing a control structure of a recursive program for realizing network creation unit 780 shown in Fig. 10. Referring to Fig. 11, this program is realized as a recursive function receiving a set of texts as an argument. In this embodiment, the set of texts as an argument is prepared as an array, and it is the address of the array that is actually passed to the program.
[0113] This program includes step 850 of generating one or more questions by repeatedly executing the question generation process of step 852 for all input texts or until the total number of answers + K × the number of questions generated (K is a positive integer) is greater than a first threshold value, step 854 of executing step 856 of searching for answers to each question using the web-based question answering system 66 for all questions generated in step 850 or until the total number (cumulative number) of answers obtained is equal to or greater than a second threshold value, step 858 of executing step 860 (described later) for each answer searched for in step 854, step 862 of branching the flow of control according to whether the total number (cumulative number) of answers obtained by the process up to that point is greater than a third threshold value after the process of step 858 is completed, and step 864 of executing a process of recursively calling itself with the set of answers searched for in step 856 as an argument when the determination in step 862 is negative, and then terminating the execution of the program and returning control to the caller. If the determination in step 862 is positive, this program terminates the execution and returns control to the caller.
[0114] Step 860 includes step 880 of determining whether the answer being processed is supporting text or contradicting text using the text verification model 802 shown in FIG. 10 and tagging the text with an indication of whether it is supporting text or contradicting text, and step 882 of adding the text tagged in step 880 to the semantic network as a child node of the node corresponding to the text that was the basis of the question, with the score obtained by the text verification model 802 and the tag obtained in step 880 attached.
[0115] The integer K used as the termination condition in step 852 has the following meaning. In this embodiment, a plurality of questions are generally obtained from one input text. Also, a plurality of answers are generally obtained for each question. As a result, a large amount of text is obtained by performing a process of searching for an answer from one input text only once. By performing recursive processing, the number of answers obtained increases exponentially. Since such processing requires a large calculation cost, it is necessary to terminate the processing at an appropriate time. In this embodiment, the condition is set as a guideline when the total number of answers obtained (the cumulative number by multiple recursive processing) exceeds a certain threshold (the third threshold). However, in the example shown in FIG. 11, when the total number of answers obtained by the processing so far is close to the third threshold, if the processing of steps 850 and 854 is completely performed, the total number of answers finally obtained may greatly exceed the third threshold, and the processing may take a long time. Therefore, in step 850, when the total number of answers obtained so far (the cumulative value) is close to the third threshold, the number of questions to be generated is limited. In this embodiment, a constant K (≧1) is assumed as a guide for the number of answers that can be obtained for one question, and the process of step 850 ends when the total number of answers so far + K × the number of input texts becomes greater than the first threshold. The third threshold may or may not be equal to the first threshold, including cases where the constant K is accurately determined (where the number of answers that can be obtained for one question is fixed).
[0116] The reason why the termination condition in step 854 is "when the total number of responses (cumulative number) becomes greater than the second threshold" is to limit the processing time. Generally, the second threshold should be equal to the third threshold, but there is no problem if the second threshold is slightly different from the third threshold.
[0117] It is not necessary to provide the termination conditions using each threshold value added in steps 850 and 854. In that case, the processing time may be longer, but these conditions may not be used when the third threshold value is small, for example.
[0118] 3 operations The dialogue system 750 according to the second embodiment operates as follows: Referring to Fig. 9, a user inputs a prompt via a text input device 60 into the large-scale language model 62. The large-scale language model 62 outputs the text following the prompt. This text is stored in the text store 130.
[0119] The text selection unit 132 first selects the first sentence (or part of the sentence) stored in the text storage unit 130, and provides this to the network creation unit 780 as the text to be processed.
[0120] 10, in the recursive text collection unit 800 of the network creation unit 780, the recursive text collection unit 800 generates one or more questions from the text to be processed, and inputs them to the question issuing unit 212. At this time, the question generating unit 810 generates both a support identifying question and a contradiction identifying question.
[0121] The question issuing unit 212 inputs each of the questions received from the question generating unit 810 to the web-based question answering system 66. The web-based question answering system 66 searches the Internet 68 for each question and outputs text determined to be appropriate as an answer to the question as an answer. The answer receiving unit 214 receives the text of these answers and provides it to the text verifying unit 804.
[0122] The text verification unit 804 combines each of the answer texts received from the answer receiving unit 214 with the processing target text input from the question generating unit 810, sandwiching a separated token therebetween, and inputs the combined text to the text verification model 802. In response to this input, the text verification model 802 outputs the probability that the answer text supports the content of the processing target text and the probability that the answer text contradicts the content of the processing target text as the score of the answer text. Based on the output of the text verification model 802, the text verification unit 804 assigns a tag to the answer text indicating whether the answer text is a supporting text or a contradicting text, and provides the tag and the score to the semantic network adding unit 806. The text verification unit 804 also inputs the answer text to the question generating unit 810.
[0123] The semantic network adding unit 806 adds the answer text received from the text verifying unit 804 as a new node to the semantic network. At this time, the semantic network adding unit 806 identifies the text that was the source of the question from which the current answer text was obtained, and adds the new node as a child node of the node corresponding to that text.
[0124] Meanwhile, the question generator 810 now generates one or more questions based on the new text received from the text verification unit 804, and provides each question to the question issuing unit 812. The question issuing unit 212 inputs each question to the Web-based question answering system 66. The answer receiving unit 214 receives the text of one or more answers output by the Web-based question answering system 66 for each question, and provides it to the text verification unit 804.
[0125] Thereafter, the network creation unit 780 repeatedly executes the recursive process described above. When the number of nodes (number of texts) added to the semantic network exceeds the third threshold, the text collection process ends. As a result, the network storage unit 772 shown in FIG. 9 stores a semantic network that includes both supporting text and contradicting text for the text to be processed.
[0126] 9, the support / contradiction determination unit 774 determines whether the text to be processed input via the text input device 60 is reliable or not based on the label and / or score assigned to each node in the semantic network stored in the network storage unit 772, and provides the result to the text editing device 108. As in the first embodiment, the text editing device 108 edits the text to be processed as necessary to clearly indicate whether the text is reliable or unreliable, or replaces the unreliable text to be processed with reliable text obtained when the semantic network is created. After completing this editing, the text editing device 108 updates the contents of the text storage unit 130 with the edited text. The text editing device 108 further instructs the text selection unit 132 to select the next text after the processed text from among the texts stored in the text storage unit 130.
[0127] In response to this instruction, the text selection unit 132 selects the text immediately following the processed text (such as the immediately following sentence) and provides it to the network creation unit 780. The network creation unit 780 creates a new semantic network using this text as a starting point. The above-mentioned process is repeated until the end of the text stored in the text storage unit 130 is reached. According to this embodiment, the semantic network includes both supporting text and contradicting text of the text to be processed. In the recursive process, even if a contradicting text is obtained from a supporting text, or a supporting text is obtained from a contradicting text, it is not necessary to distinguish them into separate networks. Compared with the first embodiment, this has the advantage that the process for collecting supporting text and contradicting text is simpler.
[0128] The recursive program shown in FIG. 11 can also be used in the first embodiment with minor modifications.
[0129] 4. Variations The process of generating a semantic network in the second embodiment (the flowchart shown in FIG. 11) is similar to the breadth-first search in a tree search. That is, in the second embodiment, the root node is generated first, then the child nodes in the second layer are generated, and then the child nodes in the third layer are generated in the form of a collection of child nodes for each node in the second layer. In this order, the semantic network is generated.
[0130] However, the creation of a semantic network in this invention is not limited to the breadth-first order as already mentioned. A semantic network may be created in a depth-first order. Figure 12 shows a schematic flowchart of a program (recursive function) corresponding to Figure 11 for realizing such a modified example.
[0131] Referring to FIG. 12, the arguments of this recursive function include a constant N (>0) that specifies the depth of the layer when adding nodes in the depth direction, and a starting text for recursively searching for supporting text and contradictory text in the text being processed.
[0132] This program includes step 910, which branches the flow of control depending on whether the value of the argument N is 0. If the determination in step 910 is positive, then the program terminates execution and returns control to the calling program.
[0133] The program further includes a step 912 of generating one or more questions based on the text of the argument when the determination in step 910 is negative. In step 912, both support-specific and contradiction-specific questions are generated.
[0134] The program further includes a step 914 for performing, for each question generated in step 912, a step 916 of searching for one or more answers to that question, and a step 918 for each answer searched for in step 914, performing a step 920, described below.
[0135] Step 920 includes step 940 of determining whether the answer to be processed is a supporting text or a contradicting text for the text to be processed using a model similar to the text verification model 802 used in the second embodiment, and tagging the answer to be processed according to the determination result. In this embodiment, in step 940, the probability that the answer to be processed is a supporting text and the probability that the answer to be processed is a contradicting text are assigned as a score to the answer together with the tag.
[0136] Step 920 further includes, following step 940, adding the scored answers tagged in step 940 to a semantic network in step 942. In step 942, the answered question to be processed is first identified, the text from which the question originates is identified, and the new text is added as a child node of the node in the semantic network that corresponds to the identified text.
[0137] Step 920 further includes step 944 of recursively calling itself with arguments being a combination of the initially received argument N minus 1, value N-1, and the answer processed in step 920 .
[0138] For simplicity, the operation of the function whose control structure is shown in FIG. 12 will be explained assuming N=2. First, this function is called. The arguments at this time are N=2 and the text to be processed (for ease of explanation, this will be called "argument text"). Since the determination in step 910 is negative, step 912 is executed. As a result, one or more questions are generated based on the argument text.
[0139] Then, in step 914, an answer search process is performed based on each of the one or more questions generated in step 912. In this process, one or more answers are obtained for each of the one or more questions, resulting in a large number of answers being obtained.
[0140] Further, the process of step 918 is executed for each answer. Specifically, the text of the first answer (referred to as the "first text") is selected, the first text is tagged in step 940, and the first text is added as a new node to the semantic network in step 942. After that, the program is recursively called with arguments N-1 (=1) and the first text.
[0141] As a result, step 910 is executed for the new argument. Since the value of the argument is 1, the determination in step 910 is negative, and step 912 and subsequent steps are executed. As a result, multiple answers are obtained based on multiple questions obtained from the first text. Here, the first text of these answers is called "1-1 text" to indicate that it is the first of the answers obtained from the first text. The 1-1 text is also tagged in step 940, and added as a new node to the semantic network in step 942. Furthermore, in step 944, this function is recursively called with the combination of N-2 (=0) and the 1-1 text as arguments.
[0142] In this recursively called function, the judgment in step 910 is executed. Since the value of the argument is 0, the judgment in step 910 is affirmative, the execution of this function is terminated, and control is returned to the calling function, i.e., step 944 when the argument N=1. Since control is returned to step 944, the execution of step 944 is terminated, and the next repetition of step 920 is started. More specifically, for the text following the 1-1 text (this will be referred to as the "1-2 text"), a process similar to the process for the 1-1 text is executed. In the same manner, when the process is completed for all the texts obtained from the 1 text (from the 1-1 text to the 1-final text), the execution of step 918 when N=1 is terminated. As a result, the execution of this function when N=1 is terminated, and control is returned to step 944 when N=2. By the end of step 944, the process of step 918 when N=2 is executed for the second text following the 1st text.
[0143] In this way, by executing the recursive program whose control structure is shown in Figure 12, a semantic network is first created in a depth-first order up to the number of levels specified by the argument N, and by repeating this process, the semantic network is further expanded along the width direction.
[0144] After creating the semantic network in this way, editing of the text to be processed is carried out in the same manner as in the second embodiment.
[0145] In any of the above embodiments, a large amount of calculations is required. However, these calculations can be performed in parallel from a certain stage. Therefore, by using a GPU, it is possible to efficiently verify the input text.
[0146] The embodiments disclosed herein are merely illustrative, and the present invention is not limited to the above-described embodiments. The scope of the present invention is defined by the claims of the appended claims, taking into consideration the detailed description of the invention, and includes all modifications within the scope and meaning equivalent to the words described therein. [Explanation of symbols]
[0147] 50 Dialogue Systems 60 Text Input Device 62,446 Large-scale language models 64 Verification Device 66 Web-based Question Answering System 68 Internet 70 Output Device 102 Semantic network creation device 104 Contradiction Network Memory 106 Support Network Storage 108 Text editing device 130 Text storage section 132 Text Selection 134 Support Network Creation Department 136 Contradiction Network Creation Department 180 Recursive Support Text Collection Unit 182 Support Text Verification Section 184 Supporting Text Verification Model 186, 358 Semantic Network Additions 210 Supporting question generator 212, 372 Questions and Answers Department 214 Response Receiving Department 300 Support Network 350 Negative form generator 352 Recursive Inconsistent Text Collection Unit 354 Conflicting Text Verification Department 356 A model for verifying contradictory text 370 Contradictory question generator 432 Text Exchange Department 436 Text Exchange Model 438 Exchanged Text 444 Prompt Creation Department 450 Text Integration Department
Claims
1. a target portion extraction means for extracting a verification target portion from an input sentence; a text collection means for collecting supporting text that supports the content of the portion to be verified and contradictory text that contradicts the content of the portion to be verified from a collection of existing texts; and selective editing means for executing a process of editing the portion to be verified in a different manner depending on whether a predetermined relationship is established between the set of supporting texts and the set of contradictory texts collected by the text collection means.
2. 2. The verification device according to claim 1, wherein the text collection means includes an answer collection means for generating supporting specific questions for obtaining answers that support the content of the portion to be verified and contradiction specific questions for obtaining answers that contradict the portion to be verified based on an expression of the portion to be verified, and for each of the questions, recursively executing a process of obtaining answers from the set of existing texts, thereby collecting the supporting text and the contradiction text.
3. The process of editing the part to be verified includes: a text selection process for selecting either the contradictory text or the supporting text according to predetermined criteria; The verification device according to claim 1 , further comprising an editing process for editing at least a part of the portion to be verified based on the contradictory text or the supporting text selected in the text selection process.
4. 4. The verification device according to claim 3, wherein the process of editing the portion to be verified further includes a text addition process of adding, following the new text, text output by the large-scale language model by inputting the new text into the large-scale language model.
5. A step in which a computer extracts a portion to be verified from an input sentence; A computer collects supporting text that supports the content of the portion to be verified and contradicting text that contradicts the content of the portion to be verified from a collection of existing texts; and a step of editing the portion to be verified according to different methods by a computer depending on whether a predetermined relationship is established between the set of supporting texts and the set of contradictory texts collected in the collecting step.
6. Computer, a target portion extraction means for extracting a verification target portion from an input sentence; a text collection means for collecting supporting text that supports the content of the portion to be verified and contradictory text that contradicts the content of the portion to be verified from a collection of existing texts; a selective editing means for editing the part to be verified in different ways depending on whether a predetermined relationship is established between the set of supporting texts and the set of contradictory texts collected by the text collecting means.
Citation Information
Patent Citations
Information processing device, information processing method, and program
JP2023083926A
Cited By
A method, system, and program for deterministic detection and physical shutdown control of output anomalies by an independent processing unit outside the formal system, based on the non-self-verification nature of artificial intelligence models.
JP7919022B1