Inference device, inference method, program, and inference system

The integration of explicit knowledge with language models in the inference device addresses the reliability and multi-step inference challenges, ensuring accurate and reliable results for complex tasks.

WO2025158507A1PCT designated stage expired Publication Date: 2025-07-31NT T INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/001709
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-22
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

The reliability of inference results from generative AI models is not guaranteed due to the black box nature of the inference process, and multi-step inference is difficult to achieve with conventional methods.

Method used

An inference device that integrates explicit knowledge, such as symbolic logic, with a language model to perform multi-step inferences by sequentially inferring relationships and rules, using a first and second inference unit to extract facts and rules, and a knowledge management unit to update and store this information.

Benefits of technology

Ensures reliable inference results for multi-step tasks by integrating explicit knowledge with language models, enabling efficient acquisition of advanced knowledge for improved inference accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024001709_31072025_PF_FP_ABST
    Figure JP2024001709_31072025_PF_FP_ABST
Patent Text Reader

Abstract

An inference device according to an aspect of the present disclosure includes an inference unit that, on the basis of a first character string, a second character string related to the first character string, knowledge having undergone inference, and a natural language processing model, sequentially infers relationship information representing relationships between specific characters or character strings in the first character string, and new knowledge representing a prescribed relationship between the relationships represented by the relationship information until relationship information satisfying a requirement represented by the second character string is inferred.
Need to check novelty before this filing date? Find Prior Art

Description

Inference device, inference method, program, and inference system

[0001] The present disclosure relates to an inference device, an inference method, a program, and an inference system.

[0002] The development of large language models (LLMs) has led to the advancement of generative artificial intelligence (AI). However, the inference process by AI is a black box, and there is a problem in that the reliability of the inference results cannot be guaranteed. To address this problem, research is being conducted into methods that can ensure the reliability of inference results by integrating knowledge with LLMs that has a clear inference process, such as knowledge expressed in symbolic logic (hereinafter also referred to as "explicit knowledge") (see, for example, Non-Patent Document 1).

[0003] Zhang, Hanlin, et al. "Improved logical reasoning of language models via differentiable symbolic programming." arXiv preprint arXiv:2305.03742 (2023).

[0004] However, conventional methods cannot acquire advanced knowledge, making multi-stage inference (multi-hop inference) difficult in some cases.

[0005] The present disclosure has been made in consideration of the above points, and aims to realize multi-stage inference that utilizes explicit knowledge.

[0006] An inference device according to one aspect of the present disclosure includes an inference unit that sequentially infers relationship information representing a relationship between specific characters or character strings in the first character string and new knowledge representing a predetermined relationship between the relationships represented by the relationship information based on a first character string, a second character string related to the first character string, inferred knowledge, and a natural language processing model, until relationship information that satisfies a requirement represented by the second character string is inferred.

[0007] Multi-stage reasoning utilizing explicit knowledge is realized.

[0008] FIG. 1 is a diagram illustrating an example of a hardware configuration of an inference device according to the present embodiment. FIG. 2 is a diagram illustrating an example of a functional configuration of an inference device according to the present embodiment. FIG. 3 is a flowchart (1 / 2) illustrating an example of an inference process according to the present embodiment. FIG. 4 is a flowchart (2 / 2) illustrating an example of an inference process according to the present embodiment. FIG. 5 is a diagram illustrating a context and a query of Example 1. FIG. 6 is a diagram illustrating facts inferred in round 0 of Example 1 and their certainty. FIG. 7 is a diagram illustrating rules inferred in round 0 of Example 1 and their certainty. FIG. 8 is a diagram illustrating a first prompt and facts in round 1 of Example 1. FIG. 9 is a diagram illustrating a first prompt and facts in round 2 of Example 1. FIG. 10 is a diagram illustrating a first prompt and facts in round 3 of Example 1. FIG. 11 is a diagram illustrating a context and a query of Example 2. FIG. 12 is a diagram illustrating facts inferred in round 0 of Example 2 and their certainty. FIG. 13 is a diagram illustrating rules inferred in round 0 of Example 2 and their certainty. FIG. 14 is a diagram illustrating a modified example of the functional configuration of an inference device according to the present embodiment.

[0009] An embodiment of the present invention will be described in detail below with reference to the drawings. In the following embodiment, an inference device 10 will be described that can acquire advanced knowledge that realizes multi-stage inference by repeating inference that integrates a language model and explicit knowledge. As a result, the inference device 10 according to this embodiment can obtain reliable inference results for problems that require multi-stage inference. Note that, hereinafter, each repetition of inference that integrates a language model and explicit knowledge by the inference device 10 according to this embodiment will be referred to as a "round," and a round will start from round r=0, where r is a variable representing the round. Furthermore, repeating inference may also be expressed as, for example, "performing inference sequentially" or "executing inference sequentially."

[0010] A language model refers to various models (machine learning models, statistical models, mathematical models, etc.) for inferring solutions to problems (hereinafter also referred to as "tasks") related to natural language processing (NLP). A language model may also be called, for example, a "natural language model," a "natural language processing model," a "natural language component," or a "natural language processing component." In the following, as an example, it is assumed that the language model is a large-scale language model (LLM) that is mainly used to realize generative AI, etc. However, the following embodiment can also be applied to language models other than large-scale language models.

[0011] Explicit knowledge refers to knowledge with a clear inference process, or a knowledge model that models such knowledge. Furthermore, knowledge or knowledge model refers to information (e.g., inference rules, etc.) or models (e.g., graphs, etc.) for performing some kind of inference on a given input and obtaining an output. Specific examples of explicit knowledge include knowledge expressed in symbolic logic, knowledge expressed in a knowledge graph, and knowledge expressed in a rule-based manner. Note that inference using knowledge expressed in symbolic logic is referred to as "symbolic inference." Similarly, inference using knowledge expressed in a knowledge graph is referred to as "knowledge graph inference" (or "knowledge graph inference"), and inference using knowledge expressed in a rule-based manner is referred to as "rule-based inference." Inference using such explicit knowledge involves a white-box inference process, so the reliability of the inference results is assured. Hereinafter, as an example, it is assumed that explicit knowledge is primarily expressed in symbolic logic, and inference using explicit knowledge is realized using symbolic inference. However, the following embodiments are also applicable to explicit knowledge other than knowledge expressed in symbolic logic.

[0012] Multi-stage inference is a term primarily used in tasks related to natural language processing, such as machine reading comprehension and question answering, and refers to a method of obtaining a solution to a task through multiple inferences. Multi-stage inference is also called, for example, "multi-hop inference" or "multiple-step inference." In contrast, a method of obtaining a solution to a task through a single inference is also called "single-step inference" or "single-hop inference."

[0013] Here, the inference device 10 according to this embodiment is primarily intended to infer solutions to tasks such as machine reading comprehension and question answering, and is given a context belonging to a certain domain and a query related to that context. However, hereinafter, we will primarily consider tasks for which it is difficult to obtain a solution that satisfies the query using single-step inference. A context refers to information such as a character string containing information necessary to obtain a solution that satisfies the query. A context is also referred to as a "context," a "sentence," a "document," an "article," etc. A query refers to information such as a character string that expresses a request, such as a "question" or "inquiry." A domain is a term that refers to an area or field. Domains can be defined in various ways with varying granularity, and specific examples include "family relationships," "medical care," "transportation," "IT," "language," "food and drink," and "entertainment."

[0014] In the following, it is assumed that both the context and the query are mainly character string information, but information other than character strings may be included in at least one of the context and the query. For example, visual information such as an image, a photograph, a diagram, a table, a graph, a slide, a list, etc. may be included in at least one of the context and the query.

[0015] <Preparation for Symbolic Logic> A brief description will be given of the grammar of symbolic logic, which is an example of a method for expressing explicit knowledge. As an example, the grammar of symbolic logic used in existing symbolic reasoning software called DataProlog (which may also be called a "symbolic reasoning program" or a "symbolic reasoning tool") will be described below. Note that such symbolic reasoning software is also called SR (Symbolic Reasoner), etc.

[0016] Term t: V|c Here, V is a variable and c is a constant. That is, term t is the variable V or the constant c. Hereinafter, the constant will be called an "entity." Note that an entity may also be called, for example, a "subject" or an "entity."

[0017] ・Atomic logical formula (Atom) α: a(t 1 , ..., t n ) where a is a symbol representing a predicate, t 1 , ..., t n is a term. Note that a predicate may also be called, for example, a "predicate."

[0018] ・Fact g: a(c 1 , ..., c n ) where a is a symbol representing a predicate, c 1 , ..., c n is a constant.

[0019] In this case, the rule is expressed as follows:

[0020] R: α:-α 1 , ..., α n Here, α, α 1 , ..., α n are all atomic formulas, α is the head part, α 1 , ..., α n is called the body part. The above rule R is 1 , ..., α n This means that if all the atomic formulas in are true, then the atomic formula α is also true.

[0021] Fact g: a(c1 , ..., c n ) is "entity c 1 , ..., c n On the other hand, the atomic formula α: a(t 1 , ..., t n ) is "term t 1 , ..., t n The rule R represents a logical formula that indicates whether there is a certain relationship between α 1 , ..., α n is true, then α is true, otherwise α is false. 1 , ..., α n This represents a clear inference process: "If all of the above are true, then α is also true." Rule R can be used as explicit knowledge. Therefore, in the following, rule R will be used as explicit knowledge as an example. Note that the relationship may also be called a "relation."

[0022] As will be described later, the inference device 10 according to this embodiment can acquire a rule R that represents the inference process of multi-stage inference as advanced knowledge by repeating rounds. This enables the inference device 10 according to this embodiment to obtain reliable inference results for problems that require multi-stage inference.

[0023] <Example of Hardware Configuration of Inference Device 10> An example of the hardware configuration of the inference device 10 according to this embodiment will be described with reference to Fig. 1. Fig. 1 is a diagram showing an example of the hardware configuration of the inference device 10 according to this embodiment.

[0024] 1, an inference device 10 according to this embodiment includes an input device 101, a display device 102, an external I / F 103, a communication I / F 104, a random access memory (RAM) 105, a read only memory (ROM) 106, an auxiliary storage device 107, and a processor 108. Each of these pieces of hardware is connected to each other via a bus 109 so as to be able to communicate with each other.

[0025] The input device 101 is, for example, a keyboard, a mouse, a touch panel, a physical button, etc. The display device 102 is, for example, a display, a display panel, etc. Note that the inference device 10 does not have to have at least one of the input device 101 and the display device 102, for example.

[0026] The external I / F 103 is an interface with an external device such as a recording medium 103a. Examples of the recording medium 103a include a CD (Compact Disc), a DVD (Digital Versatile Disk), an SD memory card (Secure Digital memory card), and a USB (Universal Serial Bus) memory card.

[0027] The communication I / F 104 is an interface for connecting to a communication network. The RAM 105 is a volatile semiconductor memory (storage device) that temporarily stores programs and data. The ROM 106 is a non-volatile semiconductor memory (storage device) that can store programs and data even when the power is turned off. The auxiliary storage device 107 is a non-volatile storage device such as a hard disk drive (HDD), a solid state drive (SSD), or a flash memory. The processor 108 is a variety of arithmetic devices such as a central processing unit (CPU) or a graphic processing unit (GPU).

[0028] 1 is an example, and the hardware configuration of the inference device 10 is not limited to this. For example, the inference device 10 may have multiple auxiliary storage devices 107 or multiple processors 108, may not have some of the hardware shown in the figure, or may have various types of hardware other than the hardware shown in the figure.

[0029] <Example of Functional Configuration of Inference Device 10> An example of the functional configuration of the inference device 10 according to this embodiment will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of the functional configuration of the inference device 10 according to this embodiment.

[0030] As shown in FIG. 2 , the inference device 10 according to this embodiment includes an input unit 201, a prompt creation unit 202, a first inference unit 203, a second inference unit 204, a knowledge management unit 205, and a termination determination unit 206. Each of these units is realized, for example, by a process in which one or more programs installed in the inference device 10 are executed by the processor 108 or the like. The inference device 10 according to this embodiment also includes a knowledge storage unit 207. The knowledge storage unit 207 is realized, for example, by a storage area of ​​the auxiliary storage device 107 or the like. However, the knowledge storage unit 207 may also be realized, for example, by a storage area of ​​a storage device (e.g., a storage device included in a database server or the like) connected to the inference device 10 so as to be able to communicate with the inference device 10.

[0031] The input unit 201 inputs a context and a query given to the inference device 10. The context and the query may be given to the inference device 10 at the same time, or one of the context and the query may be given first and then the other.

[0032] The prompt creation unit 202 creates information called a prompt to be input to the first inference unit 203. At this time, when round r≧1, the prompt creation unit 202 uses knowledge (rule R) stored in the knowledge storage unit 207 to create a prompt to be input to the first inference unit 203 in that round r. Hereinafter, the prompt created by the prompt creation unit 202 will be referred to as a "first prompt." The prompt creation unit 202 is implemented by software (which may also be referred to as a program, tool, etc.) called, for example, an Automatic Prompt Engineer (APE). A prompt is information containing instructions to be given to a language model (particularly, a large-scale language model). A prompt may also be referred to as, for example, an "instruction statement," an "instruction," or an "instruction."

[0033] The first inference unit 203 receives as input the context input by the input unit 201 and the first prompt created by the prompt creation unit 202, and infers and outputs a fact g that represents a relationship between entities in the context. At this time, the first inference unit 203 may output a certainty that represents the likelihood of the fact g. As an example, it is assumed below that the first inference unit 203 outputs the fact g and its certainty. The first inference unit 203 is realized by a language model (particularly, a general-purpose large-scale language model that is not limited to a specific domain). Note that what constitutes an entity may vary depending on the task, but typical examples of entities include named entities (e.g., person's name, organization name, place name, era, date, time, quantity, amount, etc.).

[0034] The second inference unit 204 receives as input a preset prompt (hereinafter referred to as the "second prompt") and a fact g inferred by the first inference unit 203, and infers and outputs a rule R representing a transitive relationship between the relationships represented by the fact g as knowledge. At this time, the second inference unit 204 may output a confidence level representing the likelihood of the rule R. Hereinafter, as an example, it is assumed that the second inference unit 204 outputs the rule R (knowledge) and its confidence level. The second inference unit 203 is realized by a language model (hereinafter also referred to as a "domain-specific language model") specialized for inference of the domain to which the context input by the input unit 201 belongs.

[0035] Here, a transitive relation is a relation in which the transitive law holds. That is, a rule R that expresses a transitive relation between relations expressed by facts g is 1 (c 1 , c 2 ) and g 2 (c 2 , c 3 ) (where c 1 , c 2 , c 3 is a constant) there exists a fact g 3 (c 1 , c 3 ) exists. 1Let α be an atomic formula that is true if α exists and false otherwise. 1 (t 1 , t 2 ), fact g 2 Let α be an atomic formula that is true if α exists and false otherwise. 2 (t 2 , t 3 ) then, fact g 1 The relationship and fact g 2 A rule R expressing a transitive relation between the relation expressed by 3 :-α 1 , α 2 where t 1 , t 2 , t 3 is a variable, α 3 (t 1 , t 3 ) is fact g 3 is an atomic formula that is true if g exists and false otherwise. 1 (c 1 , c 2 ) exists and fact g 2 (c 2 , c 3 ) exists, then fact g 3 (c 1 , c 3 In other words, we obtain a rule R that allows us to infer that there exists an entity c 1 and c 2 and entity c 2 and c 3 If a second relationship holds between 1 and c 3 A rule R is obtained that allows the inference that a certain third relationship holds between

[0036] It should be noted that with regard to the certainty levels output by the first inference unit 203 and the second inference unit 204, the term "certainty level" is merely an example, and may be called, for example, "reliability," "trustworthiness," "score," or "probability." The certainty level may be any information capable of expressing likelihood, and may be, for example, a continuous value, a discrete value, a value representing a label, or a value expressed as a percentage (%). Furthermore, when the certainty level is expressed as a numerical value, the range that the value can take can also be arbitrarily designed.

[0037] The knowledge management unit 205 manages the knowledge storage unit 207. That is, the knowledge management unit 205 stores the rule R (knowledge) inferred by the second inference unit 204 and its certainty in the knowledge storage unit 207, and updates the knowledge and its certainty stored in the knowledge storage unit 207 with the rule R (knowledge) inferred by the second inference unit 204. For example, when the rule R: α 3 :-α 1 , α 2 is stored in the knowledge storage unit 207, the second inference unit 204 generates a rule R′: α 5 :-α 3 , α 4 is inferred and output, the knowledge management unit 205 converts the rule R into R: α 5 :α 1 , α 2 , α 4 As a result, for example, the rule R: α 3 :-α 1 , α 2 and rule R': α 5 :-α 3 , α 4 and are both obtained by single-step inference, a rule R:α that allows two-step inference 5 :α 1 , α 2 , α 4 Therefore, by repeating round r, it is possible to obtain rule R that allows n (n≧2) step inference (that is, rule R that represents advanced knowledge).

[0038] Furthermore, the knowledge management unit 205 deletes a rule R (knowledge) stored in the knowledge storage unit 207 according to the degree of certainty of the rule R.

[0039] The termination determination unit 206 determines whether to terminate the repetition of the round. The termination determination unit 206 determines to terminate the repetition of the round if a predetermined termination condition is satisfied, and determines not to terminate the repetition of the round if the termination condition is not satisfied. Examples of the predetermined termination condition include a condition in which a fact g that satisfies a query is inferred by the first inference unit 203 from a first prompt that includes certain knowledge (rule R) stored in the knowledge storage unit 207 and a query (hereinafter referred to as a "first termination condition"), and a condition in which a fact g that satisfies the query is inferred from knowledge stored in the knowledge storage unit 207 (hereinafter referred to as a "second termination condition").

[0040] The knowledge storage unit 207 stores the knowledge (rule R) and its confidence level inferred by the second inference unit 203. That is, the knowledge storage unit 207 stores a set of information represented by a pair of the rule R and its confidence level.

[0041] 2 is an example, and is not limited to this functional configuration of the inference device 10. Various other functional configuration examples of the inference device 10 are possible, and each of these functional configurations will be described in the modified examples below.

[0042] <Inference Processing> The inference processing according to this embodiment will be described with reference to Figures 3A and 3B. Figures 3A and 3B are flowcharts showing an example of the inference processing according to this embodiment.

[0043] First, the input unit 201 inputs a given context and query (step S101).

[0044] Next, the prompt generator 202 generates a first prompt for round 0 (step S102). At this time, the prompt generator 202 generates the first prompt including instructions indicating that entities in the context input in step S101 and relationships between those entities are to be extracted from the context. For example, the prompt generator 202 generates a sentence such as "List all the entities in the context, and give the possible relationships between them." as the first prompt for round 0.

[0045] The first prompt in round 0 may be stored in advance in a prompt library that stores, for example, example prompt sentences and templates. In this case, the prompt creation unit 202 simply searches for and acquires the first prompt in round 0 from the prompt library. Such a prompt library is realized, for example, by a storage area such as the auxiliary storage device 107. The prompt library may also be called, for example, a "basic prompt library" or a "template library."

[0046] Next, the first inference unit 203 infers one or more facts g and their certainty levels using the context input in step S101 and the first prompt created in step S102 (step S103). That is, the first inference unit 203 extracts one or more entities and relationships between those entities from the context using a language model in accordance with the first prompt. As a result, the facts g representing the relationships between the entities obtained by the single-step inference and their certainty levels are output as the facts g and their certainty levels in round 0. Note that the facts g in round 0 are facts inferred directly from the context input in step S101, and therefore may be called, for example, "flat facts" or "direct facts."

[0047] Next, the second inference unit 204 infers one or more rules R and their confidence levels using the second prompt and the one or more facts g inferred in step S103 as input (step S104). That is, the second inference unit 204 infers a rule R that expresses a transitive relationship between the relationships expressed by the facts g using a domain-specific language model in accordance with the second prompt. At this time, the second inference unit 204 may infer one or more rules R and their confidence levels using a second prompt such as, for example, "Find out what are the relationships between these relationships. List them in a table." Note that "these relationships" in the second prompt "Find out what are the relationships between these relationships. List them in a table." refers to the relationships expressed by the facts g inferred in step S103. As a result, the rule R that expresses the transitive relationship between the relationships expressed by the facts g in round 0 and its confidence level are output as the rule R and its confidence level in round 0.

[0048] The second prompt may be stored in advance in a prompt library, for example. The second inference unit 204 may infer one or more rules R and their certainties without receiving the certainty of the fact g as input.

[0049] Next, the knowledge management unit 205 stores one or more rules R inferred in step S104 and their confidence levels in the knowledge storage unit 207 (step S105). As a result, each rule R obtained by single-step inference from an obvious fact g in the context and its confidence level are stored in the knowledge storage unit 207.

[0050] The above steps S102 to S105 are processes executed in round r=0. Meanwhile, the following steps S106 to S111 are processes executed in round r (where r≧1). Note that the variable r representing the round is updated to r←r+1, for example, when step S106 returns NO (or may be immediately before step S106 is executed). Below, a description is given of the case where steps S106 to S111 are executed for a certain round r (where r≧1).

[0051] The termination determination unit 206 determines whether a first termination condition is satisfied (step S106). That is, the termination determination unit 206 determines whether the first inference unit 203 infers a fact g that satisfies the query from a first prompt that includes a query and certain knowledge (rule R) stored in the knowledge storage unit 207. In particular, for example, when a query includes an entity, the termination determination unit 206 may determine whether the first inference unit 203 infers a fact g that satisfies the query from a first prompt that includes an instruction to extract a fact g that satisfies the query by referring to a rule R related to at least one entity included in the query. Note that a rule R related to an entity refers to a rule R that includes an atomic formula α that takes the entity as an argument. As a result, if the first inference unit 203 infers a fact g that satisfies the query, the repetition of the round is terminated; otherwise, the repetition of the round is continued.

[0052] If it is determined in step S106 above that the first termination condition is satisfied (YES in step S106), the inference device 10 terminates the inference process. In this case, this is because the first inference unit 203 infers a fact g that satisfies the query in a single step from the first prompt that includes a certain rule R stored in the knowledge storage unit 207 and the given query. In other words, this is because a reliable inference result has been obtained in a single step for a task that requires multi-stage inference.

[0053] On the other hand, if it is determined in step S106 that the first termination condition is not satisfied (NO in step S106), the prompt generator 202 generates a first prompt for round r (step S107). At this time, the prompt generator 202 generates the first prompt for round r using certain knowledge (rule R) stored in the knowledge storage unit 207. In particular, for example, if a query includes an entity, the prompt generator 202 uses rule R related to at least one entity included in the query to generate a first prompt including that rule R as the first prompt for round r. Note that the first prompt generator 202 may also generate a first prompt for round r that further includes at least one entity included in the query. For example, if a certain entity included in a query is c and a certain rule related to entity c is R, the prompt creation unit 202 can create a first prompt in round r that includes instructions indicating that the relationship between entity c and other entities in the context should be extracted by referring to rule R.

[0054] Next, the first inference unit 203 infers one or more facts g and their certainty levels using the context input in step S101 and the first prompt created in step S107 (step S108). That is, the first inference unit 203 extracts one or more entities and relationships between those entities from the context using a language model in accordance with the first prompt. As a result, the facts g and their certainty levels representing the relationships between the entities obtained by the single-step inference are output as the facts g and their certainty levels in round r.

[0055] In step S108, the first inference unit 203 may infer one or more facts g and their certainties using the facts g and their certainties inferred up to round r-1 as input. That is, the first inference unit 203 may extract one or more entities and relationships between those entities from the context and the inferred facts g using a language model in accordance with the first prompt.

[0056] Next, the second inference unit 204 infers one or more rules R and their confidence levels using the second prompt and the one or more facts g inferred in step S108 as input (step S109). That is, the second inference unit 204 infers a rule R that expresses a transitive relationship between the relationships expressed by the facts g using a domain-specific language model in accordance with the second prompt. For example, similar to step S104, the second inference unit 204 may infer one or more rules R and their confidence levels using a second prompt such as "Find out what are the relationships between these relationships. List them in table." Note that "these relationships" in the second prompt "Find out what are the relationships between these relationships. List them in table." refers to the relationships expressed by the facts g inferred in step S108. As a result, the rule R that expresses the transitive relationship between the relationships expressed by the facts g in round r and its confidence level are output as the rule R and its confidence level for round r.

[0057] Next, the knowledge management unit 205 uses one or more rules R and their certainty factors inferred in step S109 to update or delete the rules R and their certainty factors stored in the knowledge storage unit 207 (step S110). An example of the update or deletion method will be described below.

[0058] - Example of update / deletion method Each rule stored in the knowledge storage unit 207 is updated / deleted by R i , and the confidence level is s iIn addition, the rules inferred and output in step S109 are represented as R j ', and the confidence level is s j ', where i is an index for identifying each rule stored in the knowledge storage unit 207, and j is an index for identifying each rule inferred and output in step S109. At this time, the knowledge management unit 205 i and each rule R j Regarding ', the rule R i and the relevant rule R j Depending on the logical relationship between the rule R and the rule R, the rule R is determined as follows: i and its confidence level s i Update or delete (including cases where not updating or deleting).

[0059] (1) Rule R i and Rule R j When a new rule (new knowledge) can be derived from rule R i : α 3 :-α 1 , α 2 and rule R j ': α 5 :-α 3 , α 4 If these rules R i and R j ' to new rule R: α 5 :α 1 , α 2 , α 4 Therefore, the knowledge management unit 205 can derive the rule R i , R i ←R is overwritten and updated. This results in a rule R that represents more advanced knowledge. i is obtained.

[0060] In addition, the knowledge management unit 205 updates the confidence level s i So that the confidence s i For example, the knowledge management unit 205 updates s using a predetermined positive constant α. i ←s i +α etc. to increase confidence level iHowever, the present invention is not limited to this. For example, the knowledge management unit 205 updates s i ← (s i +s j ') / 2 gives confidence s i You can update s i ← (s i +s j ') / 2+α gives confidence level s i You can update s i ←s i +α×s j ' gives the confidence s i or by other methods, the confidence s i may be updated.

[0061] (2) Rule R i and Rule R j ' and contradict each other. For example, α 3 and α 4 Let R be an atomic formula where one is true and the other is false. i : α 3 :-α 1 , α 2 and rule R j ': α 4 :-α 1 , α 2 If these rules R i and R j ' are mutually contradictory rules. For this reason, the knowledge management unit 205, for example, i The confidence s i For example, the knowledge management unit 205 updates s using a predetermined positive constant α. i ←s i -α etc. to determine the confidence level s i However, the present invention is not limited to this. For example, the knowledge management unit 205 updates s i ←s i -α×s j ' gives the confidence s i or by other methods, the confidence s i may be updated.

[0062] Furthermore, the knowledge management unit 205 may, for example, update the confidence level s i becomes less than a predetermined threshold, the rule R i and its confidence level s i is deleted from the knowledge storage unit 207. As a result, the rule R j ' and contradictory results, confidence level s i Rule R has become lower i is deleted from the knowledge storage unit 207.

[0063] (3) Cases other than (1) and (2) above For example, Rule R i : α 3 :-α 1 , α 2 and rule R j ': α 6 :-α 4 , α 5 If these rules R i and R j Therefore, the knowledge management unit 205 determines whether the rule R j ' and its confidence level s j ' to the rule R i and its confidence level s i will not be updated and will remain as is.

[0064] The above (1) is rule R i and Rule R j ', while the above (2) is consistent with rule R i and Rule R j This applies when there is no consistency between '.

[0065] In addition, a certain rule R j ', all rules R stored in the knowledge storage unit 207 i If the logical relationship with the rule R is either (2) or (3), the knowledge management unit 205 j ' and its confidence level s j ' is stored in the knowledge storage unit 207.

[0066] Next, the termination determination unit 206 determines whether a second termination condition is satisfied (step S111). That is, the termination determination unit 206 determines whether a fact g that satisfies the query can be inferred from knowledge (rule R) stored in the knowledge storage unit 207 using symbolic inference software (SR) or the like. In particular, for example, when some entity is included in the query, the termination determination unit 206 may infer a fact g using the symbolic inference software (SR), a rule R related to at least one entity included in the query, and the at least one entity, and determine whether the fact g satisfies the query. As a result, if a fact g that satisfies the query is inferred from a certain rule R stored in the knowledge storage unit 207, the repetition of the round is terminated; otherwise, the repetition of the round is continued.

[0067] If it is determined in step S111 above that the second termination condition is met (YES in step S111), the inference device 10 terminates the inference process. In this case, the fact g that satisfies the query has been inferred in a single step by the symbolic inference software (SR) from a certain rule R stored in the knowledge storage unit 207. In other words, for a task requiring multi-stage inference, a reliable inference result has been obtained in a single step.

[0068] On the other hand, if it is not determined in step S111 above that the second termination condition is satisfied (NO in step S111), the inference device 10 returns to step S106 above, thereby starting the repetition of the next round r.

[0069] <Examples> Below, examples using existing datasets related to machine reading comprehension, question answering, etc. will be described as examples of the inference device 10 according to this embodiment. Note that in Example 1, a dataset called CLUTRR (Reference 1) was used, and in Example 2, a dataset called MetaQA (Reference 2) was used.

[0070] Example 1 In Example 1, a case will be described in which kinship relationships between people (entities) appearing in a context are inferred as a task requiring multi-stage inference.

[0071] The inference device 10 in Example 1 is provided with the context and query shown in Figure 4. This is a task to infer the kinship relationship between two people, Tracy and Louis. Note that, because the kinship relationship between Tracy and Louis is not directly stated in the context shown in Figure 4, this task is difficult to infer in a single step (i.e., a task requiring multi-stage inference).

[0072] At this time, in step S103 of FIG. 3A, the facts and their certainties shown in FIG. 5 are inferred as fact g and its certainty in round 0. In the example shown in FIG. 5, each row represents one fact and its certainty, and six facts and their certainty are inferred in Example 1. Here, each fact and its certainty in FIG. 5 represents the kinship (relationship) and its certainty between two people (person 1, person 2) described in the context shown in FIG. 4. For example, the fact and its certainty in the first row of the example shown in FIG. 5 represent the fact "Marie and Tracy are sisters" and its certainty of "90%". Similarly, for example, the fact and its certainty in the second row represent the fact "Marie and Laura are sister-in-law" and its certainty of "70%". The same applies to the facts and their certainty in the other rows.

[0073] Next, in step S104 of FIG. 3A , the rule and its certainty shown in FIG. 6 are inferred as the rule R and its certainty in round 0. Note that in the example shown in FIG. 6 , each row represents one rule and its certainty, and in Example 1, multiple rules and their certainty are inferred. Here, each rule and its certainty in FIG. 6 represents a transitive relationship between the relationships (kinship relationships) represented by each fact shown in FIG. 5 (i.e., if relationship 1 holds between person 1 and person 2 and relationship 2 holds between person 2 and person 3, then relationship 3 holds between person 1 and person 3) and its certainty. For example, when person 1 is "A," person 2 is "B," and person 3 is "C," the rule and its certainty in the first row of the example shown in FIG. 6 represent the rule "If B is A's brother and C is B's sister, then C is A's sibling," and its certainty is "100%." 6, the rule in the second line and its certainty indicate that "if B is A's Brother and C is B's Uncle, then C is A's Uncle" and its certainty is 90%. The same applies to the rules in the other lines and their certainty.

[0074] Next, in round 1, the first prompt shown in Fig. 7A is created in step S107 of Fig. 3B, and then the fact shown in Fig. 7A is inferred in step S108 of Fig. 3B. Here, the first prompt shown in Fig. 7A includes "Louis," which is an entity (person's name) included in the query shown in Fig. 4, and the statement "if B is A's daughter, and C is B's uncle, then A and C could be brothers," which represents the rule on line 9 of the example shown in Fig. 6. As a result, in step S108 of Fig. 3B in round 1, using this rule, since "Laura is Antonio's daughter" and "Louis is Laura's uncles," the fact "Louis is Antonio's brother" is inferred.

[0075] Next, in round 2, the first prompt shown in Fig. 7B is created in step S107 of Fig. 3B, and then the fact shown in Fig. 7B is inferred in step S108 of Fig. 3B. Here, the first prompt shown in Fig. 7B includes "Louis," which is an entity (person's name) included in the query shown in Fig. 4, and the statement "if B is A's brother, and C is B's brother, then A could be C's brother.", which represents the rule in the fourth line of the example shown in Fig. 6. As a result, in step S108 of Fig. 3B in round 2, using this rule, since "Antonio is Marie's brother" and "Louis is Antonio's brother," the fact "Louis is Marie's sibling" is inferred. Generally, "Marie" is a female name and "Louis" is a male name, so it is inferred that the relationship between "Marie" and "Louis" is "Siblings."

[0076] Then, in round 3, the first prompt shown in Figure 7C is created in step S107 of Figure 3B, and then the fact shown in Figure 7C is inferred in step S108 of Figure 3B. Here, the first prompt shown in Figure 7C includes "Tracy" and "Louis," which are entities (personal names) included in the query shown in Figure 4, and the sentence "if B is A's sister, and C is B's brother, then A and C are siblings.", which represents the six-line rule of the example shown in Figure 6. As a result, in step S108 of Figure 3B in round 3, using this rule, since "Marie is Tracy's sister" and "Louis is Marie's brother," the fact "Louis is a sibling of Tracy" is inferred. Since this fact satisfies the requirement represented by the query shown in Figure 4, a solution to this task has been obtained.

[0077] Second Embodiment In a second embodiment, a case will be described in which a task requiring multi-stage inference is performed to infer movies that share actors who appear in a certain movie.

[0078] The context and query shown in Fig. 8 are given to the reasoning device 10 in Example 2. This is a task to infer movies that share actors with the movie "Dil Chahta Hai" and were released in 2005. Note that the context shown in Fig. 8 does not directly mention movies that share actors with the movie "Dil Chahta Hai," so this task is difficult to infer in a single step (i.e., a task that requires multi-stage inference).

[0079] At this time, in step S103 of FIG. 3A, the facts and their certainties shown in FIG. 9 are inferred as fact g and its certainty in round 0. In the example shown in FIG. 9, each row represents one fact and its certainty, and in Example 2, 17 facts and their certainty are inferred. Here, each fact and its certainty in FIG. 9 is expressed in the format of "Entity 1 | Relationship | Entity 2, Certainty." For example, the fact and its certainty in the first row of the example shown in FIG. 9 represent the fact "The director of Dil Chahta Hai is Farhan Ankhtar" and its certainty of "100%." ​​Similarly, the fact and its certainty in the second row of the example shown in FIG. 9 represent the fact "The original author of Dil Chahta Hai is Farhan Ankhtar" and its certainty of "90%." The same applies to other facts and their certainties. However, in Example 2, each relationship is assumed to be commutative, that is, when a fact representing a relationship is g and two entities are A and B, g(A, B) = g(B, A) (therefore, atomic formula α, which is true if fact g exists and false if it does not, is also commutative).

[0080] Next, in step S104 of FIG. 3A , the rule and its confidence level shown in FIG. 10 are inferred as the rule R and its confidence level in round 0. In the example shown in FIG. 10 , each row represents one rule and its confidence level, and eight rules and their confidence levels are inferred in Example 2. Here, each rule and its confidence level in FIG. 10 represents a transitive relationship between the relationships represented by the facts shown in FIG. 9 and its confidence level. For example, assuming that the entity representing an actor's name is "A" and the entities representing movie names are "B" and "C," the rule and its confidence level in the first row of the example shown in FIG. 10 represent the rule "If A is an actor who appears in B and A is an actor who appears in C, then B and C are movies that share the same actor," with a confidence level of "90%." Similarly, if the entity representing the actor name is "A," the entity representing the movie name is "B," and the entity representing the year is "C," the rule in the second row of the example shown in Figure 10 and its certainty factor represent the rule "If A is an actor appearing in B, and the year B was released is C, then A's debut year is C," and its certainty factor is "100%." ​​The same applies to the rules and their certainty factors in the other rows.

[0081] In round 1, the first prompt shown in FIG. 11 is created in step S107 of FIG. 3B, and then the fact shown in FIG. 11 is inferred in step S108 of FIG. 3B. Here, the first prompt shown in FIG. 11 includes "Dil Chahta Hai," which is an entity (movie title) included in the query shown in FIG. 8, "2005," which is an entity (release year) included in the query shown in FIG. 8, and the sentence "Depending on the relationships above," which represents rule R inferred in round 0. As a result, in step S108 of FIG. 3B in round 1, using this rule, the fact that "Slaam Namaste" is a movie that shares actors with "Dil Chahta Hai" and was released in 2005 is inferred. Since this fact satisfies the requirement represented by the query shown in FIG. 8, a solution to this task has been obtained.

[0082] <Modifications of this embodiment> Modifications of the above embodiment will now be described.

[0083] Modification 1 A modification of the functional configuration of the inference device 10 according to this embodiment will be described with reference to Fig. 12. Fig. 12 is a diagram showing a modification of the functional configuration of the inference device 10 according to this embodiment.

[0084] 12, in the inference device 10 of the modified example, the first inference unit 203 and the second inference unit 204 are configured as a single inference unit 208. Furthermore, this inference unit 208 is realized by a language model. In this way, the first inference unit 203 and the second inference unit 204 may be configured by the inference unit 208 realized by a single language model.

[0085] Modification 2 In the above embodiment, the second inference unit 204 is realized by a domain-specific language model, but this is not limited to this. For example, the second inference unit 204 may be realized by a language model (particularly, a general-purpose language model).

[0086] Variation 3: While the above embodiment assumes that the inference device 10 is realized by a single physical device, this is not limited thereto. The inference device 10 may be realized by multiple devices connected to each other and capable of communicating with each other (including cases where communication is possible via a communication network such as the Internet). In this case, the first inference unit 203 and the second inference unit 204 may be realized by different devices that can communicate with each other via a communication network such as the Internet. In addition, in this case, the functions realized by the first inference unit 203 and the functions realized by the second inference unit 204 may be available through various APIs (Application Programming Interfaces), including Web APIs. Note that when the inference device 10 is realized by multiple devices connected to each other and capable of communicating with each other, it may be called, for example, an "inference system."

[0087] Variation 4 When creating the first prompt for round 0 in step S102 of FIG. 3A , the prompt creator 202 may create the first prompt for round 0 using entities included in the query input in step S101 of FIG. 3A . For example, the prompt creator 202 may create the first prompt including instructions to extract entities of the same or similar type as the type of entity included in the query, and the relationships between those entities. As a specific example, if the query includes a specific person's name, the prompt creator 202 may create, as the first prompt for round 0, a sentence such as, "List all the entities such like people's names in the context, and give the possible relationships between them."

[0088] 3A and 3B, both the determination of whether the first termination condition is satisfied (step S106) and the determination of whether the second termination condition is satisfied (step S111) are made. However, it is not necessary to make both determinations, and only one of the determinations may be made. Furthermore, other termination conditions may be used in addition to the first termination condition and the second termination condition, or other termination conditions may be used instead of at least one of the first termination condition and the second termination condition.

[0089] <Summary> As described above, the inference device 10 according to this embodiment can acquire advanced knowledge that realizes multi-stage inference by repeatedly performing inference that integrates a language model and explicit knowledge (i.e., sequentially performing inference that integrates a language model and explicit knowledge). This makes it possible for the inference device 10 according to this embodiment to obtain reliable inference results for problems that require multi-stage inference.

[0090] In addition to the above, the inference device 10 according to this embodiment can also use acquired advanced knowledge to infer a solution to a problem that requires multi-stage inference through single-step inference. Therefore, after acquiring advanced knowledge about a certain problem, it becomes possible to perform more efficient inference on problems that are the same as or similar to that problem.

[0091] The following supplementary notes are further disclosed with respect to the above embodiments. (Supplementary Note 1) An inference device including: a memory; and at least one processor connected to the memory, wherein the processor sequentially infers relationship information representing a relationship between specific characters or character strings in the first character string and new knowledge representing a predetermined relationship between the relationships represented by the relationship information, based on a first character string, a second character string related to the first character string, inferred knowledge, and a natural language processing model, until relationship information satisfying a requirement represented by the second character string is inferred. (Supplementary Note 2) The inference device according to Supplementary Note 1, wherein the natural language processing model includes a first natural language processing model and a second natural language processing model specialized in a domain to which the first character string belongs, and the processor infers the relationship information based on the first natural language processing model and infers the new knowledge based on the second natural language processing model. (Supplementary Note 3) The inference device according to Supplementary Note 2, wherein the processor creates an instruction for the first natural language processing model based on the second character string and the new knowledge, and infers the relationship information based on the instruction and the first natural language processing model. (Supplementary Note 4) The inference device according to any one of Supplements 1 to 3, wherein the knowledge is a rule or model with a clear inference process, and includes a knowledge graph, symbolic logic, and rules. (Supplementary Note 5) The inference device according to Supplementary Note 1, wherein the processor infers direct relationship information that represents a relationship between specific characters or character strings directly inferred from the first character string, and knowledge that represents a predetermined relationship between the relationships represented by the direct relationship information, and then sequentially infers the relationship information and the new knowledge based on the first character string, the second character string, the inferred knowledge, and the natural language model until relationship information that satisfies the requirement represented by the second character string is inferred. (Supplementary Note 6) The inference device according to Supplementary Note 1, wherein the processor updates the inferred knowledge using the new knowledge when there is consistency between the inferred new knowledge and the inferred knowledge.(Supplementary Note 7) A non-transitory storage medium storing a program executable by a computer to perform an inference process, wherein the inference process sequentially infers relationship information expressing a relationship between specific characters or character strings in the first character string and new knowledge expressing a predetermined relationship between the relationships expressed by the relationship information, based on a first character string, a second character string related to the first character string, inferred knowledge, and a natural language processing model, until relationship information satisfying a requirement expressed by the second character string is inferred. (Supplementary Note 8) An inference device including: a memory; and at least one processor connected to the memory, wherein the processor sequentially infers relationship information expressing a relationship between specific characters or character strings in the first character string, based on a first character string, a second character string related to the first character string, knowledge sequentially inferred by a second natural language processing model, and the first natural language processing model, until relationship information satisfying the requirement expressed by the second character string is inferred. (Supplementary Note 9) An inference device comprising: a memory; and at least one processor connected to the memory, wherein the processor sequentially infers new knowledge representing a predetermined relationship between the relationships represented by the relationship information, based on relationship information representing relationships between specific characters or character strings sequentially inferred by a first natural language processing model, inferred knowledge, and a second natural language processing model.

[0092] The present invention is not limited to the above-described specifically disclosed embodiments, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims.

[0093] [References] Reference 1: Koustuv Sinha, Shagun Sodhani, Jin Dong, Joelle Pineau, and William L. Hamilton. 2019. CLUTRR: A diagnostic benchmark for inductive reasoning from text. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4506-4515, Hong Kong, China. Association for Computational Linguistics. Reference 2: Zhang, Yuyu, et al. "Variational reasoning for question answering with knowledge graph." Proceedings of the AAAI conference on artificial intelligence. Vol. 32. No. 1. 2018.

[0094] 10 Inference device 101 Input device 102 Display device 103 External I / F 103a Recording medium 104 Communication I / F 105 RAM 106 ROM 107 Auxiliary storage device 108 Processor 109 Bus 201 Input unit 202 Prompt generation unit 203 First inference unit 204 Second inference unit 205 Knowledge management unit 206 End determination unit 207 Knowledge storage unit 208 Inference unit

Claims

1. An inference device having an inference unit that sequentially infers relationship information representing a relationship between specific characters or character strings in the first character string, new knowledge representing a predetermined relationship between the relationships represented by the relationship information, based on the first character string, the second character string related to the first character string, inferred knowledge, and a natural language processing model, until relationship information that satisfies the request represented by the second character string is inferred.

2. The natural language processing model includes a first natural language processing model and a second natural language processing model specialized for the domain to which the first character string belongs. The inference unit infers the relationship information based on the first natural language processing model and infers the new knowledge based on the second natural language processing model. The inference device according to claim 1.

3. A creation unit that creates an instruction for the first natural language processing model based on the second character string and the new knowledge. The inference unit infers the relationship information based on the instruction and the first natural language processing model. The inference device according to claim 2.

4. The knowledge is a rule or model with a clear inference process and includes a knowledge graph, symbolic logic, rules. The inference device according to any one of claims 1 to 3.

5. After the inference unit infers direct relationship information representing a relationship between specific characters or character strings directly inferred from the first character string and knowledge representing a predetermined relationship between the relationships represented by the direct relationship information, based on the first character string, the second character string, the inferred knowledge, and the natural language processing model, the relationship information and the new knowledge are sequentially inferred until relationship information that satisfies the request represented by the second character string is inferred. The inference device according to claim 1.

6. The inference device according to claim 1, further comprising a knowledge management unit that updates the inferred knowledge using the new knowledge when there is consistency between the new knowledge inferred by the inference unit and the inferred knowledge.

7. An inference method in which a computer executes an inference procedure of sequentially inferring relationship information representing a relationship between specific characters or character strings in the first character string, and new knowledge representing a predetermined relationship between the relationships represented by the relationship information, until relationship information that satisfies the requirement represented by the second character string is inferred, based on the first character string, the second character string related to the first character string, the inferred knowledge, and a natural language processing model.

8. A program for causing a computer to execute an inference procedure of sequentially inferring relationship information representing a relationship between specific characters or character strings in the first character string, and new knowledge representing a predetermined relationship between the relationships represented by the relationship information, until relationship information that satisfies the requirement represented by the second character string is inferred, based on the first character string, the second character string related to the first character string, the inferred knowledge, and a natural language processing model.

9. An inference device having a first inference unit that sequentially infers relationship information representing a relationship between specific characters or character strings in the first character string, based on the first character string, the second character string related to the first character string, the knowledge sequentially inferred by the second natural language processing model, and the first natural language processing model, until relationship information that satisfies the requirement represented by the second character string is inferred.

10. An inference device having a second inference unit that sequentially infers new knowledge representing a predetermined relationship between the relationships represented by the relationship information, based on the relationship information representing a relationship between specific characters or character strings sequentially inferred by the first natural language processing model, the inferred knowledge, and the second natural language processing model.

11. An inference system including a first inference device and a second inference device, wherein the first inference device sequentially infers relationship information representing the relationship between specific characters or character strings in the first character string based on the first character string, a second character string related to the first character string, knowledge sequentially inferred by a second natural language processing model, and a first natural language processing model, until relationship information satisfying the requirement represented by the second character string is inferred, and has a first inference unit; and the second inference device sequentially infers new knowledge representing a predetermined relationship between the relationships represented by the relationship information based on the relationship information sequentially inferred by the first natural language processing model, inferred knowledge, and a second natural language processing model, and has a second inference unit. Inference system.

Citation Information

Patent Citations

  • Inference device facilitating input and verification of knowledge

    JP1988241638A

  • Inference system

    JP1990220135A

  • Inferring method, inferring system, and computer program thereof

    JP2006024045A

  • Information processing method and information processing system

    WO2022009543A1

  • Processing method, processing system, and processing program

    WO2022113175A1