Large language model system combined with ontology
The ontology-assisted large-scale language model addresses semantic and logical challenges by recording vocabulary attributes and relationships, enhancing the model's ability to generate realistic and appropriate responses.
Patent Information
- Application Number
- PCT/JP2025/019784
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-04
- Filing Date
- 2025-06-02
- Publication Date
- 2025-12-11
AI Technical Summary
Conventional large-scale language models ignore the original meanings of vocabulary terms, leading to issues such as the symbol grounding problem, lack of semantic optimality, difficulty in handling inappropriate vocabulary, and unsolved frame and inference challenges.
An ontology-assisted large-scale language model system that records relational links between vocabulary terms with attributes, uses ontology query-related parts extraction, and incorporates ontology learning and insertion to enhance semantic understanding and logical bridging.
Enables realistic and tangible image generation, suppresses inappropriate vocabulary, and facilitates efficient logical inference by leveraging ontology-based attributes and relationships.
Smart Images

Figure JP2025019784_11122025_PF_FP_ABST
Abstract
Description
Ontology-based large-scale language model system
[0001] The present invention relates to an ontology-assisted large-scale language model system that assigns meaning to the input and output of a large-scale language model and to vocabulary tokens within the large-scale language model by using an ontology record that systematically organizes and records the meanings of vocabulary.
[0002] In recent years, there has been remarkable progress in artificial intelligence (AI), and in particular, large language models (LLMs) that enable inquiries and responses in natural language through deep learning, which involves multi-layering neural networks that mimic neural circuits and learning large amounts of data, have attracted attention. Prior art documents related to this application include the following:
[0003] Japanese Patent Publication No. 2023-523644 Japanese Patent Publication No. 2023-73095 Japanese Patent No. 5484317 Japanese Patent No. 7313757 Japanese Patent No. 6868860 Japanese Patent No. 6928332 Japanese Patent No. 6913308 Japanese Patent No. 6814482 Japanese Patent No. 6792751 Japanese Patent No. 7441391 US 2022 / 0405484 A1 US 2016 / 0026441 A1 Japanese Patent Publication No. 2005-157690 Japanese Patent Publication No. 2002-342146
[0004] In large-scale language models, a vocabulary term is represented by a one-hot vector, which is a long vector of zeros with the same number of dimensions as the number of vocabulary terms used, with a single one placed at the position corresponding to the term. All of the vocabulary terms in a large amount of literature are replaced with vectors of this type, and deep learning is used to determine the correlation between each vocabulary vector. Answers are then created in response to a prompt by generating and adding, one by one, vocabulary terms that are likely to appear after the prompt and any answers already generated. In this way, conventional large-scale language models completely ignore the original meanings of the vocabulary terms used within the language model, the query, and the answer, as understood by humans.
[0005] The separation of vocabulary from its original meaning leads to the following problems: (1) The vocabulary in the answer sentences ignores the various attributes of the vocabulary, making it difficult to create a realistic, tangible image (the symbol grounding problem). (2) The vocabulary indicated in the answer sentences is merely the most likely next word in the previous vocabulary sequence, and there is no guarantee that it is semantically optimal for humans. Even if we want to consider the next most likely vocabulary candidate, this is difficult because the original meaning is ignored. (3) Questions such as "How are nuclear weapons manufactured?" or answers that promote discrimination are problematic, and although humans currently deal with these manually, it is difficult to achieve a simple and complete solution. (4) When using large-scale language models to consider a situation such as removing a time bomb from a room containing one, the "frame problem" remains unsolved: the bomb explodes while the nearly infinite number of factors need to be considered. (5) When making logical inferences from one proposition to another that may appear unrelated at first glance, if there are no similar examples in the training data of a large-scale language model, the inferences are likely to be difficult or incoherent.
[0006] The present invention has been made to solve these conventional problems, and its purposes are to enable the user who made the inquiry to easily obtain a realistic and tangible image from the content of the reply, to enable comparison and verification of multiple possible reply statements rather than just one, to efficiently suppress inquiry statements and reply statements that use inappropriate vocabulary, to limit the content to be considered as much as possible without affecting the quality of the reply statement, and to enable logical bridging between propositions that are not directly related.
[0007] As a means for achieving the above-mentioned object, the ontology-assisted large scale language model system of claim 1 comprises an ontology recording means for recording relational links between vocabulary terms together with vocabulary attributes in a large scale language model (LLM) that generates a response sentence in response to a query sentence entered in a prompt, an ontology query-related part extraction means for extracting the query-related part from the recorded content of the ontology recording means, and an ontology recorded content conversion means for converting the extracted query-related part ontology into a form recognizable by the large scale language model, and is characterized in that the system comprises at least one of (1) ontology additional learning means for additionally learning the large scale language model, and (2) ontology query sentence insertion means for inserting the converted recorded content into a query sentence to the large scale language model.
[0008] The ontology-assisted large-scale language model system of claim 2 is the large-scale language model system of claim 1, characterized in that the ontology recording means further comprises parent-child relationship vocabulary attribute inheritance means for inheriting the attributes of the parent vocabulary as attributes of the child vocabulary when the relationship link between the vocabularies is a parent-child relationship.
[0009] The ontology-assisted large-scale language model system according to claim 3 is the large-scale language model system according to claim 1 or 2, characterized in that the ontology recording means is provided with a divided vocabulary group by namespace, which divides vocabulary groups using namespaces by field and eliminates interference between vocabulary fields.
[0010] The ontology-assisted large-scale language model system according to claim 4 is the large-scale language model system according to claim 1 or 2, characterized in that it further comprises a vocabulary ontology reference means for referencing, from the ontology recording means, the attributes of vocabulary contained in the query sentence or the response sentence, and further, if necessary, relational links.
[0011] The ontology-assisted large-scale language model system of claim 5 is characterized in that, in the ontology-assisted large-scale language model system of claim 1 or 2, it further comprises a step in which an operator compares and references attributes of a plurality of vocabulary candidates other than the vocabulary contained in the answer sentence, and verifies the validity of the vocabulary selection.
[0012] The ontology-assisted large-scale language model system of claim 6 is characterized in that, in the ontology-assisted large-scale language model system of claim 1 or 2, it further comprises an inappropriate vocabulary suppression means for verifying whether or not there are any inappropriate attributes in the vocabulary included in the query sentence or the answer sentence in the recorded contents of the ontology recording means, and suppressing them.
[0013] The ontology-assisted large-scale language model system of claim 7 is characterized in that, in the large-scale language model system with ontology of claim 1 or 2, when inferring a target proposition from a starting proposition in the query sentence, the system further comprises proposition-related vocabulary group information adding means for expanding vocabulary groups related to the vocabulary groups of the starting proposition, the target proposition, or both, in the recorded contents of the ontology recording means, searching for vocabulary groups related to both the starting proposition and the target proposition, and adding the related vocabulary groups to the query sentence in the ontology query sentence insertion means.
[0014] The ontology-assisted large-scale language model system of claim 8 is characterized in that, in the large-scale language model system of claim 1 or 2, when a proposition to be solved and an explanation of the situation surrounding the proposition are presented as a query statement in the query statement, the system comprises a vocabulary proximity evaluation means for evaluating proximity between a vocabulary group in the query statement and a vocabulary group in the recorded content of the ontology recording means in terms of logical proximity, spatial proximity, temporal proximity or a combination of these, and highly evaluating vocabulary groups in the recorded content of the ontology recording means that have high proximity, and the highly evaluated vocabulary group is inserted into the query statement as the ontology query statement insertion means.
[0015] The ontology-assisted large-scale language model according to claim 9 is characterized in that in the ontology-assisted large-scale language model system according to claim 1 or 2, the ontology record content conversion means in the ontology additional learning means and the ontology query statement insertion means comprises at least one of a JSON format conversion means for converting to JSON format and an XML format conversion means for converting to XML format.
[0016] The ontology-assisted large-scale language model system of claim 1 includes an ontology recording means for recording relational links between vocabulary terms along with vocabulary attributes. An ontology query-related portion extraction means is provided for extracting the query-related portion from the content recorded in the ontology recording means. An ontology record content conversion means is provided for converting the extracted query-related portion ontology into a form recognizable by the large-scale language model. An ontology record content conversion means is provided for converting any portion of the content recorded in the ontology recording means into a form recognizable by the large-scale language model. An ontology additional learning means realizes an additional learning function. An ontology query statement insertion means provides the function of inserting the converted record content into a query statement for the large-scale language model.
[0017] The ontology-assisted large-scale language model system according to claim 2 is provided with a parent-child relationship vocabulary attribute inheritance means, so that when the relational link between vocabularies is a parent-child relationship, the attributes of the parent vocabulary are inherited as attributes of the child vocabulary.
[0018] The ontology-assisted large-scale language model system according to claim 3 is provided with a vocabulary group divided by namespace, and thus the vocabulary group is divided using a namespace by field, eliminating interference between vocabulary fields.
[0019] The ontology-assisted large-scale language model system according to claim 4 includes a vocabulary ontology reference means, and therefore references vocabulary attributes and, if necessary, relational links from the ontology recording means.
[0020] The ontology-assisted large-scale language model system according to claim 5 includes a step of verifying the validity of vocabulary selection, whereby attributes are compared and referenced with a plurality of vocabulary candidates other than the vocabulary to verify the validity of vocabulary selection.
[0021] The ontology-assisted large-scale language model system according to claim 6 includes an inappropriate vocabulary suppression means, which verifies and suppresses the presence of inappropriate attributes in the vocabulary of the recorded contents of the ontology recording means.
[0022] The ontology-assisted large-scale language model system of claim 7 comprises a proposition-related vocabulary group information adding means, so that when inferring a target proposition from a starting proposition, for the vocabulary groups of the starting proposition, the target proposition, or both, the vocabulary groups related to the vocabulary groups in the recorded contents of the ontology recording means are expanded, vocabulary groups related to both the starting proposition and the target proposition are searched for, and the related vocabulary groups are added to the query statement by the ontology query statement insertion means.
[0023] In the ontology-assisted large-scale language model system according to claim 8, a lexical proximity evaluation means is provided, so that when a proposition to be solved and a description of the situation surrounding the proposition are presented as a query, the proximity between a vocabulary group in the query and a vocabulary group recorded in the ontology recording means is evaluated using one or a combination of logical proximity, spatial proximity, and temporal proximity, and vocabulary groups in the ontology record that have high proximity are highly ranked.The highly ranked vocabulary groups are then inserted into the query as the ontology query statement insertion means.
[0024] The ontology-assisted large-scale language model system according to claim 9 includes a JSON format conversion means for converting ontology record contents into the JSON format, and an XML format conversion means for converting ontology record contents into the JSON format.
[0025] This is a hardware configuration diagram of the present invention, showing an example of an ontology, illustrating the relationship between classes and instances. This diagram illustrates how, in parent-child vocabulary terms, attributes of a parent vocabulary are inherited by attributes of a child vocabulary. The diagram shows that vocabulary attributes consist of attributes inherited from the parent vocabulary and attributes newly added to the vocabulary. The recorded contents of an ontology are converted into JSON format. The recorded contents of an ontology are converted into XML format. The diagram shows that attributes in an ontology are described by referencing vocabulary defined in another ontology. This is an example of a "disease name" ontology. This is an example of the configuration of ontologies for "disease name," "symptoms," and "drugs." Using the "diabetes" class as an example, this diagram shows parent-child relationship links, inherited / added attributes for each class and instance, and vocabulary in the ontology referenced by each attribute. This is an explanatory diagram of reference links to vocabulary defined in another ontology. This is an explanatory diagram of using namespaces in an ontology to prevent vocabulary interference between different fields. This is an example of the configuration of a master table for managing namespaces. 1 is an example of the configuration of a master table for managing ontologies. FIG. 2 is an example of the configuration of a master table for managing ontology vocabulary attributes. FIG. 3 is an example of the configuration of a master table for ontology vocabulary management means. FIG. 4 is an example of a vocabulary reference means within an ontology record. It shows that matching vocabulary generated in a large-scale language model with ontology instances corresponding to that vocabulary leads to a solution to the symbol grounding problem for users. An explanatory diagram of one-hot vectors used in large-scale language models. An explanatory diagram of a vocabulary selection validity verification means. An explanatory diagram of an inappropriate vocabulary suppression means. An explanatory diagram of a proposition-related vocabulary group addition means. An explanatory diagram of a vocabulary proximity evaluation means.
[0026] The ontology-assisted large-scale language model system of the present invention includes a server device, a database, and a terminal. The server device is a known computer device and includes an arithmetic unit, a main memory device, an auxiliary memory device, an input device, an output device, and a communication device. The arithmetic unit, the main memory device, the auxiliary memory device, the input device, the output device, and the communication device are connected to each other via a bus interface. The arithmetic unit includes a known processor capable of executing an instruction set. The main memory device includes volatile memory such as RAM that can temporarily store the instruction set. The auxiliary memory device includes non-volatile data storage that can record an OS and programs. The data storage may be, for example, an HDD or an SSD. The input device is, for example, a keyboard. The output device is, for example, a display such as an LCD. The communication device includes a network interface that can be connected to a network. The server device includes means for recording ontology, extracting relevant parts of an ontology query, converting recorded content of an ontology, learning additional ontology, inserting an ontology query, inheriting parent-child relationship vocabulary attributes, referencing a vocabulary ontology, suppressing inappropriate vocabulary, adding proposition-related vocabulary group information, and evaluating lexical proximity. The processor of the server device exerts the functions and effects of these means. The database of the present invention may be configured by an auxiliary storage device of the server device, or may be configured by an auxiliary storage device separate from the server device. The database stores information handled by the ontology-assisted large-scale language model system. The terminal of the present invention, like the server device, has a known computer hardware configuration. The server device, database, and terminal of the present invention are capable of communicating via a network.
[0027] FIG. 1 is an example of a hardware configuration diagram of the present invention. Because large-scale language models require huge amounts of data and computational resources, they are typically stored on the cloud and connected to a company's local area network (LAN) via a router over the Internet. A business server is located within the company. Company staff access the business server and large-scale language models via terminals connected to the LAN. It is also possible to build part or all of the business system on the cloud, or to install part or all of the large-scale language model in a lightweight version within the company. Large-scale language models can also be accessed via the Internet using mobile devices.
[0028] Figure 2 shows an example of a user interface for a large-scale language model (LLM). LLMs are currently being rapidly developed, with numerous models being developed, including ChatGPT (a registered trademark of OpenAI), Bard, LaMDA (a registered trademark of Google), and LLaMA (a registered trademark of Meta). Naturally, user interfaces vary, but typically, as shown in Figure 2, they consist of a box for entering prompts to instruct or inquire of the LLM, a box for displaying responses to those prompts, and a box for displaying a history of prompts and responses as a usage log.
[0029] Figure 3 illustrates how parent vocabulary attributes are inherited by child vocabulary attributes in a parent-child relationship, using "organisms" as an example of ontology. "Organisms" are divided into "plants" and "animals," which are further divided into "mammals" and "birds." Common attributes of the parent vocabulary, such as "mammals" (warm-bloodedness, hairy skin, supporting front legs, supporting hind legs), are inherited by child vocabulary terms such as "human," "dog," and "cat." For "human," the inherited attribute set, "supporting front legs," is overwritten with "free front legs." This approach allows common attributes of child vocabulary terms to be grouped together as attributes of the parent vocabulary, saving storage space and clarifying the relationships between terms. Note that not all attributes of the parent vocabulary are inherited by the child vocabulary terms. When all "humans" are included in "mammals," such as when "humans" are "mammals," this is called an is-a relationship in ontology, and attribute inheritance occurs. However, attributes such as "Human A owns a car" (has-a), "Human B is the president of a company" (role-of), and "Human C is a member of Company D" (a-part-of) are limited to the vocabulary in question, and attribute inheritance to child vocabularies does not occur.
[0030] Figure 4 shows that the attributes of a vocabulary consist of attributes inherited from the parent vocabulary and attributes newly added to the vocabulary. Taking the attributes of "cat" as an example, it inherits the attributes of the parent vocabulary, "mammal" (warm-bloodedness, having fur on the skin, having supporting front legs, having supporting hind legs), and then adds attributes such as meows and movements using text, audio files, videos, etc. These files may be actual data stored on the LLM server, or they may be URLs to video software content. Furthermore, "cat" may be the parent vocabulary, and individual cats such as "Mike" and "Tama" may be child vocabulary words.
[0031] Figure 5 shows the recorded contents of the ontology shown in Figures 3 and 4 converted into JSON format (JSON format conversion means). Figure 6 shows the recorded contents of the ontology shown in Figures 3 and 4 converted into XML format (XML format conversion means). Because ontologies are composed of vocabulary and relational links, as shown in Figures 3 and 4, they cannot be used as is for training large-scale language models or for query statements. For this reason, it is necessary to convert any part of the ontology record into a format such as JSON or XML.
[0032] Using a large-scale language model as a base model, ontology record contents for a field of interest can be converted into JSON or XML format and subjected to additional learning, thereby building a large-scale language model specialized for that field (ontology additional learning means). Furthermore, related ontology record contents can be inserted into a query statement to elicit a more appropriate answer statement (ontology query statement insertion means). It should be noted that, instead of using structured expressions such as JSON or XML, it is also possible to express things as a list of individual sentences, such as "There are two types of living things: animals and plants" or "One type of animal is mammals." While using such expressions is also included in the present invention, they are lengthy, the structure is not clearly visualized, and are therefore not a desirable embodiment.
[0033] While the relationships between vocabulary terms can be obtained using large-scale language models themselves, this requires training on large amounts of text, and accurately and concisely representing complex hierarchical structures is no easy task. Ontology recording, on the other hand, directly specifies the links between vocabulary terms, making it easy to represent deep hierarchical structures without the need for extensive training. Furthermore, in hierarchical vocabulary structures, common attributes between child vocabulary terms can be extracted and aggregated into the parent vocabulary, significantly reducing recording capacity and visualizing the relationships between vocabulary terms, making them easier to understand.
[0034] Figure 7 shows that in an ontology, vocabulary defined in another ontology is referenced when describing the attributes of a vocabulary. Attributes are described for each vocabulary, but the vocabulary used to describe the attributes also references vocabulary defined in another ontology. In this way, the meanings of the vocabulary and its attributes are all defined in one of the ontologies, making it possible to express them unambiguously. Furthermore, using the ontology that defines the referenced vocabulary can greatly broaden the meaning and nuance of the vocabulary.
[0035] FIG. 8 is an example of the "disease name" ontology.
[0036] Figure 9 shows an example of the configuration of the "disease name," "symptoms / findings," and "drug" ontologies. The attributes of each disease name describe the symptoms and test findings observed with the disease, as well as the drugs used to treat the disease. However, these symptoms and drug names are defined in the "symptoms / findings" ontology and the "drug" ontology, which are separate from the disease name ontology, and the attributes of the disease name are described as references to vocabulary defined in the separate ontologies. Note that while this figure only lists the "symptoms / findings" and "drug" ontologies, other ontologies such as "tests," "treatments," and "nursing" are also used.
[0037] Using the "diabetes" class as an example, Figure 10 shows the parent-child relationship links between the parent vocabulary terms "metabolic disease" and "diabetes" itself, and the child vocabulary terms "type I diabetes" and "type II diabetes," as well as the inheritance / addition of attributes and reference links to ontology terms referenced by each attribute. The <similar disease> attribute of the disease in question references another disease name vocabulary within the same "disease name" ontology. It also shows reference links to cases of the disease in question.
[0038] Figure 11 is an explanatory diagram of a reference link to a vocabulary defined in another ontology. In Figure 10, "hyperglycemia" in the <pathological condition> of "diabetes" has a reference link to "hyperglycemia", a vocabulary in the "symptoms / findings" ontology. In the <definition> attribute of "hyperglycemia", a vocabulary in the "symptoms / findings" ontology, it is written that blood glucose level > 140 mg / dl HBa1c > 7.0.
[0039] Figure 12 is an explanatory diagram showing how namespaces are used in ontologies to prevent vocabulary interference between different fields. While various vocabulary terms are appropriate for each field, the same vocabulary is often used with completely different meanings in different fields. To prevent unexpected interference caused by newly defined vocabulary in a certain field, namespaces are used to separate vocabulary groups by field, allowing new vocabulary to be defined without considering the impact on other fields. In Figure 12, the word "fatigue" is used in various ways, such as "adrenal fatigue," "institutional fatigue," and "fatigue due to long working hours." However, by separating them using namespaces, interference can be avoided. It goes without saying that the newly defined vocabulary will not interfere with existing vocabulary within the same namespace.
[0040] 13 shows an example of the configuration of a master table for managing a group of name spaces. An ID is assigned to each name space, and records are managed using this ID.
[0041] 14 shows an example of the configuration of a master table for managing ontologies. Within the name space, IDs are assigned to ontologies and records are managed.
[0042] 15 shows an example of the configuration of a master table for managing ontology vocabulary attributes. IDs are assigned to the attributes of vocabulary used in an ontology to manage vocabulary attributes. Note that vocabulary included in the same ontology has common vocabulary attributes.
[0043] FIG. 16 shows an example of a vocabulary management unit. Along with the vocabulary ID# and the vocabulary, the unit records the namespace ID# to which the vocabulary belongs, the ontology ID#, and the path to the vocabulary within the ontology. For each vocabulary term included in a query, this diagram is used to search the ontology and extract the term and its attributes (query-related partial ontology extraction unit). The term also inherits the attributes of its parent vocabulary (attribute inheritance). If the term itself is insufficient, sibling vocabulary groups, such as "glycogen storage disease" for "diabetes" in FIG. 8, and child vocabulary groups, such as "type I diabetes" and "type II diabetes," may also be extracted as query-related partial ontologies. The query-related partial ontology extracted in this way is converted into a form recognizable by the large-scale language model using either a JSON format conversion means for converting to JSON format or an XML format conversion means for converting to XML format, as shown in Figures 5 and 6 (ontology recorded content conversion means), and inserted into a query statement for the large-scale language model (ontology query statement insertion means). Alternatively, instead of inserting into a query statement, the recorded content of the ontology can be additionally trained on the large-scale language model (ontology additional training means). While the entire recorded content of the ontology can be additionally trained, since the recorded content of the ontology can be enormous in volume, a portion of the ontology can be extracted and additionally trained depending on the purpose.
[0044] Figure 17 is an explanatory diagram of one-hot vectors used in large-scale language models. In large-scale language models, a one-hot vector (vocabulary token vector) is assigned to each vocabulary term, with only the corresponding vocabulary term set to 1 and the rest set to 0, along with a corresponding token #. The dimensionality of the vector is the total number of vocabulary terms. For documents in the training data, each vocabulary term is replaced with its corresponding vocabulary token vector. Learning is performed by solving a so-called "hole problem," in which vocabulary token vectors from a massive amount of document data are estimated from surrounding vocabulary token vectors. Based on the correlation between vocabulary token vectors, the probability of the hole vocabulary term is calculated, and the largest vocabulary term is used as the estimated vocabulary. Conventional supervised learning requires the preparation of a large number of pairs (corpora) of training data and their correct answers (teaching data), which is time-consuming and costly. In contrast, in the hole problem learning, the hidden vocabulary itself serves as the correct teaching data, making it possible to use a large corpus at once, significantly improving learning accuracy.
[0045] Because the large-scale language model abstracts all meanings of vocabulary and treats it as a vocabulary token vector with all 0s and only one 1, for example, when it outputs the vocabulary word "cat," it merely means that the probability of that vocabulary token vector is higher than others; it does not actually understand and output the specific image of "cat" that humans have. In other words, the fact that large-scale language models do not understand the relationship between the output vocabulary and its specific image (meaning) (the symbol grounding problem) is a known limitation of large-scale language models. In this invention, the vocabulary token vector "cat" is associated with the vocabulary word "cat" in the ontology, making it possible to utilize the abundant information recorded in the ontology (vocabulary ontology reference means). In other words, this is an attempt to assign meaning to the output of a large-scale language model. Large-scale language models are software within a computer and do not have physicality. Therefore, they cannot handle the texture or smell of petting a "cat." However, for humans using large-scale language models, the information, audio, video, etc. contained in the vocabulary in the ontology, combined with their own past experiences, allows them to understand it with a sufficient sense of reality. Therefore, it seems that the symbol grounding problem can be solved practically by combining the three elements of a large-scale language model, an ontology, and users.
[0046] Figure 18 shows an example of an ontology record vocabulary reference means. This figure further illustrates how matching vocabulary generated by a large-scale language model with ontology vocabulary can solve the symbol grounding problem for users. When a doctor encounters the term "hyperglycemia" in "diabetes" while using a large-scale language model, the doctor references the vocabulary record in the ontology, as shown in Figure 17. The term in the ontology is "diabetes" in the "metabolic system" and "glucose metabolism system" of the "disease name" ontology in the "medical" namespace. The attribute "hyperglycemia" in the "pathological condition" can be used to reference "hyperglycemia" in the "symptoms / findings" ontology, and information on blood glucose > 140 mg / dL can be obtained from the attribute "definition" of "hyperglycemia." This ontology record vocabulary reference means can be visually referenced by a human, or it can be referenced by other software or another large-scale language model via an API or other means and used for the relevant processing.
[0047] FIG. 19 is an explanatory diagram of the vocabulary selection validity verification means. Suppose a large-scale language model outputs "pneumonia" based on a patient's symptoms and findings. However, the probabilities of the output vocabulary are also fairly high, such as "bronchopneumonia" and "lung cancer," so the validity of "pneumonia," which had the highest probability at that time, is not overwhelmingly valid. By having the large-scale language model present not only the vocabulary with the highest probability but also multiple vocabulary with the next highest probability, and displaying the content of each using the ontology record vocabulary reference means, the operator (in this case, a doctor) can verify the possibility of vocabulary for other disease names using what is called backward inference in an expert system, and select a more appropriate vocabulary for the disease name.
[0048] Figure 20 is an explanatory diagram of an inappropriate vocabulary suppression method. The use of inappropriate vocabulary is a problem in the operation of large-scale language models. Dangerous queries such as "Tell me how to make a nuclear bomb" or "Tell me an efficient way to kill someone" should not be answered. Large-scale language models use user interactions as material for additional learning. However, some users have posted large amounts of Nazi praise and racial slander. This has led to major problems and the suspension of service for large-scale language models that have trained on these messages, generating inappropriate responses. To address this issue, large numbers of workers are currently mobilized to manually remove inappropriate vocabulary and responses. This is extremely expensive, resulting in increased service costs and delays in service launch. To address this issue, the present invention adds a "sensitive flag" to the ontology vocabulary attributes.
[0049] By setting a flag in the <Sensitive Flag> attribute of an ontology term, as shown in Figure 20, it is possible to prescreen for sensitive or potentially problematic vocabulary contained in queries, responses, and additional training data. Furthermore, by using ontology records, vocabulary related to inappropriate vocabulary can also be included in the screening. This allows for screening not only inappropriate vocabulary but also similar or related vocabulary. This makes it possible to suppress, remove, or flag sensitive vocabulary in queries, responses, and additional training data. Focusing on sensitive vocabulary can also significantly improve the efficiency of visual screening. Furthermore, by using a scale of sensitivity (sensitivity) in the <Sensitive Flag>, various responses can be made depending on the sensitivity, ranging from immediate rejection to warnings. The <Sensitive Flag> can also be set for multiple terms, making it possible to increase the sensitivity level by combining "knife" with "kill" or "injure."
[0050] Figure 21 is an explanatory diagram of the means for adding proposition-related vocabulary group information. In a certain children's story, the situation is described as follows: "An old woman living in a mountain village has died, and the neighborhood ladies are cooking something in a pot." When the teacher asks, "What are the ladies cooking?", quite a few children answer, "They are boiling the old woman's body to disinfect it," which has become a hot topic. This shows that, because there is no direct relationship between "the old woman has died" and "the neighborhood ladies are cooking something in a pot," a wrong inference can be made. In this invention, by adding related vocabulary information recorded in the ontology, such as (a) "home" in the <place of death> attribute of the ontology vocabulary "death" and "home" in the ontology attribute <funeral>, and (b) "funeral" in the <collaboration> attribute of the ontology vocabulary "life in the countryside", to the original situation description in (c), and adding proposition-related vocabulary group information such as "The funeral is held at home" and "Funerals in the countryside are a collaborative effort among neighbors" (ontology query sentence insertion means), and querying a large-scale language model, it is possible to expect the generation of a correct answer sentence such as (d) "She is probably making boiled vegetables to be served at the funeral."
[0051] When humans read text, they not only rely on the text itself, but also unconsciously supplement it with related common sense that is not apparent in the text to achieve a correct understanding. If there is a direct reference in a document trained with a large-scale language model, inference like (d) is possible, but this requires a huge amount of training data, as well as computational resources and power for processing. In this invention, by recording so-called "common sense" in an ontology and adding proposition-related vocabulary group information as appropriate, highly accurate inference is possible with a large-scale language model with a small amount of training data. Proposition-related vocabulary group information can be added to any vocabulary in the situation description, query, or answer.
[0052] FIG. 22 is an explanatory diagram of a lexical proximity evaluation method. A time bomb is sitting on a desk in a room, and its time to explode is approaching. If it is not removed quickly, it will explode. Trying to solve this problem using a large-scale language model would involve taking into account too much background information that is not necessary for solving the problem, such as "What is the structure of the ceiling and floor?", "What is the temperature of the room?", and "What is the material of the desk?", and while thinking about various things, time will run out and the bomb will explode. This problem is known as the "frame problem," and the key is how to extract the information necessary for problem solving and solve the problem quickly. This invention proposes introducing lexical proximity evaluation to eliminate as much background information as possible that does not contribute to the problem to be solved, and focus on highly relevant information for rapid problem solving.
[0053] In the ontology record shown in Figure 21, words such as "clock," "bomb," and "desk" are important because they directly relate to the situation, while words such as "banana" and "music" are hardly related. To evaluate these vocabulary terms, we use a tool such as word2vector to obtain word embeddings of each term's meaning. The cosine (cos) of these vectors represents the correlation coefficient between the two, so only words with a high correlation coefficient are considered to have high lexical proximity. Location proximity is also important. While "desk" and "bomb" have little semantic proximity, they are close in the information shown in Figure 21. Places such as "Los Angeles" and "Paris" are hardly related. Temporal proximity is also important. Words contained in information about recent minutes, hours, and dates can be said to have high proximity, but words contained in descriptions of situations from 100 years ago have little proximity and are therefore hardly related. When evaluating lexical proximity, we should at least take into account the semantic proximity, location proximity, and time proximity of the words. At least one of meaning, location, and time, or any combination of these, can be considered. By limiting the vocabulary to any of these highly proximate terms and using the ontology query insertion method to query a large-scale language model, it becomes possible to obtain an answer within a practical time frame.
[0054] Although the embodiments have been described above, the specific configuration of the present invention is not limited to the above embodiments, and the present invention also includes design changes within the scope of the invention. For example, ChatGPT and the like have been given as examples of large-scale language models, but the same applies to other systems, and the present invention also includes the use of new systems that will be developed in the future as machine learning models.
Claims
1. A large-scale language model system with ontology, characterized in that it comprises a large-scale language model (LLM) that generates an answer sentence in response to a query sentence entered in a prompt, the system comprising: an ontology recording means for recording the relational links between each vocabulary along with the attributes of the vocabulary; an ontology query-related part extraction means for extracting the query-related part from the recorded content of the ontology recording means; and an ontology recorded content conversion means for converting the extracted query-related part ontology into a form recognizable by the large-scale language model, and using the converted recorded content, it comprises at least one of: (1) ontology additional learning means for additional learning the large-scale language model; and (2) ontology query sentence insertion means for inserting into a query sentence to the large-scale language model.
2. The ontology-assisted large-scale language model system according to claim 1, characterized in that said ontology recording means is provided with a parent-child relationship vocabulary attribute inheritance means for inheriting the attributes of the parent vocabulary as attributes of the child vocabulary when the relationship link between each of said vocabulary is a parent-child relationship.
3. The ontology-assisted large-scale language model system according to claim 1 or 2, characterized in that the ontology recording means is provided with a divided vocabulary group by namespace, which divides vocabulary groups using namespaces by field and eliminates interference between vocabulary fields.
4. A large-scale ontology-assisted language model system as claimed in claim 1 or 2, characterized in that it comprises a vocabulary ontology reference means for referencing the attributes of vocabulary contained in the query sentence or response sentence, and further, if necessary, relational links, from the ontology recording means.
5. The ontology-assisted large-scale language model system according to claim 1 or 2, characterized in that it comprises a step in which an operator compares and references attributes of multiple vocabulary candidates other than the vocabulary contained in the answer sentence, and verifies the validity of the vocabulary selection.
6. The ontology-assisted large-scale language model system according to claim 1 or 2, characterized in that it comprises an inappropriate vocabulary suppression means for verifying whether or not there are any inappropriate attributes in the vocabulary contained in the query sentence or the response sentence in the contents recorded in the ontology recording means, and suppressing them.
7. The ontology-assisted large-scale language model system according to claim 1 or 2, characterized in that, when inferring a target proposition from a starting proposition in the query sentence, the system further comprises a proposition-related vocabulary group information adding means for expanding vocabulary groups related to the vocabulary groups of the starting proposition, the target proposition, or both, recorded in the ontology recording means, searching for vocabulary groups related to both the starting proposition and the target proposition, and adding the related vocabulary groups to the query sentence in the ontology query sentence insertion means.
8. The ontology-assisted large-scale language model system according to claim 1 or 2, characterized in that, when a proposition to be solved and an explanation of the situation surrounding the proposition are presented as a query statement, the system is provided with a vocabulary proximity evaluation means which evaluates the proximity between a vocabulary group in the query statement and a vocabulary group in the recorded content of the ontology recording means using either logical proximity, spatial proximity, or temporal proximity, or a combination of these, and highly rates a vocabulary group in the recorded content of the ontology recording means which has high proximity, and which inserts the highly rated vocabulary group into the query statement as the ontology query statement insertion means.
9. A large-scale ontology-compatible language model system as described in claim 1 or 2, characterized in that the ontology record content conversion means in the ontology additional learning means and the ontology query statement insertion means is provided with at least one of a JSON format conversion means for converting to JSON format and an XML format conversion means for converting to XML format.
Citation Information
Patent Citations
Intelligent trademark law question and answer method for enhancing big language model reasoning through knowledge graph
CN117807202A
Method and device for monitoring electronic bulletin board information on homepage, method and device for monitoring alteration inhibition information, method and device for monitoring particular term or particular sentence
JP2002342146A
Electronic equipment and information providing method
JP2005157690A
Knowledge management system
JP2020194520A
Recursive ontology-based systems engineering
US20160026441A1