Ontology-integrated large language model server
The ontology-assisted large-scale language model addresses the limitations of conventional models by integrating vocabulary attributes and relational links to enhance semantic understanding and logical coherence, providing realistic and efficient responses.
Patent Information
- Application Number
- JP2024151854
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-04
- Filing Date
- 2024-09-04
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2044-09-04
AI Technical Summary
Conventional large-scale language models ignore the original meaning of vocabulary, leading to unrealistic and semantically suboptimal responses, difficulty in forming coherent logical inferences, and challenges in handling sensitive queries and complex situations.
An ontology-assisted large-scale language model that records relational links and attributes of vocabulary, performs additional learning, and inserts ontology content into query sentences to enhance semantic understanding and logical bridging.
Enables realistic and tangible responses, suppresses inappropriate content, and facilitates efficient logical inference by leveraging ontology records to improve the quality and relevance of language model outputs.
Smart Images

Figure 2025183131000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a large-scale language model with ontology that assigns meaning to the input and output of a large-scale language model and to vocabulary tokens within the large-scale language model by using an ontology record that systematically organizes and records the meanings of vocabulary. [Background technology]
[0002] In recent years, there has been remarkable progress in artificial intelligence (AI), and in particular, large-scale language models (LLMs) that enable inquiries and responses in natural language through deep learning, which involves layering neural networks that mimic neural circuits and training them with large amounts of data, have attracted attention. Prior art documents relevant to this application include the following: [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Special Publication No. 2023-523644 [Patent Document 2] Japanese Patent Publication No. 2023-73095 [Patent Document 3] Patent No. 5484317 [Patent Document 4] Patent No. 7313757 [Patent Document 5] Patent No. 6868860 [Patent Document 6] Patent No. 6928332 [Patent Document 7] Patent No. 6913308 [Patent Document 8] Patent No. 6814482 [Patent Document 9] Patent No. 6792751 [Patent Document 10] Patent No. 7441391 [Patent Document 11] US 2022 / 0405484 A1 [Patent Document 12] US 2016 / 0026441 A1 [Patent Document 13] Japanese Patent Application Laid-Open No. 2005-157690 [Patent Document 14] Japanese Patent Application Laid-Open No. 2002-342146 Summary of the Invention [Problem to be solved by the invention]
[0004] In large-scale language models, to represent a certain vocabulary, a one-hot vector is used, which is a long vector consisting of zeros with the same number of dimensions as the number of types of vocabulary used, with a single 1 placed in the position corresponding to the vocabulary in question. All the vocabulary in a large amount of literature is replaced with vectors of this format, and deep learning is used to determine the correlation between each vocabulary vector.In response to a query (prompt), vocabulary that is likely to appear next to the query and the answer sentences that have already been generated is generated and added one by one to create an answer sentence. In this way, in conventional large-scale language models, the original meaning of the vocabulary used in the language model, query sentences, and response sentences is completely ignored as understood by humans.
[0005] The following problems arise when vocabulary is separated from its original meaning: (1) The vocabulary in the answers ignores the various attributes of the vocabulary in question, making it difficult to form a realistic image with a sense of reality (symbol grounding problem). (2) The vocabulary indicated in the answer sentence is merely the most likely next word in the previous vocabulary string, and there is no guarantee that it is semantically optimal for humans to understand. Even if we want to consider the next most likely vocabulary candidate, it is difficult to do so because the original meaning is ignored. (3) Questions such as "How are nuclear weapons manufactured?" or answers that encourage discrimination are becoming a problem. Currently, these are handled manually by humans, but it is difficult to provide a simple and complete response. (4) When using a large-scale language model to consider how to respond in a situation such as using a robot to remove a time bomb from a room containing one, there are an almost infinite number of items to consider, and the ``frame problem'' remains unsolved; the bomb may explode while the robot is still considering them. (5) When making logical inferences from one proposition to another that may seem unrelated at first glance, if there are no similar examples in the training data of a large-scale language model, the inferences are likely to be difficult or incoherent.
[0006] The present invention has been made to solve these conventional problems, and its purposes are to enable the user who made the inquiry to easily obtain a realistic and tangible image from the content of the reply, to enable comparison and verification of multiple possible reply statements rather than just one, to efficiently suppress inquiry statements and reply statements that use inappropriate vocabulary, to limit the content to be considered as much as possible without affecting the quality of the reply statement, and to enable logical bridging between propositions that are not directly related. [Means for solving the problem]
[0007] As a means for achieving the above-mentioned object, the ontology-assisted large-scale language model described in claim 1 is characterized in that, in a large-scale language model (LLM) that generates an answer sentence in response to a query sentence entered in a prompt, it comprises an ontology recording means that records the relational links between each vocabulary along with the attributes of the vocabulary, an ontology query-related part extraction means that extracts the query-related part from the recorded content of the ontology recording means, and an ontology recorded content conversion means that converts the extracted query-related part ontology into a form recognizable by the large-scale language model, and that it comprises at least one of (1) an ontology additional learning means that performs additional learning on the large-scale language model using the converted recorded content, and (2) an ontology query sentence insertion means that inserts the converted recorded content into a query sentence to the large-scale language model.
[0008] The ontology-assisted large-scale language model of claim 2 is characterized in that in the ontology-assisted large-scale language model of claim 1, the ontology recording means comprises a parent-child relationship vocabulary attribute inheritance means for inheriting the attributes of the parent vocabulary as attributes of the child vocabulary when the relationship link between the vocabulary is a parent-child relationship.
[0009] The ontology-assisted large-scale language model of claim 3 is characterized in that in the ontology-assisted large-scale language model of claim 1 or 2, the ontology recording means is provided with a divided vocabulary group by namespace, which divides vocabulary groups using namespaces by field and eliminates interference between vocabulary fields.
[0010] The ontology-assisted large-scale language model of claim 4 is characterized in that, in the ontology-assisted large-scale language model of claim 1 or 2, it is provided with a vocabulary ontology reference means that references the attributes of vocabulary contained in the query sentence or answer sentence, and further, if necessary, relational links, from the ontology recording means.
[0011] The ontology-assisted large-scale language model of claim 5 is characterized in that, in the ontology-assisted large-scale language model of claim 1 or 2, it comprises a step of comparing and referencing attributes of vocabulary included in the answer sentence with multiple vocabulary candidates other than the vocabulary itself, and verifying the validity of vocabulary selection.
[0012] The ontology-assisted large-scale language model of claim 6 is characterized in that, in the ontology-assisted large-scale language model of claim 1 or 2, it is provided with an inappropriate vocabulary suppression means for verifying whether or not there are any inappropriate attributes in the vocabulary included in the query sentence or the answer sentence in the recorded contents of the ontology recording means, and suppressing them.
[0013] The ontology-assisted large-scale language model of claim 7 is characterized in that, in the large-scale language model with ontology of claim 1 or 2, when inferring a target proposition from a starting proposition in the query sentence, the ontology-assisted large-scale language model further comprises a proposition-related vocabulary group information addition means which, for the vocabulary groups of the starting proposition, the target proposition, or both, expands vocabulary groups related to the vocabulary groups in the recorded contents of the ontology recording means, searches for vocabulary groups related to both the starting proposition and the target proposition, and adds the related vocabulary groups to the query sentence in the ontology query sentence insertion means.
[0014] The ontology-assisted large-scale language model of claim 8 is characterized in that, in the ontology-assisted large-scale language model of claim 1 or 2, when a proposition to be solved and an explanation of the situation surrounding the proposition are presented as a query statement in the query statement, the model is provided with a vocabulary proximity evaluation means that evaluates proximity between a vocabulary group in the query statement and a vocabulary group recorded in the ontology recording means using either logical proximity, spatial proximity, or temporal proximity, or a combination of these, and highly evaluates vocabulary groups in the ontology record that have high proximity, and that the highly evaluated vocabulary group is inserted into the query statement as the ontology query statement insertion means.
[0015] The ontology-assisted large-scale language model of claim 9 is characterized in that in the ontology-assisted large-scale language model of claim 1 or 2, the ontology additional learning means and the ontology record content conversion means in the ontology query statement insertion means are provided with at least one of a JSON format conversion means for converting into JSON format and an XML format conversion means for converting into XML format. [Effects of the Invention]
[0016] The ontology-assisted large-scale language model according to claim 1 includes ontology recording means, which records the relational links between each vocabulary along with the attributes of the vocabulary. The ontology query related portion extracting means is provided, and the query related portion is extracted from the recorded contents of the ontology recording means. The ontology record content conversion means converts the extracted query-related partial ontology into a form that can be recognized by the large-scale language model. The ontology record content conversion means converts any part of the record content of the ontology record means into a format that can be recognized by the large-scale language model. The ontology additional learning means realizes the additional learning function. The ontology query insertion means provides the ability to insert the transformed record content into a query to a large-scale language model.
[0017] The ontology-assisted large-scale language model according to claim 2 is provided with a parent-child relationship vocabulary attribute inheritance means, so that when the relational link between vocabularies is a parent-child relationship, the attributes of the parent vocabulary are inherited as attributes of the child vocabulary.
[0018] The ontology-assisted large-scale language model according to claim 3 includes a vocabulary group divided by namespace, and thus the vocabulary group is divided using a namespace by field, eliminating interference between vocabulary fields.
[0019] The ontology-assisted large-scale language model according to claim 4 includes a vocabulary ontology reference means, and therefore references vocabulary attributes and, if necessary, relational links from the ontology recording means.
[0020] The ontology-assisted large-scale language model according to claim 5 includes a step of verifying the validity of vocabulary selection, whereby attributes are compared and referenced with a plurality of vocabulary candidates other than the vocabulary to verify the validity of vocabulary selection.
[0021] The ontology-assisted large-scale language model according to claim 6 includes an inappropriate vocabulary suppression means, which verifies and suppresses the presence of inappropriate attributes in the vocabulary in the recorded content of the ontology recording means.
[0022] The ontology-combined large-scale language model as set forth in claim 7 is provided with a proposition-related vocabulary group information adding means, so that when inferring a target proposition from a starting proposition, for the vocabulary groups of the starting proposition, the target proposition, or both, the vocabulary groups related to the vocabulary groups in the recorded contents of the ontology recording means are expanded, vocabulary groups related to both the starting proposition and the target proposition are searched for, and the related vocabulary groups are added to the query statement by the ontology query statement insertion means.
[0023] The ontology-assisted large-scale language model of claim 8 is equipped with a vocabulary proximity evaluation means, so that when a proposition to be solved and a description of the situation surrounding the proposition are presented as a query statement, the proximity between the vocabulary groups in the query statement and the vocabulary groups recorded in the ontology recording means is evaluated using either logical proximity, spatial proximity, or temporal proximity, or a combination of these, and vocabulary groups in the ontology record that have high proximity are highly evaluated. Then, the vocabulary group that has been highly evaluated is inserted into the query statement by the ontology query statement insertion means.
[0024] The ontology-assisted large-scale language model according to claim 9 includes a JSON format conversion means, and converts ontology record contents into JSON format. It also has an XML format conversion means, which converts ontology record contents into JSON format. [Brief explanation of the drawings]
[0025] [Figure 1] 1 is a diagram showing the hardware configuration of the present invention. [Figure 2] This is an example of the screen layout when using a large-scale language model. [Figure 3] FIG. 1 is an explanatory diagram showing the relationship between classes and instances in an example of an ontology, in which attributes of a parent vocabulary are inherited by attributes of a child vocabulary in parent-child relationships. [Figure 4] This indicates that the attributes of a vocabulary consist of attributes inherited from the parent vocabulary and attributes newly added in the vocabulary. [Figure 5] The recorded contents of the ontology are converted into JSON format. [Figure 6] The recorded contents of the ontology are converted into XML format. [Figure 7] Indicates that the description of an attribute in an ontology refers to and uses a vocabulary defined in another ontology. [Figure 8] This is an example of the "disease name" ontology. [Figure 9] This is an example of the structure of the "disease name," "symptoms and findings," and "drug" ontology. [Figure 10] Taking the "Diabetes" class as an example, the following shows parent-child relationship links, inherited / added attributes for each class and instance, and ontology vocabulary referenced by each attribute. [Figure 11] FIG. 10 is an explanatory diagram of a reference link to a vocabulary defined in another ontology. [Figure 12] This is an explanatory diagram showing how namespaces are used in ontologies to prevent vocabulary interference between different fields. [Figure 13] 10 is a configuration example of a master table that manages namespace groups. [Figure 14] 10 is a diagram showing an example of the configuration of a master table for managing ontologies. [Figure 15]10 is a diagram illustrating an example of the configuration of a master table for managing ontology vocabulary attributes. [Figure 16] 10 is a diagram showing an example of the configuration of a master table of an ontology vocabulary management means. [Figure 17] This is an example of a vocabulary reference mechanism within an ontology record. It shows that matching vocabulary generated by a large-scale language model with the corresponding ontology instances can help solve the symbol grounding problem for users. [Figure 18] FIG. 1 is an explanatory diagram of one-hot vectors used in a large-scale language model. [Figure 19] FIG. 10 is an explanatory diagram of a vocabulary selection validity verification unit. [Figure 20] FIG. 10 is an explanatory diagram of an inappropriate vocabulary suppression means. [Figure 21] FIG. 10 is an explanatory diagram of a proposition-related vocabulary group adding means. [Figure 22] FIG. 10 is an explanatory diagram of a lexical proximity assessment means. DETAILED DESCRIPTION OF THE INVENTION
[0026] FIG. 1 is a diagram showing an example of a hardware configuration according to the present invention. Because large-scale language models require huge amounts of data and computational resources, they are usually stored on the cloud and connected to a company's LAN (Local Area Network) via a router over the Internet. The company has a business server. Company staff use business servers and large-scale language models via terminals connected to a LAN. It is also possible to build part or all of a business system on the cloud, or in the case of a lightweight version, to install part or all of a large-scale language model within a company. It is also possible to use large-scale language models via the Internet using mobile devices.
[0027] Figure 2 shows an example of a user interface for a large-scale language model (LLM). Development of LLM is currently progressing rapidly, with numerous models being developed, including ChatGPT (a registered trademark of OpenAI), Bard, LaMDA (a registered trademark of Google), and LLaMA (a registered trademark of Meta). Naturally, the user interface will differ, but typically, as shown in Figure 2, it consists of a box for entering prompts to instruct or inquire of the LLM, a box for displaying the answers to those prompts, and a box for displaying the history of prompts and answers as a usage log.
[0028] FIG. 3 is an explanatory diagram showing how, in vocabularies in a parent-child relationship, the attributes of a parent vocabulary are inherited by the attributes of a child vocabulary, using "organisms" as an example of an ontology. "Living things" are divided into "plants" and "animals," and "animals" are further divided into "mammals" and "birds," etc. The attributes of the parent vocabulary, such as the common attributes of "mammals" (warm-bloodedness, having hair on the skin, having supporting front legs, having supporting hind legs), are inherited by the child words "human," "dog," "cat," etc. In humans, the inherited attributes of the supporting front legs are overwritten by the free front legs. By doing this, the common part of the attributes of the child vocabulary can be grouped as the attribute of the parent vocabulary, which has the advantage of saving recording capacity and making the relationship between vocabulary clearer. However, not all attributes of the parent vocabulary are inherited by the child vocabulary. In ontology, when all "humans" are included in "mammals," such as when "humans" are "mammals," this is called an is-a relationship, and attribute inheritance occurs. However, attributes such as Person A owns a car (has-a), Person B is the president of a company (role-of), and Person C is a member of Company D (a-part-of) are limited to the vocabulary in question, and are not inherited by child vocabularies.
[0029] Figure 4 shows that the attributes of a vocabulary consist of attributes inherited from the parent vocabulary and attributes newly added in the vocabulary. For example, looking at the attributes of a "cat," it inherits the attributes of its parent vocabulary, "mammal" (warm-bloodedness, having hair on the skin, having supporting front legs, having supporting hind legs), and then adds attributes such as meows and movements via text, audio files, videos, etc. These files may be actual data stored on the LLM server, or they may be URLs to video software content. Also, "cat" may be the parent vocabulary, and individual cats such as "Mike" and "Tama" may be child vocabulary words.
[0030] FIG. 5 shows the recorded contents of the ontology shown in FIGS. 3 and 4 converted into JSON format (JSON format conversion means). FIG. 6 shows the recorded contents of the ontology shown in FIGS. 3 and 4 converted into XML format (XML format conversion means). As shown in Figures 3 and 4, an ontology is composed of vocabulary and relational links, and therefore cannot be used as is for training large-scale language models or for query statements. This requires converting any part of an ontology record into a format such as JSON or XML.
[0031] Using a large-scale language model as a base model, ontology records in a field of interest can be converted into JSON or XML format and subjected to additional learning, thereby constructing a large-scale language model specialized for that field (ontology additional learning means). Furthermore, by inserting related ontology record contents into a query, a more appropriate answer can be obtained (ontology query insertion means). It is also possible to express the data as a list of individual sentences, such as "There are two types of living things: animals and plants" or "One type of animal is mammals," without using a structured representation such as JSON or XML. This representation is also included in the present invention, but it is redundant, the structure is not clearly visualized, and it is not a desirable embodiment.
[0032] The relationships between vocabulary words can be obtained using large-scale language models themselves, but this requires training on a large amount of text, and it is not easy to accurately and concisely represent complex hierarchical structures. In contrast, when ontology records are used, links between vocabularies are directly specified, making it possible to easily express deep hierarchical structures without requiring extensive learning. Furthermore, in a hierarchical vocabulary structure, common attributes between child vocabulary elements can be grouped and aggregated into parent vocabulary elements, which allows for significant savings in storage capacity and makes the relationships between vocabulary elements visible, making them easier to understand.
[0033] FIG. 7 shows that in an ontology, vocabulary attributes are described by referring to vocabulary defined in another ontology. Attributes are described for each vocabulary, but the vocabulary used to describe the attributes is itself referenced from vocabulary defined in another ontology. In this way, the meanings of the vocabulary and the attributes are all defined in one of the ontologies, making it possible to express them unambiguously. Furthermore, using the ontologies in which the referenced vocabulary is defined can potentially provide a greater breadth of meaning and nuance to the vocabulary.
[0034] Figure 8 is an example of the "disease name" ontology.
[0035] Figure 9 shows an example of the structure of the "disease name," "symptoms / findings," and "drug" ontologies. The attributes of each disease name include the symptoms and test findings observed with the disease, and the drugs used to treat the disease. However, these symptoms and drug names are defined in the "symptoms / findings" ontology and the "drug" ontology, which are separate from the disease name ontology, and the attributes of the disease name are described as references to vocabularies defined in the separate ontologies. Although this diagram only lists the "symptoms / findings" and "drug" ontologies, other ontologies such as "examination," "treatment," and "nursing" are also used.
[0036] Figure 10 takes the "diabetes" class as an example and shows the parent-child relationship links between the parent vocabulary terms "metabolic disease" and "diabetes" itself, and the child vocabulary terms "type I diabetes" and "type II diabetes," as well as the inheritance / addition of attributes and reference links to ontology terms referenced by each attribute. The <similar diseases> attribute of the disease in question references another disease name vocabulary within the same "disease name" ontology. It also shows reference links to cases of the disease in question.
[0037] FIG. 11 is an explanatory diagram of a reference link to a vocabulary defined in another ontology. In Figure 10, "hyperglycemia" in the <pathological condition> of "diabetes" has a reference link to "hyperglycemia," a vocabulary in the "symptoms / findings" ontology. The <definition> attribute of the "hyperglycemia" vocabulary in the "symptoms / findings" ontology describes blood glucose level >140 mg / dl and HBa1c >7.0.
[0038] FIG. 12 is an explanatory diagram showing how namespaces are used in ontology to prevent vocabulary interference between different fields. Different fields use different vocabulary, and often the same words are used with completely different meanings in different fields. In order to prevent unexpected interference caused by newly defined vocabulary in a certain field, namespaces are used to separate vocabularies for each field, allowing new vocabulary to be defined without considering the impact on other fields. In Figure 12, the term "fatigue" is used in various ways, such as "adrenal fatigue," "institutional fatigue," and "fatigue due to long working hours," but by separating them in namespaces, interference can be avoided. It goes without saying that the newly defined vocabulary does not interfere with existing vocabulary in the same name space.
[0039] Figure 13 shows an example of the configuration of a master table that manages a group of namespaces. An ID is assigned to each namespace, and records are managed using this ID.
[0040] Figure 14 shows an example of the configuration of a master table for managing ontologies. Within the namespace, IDs are assigned to ontologies and records are managed.
[0041] Figure 15 shows an example of the configuration of a master table for managing ontology vocabulary attributes. IDs are assigned to the attributes of vocabulary used in an ontology to manage vocabulary attributes. Note that vocabulary included in the same ontology share common vocabulary attributes.
[0042] Figure 16 shows an example of a vocabulary management means. Along with the vocabulary ID# and the vocabulary, the namespace ID# to which the vocabulary belongs, the ontology ID#, and the path to the vocabulary within the ontology are recorded. For each vocabulary term included in a query, this diagram is used to search the ontology and extract the term and its attributes (query-related partial ontology extraction means). The term also inherits the attributes of the higher-level vocabulary term (attribute inheritance). If the vocabulary term alone is insufficient, appropriate sibling vocabulary groups, such as "glycogen storage disease" in the case of "diabetes" in Figure 8, and child vocabulary groups, such as "type I diabetes" and "type II diabetes," may also be extracted as query-related partial ontologies. The query-related partial ontology extracted in this way is converted into a form recognizable by the large-scale language model using either a JSON format conversion means for converting it into JSON format or an XML format conversion means for converting it into XML format, as shown in Figures 5 and 6 (ontology record content conversion means), and is inserted into a query statement for the large-scale language model (ontology query statement insertion means). Furthermore, instead of inserting the ontology into the query sentence, it is also possible to additionally train the recorded contents of the ontology into a large-scale language model (ontology additional training means). Although the entire recorded contents of the ontology can be additionally trained, since the recorded contents of the ontology are enormous in volume, it is also possible to extract a portion of the ontology and additionally train it depending on the purpose.
[0043] FIG. 17 is an explanatory diagram of one-hot vectors used in large-scale language models. In a large-scale language model, for each vocabulary word, a one-hot vector is assigned in which only the part of the vocabulary word is 1 and the rest is 0, and the corresponding token# is assigned (vocabulary token vector). The number of dimensions of the vector is the total number of vocabulary words. For the documents in the training data, an operation is performed to replace each vocabulary word with the corresponding vocabulary token vector. For vocabulary token vectors of a huge amount of document data, learning is performed by solving a so-called hole problem, in which a vocabulary token vector is estimated from a group of surrounding vocabulary token vectors. Based on the correlation between vocabulary token vectors, the probability of the missing vocabulary is calculated, and the maximum vocabulary is taken as the estimated vocabulary. Previously, supervised learning required the preparation of a large amount of pairs (corpora) of training data and their correct answers (teaching data), which was time-consuming and costly. In contrast, with learning from the above-mentioned gap-in-the-blank problem, the hidden vocabulary itself becomes the correct teaching data, making it possible to use a large amount of corpora all at once, significantly improving learning accuracy.
[0044] All meanings of vocabulary are ignored and the vocabulary is treated as a vocabulary token vector where all vocabulary tokens are 0 and only one 1. For example, when a large-scale language model outputs the vocabulary word "cat," it simply means that the probability of that vocabulary token vector is higher than others; it does not represent an understanding of the specific image of "cat" that humans have. In other words, a known limitation of large-scale language models is that they do not understand the relationship between the vocabulary they output and its concrete image (meaning) (the symbol grounding problem). In the present invention, the vocabulary token vector "cat" is associated with the vocabulary "cat" in the ontology, making it possible to use the abundant information recorded in the ontology (vocabulary ontology referencing means). In other words, it is an attempt to give meaning to the output of a large-scale language model. Large-scale language models are software within a computer and have no physicality. For this reason, it is not possible to experience the feel and smell of a cat when petting it. However, for people who use large-scale language models, the information, audio, video, etc. contained in the vocabulary within the ontology can be understood with a sense of reality, combined with their own past experiences. Therefore, it seems that the symbol grounding problem can be solved practically by combining the three elements of a large-scale language model, an ontology, and users.
[0045] FIG. 18 is an example of a vocabulary reference means within an ontology record. We further demonstrate that matching the vocabulary generated by a large-scale language model with the vocabulary of an ontology leads to a solution to the symbol grounding problem for users. When a doctor encounters the vocabulary term "hyperglycemia" in "diabetes" while using a large-scale language model, he or she refers to the vocabulary record in the ontology as shown in Figure 17. The relevant vocabulary in the ontology is "diabetes" in the "metabolic system" and "glucose metabolism system" of the "disease name" ontology in the namespace "medical care", and by referencing "hyperglycemia" in the "symptoms / findings" ontology from the attribute "hyperglycemia" in <pathological condition>, it is possible to obtain the information that blood glucose is 140 mg / dl from the attribute <definition> of "hyperglycemia". This vocabulary reference means within the ontology record may be referenced visually by a human, or may be referenced by other software or other large-scale language model via an API or the like and utilized in the processing.
[0046] FIG. 19 is an explanatory diagram of the vocabulary selection validity verification means. Let's say a large-scale language model outputs "pneumonia" based on a patient's symptoms and findings. However, the probabilities of the output vocabulary words such as "bronchopneumonia" and "lung cancer" are also relatively high, and the validity of "pneumonia," which had the highest probability at that time, is not overwhelmingly valid. By having the large-scale language model present not only the vocabulary with the highest probability, but also multiple vocabulary words with the next highest probability, and displaying the content of each using the vocabulary reference means within the ontology record, the person operating the model (in this case, a doctor) can verify the possibility of vocabulary words for other disease names in what is called backward inference in an expert system, and select a more appropriate vocabulary word for a disease name.
[0047] FIG. 20 is an explanatory diagram of the inappropriate vocabulary suppression means. Inappropriate vocabulary usage is a problem in the operation of large-scale language models. Dangerous queries such as "Tell me how to make a nuclear bomb" or "Tell me how to kill people efficiently" should not be answered. Large-scale language models use user interactions as material for additional learning, but there have been cases where some users have posted large amounts of praise for the Nazis and slander against specific races, causing the large-scale language model that learned from these posts to generate inappropriate replies, which caused major problems and forced the service to be discontinued. To address this issue, the current practice is to mobilize large numbers of workers to manually remove inappropriate vocabulary and answer sentences, which is extremely expensive, increases the cost of providing services, and delays the launch of services. To solve this problem, the present invention provides a <sensitive flag> as an attribute of the ontology vocabulary.
[0048] By setting a flag in the <sensitive flag> attribute of ontology vocabulary as shown in Figure 20, it is possible to pre-screen sensitive vocabulary and vocabulary that may cause problems contained in queries, answers, and additional training data. Furthermore, by using ontology records, it is possible to include vocabulary in the related links of inappropriate vocabulary in the screening of inappropriate vocabulary. This makes it possible to screen not only inappropriate vocabulary, but also similar or related vocabulary. This makes it possible to suppress, remove, or draw attention to sensitive vocabulary in inquiries, responses, and additional training data. Even when screening by human eyes, it is possible to significantly improve efficiency by focusing on sensitive vocabulary. Furthermore, if a scale of sensitivity (sensitivity) is used for the <Sensitivity Flag>, various responses can be made depending on the sensitivity, ranging from immediate refusal to issuing a warning. The <Sensitive Flag> can also be set for multiple words, making it possible to increase the sensitivity of combinations such as "kill" or "hurt" for "knife."
[0049] FIG. 21 is an explanatory diagram of the proposition-related vocabulary group information adding means. In one children's story, an old woman living in a mountain village dies, and the neighborhood ladies are boiling her body in a pot. When the teacher asks, "What are the ladies boiling?", quite a few children answer, "They are boiling the old woman's body to disinfect it," which has become a hot topic. This shows that the reader makes a mistaken inference because there is no direct connection between "the old lady has passed away" and "the neighborhood ladies are cooking a pot." In this invention, by adding related vocabulary information recorded in the ontology, such as (a) "home" in the <place of death> attribute of the ontology vocabulary "death" and "home" in the ontology attribute <funeral>, and (b) "funeral" in the <collaboration> attribute of the ontology vocabulary "life in the countryside", to the original text describing the situation in (c) (ontology query sentence insertion means), and adding proposition-related vocabulary group information such as "The funeral is held at home" and "Funerals in the countryside are a collaborative effort among neighbors", and then querying the large-scale language model, it is possible to expect the generation of a correct answer sentence such as (d) "She is probably making boiled vegetables to be served at the funeral".
[0050] When humans read a text, they not only pay attention to the text itself, but also unconsciously supplement it with related common sense that is not expressed in the text to gain a correct understanding. If there is a direct reference in a document trained with a large-scale language model, inferences like (d) are possible, but this requires a huge amount of training data and the computational resources and power required to process it. In the present invention, by recording so-called "common sense" in an ontology and adding proposition-related vocabulary group information as appropriate, highly accurate inference becomes possible using a large-scale language model with a small amount of training data. Proposition-related vocabulary group information can be added to any of the vocabulary of the situation description, the inquiry sentence, and the answer sentence.
[0051] FIG. 22 is an explanatory diagram of the lexical proximity evaluation means. There is a time bomb on the desk in the room, and the time to explode is approaching. If we don't remove it quickly, it will explode. If we try to solve this problem with a large-scale language model, we will take into account too much background information that is not necessary for solving the problem, such as "What is the structure of the ceiling and floor?", "What is the temperature of the room?", "What is the material of the desk?", etc., and while we are thinking about various things, time will run out and the model will explode. This problem is known as the "frame problem," and the key is how to extract the information necessary to solve the problem and solve it quickly. In this invention, we propose to introduce lexical proximity evaluation to eliminate as much background information as possible that does not contribute to the problem to be solved, and to focus on highly relevant information to achieve rapid problem solving.
[0052] Among the ontology records in Figure 21, clock, bomb, desk, etc. are important because they directly relate to the situation, but banana, music, etc. are hardly related. To evaluate these vocabularies, we use tools such as word2vector to obtain word embeddings of the meaning of each vocabulary. The cosine (cos) between these vectors represents the correlation coefficient between the two, so only vocabulary with a large correlation coefficient is evaluated as having high lexical proximity and taken into account. Proximity of location is also important. The desk and the bomb are not closely related in meaning, but they are close in the information in Figure 21. Los Angeles, Paris, etc. are not related at all. Furthermore, temporal proximity is also important. Vocabulary contained in information related to recent minutes, hours, and dates can be said to have a high degree of proximity, but vocabulary contained in descriptions of situations from 100 years ago has a low degree of proximity and can be said to have little relevance. Lexical proximity evaluation should take into account at least the semantic proximity, location proximity, and temporal proximity of vocabulary. At least one of semantic proximity, location proximity, and time proximity, or any combination of these, is possible. By limiting the vocabulary to terms with high proximity and using ontology query insertion to query a large-scale language model, it is possible to obtain answers within a practical time frame.
[0053] Although the embodiments have been described above, the specific configuration of the present invention is not limited to the above-described embodiments, and the present invention also includes design changes and the like that do not deviate from the gist of the invention. For example, ChatGPT and the like have been given as examples of large-scale language models, but the same applies to other systems, and the present invention also includes the use of new systems that will be developed in the future as machine learning in general.
Claims
1. In a large-scale language model (LLM) that generates an answer sentence in response to a query sentence entered in a prompt, An ontology-assisted large-scale language model comprising: an ontology recording means for recording relational links between vocabulary terms together with vocabulary attributes; an ontology query-related part extraction means for extracting the query-related part from the recorded content of the ontology recording means; and an ontology recorded content conversion means for converting the extracted query-related part ontology into a form recognizable by the large-scale language model, wherein the converted recorded content is used to provide at least one of: (1) ontology additional learning means for additionally learning the large-scale language model; and (2) ontology query statement insertion means for inserting the ontology query statement into a query statement for the large-scale language model.
2. 2. The ontology-assisted large-scale language model according to claim 1, further comprising a parent-child relationship vocabulary attribute inheritance means for inheriting the attributes of a parent vocabulary as attributes of a child vocabulary when the relationship link between each of the vocabulary terms is a parent-child relationship in the ontology recording means.
3. In the ontology recording means, 3. The ontology-assisted large-scale language model according to claim 1, further comprising a vocabulary group divided by namespace, which is divided using a domain-specific namespace to eliminate interference between vocabulary domains.
4. 3. The ontology-assisted large-scale language model according to claim 1, further comprising a vocabulary ontology reference means for referencing the attributes of vocabulary contained in the query sentence or the response sentence, and furthermore, as necessary, relational links, from the ontology recording means.
5. In the vocabulary contained in the answer, 3. The ontology-assisted large-scale language model according to claim 1, further comprising a step of comparing and referencing attributes of a plurality of vocabulary candidates other than the selected vocabulary to verify the validity of the vocabulary selection.
6. In the vocabulary contained in the inquiry sentence or the reply sentence, 3. The ontology-assisted large-scale language model according to claim 1, further comprising an inappropriate vocabulary suppression means for verifying whether or not there is an inappropriate attribute in the vocabulary recorded in the ontology recording means and suppressing it.
7. In the query, 3. The ontology-assisted large-scale language model according to claim 1, further comprising a proposition-related vocabulary group information addition means for, when inferring a target proposition from a starting proposition, expanding vocabulary groups related to the vocabulary groups of the starting proposition, the target proposition, or both, in the recorded contents of said ontology recording means, searching for vocabulary groups related to both the starting proposition and the target proposition, and adding the related vocabulary groups to the query statement in said ontology query statement insertion means.
8. In the query, 3. The ontology-assisted large-scale language model according to claim 1 or 2, further comprising a vocabulary proximity evaluation means for evaluating, when a proposition to be solved and an explanation of the situation surrounding the proposition are presented as a query statement, proximity between a vocabulary group in the query statement and a vocabulary group recorded in the ontology recording means, either logical proximity, spatial proximity, or temporal proximity, or a combination thereof, and highly evaluating vocabulary groups in the ontology record that have high proximity, and wherein the highly evaluated vocabulary groups are inserted into the query statement as the ontology query statement insertion means.
9. The ontology-assisted large-scale language model according to claim 1 or 2, characterized in that the ontology record content conversion means in the ontology additional learning means and the ontology query statement insertion means comprises at least one of a JSON format conversion means for converting to JSON format and an XML format conversion means for converting to XML format.
Citation Information
Patent Citations
Intelligent trademark law question and answer method for enhancing big language model reasoning through knowledge graph
CN117807202A
Method and device for monitoring electronic bulletin board information on homepage, method and device for monitoring alteration inhibition information, method and device for monitoring particular term or particular sentence
JP2002342146A
Electronic equipment and information providing method
JP2005157690A
Knowledge management system
JP2020194520A
Recursive ontology-based systems engineering
US20160026441A1