A method for evaluating the ability of large language models to judge the relationship between hyponyms and hyponyms

By designing different prompt words and building classification diagrams, we evaluate the ability to judge the upper and lower relationships of large language models, and solve the problem of lack of special evaluation methods in the existing technology, and improve the recognition accuracy and generalization ability.

CN119397228BActive Publication Date: 2025-06-06NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510015443.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-06-06
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

The existing technology lacks a special evaluation method for judging the upper and lower relationships of large language models, resulting in limitations in accuracy and generalization capabilities when identifying and utilizing upper and lower relationships.

Method used

By designing different prompt words, combining the ability to judge the relationship between the upper and lower relationships and external knowledge of the large language model, a classification diagram and test cases are constructed to evaluate the ability to judge the relationship between the upper and lower relationships of the large language model.

Benefits of technology

Improve the accuracy and generalization ability of large language models in upper and lower relationship recognition, and overcome the limitations of existing methods in understanding context nuances and cross-domain applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119397228B_ABST
    Figure CN119397228B_ABST
Patent Text Reader

Abstract

The invention discloses an evaluation method for the ability of a large language model to judge the relationship between a word and a sub-word, and belongs to the field of natural language processing. The method comprises: constructing a classification diagram according to a plurality of data sets containing the relationship between a word and a sub-word; constructing a test case according to the classification diagram and the structural information in the data set; combining the test case with a plurality of pre-designed prompt words, inputting the test case into the large language model to be tested, and obtaining a return result; and evaluating the ability of the large language model to judge the relationship between a word and a sub-word according to the return result of the large language model. The invention fully combines the ability of the large language model to understand the meaning of words in different contexts, and overcomes the shortcomings of the previous method. The method not only evaluates the ability of the large language model to judge the relationship between a word and a sub-word by designing different prompt words, but also improves the reasoning ability of the large language model and the recognition accuracy of the relationship between a word and a sub-word by injecting external knowledge.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for evaluating the superior-female relationship judgment ability of a large language model, and belongs to the field of natural language processing. Background Art

[0002] Among many linguistic concepts, the hierarchical relationship is a basic semantic relationship that expresses the inclusion relationship between things. It is crucial to many fields such as semantic understanding, knowledge representation, and information retrieval. For example, in the field of natural language processing, the hierarchical relationship can promote contextual understanding and help the model understand the specific meaning of words or phrases in different contexts; when building knowledge graphs or ontologies, the hierarchical relationship provides a hierarchical framework for concepts and entities. This structure is not only conducive to human understanding and organization of knowledge, but also facilitates computers to perform efficient information retrieval and logical reasoning; in information retrieval, the hierarchical relationship can be used for query expansion, that is, automatically including broader or more specific related information based on the user's initial query to increase the relevance and coverage of the returned results. It can be seen that accurately identifying and utilizing the hierarchical relationship can effectively improve the performance and accuracy of many applications.

[0003] Existing methods for identifying hyponymy and hyponymy are mainly divided into template-based methods, distribution-based methods, and hybrid methods. Template-based methods rely on predefined language patterns (such as vocabulary templates) to identify hyponymy and hyponymy. Although such methods have a relatively high accuracy in identifying hyponymy and hyponymy, the generalization of the pattern is poor. New or uncommon expressions may not be captured by existing patterns, and the process of formulating patterns is time-consuming and error-prone. Distribution-based methods use the co-occurrence information of words in large-scale text corpora to identify hyponymy and hyponymy. According to the "distribution hypothesis", semantically similar or related words have similar distribution patterns in text. Although such methods perform well in the fields covered by the training data, their generalization ability may be limited when encountering new fields or tasks that are significantly different from the distribution of the training data, and they are still limited in understanding subtle semantic differences in specific contexts. Hybrid methods refer to the fusion of template-based methods and distribution-based methods. Although hybrid methods aim to reduce the dependence on large-scale annotated data, their performance still depends largely on the quality and diversity of available data, and when there is bias or insufficient coverage in the dataset used, the effect and generalization ability of the model may be limited.

[0004] With the rapid development of artificial intelligence technology, large language models (such as LLaMa large language models) have become the core of natural language processing technology, and also provide new methods for the identification of hierarchical relationships. Large language models are natural language processing models with large-scale parameters and computing power, designed to understand and generate human language. They have mastered rich language knowledge through pre-training on large-scale data sets, thus showing excellent performance in a variety of natural language processing tasks (such as machine translation, question-answering systems, text summarization, etc.). Large language models can also capture language patterns and semantic information from a large amount of text. Through pre-trained rich corpus and fine-tuning tasks, they can understand and extract hierarchical relationships, becoming an emerging method for extracting hierarchical relationships. Therefore, it is very important to select or develop a large language model suitable for extracting hierarchical relationships. In this process, first of all, there is a need for a comprehensive and accurate evaluation method of the large language model in terms of its ability to judge hierarchical relationships.

[0005] Existing large language model evaluation methods mostly focus on the language generation quality (such as fluency, coherence and diversity), task performance (such as question-answering, text generation and reasoning ability), knowledge and factuality (such as accuracy and knowledge coverage), and dialogue and interaction capabilities (such as context understanding and multi-round dialogue coherence) of large language models, but lack specialized evaluation of the ability to judge hierarchical relationships. Summary of the invention

[0006] The purpose of the present invention is to overcome the deficiencies in the prior art and to provide an evaluation method for the ability of a large language model to judge the hyponym and hyponym relationships. By designing different prompt words, the ability of the large language model to judge the hyponym and hyponym relationships is evaluated, and the reasoning ability of the large language model and the recognition accuracy of the hyponym and hyponym relationships are improved by injecting external knowledge.

[0007] To achieve the above object, the present invention is implemented by adopting the following technical solutions:

[0008] In a first aspect, the present invention provides a method for evaluating the ability of a large language model to judge the relationship between a high and a low level, comprising:

[0009] Construct classification graphs based on multiple datasets containing hyponymy;

[0010] Construct test cases based on the classification diagram and structural information in the data set;

[0011] Combine the test case with a variety of pre-designed prompt words, input them into the large language model under test, and obtain the return result;

[0012] Based on the results returned by the large language model, evaluate the large language model's ability to judge the relationship between superiors and subordinates.

[0013] Further, the data set includes a data set with concept definition, a data set with instance layer information, a data set with pattern layer information, and a data set with both instance layer and pattern layer information;

[0014] The dataset with concept definitions represents a hyponymy recognition dataset; it includes a taxo file, a terms file, and a term2def file, wherein each line in the taxo file describes a child node number and a parent node number thereof, each line in the terms file describes a node number and a concept corresponding to it, and each line in the term2def file is a concept and its definition from a corresponding corpus; the node numbers in the taxo file are replaced with the corresponding concepts in the terms file to construct concept pairs with hyponymy relationships;

[0015] The data set with instance layer information represents a knowledge graph containing upper and lower relationships. The knowledge graph is a structured data model that represents entities and their relationships through nodes and edges, and is used to organize and associate information to support intelligent analysis and applications, wherein nodes represent instances, concepts or specific values, and edges represent semantic relationships between nodes;

[0016] The data set with pattern layer information represents an ontology containing a hypernym relationship, and the ontology is used to describe the semantic relationship between concepts, attributes and instances; when selecting the hypernym and hyponym relationships for evaluation, hypernyms and hyponyms with pattern layer information definitions are selected, and the pattern layer information definitions include the non-intersection axioms of concepts, the value taking methods of concepts on certain attributes or the definition of constraints on the number of instances, and attributes associated with concepts;

[0017] The data set that has both instance layer and pattern layer information is a knowledge graph that contains both hyponymy and pattern layer information definitions, or an ontology that contains both hyponymy and instance layer information. When selecting hyponymy for evaluation, the hypernyms and hyponyms have both relevant instances and relevant pattern definitions.

[0018] Furthermore, the classification diagram is constructed based on a plurality of data sets containing upper and lower relationships, including:

[0019] Construct an empty directed graph and add nodes to the directed graph. Each node represents a hypernym or hyponym. Each hypernym or hyponym is a concept in the knowledge graph or ontology.

[0020] Add edges to the directed graph based on the hypernymy relationship extracted from the data set. If there is a hypernymy relationship between two nodes, add a directed edge from the hypernym to the hyponym.

[0021] After adding all the edges to the directed graph, we get a new classification graph;

[0022] The nodes in the classification graph are divided into root nodes, leaf nodes and non-leaf nodes. The root node refers to a node with only outgoing edges but no incoming edges, the leaf node refers to a node with only incoming edges but no outgoing edges, and the non-leaf node refers to a node with both incoming and outgoing edges.

[0023] Among them, the path node set anchors(q) consisting of all paths from leaf nodes to reachable root nodes is calculated as follows:

[0024] ;

[0025] Among them, q is a leaf node, p is a path, roots are all root nodes, r is a root node, pathNodes(q, p) represents the set of all nodes except q on a path p of leaf node q, and paths(q, r) represents all paths from q to r.

[0026] Furthermore, constructing a test case according to the classification diagram and the structural information in the data set includes:

[0027] Construct positive examples with a hierarchical relationship between concept pairs or node pairs tested according to the classification graph;

[0028] When constructing positive examples, first calculate the path node set for each leaf node in the classification graph, and then construct all possible parent-child node pairs for each leaf node. The formula is as follows:

[0029] ;

[0030] in, Represents all parent-child node pairs of leaf node q; anchor is any element in the set anchors(q), representing the set of all nodes on a path from node q to a root node that q can reach; sup is an element in anchor, representing a node;

[0031] AllNodesPairs is used to represent the set of parent-child node pairs consisting of all leaf nodes. The formula is as follows:

[0032] ;

[0033] in, Represents all leaf nodes in a directed graph; Represents a leaf node All parent-child node pairs of ;

[0034] Take allNodesPairs as all positive examples constructed;

[0035] According to the classification graph and the structural information in the data set, negative examples are constructed in which there is no upper-lower relationship between the concept pairs or node pairs tested, including:

[0036] According to the classification graph, find out the nodes with sibling relationships, and establish node pairs between every two nodes as negative examples;

[0037] For datasets with pattern layer information, we use the disjoint relationships in the structural information to construct negative examples.

[0038] Combine the constructed positive examples with the negative examples to obtain the constructed test cases.

[0039] Furthermore, the prompt words include prompt words that are directly expressed, prompt words that utilize concept definitions, prompt words that utilize structure, and prompt words that utilize both concept definitions and structure.

[0040] The direct expression of the prompt word indicates that for each test case (q m , a n ) Design prompt words that do not inject any external knowledge or use any information in the dataset, and directly let the large language model judge q m with a n Whether there is a hierarchical relationship, where for the test case (q m , a n ), if it is a positive example, it means a n Yes m The parent node of ; if it is a negative example, it means a n Not q m The parent node of

[0041] The hint words defined by the concept indicate that for each test case (q m , a n ) Add node definitions, where node definitions are divided into the following three types: definitions from the dataset itself; definitions from a large language model; definitions from external data resources;

[0042] The hint word of the utilization structure indicates that for each test case (q m , a n ) Consider q m with a n The structural information is as follows:

[0043] For a dataset with concept definitions, get q m with a n The concept hierarchy information of each node includes sibling nodes, father nodes, child nodes, ancestor nodes, and descendant nodes. For data sets with instance-level information, obtain q m with an Each of the associated instances, and then obtain other related instances or specific values ​​through these instances; for data sets with pattern layer information, obtain q m with a n The directly related pattern layer information, directly related means the appearance of q m or a n The pattern layer axiom of q m with a n Respectively related instance layer and pattern layer information;

[0044] The concept definition and structure prompt words are used simultaneously to indicate that for each test case (q m , a n ) Both inject q m with a n The definition of the node injects the corresponding structural information.

[0045] Furthermore, the return result of the tested large language model includes: a clear judgment on whether there is a superior-female relationship, and a basis for the judgment;

[0046] The returned results are divided into the following four situations:

[0047] For the prompt words in the positive example, the returned result correctly identifies the hyponym relationship and the corresponding explanation is also correct. All inputs with this situation are marked as a set. ;

[0048] For the negative example, the returned result does not identify the upper and lower relationship, and the corresponding explanation is also correct. All inputs with this situation are marked as a set. ;

[0049] Make a correct judgment on the hyponymy relationship, but the corresponding explanation is incorrect or incompletely correct. Mark all inputs with this situation as a set ;

[0050] The judgment of the upper and lower relationship is incorrect. All inputs with this situation are marked as a set. .

[0051] Furthermore, the step of evaluating the judgment capability of the large language model for the upper and lower relationships according to the return result of the large language model includes:

[0052] Evaluate the ability of the large language model under test to judge the relationship between the upper and lower words from the perspective of the correctness of judgment and the correctness of interpretation, and define the accuracy rate of identifying the relationship between the upper and lower words , the accuracy rate of judging the relationship between superior and subordinate , Error rate when judging the relationship between superior and subordinate and the probability of hallucinations when judging , the calculation formula is as follows:

[0053] ;

[0054] ;

[0055] ;

[0056] ;

[0057] in, and They represent the node pair sets of positive examples and the node pair sets of negative examples respectively. Representing a collection The number of elements in , and , The larger the value of , the more likely the corresponding large language model will hallucinate when judging the hyponym relationship.

[0058] In a second aspect, the present invention provides an evaluation device for the ability to judge the superior-subordinate relationship of a large language model, comprising:

[0059] A classification map construction module is used to construct classification maps based on a variety of data sets containing hierarchical relationships;

[0060] The test case construction module is used to construct test cases based on the classification graph and the structural information in the data set;

[0061] The model testing module is used to combine the test case with a variety of pre-designed prompt words, input them into the large language model under test, and obtain the return result;

[0062] The performance evaluation module is used to evaluate the judgment ability of the large language model on the upper and lower relationships according to the return results of the large language model.

[0063] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the aforementioned methods.

[0064] In a fourth aspect, the present invention provides a computer device, comprising:

[0065] Memory, for storing computer programs / instructions;

[0066] A processor is used to execute the computer program / instructions to implement the steps of any of the aforementioned methods.

[0067] In a fifth aspect, the present invention provides a computer program product, comprising a computer program / instruction, which implements the steps of any one of the aforementioned methods when executed by a processor.

[0068] Compared with the prior art, the present invention has the following beneficial effects:

[0069] The present invention provides an evaluation method for the ability of a large language model to judge the relationship between the upper and lower words. This method fully combines the ability of the large language model to understand the meaning of words in different contexts, including the subtle differences in the relationship between the upper and lower words, and the characteristics of the large language model with extensive cross-domain knowledge and strong generalization ability, thus overcoming the shortcomings of previous methods. In addition, this method not only evaluates the ability of the large language model to judge the relationship between the upper and lower words by designing different prompt words, but also improves the reasoning ability of the large language model and the recognition accuracy of the relationship between the upper and lower words by injecting external knowledge. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 It is a flow chart of a method for evaluating the ability of a large language model to judge the relationship between a hyponym and a hypernym provided by an embodiment of the present invention;

[0071] Figure 2 It is a schematic diagram of an evaluation device for the ability to judge the hyponym and hyponym relationships of a large language model provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0072] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.

[0073] Embodiment 1: This embodiment introduces a method for evaluating the ability of a large language model to judge the relationship between the upper and lower levels, including:

[0074] Construct a classification diagram based on multiple datasets containing hierarchical relationships; involving datasets with concept definitions, datasets with instance layer information, datasets with pattern layer information, and datasets with both instance layer and pattern layer information;

[0075] Construct test cases, including positive and negative examples, based on the classification graph and structural information in the data set;

[0076] Four types of prompt words are designed for testing large language models, including prompt words of direct expression, prompt words using concept definition, prompt words using structure, and prompt words using both concept definition and structure;

[0077] Combine the test case with a variety of pre-designed prompt words, input them into the large language model under test, and obtain the return result;

[0078] Based on the results returned by the large language model, evaluate the large language model's ability to judge the relationship between superiors and subordinates.

[0079] like Figure 1 As shown, the method for evaluating the ability of judging the hyponymy of a large language model provided in this embodiment specifically involves the following steps in its application process:

[0080] Step 1: Construct a classification map based on four data sets containing hierarchical relationships, including extracting hierarchical relationships that meet certain conditions from different types of data sets, and then constructing a classification map based on the hierarchical relationships.

[0081] In this embodiment, step 1 specifically includes the following steps:

[0082] Step 1.1: Extract the hierarchical relationships from four datasets containing hierarchical relationships, including datasets with concept definitions, datasets with instance layer information, datasets with pattern layer information, and datasets with both instance layer and pattern layer information.

[0083] The dataset with concept definitions refers to a common hyponym recognition dataset. For example, the MAG-WiKi dataset, the SemEval dataset, and the SE16 dataset come from different fields such as computer science, psychology, environment, and food. This type of dataset mainly includes taxo files, terms files, and term2def files. Each line in the taxo file describes a child node number and its parent node number, each line in the terms file describes a node number and its corresponding concept, and each line in the term2def file is a concept and its definition from the corresponding corpus. The extraction of hyponym relationships based on a dataset with concept definitions refers to the construction of concept pairs with hyponym relationships through taxo files and terms files, that is, replacing the node numbers in taxo with the corresponding concepts in the terms file.

[0084] The dataset with instance layer information refers to a knowledge graph containing hypernyms and hyponyms. A knowledge graph is a structured data model that represents entities and their relationships through nodes and edges, and is used to organize and associate information to support intelligent analysis and applications, where nodes can be instances, concepts, or specific values, and edges represent semantic relationships between nodes. Extracting hypernyms and hyponyms based on a dataset with instance layer information (i.e., a knowledge graph) means that when selecting hypernyms and hyponyms for evaluation, it is necessary to select hypernyms and hyponyms that have relevant instance definitions. The relevant instance definitions here include instances of hypernyms or hyponyms, as well as other instances or specific values ​​associated with these instances.

[0085] The data set with pattern layer information refers to an ontology containing hierarchical relationships. Ontology is a formal definition of knowledge in a certain field, which is used to describe the semantic relationship between concepts, attributes and instances. Extracting hierarchical relationships based on a data set with pattern layer information (i.e., ontology) means that when selecting hierarchical relationships for evaluation, it is necessary to select hypernyms and hyponyms with pattern layer information definitions. The pattern layer information definition here includes the non-intersection axioms of concepts, the definition of the way concepts take values ​​on certain attributes or the constraints on the number of instances, attributes associated with concepts, etc.

[0086] The data set having both instance layer and pattern layer information is a knowledge graph containing both the hypernymy and the pattern layer information definition, or an ontology containing both the hypernymy and the instance layer information. Extracting the hypernymy from the data set having both instance layer and pattern layer information means that when selecting the hypernymy for evaluation, the hypernym and the hyponym need to have both relevant instances and relevant pattern definitions.

[0087] Step 1.2: Construct a classification graph based on the hyponymy relationships extracted from the dataset.

[0088] The classification graph is constructed using the complex network creation and analysis tool NetworkX, and its construction steps are as follows: first, an empty directed graph is constructed; then nodes are added to the directed graph, each node represents a hypernym or hyponym, and each hypernym or hyponym is an atomic concept (hereinafter referred to as concept) in the knowledge graph or ontology; finally, edges are added to the directed graph according to the hypernym relationship extracted from the data set, that is, if there is a hypernym relationship between two nodes, a directed edge is added from the hypernym to the hyponym. Among them, NetworkX is a graph theory and complex network modeling tool developed in Python language, which has built-in commonly used graph and complex network analysis algorithms, and can easily perform complex network data analysis, simulation modeling and other tasks.

[0089] The nodes of the classification graph are divided into root nodes, leaf nodes and non-leaf nodes. The root node refers to a node with only outgoing edges but no incoming edges, the leaf node refers to a node with only incoming edges but no outgoing edges, and the non-leaf node refers to a node with both incoming and outgoing edges. Assume that all leaf nodes in a directed graph are represented by queries, and all root nodes are represented by roots. If there is a path from a leaf node q to a root node r, paths(q, r) are used to represent all paths from q to r, and pathNodes(q, p) is used to represent the set of all nodes except q on a path p of leaf node q (called path nodes). For a path p of q, each node in pathNodes(q, p) is a hypernym of q, and q is a hyponym. Then, anchors(q) is used to represent the set of path nodes consisting of all paths from q to all reachable root nodes, that is:

[0090] .

[0091] Step 2: Construct test cases based on the classification graph and structural information, that is, construct test concept pairs or node pairs for testing.

[0092] In this embodiment, step 2 specifically includes the following steps:

[0093] Step 2.1: Construct positive examples based on the classification graph, that is, the concept pairs or node pairs tested do have a hierarchical relationship. When constructing positive examples, first calculate the path node set for each leaf node in the classification graph, and then construct all possible parent-child node pairs for each leaf node, which is formally defined as:

[0094] ;

[0095] Right now Represents all parent-child node pairs of leaf node q. Where anchor is any element in the set anchors(q), representing the set of all nodes on a path from node q to a root node that q can reach; sup is an element in anchor, representing a node.

[0096] AllNodesPairs represents the set of parent-child node pairs consisting of all leaf nodes, and its formal definition is:

[0097] ;

[0098] in, Represents a leaf node All parent-child node pairs of ;

[0099] Therefore, allNodesPairs are all the positive examples constructed.

[0100] Step 2.2: Construct negative examples based on the classification graph and the structural information in the data set, that is, there is no hierarchical relationship between the tested concept pairs or node pairs. First, it can be constructed based on the hierarchical structure of concepts. In existing ontology construction work, it is often assumed that sibling concepts are disjoint. Here, sibling concepts refer to concepts with a common direct parent. Therefore, nodes with sibling relationships can be found based on the classification graph, and node pairs can be established between each two nodes as negative examples. In order to ensure the correctness of negative examples, manual inspection is added here to filter out incorrect negative examples. Secondly, for data sets with pattern layer information, negative examples can also be constructed using disjoint relationships in structural information. Specifically, for a node n in a directed graph, a concept c with a disjoint relationship with the node is selected from the structural information of the data set. If c has the definition of relevant axioms, negative examples (n, c) and (c, n) can be constructed. In addition, all ancestor nodes of n are obtained from the directed graph to form a set An(n), then negative examples can be constructed for c and each node in An(n).

[0101] Step 3: Design four types of prompt words for testing the large language model, including prompt words of direct expression, prompt words using concept definition, prompt words using structure, and prompt words using both concept definition and structure.

[0102] The direct expression of the prompt word refers to each test case (q m , a n ) Design prompt words that do not inject any external knowledge or use any information in the dataset, and directly let the large language model judge q m with a n Whether there is a hierarchical relationship. Among them, for the test case (q m , a n ), if it is a positive example, it means a n Yes m The parent node of ; if it is a negative example, it means a n Not q m The parent node of .

[0103] The hint words defined by the concept refer to the following: m , a n) Add node definitions, where node definitions are divided into the following three types: definitions come from the dataset itself; definitions come from a large language model. The definitions from the dataset itself are given based on the characteristics of the dataset itself. Specifically, for datasets with concept definitions, the node definitions can be obtained directly, that is, from the term2def file; for other datasets, the corresponding annotation information can be obtained as the node definition, such as using annotation information such as rdfs:comment and rdfs:label, or by converting the structural information directly related to the node into text as the node definition. Among them, rdfs represents a prefix, indicating that comment and label are defined from the resource description framework model language RDFS; the structural information can be converted into text by using a large language model, or by using the methods provided by existing conversion tools (such as NaturalOWL, a tool for converting axioms to natural languages).

[0104] The use of structured prompts refers to the use of each test case (q m , a n ) Consider q m with a n Specifically, for a dataset with concept definitions, obtain q m with a n The concept hierarchy information of each node includes sibling nodes, father nodes, child nodes, ancestor nodes, and descendant nodes. For data sets with instance-level information, obtain q m with a n Each of the associated instances, and then obtain other related instances or specific values ​​through these instances; for data sets with pattern layer information, obtain q m with a n The directly related pattern layer information, where direct correlation refers to the occurrence of q m or a n The pattern layer axiom of q m with a n Respectively related instance layer and model layer information.

[0105] For example, take the hyponym q in the environment dataset SE16-environment in the 2016 semantic evaluation dataset m : marine ecosystem and a n: Take the physical environment as an example to illustrate the acquisition of conceptual hierarchical information: the brother nodes of physical environment are pollution control measures, waste management and climate change policy, and the father node is environment; the child nodes are aquatic environment, ecosystem and ecological balance, and the descendant nodes are marine environment, ozone and biodiversity.

[0106] Take the hyponyms q in the DBpedia dataset m :Artist and a n : Taking Person as an example, we can illustrate the acquisition of instance layer structure: First, we can query the standard query language SPARQL of the resource description framework to obtain the instances of the concept Artist from the dataset, such as Pablo_Picasso, Vincent_van_Gogh, and Leonardo_da_Vinci, and the instances of Person include Albert_Einstein, Charles_Darwing, and Isaac_Newton. Here, SPARQL query is commonly used in the field of semantic web and knowledge graph. Through these instances, we can further obtain more relevant information. For example, through the instance Pablo_Picasso, we can know that the instance's date of birth is October 25, 1881, his nationality is Spanish, and his famous work is Guernica. The following are the corresponding triples in the dataset:

[0107] Pablo_Picasso birthDate "1881-10-25"^^xsd:date;

[0108] Pablo_Picasso nationality Spain;

[0109] Pablo_Picasso works Guernica;

[0110] Among them, xsd represents the namespace represented by the XML schema definition, date is the attribute of date type under the xsd namespace; birthDate represents date of birth; nationality represents nationality; Spain represents Spain; works represents work, and Guernica represents Guernica.

[0111] Then take the hyponyms a in the DBpedia dataset n : PopulatedPlace (inhabited place) with class q m : Community is used as an example to illustrate the acquisition of pattern layer information. Specifically, the former is used as an example. Through SPARQL query, it can be obtained that the class PopulatedPlace has a sub-concept Region, which is the value domain of the object attribute largestMetro and the definition domain of the data type attribute areaMetro.

[0112] The simultaneous use of concept definitions and structural prompts means that for each test case (q m , a n ) Both inject q m with a n The definition of the node injects the corresponding structural information.

[0113] Step 4: Combine each test case with each prompt word, input it into the large language model and obtain the result.

[0114] The return result of the tested large language model is divided into two parts: whether there is a clear judgment of the upper and lower relationship, and the basis for the judgment. Among them, the tested large language model can be any large language model, such as the ChatGPT large language model, the LLaMA large language model or the PaLM large language model. Due to the well-known hallucination problem of large language models, even if the large language model can correctly judge the upper and lower relationship, the corresponding interpretation may not be correct. Therefore, the explanation in the result returned by the large language model may contain some fictitious knowledge or partially wrong interpretation.

[0115] Step 5: Based on the results returned by the large language model, evaluate the ability of the large language model to judge the relationship between the upper and lower levels.

[0116] The judgment ability is evaluated from two aspects: the correctness of the judgment and the correctness of the interpretation of the large language model under test. Various metrics are formulated according to the returned results to measure the large language model from different dimensions. The returned results are divided into four situations: 1) For the prompt words of the positive example, the returned results correctly identify the upper and lower relationship, and the corresponding interpretation is also correct (mark all inputs with this situation to form a set ); 2) For the negative example, the returned result does not identify the upper and lower relationship, and the corresponding explanation is also correct (mark all inputs with this situation to form a set ); 3) Make a correct judgment on the hierarchical relationship, but the corresponding explanation is incorrect or not completely correct (mark all inputs with this situation to form a set ); 4) Incorrect judgment of the upper and lower relationships (mark all inputs with this situation to form a set ). Among them, the first case is the key to evaluate the ability of the large language model to identify the relationship between the upper and lower levels; the first and second cases can be combined to evaluate the ability of the large language model to judge the relationship between the upper and lower levels; the third case can be used to quantitatively evaluate the situation of hallucinations when the large language model recognizes the relationship between the upper and lower levels; for the fourth case, when the large language model fails to make a correct judgment on the relationship between the upper and lower levels, its explanation actually loses its meaning, so there is no need to care about whether its explanation is right or wrong.

[0117] Whether the large language model can make a correct judgment on the upper-lower relationship depends on whether the test case is a positive example or a negative example. Specifically, when a positive example is used as part of the input of the large language model, if the returned result determines that there is a upper-lower relationship between the input node pairs, it means that the large language model has made a correct judgment, otherwise it means that the judgment is incorrect; when a negative example is used as part of the input of the large language model, if the returned result determines that there is no upper-lower relationship between the input node pairs, it means that the large language model has made a correct judgment, otherwise it means that the judgment is incorrect.

[0118] The formulation of various metrics refers to defining them based on the four situations of the results returned by a large language model from four aspects: the accuracy of identifying the upper and lower relationships, the accuracy of judging the upper and lower relationships, the error rate of judging the upper and lower relationships, and the probability of hallucination. Among them, the accuracy of a large language model in identifying the upper and lower relationships can be formally defined as follows:

[0119] ;

[0120] The accuracy of a large language model in determining the relationship between the upper and lower levels can be formally defined as follows:

[0121] ;

[0122] The error rate of a large language model in determining the relationship between the upper and lower levels can be formally defined as follows:

[0123] ;

[0124] The probability of a large language model hallucinating when judging the relationship between the upper and lower levels can be defined as follows:

[0125] ;

[0126] in, and They represent the node pair sets of positive examples and the node pair sets of negative examples respectively. Representing a collection The number of elements in , and , The larger the value of , the more likely the corresponding large language model will hallucinate when judging the hyponym relationship.

[0127] This embodiment provides an evaluation method for the ability of a large language model to judge the relationship between the upper and lower rank. This method fully combines the ability of the large language model to understand the meaning of words in different contexts, including the subtle differences in the relationship between the upper and lower rank, and the characteristics of the large language model with extensive cross-domain knowledge and strong generalization ability, and overcomes the shortcomings of previous methods. In addition, this method evaluates the ability of the large language model to judge the relationship between the upper and lower rank by designing different prompt words, and also improves the reasoning ability of the large language model and the recognition accuracy of the relationship between the upper and lower rank by injecting external knowledge.

[0128] Embodiment 2, as Figure 2 As shown, this embodiment provides an evaluation device for the ability to judge the superior-subordinate relationship of a large language model, including:

[0129] A classification map construction module is used to construct classification maps based on a variety of data sets containing hierarchical relationships;

[0130] The test case construction module is used to construct test cases based on the classification graph and the structural information in the data set;

[0131] The model testing module is used to combine the test case with a variety of pre-designed prompt words, input them into the large language model under test, and obtain the return result;

[0132] The performance evaluation module is used to evaluate the judgment ability of the large language model on the upper and lower relationships according to the return results of the large language model.

[0133] The specific functional implementation of each of the above modules can be found in the relevant contents of the method in Example 1 and will not be elaborated here.

[0134] Embodiment 3, this embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods described in Embodiment 1 are implemented.

[0135] Embodiment 4: This embodiment provides a computer device, including:

[0136] Memory, for storing computer programs / instructions;

[0137] A processor, configured to execute the computer program / instructions to implement the steps of any one of the methods described in Example 1.

[0138] Embodiment 5: This embodiment provides a computer program product, including a computer program / instruction, which implements the steps of any method described in Embodiment 1 when executed by a processor.

[0139] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

[0140] It should be understood by those skilled in the art that the embodiments of the present disclosure may be provided as methods, systems or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0141] The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0142] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0144] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure rather than to limit its protection scope. Although the present disclosure has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that after reading the present disclosure, those skilled in the art can still make various changes, modifications or equivalent substitutions to the specific implementation methods of the invention, but these changes, modifications or equivalent substitutions are all within the protection scope of the disclosed claims to be approved.

Claims

1. A method for evaluating the ability of a large language model to judge the relationship between the upper and lower levels, characterized in that: include: Construct classification graphs based on multiple datasets containing hyponymy; Construct test cases based on the classification diagram and structural information in the data set; Combine the test case with a variety of pre-designed prompt words, input them into the large language model under test, and obtain the return result; Based on the results returned by the large language model, evaluate the ability of the large language model to judge the relationship between the upper and lower levels; The data set includes a data set with concept definition, a data set with instance layer information, a data set with pattern layer information, and a data set with both instance layer and pattern layer information; The dataset with concept definitions represents a hyponymy recognition dataset; it includes a taxo file, a terms file, and a term2def file, wherein each line in the taxo file describes a child node number and a parent node number thereof, each line in the terms file describes a node number and a concept corresponding to it, and each line in the term2def file is a concept and its definition from a corresponding corpus; the node numbers in the taxo file are replaced with the corresponding concepts in the terms file to construct concept pairs with hyponymy relationships; The data set with instance layer information represents a knowledge graph containing upper and lower relationships. The knowledge graph is a structured data model that represents entities and their relationships through nodes and edges, and is used to organize and associate information to support intelligent analysis and applications, wherein nodes represent instances, concepts or specific values, and edges represent semantic relationships between nodes; The data set with pattern layer information represents an ontology containing a hypernym relationship, and the ontology is used to describe the semantic relationship between concepts, attributes and instances; when selecting the hypernym and hyponym relationships for evaluation, hypernyms and hyponyms with pattern layer information definitions are selected, and the pattern layer information definitions include the non-intersection axioms of concepts, the value taking methods of concepts on certain attributes or the definition of constraints on the number of instances, and attributes associated with concepts; The data set having both instance layer and pattern layer information is a knowledge graph containing both hyponymy and pattern layer information definitions, or an ontology containing both hyponymy and instance layer information. When selecting the hyponymy for evaluation, the hypernym and the hyponym have both relevant instances and relevant pattern definitions. The classification diagram is constructed based on a plurality of data sets containing upper and lower relationships, including: Construct an empty directed graph and add nodes to the directed graph. Each node represents a hypernym or hyponym. Each hypernym or hyponym is a concept in the knowledge graph or ontology. Add edges to the directed graph based on the hypernymy relationship extracted from the data set. If there is a hypernymy relationship between two nodes, add a directed edge from the hypernym to the hyponym. After adding all the edges to the directed graph, we get a new classification graph; The nodes in the classification graph are divided into root nodes, leaf nodes and non-leaf nodes. The root node refers to a node with only outgoing edges but no incoming edges, the leaf node refers to a node with only incoming edges but no outgoing edges, and the non-leaf node refers to a node with both incoming and outgoing edges. Among them, the path node set anchors(q) consisting of all paths from leaf nodes to reachable root nodes is calculated as follows: ; Among them, q is a leaf node, p is a path, roots are all root nodes, r is a root node, pathNodes(q, p) represents the set of all nodes except q on a path p of leaf node q, and paths(q, r) represents all paths from q to r.

2. The method for evaluating the ability of judging the hyponymy of a large language model according to claim 1, characterized in that: The step of constructing a test case based on the classification diagram and the structural information in the data set includes: Construct positive examples with a hierarchical relationship between concept pairs or node pairs tested according to the classification graph; When constructing positive examples, first calculate the path node set for each leaf node in the classification graph, and then construct all possible parent-child node pairs for each leaf node. The formula is as follows: ; in, Represents all parent-child node pairs of leaf node q; anchor is any element in the set anchors(q), representing the set of all nodes on a path from node q to a root node that q can reach; sup is an element in anchor, representing a node; AllNodesPairs is used to represent the set of parent-child node pairs consisting of all leaf nodes. The formula is as follows: ; in, Represents all leaf nodes in a directed graph; Represents a leaf node All parent-child node pairs of ; Take allNodesPairs as all positive examples constructed; According to the classification graph and the structural information in the data set, negative examples are constructed in which there is no upper-lower relationship between the concept pairs or node pairs tested, including: According to the classification graph, find out the nodes with sibling relationships, and establish node pairs between every two nodes as negative examples; For datasets with pattern layer information, we use the disjoint relationships in the structural information to construct negative examples. Combine the constructed positive examples with the negative examples to obtain the constructed test cases.

3. The method for evaluating the ability of judging the hyponymy of a large language model according to claim 1, characterized in that: The prompt words include prompt words that are directly expressed, prompt words that use concept definitions, prompt words that use structure, and prompt words that use both concept definitions and structure. The direct expression of the prompt word indicates that for each test case (q m , a n ) Design prompt words that do not inject any external knowledge or use any information in the dataset, and directly let the large language model judge q m with a n Whether there is a hierarchical relationship, where for the test case (q m , a n ), if it is a positive example, it means a n Yes m The parent node of ; if it is a negative example, it means a n Not q m The parent node of The hint words defined by the concept indicate that for each test case (q m , a n ) Add node definitions, where node definitions are divided into the following three types: definitions from the dataset itself; definitions from a large language model; definitions from external data resources; The hint word of the utilization structure indicates that for each test case (q m , a n ) Consider q m with a n The structural information is as follows: For a dataset with concept definitions, get q m with a n The concept hierarchy information of each node includes sibling nodes, father nodes, child nodes, ancestor nodes, and descendant nodes. For data sets with instance-level information, obtain q m with a n Each of the associated instances, and then obtain other related instances or specific values ​​through these instances; for data sets with pattern layer information, obtain q m with a n The directly related pattern layer information, directly related means the appearance of q m or a n The pattern layer axiom of q m with a n Respectively related instance layer and pattern layer information; The concept definition and structure prompt words are used simultaneously to indicate that for each test case (q m , a n ) Both inject q m with a n The definition of the node injects the corresponding structural information.

4. The method for evaluating the ability of judging the hyponymy of a large language model according to claim 1, characterized in that: The return result of the tested large language model includes: a clear judgment on whether there is a superior-female relationship, and a basis for the judgment; The returned results are divided into the following four situations: For the prompt words in the positive example, the returned result correctly identifies the hyponym relationship and the corresponding explanation is also correct. All inputs with this situation are marked as a set. ; For the negative example, the returned result does not identify the upper and lower relationship, and the corresponding explanation is also correct. All inputs with this situation are marked as a set. ; Make a correct judgment on the hyponymy relationship, but the corresponding explanation is incorrect or incompletely correct. Mark all inputs with this situation as a set ; The judgment of the upper and lower relationship is incorrect. All inputs with this situation are marked as a set. .

5. The method for evaluating the ability of judging the hyponymy of a large language model according to claim 4, characterized in that: The step of evaluating the ability of the large language model to judge the superior-subordinate relationship based on the result returned by the large language model includes: Evaluate the ability of the large language model under test to judge the relationship between the upper and lower words from the perspective of the correctness of judgment and the correctness of interpretation, and define the accuracy rate of identifying the relationship between the upper and lower words , the accuracy rate of judging the relationship between superior and subordinate , Error rate when judging the relationship between superior and subordinate and the probability of hallucinations when judging , the calculation formula is as follows: ; ; ; ; in, and They represent the node pair sets of positive examples and the node pair sets of negative examples respectively. Representing a collection The number of elements in , and , The larger the value of , the more likely the corresponding large language model will hallucinate when judging the hyponym relationship.

6. An evaluation device for the ability to judge the relationship between a large language model and a low-level relationship, using the evaluation method for the ability to judge the relationship between a large language model and a low-level relationship according to claim 1, characterized in that: include: A classification map construction module is used to construct classification maps based on a variety of data sets containing hierarchical relationships; The test case construction module is used to construct test cases based on the classification graph and the structural information in the data set; The model testing module is used to combine the test case with a variety of pre-designed prompt words, input them into the large language model under test, and obtain the return result; The performance evaluation module is used to evaluate the judgment ability of the large language model on the upper and lower relationships according to the return results of the large language model.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 5 are implemented.

8. A computer device, characterized in that: include: Memory, for storing computer programs / instructions; A processor, configured to execute the computer program / instructions to implement the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Tobacco enterprise intelligent information question and answer method based on knowledge graph and large language model

    CN117216227A

  • Data asset identification and risk early warning system and method based on financial knowledge graph and large language model

    CN118247057A