Method for establishing knowledge database based on language model
By building a domain knowledge graph and using graph prediction model, combined with dynamic optimization of user Q&A data, the problem of insufficient coverage of the language model in niche industries or specific fields is solved, and the generation and optimization of high-quality knowledge graphs are achieved.
Patent Information
- Application Number
- CN202510055807.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-14
AI Technical Summary
The lack of coverage of knowledge databases in niche industries or specific fields leads to poor performance in related application scenarios and affects user experience.
By constructing a domain knowledge graph based on common domain data, establishing a common domain knowledge graph, generating an initial knowledge graph using a graph prediction model, and dynamically optimizing the knowledge graph through user Q&A data until the preset similarity threshold is reached.
It realizes the rapid generation of high-quality knowledge graphs in niche fields or data scarce scenarios, improves the accuracy and applicability of the knowledge graphs, and significantly enhances the practical value and expansion capabilities of the language model.
Smart Images

Figure CN119990291A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more specifically, to a method for establishing a knowledge database based on a language model. Background Art
[0002] With its knowledge database built based on massive Internet data, the generative language model has demonstrated significant universal advantages in a wide range of cross-domain and cross-industry applications, effectively improving the work efficiency of practitioners in multiple fields and industries. However, despite the excellent performance of the generative language model in terms of universality, its knowledge database does not cover niche industries or specific fields. Due to the scarcity of professional data in niche industries or fields, it is difficult for language models to fully acquire and accurately understand the knowledge in these specific fields, resulting in limited performance in related application scenarios. Therefore, in actual applications, users in niche industries or fields may face problems such as inaccurate answers and difficulty in understanding due to the lack of knowledge in the language model, which affects the user experience. Summary of the invention
[0003] The present invention provides a method for establishing a knowledge database based on a language model, which solves the technical problems raised in the background technology.
[0004] The present invention provides a method for establishing a knowledge database based on a language model, comprising:
[0005] Step 1: construct domain knowledge graphs of several domains based on public domain data; the domain knowledge graphs include root nodes, child nodes, node features of child nodes, and edge relationships of child nodes;
[0006] Step 2: According to the similarity of the root nodes of different domain knowledge graphs, edge relationships are constructed between different domain knowledge graphs to obtain a public domain knowledge graph;
[0007] Step 3: Based on the public domain knowledge graph as training data, a graph prediction model is constructed, and an initial knowledge graph is generated for the target private domain through the graph prediction model;
[0008] Step 4: configure the initial knowledge graph as the knowledge database of the target private domain, match several private domain users for question and answer, so as to obtain the question and answer knowledge graphs of several private domain users, and generate the confidence graph according to the question and answer knowledge graphs of several private domain users;
[0009] Step 5: Replace and update the sub-nodes of the initial knowledge graph with the sub-nodes of the confidence graph to obtain an updated knowledge graph;
[0010] Step 6, loop step 4 and step 5 until the similarity between at least one updated knowledge graph in the Mth cycle and at least one updated knowledge graph in the M-1th cycle is greater than a preset similarity threshold, thereby obtaining the private domain knowledge graph of the target private domain.
[0011] Furthermore, domain knowledge graphs of several fields are constructed based on public domain data, including:
[0012] Clean, identify parts of speech, and segment public domain data to obtain several domain names and corresponding knowledge entities; map the domain names to the root nodes of the domain knowledge graph, map the knowledge entities to the child nodes of the domain knowledge graph, and use the interpretations of the knowledge entities as the features of the child nodes;
[0013] Other knowledge entities in the interpretation of the knowledge entity are extracted to obtain the dependency relationship between the knowledge entities; and edge relationships are established between child nodes based on the dependency relationship between the knowledge entities.
[0014] Furthermore, according to the similarity of the root nodes of different domain knowledge graphs, edge relationships are constructed between different domain knowledge graphs, including:
[0015] Map the domain names corresponding to the root nodes of different domain knowledge graphs to the word vector space to obtain the word vector representation corresponding to the domain name;
[0016] Calculate the cosine similarity of the word vectors corresponding to the domain name. If the cosine similarity is greater than the corresponding preset threshold, an exogenous edge relationship is constructed between the child nodes of the corresponding domain knowledge graph, as follows:
[0017] The sub-nodes in the domain knowledge graph are divided according to the degree with the root node, and the sub-nodes with the same degree are divided into a group to obtain 1 to N groups of sub-nodes;
[0018] Based on the mathematical operation tree algorithm, the dependency relationship of the a-th child node of the i-th group of child nodes in the domain knowledge graph A and the dependency relationship of the b-th child node of the i-th group of child nodes in the domain knowledge graph B are formulated to obtain mathematical operation trees a and b;
[0019] Calculate the structural similarity between mathematical operation tree a and mathematical operation tree b. The calculation formula is as follows:
[0020]
[0021] Among them, Sim TED (a,b) represents the structural similarity between mathematical operation tree a and mathematical operation tree b, ED(a,b) represents the minimum number of edit operations from formula a to formula b, |a| represents the total number of tree nodes of mathematical operation tree a, |b| represents the total number of tree nodes of mathematical operation tree a, and denote the first linear weight, the second linear weight and the third linear weight respectively, and η denotes the exponential weight;
[0022] If the structural similarity between formula a and formula b is greater than the corresponding preset threshold, an exogenous edge relationship is established between the ath child node of the i-th group of child nodes in the domain knowledge graph A and the bth child node of the i-th group of child nodes in the domain knowledge graph B.
[0023] Furthermore, based on the public domain knowledge graph as training data, a graph prediction model is constructed, including:
[0024] Step 7: randomly select a domain knowledge graph from the public domain knowledge graph as a sample graph, use the features of the sub-nodes in the sample graph as sample labels, and use the public domain knowledge graph excluding the sample graph as a training sample;
[0025] Step 8, repeat step 7 to obtain several training samples and corresponding sample labels;
[0026] Step 9: The graph prediction model inputs the training sample and outputs the prediction features of the corresponding sub-nodes in the sample graph through the exogenous edge relationship;
[0027] Step 10: Based on the prediction features and sample labels, the hyperparameters of the hidden layer of the graph prediction model are reversely optimized and updated through the mean error variance loss function.
[0028] Furthermore, the calculation formula of the hidden layer of the graph prediction model is as follows:
[0029]
[0030]
[0031] β (K) =Relu(β (K-1) ×W (K) ), and β (1) =I;
[0032] Among them, Z m represents the prediction feature of the hidden layer for the mth child node in the sample graph, α mn represents the child node interaction weight correction factor, β (K) represents the correction matrix of the Kth hidden unit in the hidden layer, represents the feature set of the child nodes that establish exogenous edge relationships with the mth child node in the sample graph, k represents the kth child node in the feature set, n represents the nth child node in the feature set, deg(m) represents the sum of the number of exogenous edge relationships and edge relationships of the mth child node in the sample graph, deg(n) represents the sum of the number of exogenous edge relationships and edge relationships of the nth child node in the feature set, Represents the feature of the K-1th hidden unit in the hidden layer of the nth child node in the feature set, represents the feature of the K-1th hidden unit in the hidden layer of the mth child node in the sample graph, W (K) represents the weight matrix of the Kth hidden unit in the hidden layer, exp(·) = e (·) , e represents the natural base, Indicates the L2 norm square operation, β (1) represents the correction matrix of the first hidden unit in the hidden layer, I represents the identity matrix of size n×n, with diagonal elements of 1 and other elements of 0, Relu represents the Relu activation function, and softmax represents the softmax activation function.
[0033] Furthermore, an initial knowledge graph is generated for the target private domain through a graph prediction model, including:
[0034] Obtain the private domain name of the target private domain, match the private domain name with several domain names in the public domain knowledge graph based on the cosine similarity of the word vectors, obtain the domain name with the largest cosine similarity value with the private domain name, and use the structure of the domain knowledge graph corresponding to the domain name as the basic structure of the initial knowledge graph of the target private domain, and the exogenous edge relationship of the corresponding domain knowledge graph as the basic exogenous edge relationship of the initial knowledge graph;
[0035] The knowledge entity of the first-degree sub-node of the corresponding domain knowledge graph is used as the knowledge entity of the first-degree sub-node of the basic structure of the initial knowledge graph. Based on the basic exogenous edge relationship and the edge relationship contained in the basic structure of the initial knowledge graph, the prediction features of the first-degree sub-node to the P-degree sub-node of the initial knowledge graph are sequentially generated through the graph prediction model to obtain the initial knowledge graph; wherein P degree represents the maximum degree of the sub-nodes in the basic structure of the initial knowledge graph.
[0036] Furthermore, a question-and-answer knowledge graph of several private domain users is obtained, and a confidence graph is generated according to the question-and-answer knowledge graph of several private domain users, including:
[0037] Clean, identify parts of speech, and segment user answers to obtain question-answer knowledge entities and corresponding question-answer interpretations in user answers;
[0038] Extract other question-and-answer knowledge entities in the question-and-answer interpretation of the question-and-answer knowledge entity to obtain the dependency relationship between the question-and-answer knowledge entities;
[0039] Map the question-answering knowledge entities to question-answering nodes of the question-answering knowledge graph, use the corresponding question-answering interpretations as the features of the question-answering nodes, and establish question-answering edge relationships between question-answering nodes based on dependency relationships to obtain a question-answering knowledge graph;
[0040] Get the maximum degree of each question-answer knowledge graph;
[0041] Combine the question-answer knowledge graphs of several private domain users to make voting decisions, including:
[0042] Get the question-answer knowledge graph including the yth question-answer node and calculate the weighted score y of the yth question-answer node S :
[0043] Where L represents the total number of question-answer knowledge graphs including the yth question-answer node, q represents the index of L, and F q represents the maximum degree of the qth question-answer knowledge graph;
[0044] Judgment weighted score y S > Preset the node score threshold, then retain the yth question-answer node, and get all the retained question-answer nodes;
[0045] Based on all retained question-answer nodes, a confidence graph is established.
[0046] Furthermore, the sub-nodes of the confidence graph are replaced with the sub-nodes of the initial knowledge graph, including:
[0047] Map the features of the cth child node of the confidence graph and the rth child node of the initial knowledge graph to the word vector space respectively, and obtain the embedding vectors X1 and X2, X1 = {g1, g2, … g V}, X2={g1,g2,…g Q}, V represents the total number of word segments of the text of the feature of the c-th child node, Q represents the total number of word segments of the text of the feature of the r-th child node, Q ≥ V;
[0048] For the subvector g in the embedding vector X1 d , calculate g d The scaled dot product attention weights for each sub-vector in the embedding vector X2 are given by:
[0049] Among them, δ dc represents g in the embedding vector X1 d With the embedding vector g in X2 cThe scaled dot product attention weights, ω represents the learning weight matrix, (·) T Represents the transpose of a vector;
[0050] By scaling the dot product attention weights, the weighted vector is calculated for each sub-vector of the embedding vector X2 to obtain the comprehensive attention vector. The formula of the comprehensive attention vector is as follows:
[0051] Among them, G d represents the comprehensive attention vector;
[0052] Calculate the subvector A in the embedding vector X1 d The cosine similarity with the comprehensive attention vector is used to obtain the sub-vector A in the embedding vector X1 d The matching score S with the embedding vector X2 d ;
[0053] Based on the subvector A d , calculate the correlation between the embedding vector X1 and the embedding vector X2. The correlation calculation formula is as follows:
[0054] Among them, S totle Represents the correlation between the embedding vector X1 and the embedding vector X2;
[0055] If the correlation is greater than the preset correlation threshold, the rth child node of the initial knowledge graph is included in the replacement target of the cth child node of the confidence graph, and the replacement target set of the cth child node is obtained;
[0056] According to the replacement target set of each child node in the confidence graph, several replacement schemes are obtained, and each replacement scheme is used to replace the initial knowledge graph to generate several replacement knowledge graphs.
[0057] Furthermore, the sub-nodes of the confidence graph are updated to the nodes of the initial knowledge graph, including:
[0058] Each replacement knowledge graph is updated through the graph prediction model. The update includes:
[0059] Update the unreplaced child nodes in the replacement knowledge graph;
[0060] The replaced sub-nodes, edge relationships and exogenous edge relationships in the replacement knowledge graph are used as inputs of the graph prediction model to obtain the prediction features of the graph prediction model for the non-replaced sub-nodes, and the prediction features are used as update features of the replaced sub-nodes in the replacement knowledge graph to obtain the updated knowledge graph.
[0061] Furthermore, the private domain knowledge graph of the target private domain is obtained, including:
[0062] When the Mth cycle is reached, the graph tool is used to analyze the largest common subgraph between the γth updated knowledge graph obtained in the Mth cycle and any updated knowledge graph obtained in the M-1th cycle;
[0063] If there is at least one updated knowledge graph in the updated knowledge graph obtained in the M-1th cycle that is exactly the same as the largest common subgraph, the largest common subgraph is used as the private domain knowledge graph of the target private domain.
[0064] The beneficial effect of the present invention is that by combining the public domain knowledge graph with private domain needs, the private domain knowledge graph is generated by dynamically optimizing the graph prediction model and user feedback, thus achieving efficient cross-domain knowledge integration and private domain customized construction. Its most significant beneficial effect is that it can quickly generate high-quality knowledge graphs in niche fields or data-scarce scenarios, while improving the accuracy and applicability of knowledge graphs through iterative updates, significantly enhancing the practical value and expansion capabilities of language models. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 It is a flow chart of a method for establishing a knowledge database based on a language model of the present invention;
[0066] Figure 2 It is a construction diagram of the graph prediction model of the present invention. DETAILED DESCRIPTION
[0067] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that the discussion of these embodiments is only to enable those skilled in the art to better understand and implement the subject matter described herein, and the functions and arrangements of the elements discussed may be changed without departing from the scope of protection of the contents of this specification. Each example may omit, replace or add various processes or components as needed. In addition, the features described relative to some examples may also be combined in other examples.
[0068] like Figure 1-2 As shown, a method for establishing a knowledge database based on a language model includes:
[0069] Step 1: construct domain knowledge graphs of several domains based on public domain data; the domain knowledge graphs include root nodes, child nodes, node features of child nodes, and edge relationships of child nodes;
[0070] Step 2: According to the similarity of the root nodes of different domain knowledge graphs, edge relationships are constructed between different domain knowledge graphs to obtain a public domain knowledge graph;
[0071] Step 3: Based on the public domain knowledge graph as training data, a graph prediction model is constructed, and an initial knowledge graph is generated for the target private domain through the graph prediction model;
[0072] Step 4: Configure the initial knowledge graph as the knowledge database of the target private domain, and match several private domain users for Q&A to obtain the Q&A knowledge graphs of several private domain users, and generate a confidence graph based on the Q&A knowledge graphs of several private domain users;
[0073] Step 5: Replace and update the child nodes of the initial knowledge graph with the child nodes of the confidence graph to obtain an updated knowledge graph;
[0074] Step 6: Loop Step 4 and Step 5 until the similarity between at least one updated knowledge graph in the M-th loop and at least one updated knowledge graph in the (M - 1)-th loop > the preset similarity threshold, then obtain the private domain knowledge graph of the target private domain.
[0075] In an embodiment of the present invention, domain knowledge graphs in several domains are constructed based on public domain data, including:
[0076] Clean, identify the part of speech, and segment and cut the public domain data to obtain several domain names and several corresponding knowledge entities; map the domain names to the root nodes of the domain knowledge graphs, map the knowledge entities to the child nodes of the domain knowledge graphs, and use the definitions of the knowledge entities as the features of the child nodes;
[0077] Extract other knowledge entities in the definitions of the knowledge entities to obtain the dependency relationships between the knowledge entities; and establish edge relationships between the child nodes based on the dependency relationships between the knowledge entities.
[0078] Specifically, remove redundant, irrelevant, or non-standard data from the public domain data to improve the efficiency and accuracy of subsequent processing, including: deleting HTML tags, special characters (such as @#$%^, etc.) and invisible characters (such as line breaks). Unify the case (such as converting all to lowercase), and delete redundant spaces. Remove stop words (such as "de", "shi", "le") to retain words with actual semantics. Clear duplicate records to ensure data uniqueness. Identify and correct possible spelling mistakes (for pinyin, abbreviations, etc.). Tools used include: filtering irrelevant words through the stop word library provided by spaCy in Python, and cleaning noise using the re (regular expression) module. Removing duplicates from the data table using Pandas. Detecting and filtering non-target language data (such as non-Chinese content in multi-language mixed data) using LangDetect. Identify the part of speech of each word in the sentence to further extract key information such as domain names and entities, including: loading the segmented text data. Assigning the corresponding part of speech tags (such as nouns, verbs, etc.) to each word. Extracting specific parts of speech according to requirements, such as nouns (domain names) and entity names. Tools used include: HanLP. Implementing text segmentation into words or phrases through Jieba to facilitate the construction of node and edge relationships in the knowledge graph.
[0079] Specifically, the interpretation of knowledge entity 1 includes knowledge entity 2 and knowledge entity 3. Then, edge relationships are established between knowledge entity 2 and knowledge entity 3 and knowledge entity 1 respectively.
[0080] In one embodiment of the present invention, according to the similarity of the root nodes of different domain knowledge graphs, edge relationships are constructed between different domain knowledge graphs, including:
[0081] Map the domain names corresponding to the root nodes of different domain knowledge graphs to the word vector space to obtain the word vector representation corresponding to the domain name;
[0082] Calculate the cosine similarity of the word vectors corresponding to the domain name. If the cosine similarity is greater than the corresponding preset threshold, an exogenous edge relationship is constructed between the child nodes of the corresponding domain knowledge graph, as follows:
[0083] The sub-nodes in the domain knowledge graph are divided according to the degree with the root node, and the sub-nodes with the same degree are divided into a group to obtain 1 to N groups of sub-nodes;
[0084] Based on the mathematical operation tree algorithm, the dependency relationship of the a-th child node of the i-th group of child nodes in the domain knowledge graph A and the dependency relationship of the b-th child node of the i-th group of child nodes in the domain knowledge graph B are formulated to obtain mathematical operation trees a and b;
[0085] Calculate the structural similarity between mathematical operation tree a and mathematical operation tree b. The calculation formula is as follows:
[0086]
[0087] Among them, Sim TED (a,b) represents the structural similarity between mathematical operation tree a and mathematical operation tree b, ED(a,b) represents the minimum number of edit operations from formula a to formula b, |a| represents the total number of tree nodes of mathematical operation tree a, |b| represents the total number of tree nodes of mathematical operation tree a, and denote the first linear weight, the second linear weight and the third linear weight respectively, and η denotes the exponential weight;
[0088] If the structural similarity between formula a and formula b is greater than the corresponding preset threshold, an exogenous edge relationship is established between the ath child node of the i-th group of child nodes in the domain knowledge graph A and the bth child node of the i-th group of child nodes in the domain knowledge graph B.
[0089] It should be noted that the operation of mapping to the word vector space can also be implemented by word2vec.
[0090] Specifically, the root nodes (i.e., domain names) of knowledge graphs in different fields are mapped to the word vector space, and the cosine similarity is calculated to determine the semantic relevance between the fields. If the similarity is higher than the preset threshold, an exogenous edge relationship is established between the child nodes of the corresponding graph. Furthermore, the child nodes are grouped according to the degree of connection, the dependency relationship between the child nodes is formulated through a mathematical operation tree, and the structural similarity of the tree is calculated (based on the minimum number of editing operations). If the structural similarity of the two trees exceeds the threshold, a cross-graph association edge is established between the child nodes, thereby realizing the fusion and integration of knowledge in different fields.
[0091] For example, for knowledge entity 1, knowledge entity 2 and knowledge entity 3, knowledge entity 2 plus knowledge entity 3 equals knowledge entity 1. Then map knowledge entity 1, knowledge entity 2, knowledge entity 3 and the plus sign to nodes in the mathematical operation tree, connect knowledge entity 2 and knowledge entity 3 to the plus sign, and connect the plus sign to knowledge entity 1, and obtain the mathematical operation tree of knowledge entity 1, knowledge entity 2 and knowledge entity 3.
[0092] In one embodiment of the present invention, a graph prediction model is constructed based on a public domain knowledge graph as training data, including:
[0093] Step 7: randomly select a domain knowledge graph from the public domain knowledge graph as a sample graph, use the features of the sub-nodes in the sample graph as sample labels, and use the public domain knowledge graph excluding the sample graph as a training sample;
[0094] Step 8, repeat step 7 to obtain several training samples and corresponding sample labels;
[0095] Step 9: The graph prediction model inputs the training sample and outputs the prediction features of the corresponding sub-nodes in the sample graph through the exogenous edge relationship;
[0096] Step 10: Based on the prediction features and sample labels, the hyperparameters of the hidden layer of the graph prediction model are reversely optimized and updated through the mean error variance loss function.
[0097] Specifically, a domain knowledge graph is selected from the public domain knowledge graph as a sample in the training data so as to extract the features of the subnodes for model training. A knowledge graph is randomly selected from the public domain knowledge graph set, such as the "medical domain knowledge graph" or the "physical domain knowledge graph". The subnode set in the domain knowledge graph is determined. The features of these subnodes are extracted. Output: The extracted subnode features are used as labels for training data (for supervised learning). By repeatedly selecting domain knowledge graphs, a sufficient number of sample sets are constructed to ensure the diversity and coverage of model training. Different domain knowledge graphs are continuously randomly selected from the public domain knowledge graph. For each selected domain knowledge graph, subnode features are extracted and corresponding training labels are generated. These training samples and labels are accumulated to form a complete set of training data sets. Through multiple random selections, as many fields as possible in the public domain knowledge graph are covered to improve the generalization ability and scope of application of the model. Through the graph prediction model, the training samples are used to predict the features of the subnodes and generate predicted features corresponding to the actual labels. The training samples are input into the graph prediction model. The model generates features based on exogenous edge relationships (i.e., dependencies between child nodes) through internal calculations. Output the predicted features of the child nodes of each sample graph, such as the category, attribute, or relationship description of the node. The prediction model needs to accurately restore the true characteristics of the child nodes based on the input training sample features so as to compare with the label for optimization. By comparing the model's predicted value with the actual label value, the error is calculated and the model's hyperparameters are adjusted to gradually improve the model's prediction accuracy. Compare the predicted features of the child nodes generated by the model with the actual sample labels and calculate the error between the two. Use the mean square error loss function to quantify the error value. Based on the size of the error, adjust the model's hidden layer hyperparameters (such as node weights, connection relationship weights, etc.) to reduce the error. Repeat the training and optimization process until the model's prediction error reaches an acceptable range. By continuously optimizing the model, it can more accurately predict the characteristics of child nodes, providing a reliable foundation for building a new knowledge graph. This method achieves accurate prediction of child node characteristics in knowledge graphs through random sampling and feature generation based on exogenous edge relationships, combined with supervised learning and loss optimization. This process lays the technical foundation for the subsequent construction and dynamic updating of private domain knowledge graphs.
[0098] In one embodiment of the present invention, the calculation formula of the hidden layer of the graph prediction model is as follows:
[0099]
[0100] Among them, Z m represents the prediction feature of the hidden layer for the mth child node in the sample graph, α mn represents the child node interaction weight correction factor, β (K) represents the correction matrix of the Kth hidden unit in the hidden layer, represents the feature set of the child nodes that establish exogenous edge relationships with the mth child node in the sample graph, k represents the kth child node in the feature set, n represents the nth child node in the feature set, deg(m) represents the sum of the number of exogenous edge relationships and edge relationships of the mth child node in the sample graph, deg(n) represents the sum of the number of exogenous edge relationships and edge relationships of the nth child node in the feature set, Represents the feature of the K-1th hidden unit in the hidden layer of the nth child node in the feature set, represents the feature of the K-1th hidden unit in the hidden layer of the mth child node in the sample graph, W (K) represents the weight matrix of the Kth hidden unit in the hidden layer, exp(·)=e(·), where e represents the natural base number, Indicates the L2 norm square operation, β (1) represents the correction matrix of the first hidden unit in the hidden layer, I represents the identity matrix of size n×n, with diagonal elements of 1 and other elements of 0, Relu represents the Relu activation function, and softmax represents the softmax activation function.
[0101] Specifically, by aggregating the features of child nodes and their neighboring nodes, the model can simulate the dependency relationship between child nodes and their context. The introduction of weight matrix and correction matrix enables the model to automatically optimize the strength of feature interaction according to specific tasks. Activation function and feature normalization ensure the nonlinearity and numerical stability of the model. The actual application effect evaluation table based on the graph prediction model is as follows:
[0102]
[0103] Among them, the test data set represents the knowledge graph data set in different fields; the number of nodes represents the number of nodes contained in the graph; the number of edges represents the number of edges connecting nodes in the graph; the prediction accuracy represents the degree of matching between the model's predicted features of sub-nodes and actual labels; the number of hidden layers of the graph prediction model represents the depth and complexity of the model; the comparative improvement represents the improvement in prediction accuracy compared with traditional graph prediction methods.
[0104] In one embodiment of the present invention, generating an initial knowledge graph for a target private domain through a graph prediction model includes:
[0105] Obtain the private domain name of the target private domain, match the private domain name with several domain names in the public domain knowledge graph based on the cosine similarity of the word vectors, obtain the domain name with the largest cosine similarity value with the private domain name, and use the structure of the domain knowledge graph corresponding to the domain name as the basic structure of the initial knowledge graph of the target private domain, and the exogenous edge relationship of the corresponding domain knowledge graph as the basic exogenous edge relationship of the initial knowledge graph;
[0106] The knowledge entity of the first-degree sub-node of the corresponding domain knowledge graph is used as the knowledge entity of the first-degree sub-node of the basic structure of the initial knowledge graph. Based on the basic exogenous edge relationship and the edge relationship contained in the basic structure of the initial knowledge graph, the prediction features of the first-degree sub-node to the P-degree sub-node of the initial knowledge graph are sequentially generated through the graph prediction model to obtain the initial knowledge graph; wherein P degree represents the maximum degree of the sub-nodes in the basic structure of the initial knowledge graph.
[0107] Specifically, the closest domain is found by calculating the cosine similarity of the word vectors of the target private domain name and several domain names in the public domain knowledge graph. Cosine similarity is used to quantify the semantic similarity between the private domain name and each domain name. The domain name with the highest semantic similarity is selected, and the target private domain knowledge graph is initialized based on the knowledge graph structure of the domain. "First-degree child nodes" (child nodes directly connected to the root node) are extracted from the selected domain knowledge graph as the basic nodes of the initial graph. At the same time, all edge relationships (including exogenous edge relationships) in the basic structure are retained. An initial knowledge graph is obtained, which contains part of the structure and association relationships of the selected domain. Using the graph prediction model, based on the sub-node features and exogenous edge relationships of the initial structure, the features of deeper nodes are generated layer by layer. From the first-degree child node to the multi-degree child node (such as the second-degree and third-degree nodes). For example, the features of the first-degree child node are obtained based on the graph prediction model, and then the features of the first-degree child node are used as the input of the graph prediction model to obtain the features of the second-degree child node, and so on. A complete initial knowledge graph is obtained, containing all predicted node features.
[0108] For example, when generating a knowledge graph for the private domain "electric vehicle", the domain name "automobile design" has the highest similarity, so its corresponding knowledge graph is selected as the basis. The structure of the knowledge graph of "automobile design" is obtained. For example, the root node is "automobile design", and the first-degree child nodes are "the meaning of automobile design", "the concept of automobile design", "the actual role of automobile" and "the prospect of automobile design". Correspondingly, "electric vehicle" is the root node of the private domain knowledge graph, and the first-degree child nodes are "the meaning of electric vehicles", "the concept of electric vehicles", "the role of electric vehicles" and "the prospect of electric vehicles".
[0109] In one embodiment of the present invention, a question-answer knowledge graph of several private domain users is obtained, and a confidence graph is generated according to the question-answer knowledge graph of several private domain users, including:
[0110] Clean, identify parts of speech, and segment user answers to obtain question-answer knowledge entities and corresponding question-answer interpretations in user answers;
[0111] Extract other question-and-answer knowledge entities in the question-and-answer interpretation of the question-and-answer knowledge entity to obtain the dependency relationship between the question-and-answer knowledge entities;
[0112] Map the question-answering knowledge entities to question-answering nodes of the question-answering knowledge graph, use the corresponding question-answering interpretations as the features of the question-answering nodes, and establish question-answering edge relationships between question-answering nodes based on dependency relationships to obtain a question-answering knowledge graph;
[0113] Get the maximum degree of each question-answer knowledge graph;
[0114] Combine the question-answer knowledge graphs of several private domain users to make voting decisions, including:
[0115] Get the question-answer knowledge graph including the yth question-answer node and calculate the weighted score y of the yth question-answer node S :
[0116] Where L represents the total number of question-answer knowledge graphs including the yth question-answer node, q represents the index of L, and F q represents the maximum degree of the qth question-answer knowledge graph;
[0117] Judgment weighted score y S > Preset the node score threshold, then retain the yth question-answer node, and get all the retained question-answer nodes;
[0118] Based on all retained question-answer nodes, a confidence graph is established.
[0119] Specifically, through this method, the question and answer content of multiple private domain users can be structured. First, the user answers are cleaned, segmented and identified, the core knowledge entities and their interpretations in the questions and answers are extracted, and then the dependencies between the knowledge entities are extracted to generate the question and answer knowledge graphs corresponding to each user. In these graphs, nodes represent knowledge entities, and edges represent the logical or semantic associations between entities. Subsequently, by traversing each question and answer knowledge graph, the importance of the nodes is calculated and their contribution in each graph is quantified. Combined with the statistical data of multi-user questions and answers, this method adopts a voting strategy to screen the nodes with high frequency in multiple graphs, eliminate the nodes with low confidence, and only retain the nodes with high confidence. Finally, based on the high-confidence nodes screened out, a comprehensive confidence graph is constructed. This confidence graph can fully integrate the knowledge feedback of multiple users, ensure that the expressed knowledge structure is more accurate, authoritative and widely applicable, and provide a reliable foundation for the optimization of private domain knowledge graphs.
[0120] In one embodiment of the present invention, sub-nodes of the confidence graph are replaced with sub-nodes of the initial knowledge graph, including:
[0121] Map the features of the cth child node of the confidence graph and the rth child node of the initial knowledge graph to the word vector space respectively, and obtain the embedding vectors X1 and X2, X1 = {g1, g2, … g V}, X2={g1,g2,…g Q}, V represents the total number of word segments of the text of the feature of the c-th child node, Q represents the total number of word segments of the text of the feature of the r-th child node, Q ≥ V;
[0122] For the subvector g in the embedding vector X1 d , calculate g d The scaled dot product attention weights for each sub-vector in the embedding vector X2 are given by:
[0123] Among them, δ dc represents g in the embedding vector X1 d With the embedding vector g in X2 c The scaled dot product attention weights, ω represents the learning weight matrix, (·) T Represents the transpose of a vector;
[0124] By scaling the dot product attention weights, the weighted vector is calculated for each sub-vector of the embedding vector X2 to obtain the comprehensive attention vector. The formula of the comprehensive attention vector is as follows:
[0125] Among them, G d represents the comprehensive attention vector;
[0126] Calculate the subvector A in the embedding vector X1 d The cosine similarity with the comprehensive attention vector is used to obtain the sub-vector A in the embedding vector X1 d The matching score S with the embedding vector X2 d ;
[0127] Based on the subvector A d , calculate the correlation between the embedding vector X1 and the embedding vector X2. The correlation calculation formula is as follows:
[0128] Among them, S totle Represents the correlation between the embedding vector X1 and the embedding vector X2;
[0129] If the correlation is greater than the preset correlation threshold, the rth child node of the initial knowledge graph is included in the replacement target of the cth child node of the confidence graph, and the replacement target set of the cth child node is obtained;
[0130] According to the replacement target set of each child node in the confidence graph, several replacement schemes are obtained, and each replacement scheme is used to replace the initial knowledge graph to generate several replacement knowledge graphs.
[0131] Specifically, by gradually integrating high-confidence nodes in the confidence graph into the initial knowledge graph, dynamic optimization of the initial graph is achieved, including:
[0132] The sub-node features in the confidence graph and the sub-node features in the initial knowledge graph are mapped to the high-dimensional word vector space to generate the corresponding embedding vectors. By converting text features into embedding vectors, it is convenient to calculate the similarity and correlation between nodes in the high-dimensional space. Each node feature is represented as an embedding vector, which contains the semantic information of the feature.
[0133] For each sub-part of the embedding vector, the attention mechanism is used to calculate its importance in the overall embedding vector. The attention mechanism determines which parts are most critical to the node feature by weighting the sub-features. The semantic importance of each sub-feature is highlighted to generate a comprehensive attention vector. Each sub-feature of the embedding vector is assigned an attention weight to obtain a comprehensive semantic representation.
[0134] Compare the nodes in the confidence graph with the nodes in the initial knowledge graph in the embedding space and calculate their cosine similarity as the matching score. Through the similarity measurement, determine whether the nodes in the confidence graph are suitable to replace the nodes in the initial knowledge graph. Generate a matching score for each node to be compared to evaluate their semantic consistency.
[0135] The node association is further calculated based on the matching score. The association reflects the overall correlation between nodes in the graph, not just the individual feature similarity. The semantic similarity of the nodes and the dependency of the graph structure are comprehensively considered to improve the accuracy of the replacement operation. A correlation value is generated for each pair of nodes. The higher the correlation, the more reasonable the replacement.
[0136] If the correlation between the nodes of the confidence graph and the nodes of the initial knowledge graph exceeds the preset threshold, the nodes of the initial knowledge graph are included in the replacement target set of the confidence graph nodes. The initial knowledge graph nodes suitable for replacement are screened out to provide a basis for updating the graph. A replacement target set is generated, which contains all nodes that meet the replacement conditions.
[0137] Based on the replacement target set, multiple replacement schemes are constructed, each of which corresponds to a replacement combination of the initial knowledge graph nodes. Through diverse replacement schemes, different node update paths are explored. Several replacement knowledge graphs are obtained, each of which is based on a different replacement scheme.
[0138] The replacement scheme is applied to the initial knowledge graph to generate an updated knowledge graph. Through iterative replacement, the node features of the knowledge graph are dynamically optimized to make it more accurate and comprehensive. The new knowledge graph integrates the high-confidence node features in the confidence graph and expands and optimizes the initial graph through the replacement scheme.
[0139] In one embodiment of the present invention, updating the node of the initial knowledge graph with the child node of the confidence graph includes:
[0140] Each replacement knowledge graph is updated through the graph prediction model. The update includes:
[0141] Update the unreplaced child nodes in the replacement knowledge graph;
[0142] The replaced sub-nodes, edge relationships and exogenous edge relationships in the replacement knowledge graph are used as inputs of the graph prediction model to obtain the prediction features of the graph prediction model for the non-replaced sub-nodes, and the prediction features are used as update features of the replaced sub-nodes in the replacement knowledge graph to obtain the updated knowledge graph.
[0143] In one embodiment of the present invention, obtaining a private domain knowledge graph of a target private domain includes:
[0144] When the Mth cycle is reached, the graph tool is used to analyze the largest common subgraph between the γth updated knowledge graph obtained in the Mth cycle and any updated knowledge graph obtained in the M-1th cycle;
[0145] If there is at least one updated knowledge graph in the updated knowledge graph obtained in the M-1th cycle that is exactly the same as the largest common subgraph, the largest common subgraph is used as the private domain knowledge graph of the target private domain.
[0146] Specifically, through this method, in each round of replacement process, multiple updated knowledge graphs are generated, in which the correct replacement scheme must be included in all replacement schemes. Because the confidence graph provided by the user is smaller than the initial knowledge graph, the replacement operation is mainly based on the local high-confidence content to gradually adjust the initial knowledge graph. Therefore, the correct replacement scheme is the first to achieve the stability of the result (i.e., "fastest convergence") in the generated updated knowledge graph because it conforms to the logic of the confidence graph. In this process, the updated knowledge graph generated by the replacement scheme of each cycle will be compared with its parent. When the maximum common subgraph generated by the replacement scheme of a certain cycle is completely consistent with the updated knowledge graph of the previous cycle, it indicates that the current replacement operation is consistent with the correct replacement scheme. Since the adjustment logic of the correct replacement scheme is optimal and consistent with the confidence graph, it will preferentially form a stable structure. Therefore, the first time that the maximum common subgraph is completely consistent with any updated knowledge graph of a certain cycle, it means that we have found the correct replacement scheme and successfully converged to the target private domain knowledge graph.
[0147] The above describes an embodiment of the present embodiment, but the present embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present embodiment, ordinary technicians in this field can also make many forms, all of which are within the protection of the present embodiment.
Claims
1. A method for establishing a knowledge database based on a language model, characterized in that: include: Step 1: construct domain knowledge graphs of several domains based on public domain data; the domain knowledge graphs include root nodes, child nodes, node features of child nodes, and edge relationships of child nodes; Step 2: According to the similarity of the root nodes of different domain knowledge graphs, edge relationships are constructed between different domain knowledge graphs to obtain a public domain knowledge graph; Step 3: Based on the public domain knowledge graph as training data, a graph prediction model is constructed, and an initial knowledge graph is generated for the target private domain through the graph prediction model; Step 4: configure the initial knowledge graph as the knowledge database of the target private domain, match several private domain users for question and answer, so as to obtain the question and answer knowledge graphs of several private domain users, and generate the confidence graph according to the question and answer knowledge graphs of several private domain users; Step 5: Replace and update the sub-nodes of the initial knowledge graph with the sub-nodes of the confidence graph to obtain an updated knowledge graph; Step 6, loop step 4 and step 5 until the similarity between at least one updated knowledge graph in the Mth cycle and at least one updated knowledge graph in the M-1th cycle is greater than a preset similarity threshold, thereby obtaining the private domain knowledge graph of the target private domain.
2. The method for establishing a knowledge database based on a language model according to claim 1, characterized in that: Construct domain knowledge graphs in several fields based on public domain data, including: Clean, identify parts of speech, and segment public domain data to obtain several domain names and corresponding knowledge entities; map the domain names to the root nodes of the domain knowledge graph, map the knowledge entities to the child nodes of the domain knowledge graph, and use the interpretations of the knowledge entities as the features of the child nodes; Other knowledge entities in the interpretation of the knowledge entity are extracted to obtain the dependency relationship between the knowledge entities; and edge relationships are established between child nodes based on the dependency relationship between the knowledge entities.
3. The method for establishing a knowledge database based on a language model according to claim 2, characterized in that: According to the similarity of the root nodes of different domain knowledge graphs, edge relationships are constructed between different domain knowledge graphs, including: Map the domain names corresponding to the root nodes of different domain knowledge graphs to the word vector space to obtain the word vector representation corresponding to the domain name; Calculate the cosine similarity of the word vectors corresponding to the domain name. If the cosine similarity is greater than the corresponding preset threshold, an exogenous edge relationship is constructed between the child nodes of the corresponding domain knowledge graph, as follows: The sub-nodes in the domain knowledge graph are divided according to the degree with the root node, and the sub-nodes with the same degree are divided into a group to obtain 1 to N groups of sub-nodes; Based on the mathematical operation tree algorithm, the dependency relationship of the a-th child node of the i-th group of child nodes in the domain knowledge graph A and the dependency relationship of the b-th child node of the i-th group of child nodes in the domain knowledge graph B are formulated to obtain mathematical operation trees a and b; Calculate the structural similarity between mathematical operation tree a and mathematical operation tree b. The calculation formula is as follows: Among them, Sim TED (a,b) represents the structural similarity between mathematical operation tree a and mathematical operation tree b, ED(a,b) represents the minimum number of edit operations from formula a to formula b, |a| represents the total number of tree nodes of mathematical operation tree a, |b| represents the total number of tree nodes of mathematical operation tree a, and denote the first linear weight, the second linear weight and the third linear weight respectively, and η denotes the exponential weight; If the structural similarity between formula a and formula b is greater than the corresponding preset threshold, an exogenous edge relationship is established between the ath child node of the i-th group of child nodes in the domain knowledge graph A and the bth child node of the i-th group of child nodes in the domain knowledge graph B.
4. The method for establishing a knowledge database based on a language model according to claim 3, characterized in that: Based on the public domain knowledge graph as training data, a graph prediction model is constructed, including: Step 7: randomly select a domain knowledge graph from the public domain knowledge graph as a sample graph, use the features of the sub-nodes in the sample graph as sample labels, and use the public domain knowledge graph excluding the sample graph as a training sample; Step 8, repeat step 7 to obtain several training samples and corresponding sample labels; Step 9: The graph prediction model inputs the training sample and outputs the prediction features of the corresponding sub-nodes in the sample graph through the exogenous edge relationship; Step 10: Based on the prediction features and sample labels, the hyperparameters of the hidden layer of the graph prediction model are reversely optimized and updated through the mean error variance loss function.
5. The method for establishing a knowledge database based on a language model according to claim 4, characterized in that: The calculation formula of the hidden layer of the graph prediction model is as follows: Among them, Z m represents the prediction feature of the hidden layer for the mth child node in the sample graph, α mn represents the modification factor of the interaction weight of child nodes, β (K) represents the correction matrix of the Kth hidden unit in the hidden layer, represents the feature set of the child nodes that establish exogenous edge relationships with the mth child node in the sample graph, k represents the kth child node in the feature set, n represents the nth child node in the feature set, deg(m) represents the sum of the number of exogenous edge relationships and edge relationships of the mth child node in the sample graph, deg(n) represents the sum of the number of exogenous edge relationships and edge relationships of the nth child node in the feature set, Represents the feature of the K-1th hidden unit in the hidden layer of the nth child node in the feature set, represents the feature of the K-1th hidden unit in the hidden layer of the mth child node in the sample graph, W (K) represents the weight matrix of the Kth hidden unit in the hidden layer, exp(·) = e (·) , e represents the natural base, Indicates the L2 norm square operation, β (1) represents the correction matrix of the first hidden unit in the hidden layer, I represents the identity matrix of size n×n, with diagonal elements of 1 and other elements of 0, Relu represents the Relu activation function, and softmax represents the softmax activation function.
6. The method for establishing a knowledge database based on a language model according to claim 5, characterized in that: Generate an initial knowledge graph for the target private domain through the graph prediction model, including: Obtain the private domain name of the target private domain, match the private domain name with several domain names in the public domain knowledge graph based on the cosine similarity of the word vectors, obtain the domain name with the largest cosine similarity value with the private domain name, and use the structure of the domain knowledge graph corresponding to the domain name as the basic structure of the initial knowledge graph of the target private domain, and the exogenous edge relationship of the corresponding domain knowledge graph as the basic exogenous edge relationship of the initial knowledge graph; The knowledge entity of the first-degree sub-node of the corresponding domain knowledge graph is used as the knowledge entity of the first-degree sub-node of the basic structure of the initial knowledge graph. Based on the basic exogenous edge relationship and the edge relationship contained in the basic structure of the initial knowledge graph, the prediction features of the first-degree sub-node to the P-degree sub-node of the initial knowledge graph are sequentially generated through the graph prediction model to obtain the initial knowledge graph; wherein P degree represents the maximum degree of the sub-nodes in the basic structure of the initial knowledge graph.
7. The method for establishing a knowledge database based on a language model according to claim 1, characterized in that: A question-and-answer knowledge graph of several private domain users is obtained, and a confidence graph is generated based on the question-and-answer knowledge graph of several private domain users, including: Clean, identify parts of speech, and segment user answers to obtain question-answer knowledge entities and corresponding question-answer interpretations in user answers; Extract other question-and-answer knowledge entities in the question-and-answer interpretation of the question-and-answer knowledge entity to obtain the dependency relationship between the question-and-answer knowledge entities; Map the question-answering knowledge entities to question-answering nodes of the question-answering knowledge graph, use the corresponding question-answering interpretations as the features of the question-answering nodes, and establish question-answering edge relationships between question-answering nodes based on dependency relationships to obtain a question-answering knowledge graph; Get the maximum degree of each question-answer knowledge graph; Combine the question-answer knowledge graphs of several private domain users to make voting decisions, including: Get the question-answer knowledge graph including the yth question-answer node and calculate the weighted score y of the yth question-answer node S : Where L represents the total number of question-answer knowledge graphs including the yth question-answer node, q represents the index of L, and F q represents the maximum degree of the qth question-answer knowledge graph; Judgment weighted score y S > Preset the node score threshold, then retain the yth question-answer node, and get all the retained question-answer nodes; Based on all retained question-answer nodes, a confidence graph is established.
8. The method for establishing a knowledge database based on a language model according to claim 7, characterized in that: Replace the sub-nodes of the confidence graph with the sub-nodes of the initial knowledge graph, including: Map the features of the cth child node of the confidence graph and the rth child node of the initial knowledge graph to the word vector space respectively, and obtain the embedding vectors X1 and X2, X1 = {g1, g2, … g V }, X2={g1,g2,…g Q }, V represents the total number of word segments of the text of the feature of the c-th child node, Q represents the total number of word segments of the text of the feature of the r-th child node, Q ≥ V; For the subvector g in the embedding vector X1 d , calculate g d The scaled dot product attention weights for each sub-vector in the embedding vector X2 are given by: Among them, δ dc represents g in the embedding vector X1 d With the embedding vector g in X2 c The scaled dot product attention weights, ω represents the learning weight matrix, (·) T Represents the transpose of a vector; By scaling the dot product attention weights, the weighted vector is calculated for each sub-vector of the embedding vector X2 to obtain the comprehensive attention vector. The formula of the comprehensive attention vector is as follows: Among them, G d represents the comprehensive attention vector; Calculate the subvector A in the embedding vector X1 d The cosine similarity with the comprehensive attention vector is used to obtain the sub-vector A in the embedding vector X1 d The matching score S with the embedding vector X2 d ; Based on the subvector A d , calculate the correlation between the embedding vector X1 and the embedding vector X2. The correlation calculation formula is as follows: Among them, S totle Represents the correlation between the embedding vector X1 and the embedding vector X2; If the correlation is greater than the preset correlation threshold, the rth child node of the initial knowledge graph is included in the replacement target of the cth child node of the confidence graph, and the replacement target set of the cth child node is obtained; According to the replacement target set of each child node in the confidence graph, several replacement schemes are obtained, and each replacement scheme is used to replace the initial knowledge graph to generate several replacement knowledge graphs.
9. The method for establishing a knowledge database based on a language model according to claim 5 or 8, characterized in that: Update the nodes of the initial knowledge graph with the child nodes of the confidence graph, including: Each replacement knowledge graph is updated through the graph prediction model. The update includes: Update the unreplaced child nodes in the replacement knowledge graph; The replaced sub-nodes, edge relationships and exogenous edge relationships in the replacement knowledge graph are used as inputs of the graph prediction model to obtain the prediction features of the graph prediction model for the non-replaced sub-nodes, and the prediction features are used as update features of the replaced sub-nodes in the replacement knowledge graph to obtain the updated knowledge graph.
10. The method for establishing a knowledge database based on a language model according to claim 9, characterized in that: Get the private domain knowledge graph of the target private domain, including: When the Mth cycle is reached, the graph tool is used to analyze the largest common subgraph between the γth updated knowledge graph obtained in the Mth cycle and any updated knowledge graph obtained in the M-1th cycle; If there is at least one updated knowledge graph in the updated knowledge graph obtained in the M-1th cycle that is exactly the same as the largest common subgraph, the largest common subgraph is used as the private domain knowledge graph of the target private domain.
Citation Information
Patent Citations
Large model training deployment method and system based on knowledge graph
CN117332852A
Financial industry knowledge graph construction method based on generative large language model
CN117454985A
Software development data processing system based on data analysis
CN118152221A
Digital human interaction system based on knowledge graph and large model
CN118535022A
Knowledge retrieval-based vertical type government affair large model service method and system
CN118916470A
Cited By
Bidding document analysis method based on engineering cost knowledge graph
CN122364690A