A knowledge base based inference comparison method
Patent Information
- Application Number
- CN202511237382.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2045-09-01
AI Technical Summary
以金融市场风险预测为例,市场动态、政策法规的变化瞬息万变,传统方法无法及时捕捉新知识,导致推理结果滞后,无法满足实时性要求极高的实际应用需求
[0099](1)本发明首先对交互数据集、领域数据集、通用数据集进行数据清洗、标准化、去重及分词预处理,一方面通过特征阈值判定噪声数据并剔除,确保数据准确性;另一方面将多源数据统一格式与尺度,消除量纲差异,同时进行去重,减少数据冗余,分词处理则将文本转化为结构化词语序列,为后续知识图谱嵌入提供高质量的数据基础;
Smart Images

Figure CN121119142B_ABST
Abstract
Description
[0001] This invention relates to the fields of knowledge graphs, large models, and intelligent decision-making, and specifically to a reasoning comparison method based on a knowledge base. Background Technology
[0002] With the exponential development of artificial intelligence and knowledge graph technologies, knowledge base-based reasoning and comparison methods have become a hot topic in research and application across numerous fields. From credit risk assessment in the financial industry to assisted diagnosis in the medical field; from precise responses in intelligent customer service to decision support systems in smart cities, these technologies are deeply penetrating all aspects of society. However, traditional reasoning methods, such as rule-based reasoning and simple machine learning reasoning, reveal many significant limitations when dealing with complex knowledge and large-scale data.
[0003] While simple machine learning reasoning can learn patterns from massive amounts of data, it lacks a deep understanding of knowledge structures and semantics. Taking intelligent question-answering systems as an example, when users ask questions involving the intersection of multiple disciplines, such as "How to improve cardiovascular health through diet," traditional machine learning models often only match surface keywords and cannot perform deep reasoning by combining knowledge from multiple disciplines such as nutrition and physiology, making it difficult to provide a systematic and comprehensive answer. When dealing with tasks that require integrating multi-source knowledge and uncovering deep logical relationships, these methods often suffer from a lack of semantic understanding, leading to reasoning results that deviate from the actual needs.
[0004] With the continuous accumulation and updating of knowledge across various fields, the scale of knowledge bases is expanding at an astonishing rate. Taking the biomedical field as an example, the PubMed database adds over 2,000 new articles daily, the number of knowledge graph nodes and edges is growing exponentially, and the diversity and complexity of the data are also increasing dramatically. Existing reasoning and comparison methods have significant shortcomings in integrating multi-source heterogeneous knowledge, extracting key information, and responding to dynamically changing knowledge graphs.
[0005] At the knowledge fusion level, knowledge from different sources differs significantly in its representation and semantic interpretation. For example, internal corporate databases use structured data storage, while online encyclopedia information is mostly unstructured text, and open-source research literature contains semi-structured tables and formulas. This difference makes knowledge fusion difficult, hindering the formation of a unified reasoning foundation. At the knowledge update level, traditional methods typically suffer from significant delays in updating rapidly iterating knowledge graphs. Taking financial market risk prediction as an example, market dynamics and policy regulations change rapidly, and traditional methods cannot capture new knowledge in a timely manner, resulting in lagging reasoning results that fail to meet the extremely high real-time requirements of practical applications.
[0006] Against this backdrop, researching a reasoning comparison method that can effectively integrate multi-source knowledge and deeply mine the structural and semantic information of knowledge graphs is of significant practical importance. The reasoning comparison method based on a knowledge base proposed in this invention aims to address the problems existing in current technologies. By integrating a reasoning graph model and the Mingtu general large model, it fully leverages the advantages of both to achieve accurate knowledge reasoning and analysis, providing more reliable and efficient support for applications in various fields. Summary of the Invention
[0007] In view of this, the purpose of this invention is to provide a reasoning comparison method based on a knowledge base, which mainly integrates the reasoning model of a domain database with the general large model of a general database to achieve accurate knowledge reasoning and analysis.
[0008] The objective of this invention is achieved through the following technical solution:
[0009] A knowledge base-based reasoning comparison method includes the following steps:
[0010] Step S1: Establish interactive datasets, domain datasets, and general datasets, and perform data preprocessing on the three databases respectively;
[0011] Step S2: Embed the preprocessed data using a knowledge graph and construct a ternary embedding matrix;
[0012] Step S3: Use the inference graph model to traverse the domain knowledge graph and search for candidate entities related to the interactive content;
[0013] Step S4: Using a multi-layer self-focusing network and a feedforward neural network, capture the global dependencies and local feature representations of the input data to build a large model and pre-train the large model;
[0014] Step S5: Utilize the implicit knowledge within the pre-trained large model and the explicit knowledge within the domain knowledge graph, and gradually generate accurate prediction or analysis results through multi-dimensional feature fusion and optimization iteration.
[0015] Further, step S1 specifically includes:
[0016] Step S101: Establish an interactive dataset, a domain dataset, and a general dataset. The interactive dataset provides dynamic scene data directly related to the analysis task to clarify the task context; the domain dataset provides structured knowledge from a specific professional field to support the construction of a domain knowledge graph and professional reasoning; the general dataset provides cross-domain common sense and basic logic to ensure the model's basic understanding and generalization ability. These three datasets work together to provide the data foundation for subsequent processes. The preprocessing includes data cleaning of the domain dataset D1, the general dataset D2, and the interactive dataset D3 to remove noisy and invalid data. Let the domain dataset be... The general dataset is Interactive dataset is Where d i d j and d k These represent data instances in the domain dataset, the general dataset, and the interactive dataset, respectively, with I, J, and K representing the total number of data instances in the three datasets.
[0017] Define a unified noise data judgment function f(d). When f(d) = 1, d is judged as noise data; when f(d) = 0, d is considered as valid data. f(d) is set based on a feature threshold. If x l If the value is less than θ1 or θ2 (θ1 and θ2 are pre-set feature thresholds), then f(d) = 1, and the numerical feature x... l (l is the feature dimension index) is a numerical indicator used to describe a certain aspect of an attribute of a data instance; for each data instance d, multiple numerical features x can be used. l To represent its different dimensions of characteristics;
[0018] After cleaning, the domain dataset and the general dataset were updated as follows: and Among them, θ1 and θ2 are used to define the reasonable range of values for data features. Data that exceeds this range will be judged as noise data and removed, thereby ensuring the accuracy and effectiveness of data processing.
[0019] Step S102: Process the cleaned domain dataset General Dataset and interactive datasets Standardization is performed to unify the data to the same format and metric, resulting in a standardized domain dataset. General Dataset and interactive datasets For the numerical features in all three datasets, a standardization formula was used, as shown below:
[0020]
[0021] Where μ represents the feature in the merged dataset (where μ is the sum of the features in the merged dataset). and The mean of the merged dataset (x) is used to characterize the central tendency of the data. i x j x kThese correspond to the numerical features of a certain attribute of a certain data instance in the three datasets. Through this standardization formula, data of different magnitudes and distribution characteristics can be transformed to the same scale, which is convenient for subsequent data analysis and model training. σ is the standard deviation of the feature in the merged dataset, which is used to measure the dispersion of the data. x' is a numerical feature, which represents the standardized numerical index used to describe a certain attribute of a certain data instance.
[0022] Step S103: Process the standardized domain dataset and general datasets A deduplication operation is performed to eliminate duplicate data records. A hash function h(d) is used to hash the data records in both datasets, mapping each data record d to a fixed-length hash value. If h(d1) = h(d2) (where d1 and d2 come from the domain dataset or the general dataset respectively, and d1 = d2), and further comparison confirms that the key features of d1 and d2 are completely identical, then these two data records are determined to be duplicates, and only one is retained while the other is deleted. After deduplication, the domain dataset and the general dataset are updated to... and
[0023] Step S104: Process the deduplicated domain dataset General Dataset and the standardized interactive dataset Word segmentation is performed to convert the text data into a word sequence; a statistical word segmentation method is adopted, and the probability function of the word segmentation model is known to be P(w1,w2,…,w…). s |T), where T is the input text data, w1, w2, ..., w s The sequence of words obtained after word segmentation is s, where s is the length of the word sequence. The probability function is statistical information obtained from training on a large corpus, which helps the model determine the segmentation method of words in the text, transforming the text data into a computer-processable word sequence form, and is used to calculate the word sequence w1, w2, ..., w3 given a text T. s The probability is calculated, and the optimal word segmentation result is obtained by maximizing this probability function. After word segmentation, the domain dataset, general dataset, and interactive dataset are updated to D1', D2', and D3', respectively.
[0024] Further, step S2 specifically includes:
[0025] Step S201: Perform knowledge graph embedding processing on the deduplicated and segmented domain dataset D1', general dataset D2', and interactive dataset D3' from S104. First, extract the entity set E1 from the domain dataset D1' using part-of-speech tagging, so that the entities in E1 cover the key concepts and objects in the domain. Then, use a knowledge graph embedding model to represent the triples (h,r,t) in the knowledge graph (where h is the head entity, r is the relation, and t is the tail entity) as relations in the vector space, i.e., h⊕r≈t, thereby obtaining the domain data triple set S1 and the general data triple set S2. Optimize the entity vectors by minimizing the objective function, and then use a scoring function and undergo multiple rounds of iterative training to obtain the entity vector representation E1' of the domain dataset.
[0026] For the general dataset D2' and the interactive dataset D3', after extracting the entity sets E2 and E3 using a similar method, the entities in the general dataset are embedded using the above embedding method and optimization strategy to obtain the entity vector representation E2' of the general dataset and the entity vector representation E3' of the interactive dataset.
[0027] Step S202: Using the same model as the entity embedding in S201, extract relation sets R1 and R2 from the domain dataset D1' and the general dataset D2' respectively. Map the relations to a vector space of the same dimension L as the entity vectors to obtain the relation vector r∈R L During the mapping process, the model learns the semantic features of the relationship, so that relationships with similar semantics are closer in the vector space;
[0028] For each triple (h, r, t) in the domain knowledge graph, the tensor product operation M is performed. h,r,t =e h ⊙r⊙e t Construct its embedding matrix representation, where e h and e t For the elements in the entity vector representation E1' of the domain dataset or the entity vector representation E2' of the general dataset, the head entity vector, relation vector, and tail entity vector are fused using the tensor product operation to form a multidimensional tensor. The semantic information of the triples is then encoded into a matrix, ultimately constructing the triple embedding matrix M1∈R of the domain dataset. L ×L×L Similarly, the same operation is performed on the triples in the general knowledge graph to generate the triple embedding matrix M2∈R of the general dataset. L×L×L ;
[0029] Step S203: Perform L2 normalization on the embedding matrix M1 generated from the domain dataset and the embedding matrix M2 generated from the general dataset to obtain the normalized embedding matrix M1' of the domain dataset and the embedding matrix M2' of the general dataset. This is done by dividing each element of the matrices M1 and M2 by their norm, so that the length (modulus) of each vector in the matrix becomes 1.
[0030] The calculation formula in step S201 is as follows:
[0031] f(h,r,t)=-||h+rt||2
[0032] ξ=∑ (h,r,t∈S) ∑ (h,r,t∈S') |γ+f(h,r,t)-f(h',r,t')|
[0033] In the above formula, f(h,r,t) represents the scoring function, which is used to evaluate the rationality of the triples. The smaller the value, the more reasonable the relationship of the triples in the vector space. ξ represents the minimization objective function, which is used to optimize the entity vectors. S represents the domain data triple set S1 or the general dataset triple set S2. S' represents the set of negative triples generated by randomly replacing the head or tail entity in the positive triples. h' is the replaced head entity vector in the negative triples. t' is the replaced tail entity vector in the negative triples. γ represents the marginal hyperparameter, which is used to control the interval between positive and negative examples in the vector space to avoid overfitting of the model.
[0034] The calculation formula in step S203 is as follows:
[0035]
[0036] In the above formula, ||M|| represents the L2 norm of the matrix, M represents the embedding matrix M1 generated from the domain dataset or the embedding matrix M2 generated from the general dataset, L represents the size of the ternary matrix M in a certain dimension, i.e., the order of the matrix, i, j, k represent the index variables for each triple summation, i.e., the element positions in the three dimensions of the ternary matrix M, M ijk M represents the element value of the ternary matrix M in the i-th row, j-th column, and k-th layer; M' represents the normalized domain dataset embedding matrix M1' or the general dataset embedding matrix M2'.
[0037] Further, step S3 specifically includes:
[0038] Step S301: Embed the normalized domain knowledge embedding matrix M1'∈R from S203 L×L×L Based on this, a directed weighted inference graph model G = (V, B, W) is constructed, and the specific construction methods of each component are as follows:
[0039] (1) Node set V: Consists of entities in the domain knowledge graph, denoted as V = E1', where each node v i Each node in V corresponds one-to-one with an entity in the domain knowledge graph. If there are |E1'| entities in the domain knowledge graph, then the set of nodes can be represented as V = {v1, v2, ..., v...} |E1'|};
[0040] (2) Edge set B: Edges in the graph (v i ,v j )∈B represents entity v i With v j The semantic association between entities v is encoded in the normalized domain knowledge graph matrix M1; in the domain knowledge graph, if there exists a relationship between entity v and entity v... i Point to v j The relationship path is defined in the graph model (v i ,v j )∈B;
[0041] (3) Weight matrix W: To accurately measure the semantic strength of an edge, a weight function f is introduced. For an edge (v i ,v j ), its weight element w ij via f(M1'[h i ,r ij ,t j ]) calculate, where h i and t j The head entity v i With tail entity v j The vector representation in M1', r ij The corresponding relation vector is used; the weight function f adopts a proportional weighting strategy, which determines the weight of the edge by the proportion of the relation vector's magnitude to the sum of all relation vectors' magnitudes. That is, the larger the magnitude of the relation vector, the higher the weight of the corresponding edge, and the more important the relation is in the knowledge graph.
[0042] Step S302: Locate the corresponding node v in the inference graph model G of the key entity set E3 extracted from the interactive dataset in S201. c Starting from this point, a breadth-first search algorithm is executed, as follows:
[0043] (1) Initialize an empty queue Q, enqueue the starting node, and mark v. c For items that have been visited, denoted as visited(v c =True;
[0044] (2) When queue Q is not empty, perform the following operations:
[0045] 1) Retrieve the head node of the queue, denoted as u;
[0046] 2) Traverse all unvisited adjacent nodes v of node u. For each adjacent node v:
[0047] i) Calculate the relation vector r uv Relationship vector r with interactive content c Cosine similarity;
[0048] ii) Update the association score of node v according to the association score formula;
[0049] iii) Mark v as visited, i.e., visited(v) c = True, and enqueue v;
[0050] To effectively improve computational efficiency, a pruning optimization strategy is adopted. Constraints are constructed by setting an association score threshold β and an upper limit on the search depth: when the association score s(v) of node v is less than the threshold β, or when the association score ... the association score of node v is less than the threshold β. c When the path length path_length(v) to node v exceeds the upper limit, the expansion operation on that node and its subsequent paths is terminated immediately.
[0051] The path length path_length(v) is precisely calculated using a breadth-first search algorithm. Specifically, during the node enqueue phase, its hierarchical information is recorded, and the initial node v... c The path length is defined as 0, and the path length is incremented by 1 for each level of nodes traversed.
[0052] Step S303: After completing the traversal operation of the domain knowledge graph G = (V, E1', W), obtain the set of association scores {S(v)}v of all nodes in the node set V. ∈V To eliminate the impact of differences on the comparison of results and enhance the comparability between scores at different nodes, a linear normalization method is used to standardize the score set.
[0053] Based on the known threshold N for the number of candidate entities, the top N nodes in terms of association score are selected, and their corresponding entities form the candidate entity set CE = {e...} i} N i=1 Where N represents the number of candidate entities to be selected, and e i Let i represent the i-th entity in the candidate entity set (i = 1, 2, ..., N). Then, evaluate the reliability of the candidate entities and introduce execution degree calculation. Combined with the frequency of the entity's appearance in the knowledge graph, use the sum confidence calculation formula to sort the candidate entities by confidence degree and perform a second screening to ensure that the final output candidate entities are highly relevant to the interactive content.
[0054] The calculation formula in step S301 is as follows:
[0055]
[0056] In the above formula, w ij Represents an edge (v) i ,v j The weights of |r ij | represents the relation vector r ij The larger the modulus, the higher the importance of the relation in the semantic space; This represents the sum of the magnitudes of all relation vectors, used as the denominator for normalization, so that the weight of each edge is between 0 and 1;
[0057] The calculation formula in step S302 is as follows:
[0058]
[0059] S'(v)=S(v)+α*sim(r uv ,r c )*S(u)*w uv
[0060] In the above formula, sim(r) uv ,r c ) represents the relation vector r uv With relation content vector r c The cosine similarity; S(v) represents a certain score or weight value of node v, which is dynamically updated based on subsequent calculation results. S(v) is the value before the update, and S'(v) is the value after the update; α is a hyperparameter used to control the degree of influence of the entire update item; S(u) represents the score or weight value of node u, which is used to influence the update of node v, reflecting the ability of node u to influence node v; w uv It is the connection weight from node u to node v, reflecting the strength or importance of the relationship between nodes u and v. The larger the weight, the closer the connection between the two.
[0061] The calculation formula in step S303 is as follows:
[0062]
[0063] In the above formula, S'(v) represents the normalized association score. The normalized score is in the interval [0,1], which is a difference distribution of the original scores that are too large, so that the association scores of different nodes are comparable. S(v) represents the original association score of node v mentioned above. min({S(v)}) represents the minimum value among all node association scores in the set S(v), and max({S(v)}) represents the maximum value among all node association scores in the set S(v).
[0064]
[0065] In the above formula, CF(e i ) represents entity e i The overall confidence score is quantified by weighted fusion of three different evaluation indicators to determine the reliability and importance of entities in knowledge base reasoning scenarios; γ1, γ2, γ3∈[0,1] and γ1+γ2+γ3=1, which are weight parameters used to balance the influence of normalized association score, node degree, and occurrence frequency on confidence score; S'(v) represents the confidence score of node v. i The normalized correlation score, H(v) i ) is node v i In a reasoning graph, the degree reflects the richness of the connections between entities, f(e i ) represents entity e i The frequency of an entity's appearance in a domain knowledge graph indicates its importance and prevalence within the domain knowledge.
[0066] Further, step S4 specifically includes:
[0067] Step S401: The normalized general knowledge embedding matrix M2' from S203 is directly used as input data to construct the input layer of the neural network. This directly affects the subsequent network's data processing efficiency and feature extraction capability. The dimension of M2' is R. L×L×L This three-dimensional structure utilizes the multi-relation embedding characteristics of knowledge graphs, where the three dimensions correspond to different semantic association dimensions of knowledge. To match the design principles of neural networks, the number of neurons in the input layer must match the dimensions of the input data; therefore, the number of neurons in this input layer is L. 3 To meet the input data format requirements of subsequent network layers, the tensor flattening operation function is used to transform the dimension of M2', reshaping it from a three-dimensional matrix into a one-dimensional vector X∈R in sequence. L3 That is, the operation X = Flatten(M2') is performed, and the generated X will be used as the input data for subsequent network layers;
[0068] Step S402: After the input layer, a multi-layer self-focusing network consisting of Z cascaded self-focusing layers is constructed. This network adaptively focuses on key regions of the input data, effectively extracting and representing local features. For the z-th self-focusing layer (z∈{1,2,…,Z}), let the input be the output X of the previous layer. (z-1) The output is X (z) Its calculation process follows these steps:
[0069] First, the input feature matrix X is transformed through three independent linear transformation operations. (z-1) Mapped to query vector Q respectively (z) Key vector K (z) Sum vector U (z) The specific formula is shown below:
[0070]
[0071] In the above formula, These are all trainable weight matrices, whose function is to perform a linear transformation on the input data. The bias vector is used to adjust the result after linear transformation, thereby improving the model's expressive power; {·} represents matrix multiplication.
[0072] Then, the attention score matrix A is calculated using the scaled dot product attention mechanism. (z) The specific formula is shown below:
[0073]
[0074] In the above formula, d k For the key vector K (z) In terms of dimensions, the Softmax function normalizes the scores, making A... (z) The sum of each row of elements is 1, thus quantifying the importance weight of features at different positions;
[0075] Finally, the value vectors are weighted and summed using attention weights to obtain the output feature matrix of this layer, as shown in the following formula:
[0076] X (z) =A (z) ·U (z)
[0077] In the above formula, X (z) This represents the output of the self-focusing layer at layer z, which is the final vector obtained after weighted summation. It will be used as input for subsequent calculations or other layers of the model; A (z)U is the attention weight matrix of the z-th layer. Each element in this matrix represents the importance of the corresponding feature in the weighted summation; the larger the value, the greater the contribution of the corresponding feature to the final output. (z) Let A be the feature vector matrix of the z-th layer, containing various feature information extracted from the input data at this layer. This is then combined with the attention weight matrix A. (z) Multiplication enables weighted combinations of different features;
[0078] Step S403: After the multi-layer self-focusing network, a general feedforward neural network is connected. This network architecture consists of several fully connected layers, designed to capture the global dependencies of the input data and optimize the output X of the self-focusing network. (z) The data is sequentially fed into n fully connected layers for processing.
[0079] t (1) =ReLU(w (1) ·X (z) +b (1) )
[0080] t (2) =ReLU(w (2) ·X (z) +b (2) ...
[0081] t (n) =ReLU(w (n) ·X (z) +b (n) )
[0082] In the above formula, t (i) (i = 1, 2, ..., n) represents the output of the i-th fully connected layer, t (1) It is the first fully connected layer to output X of the self-focusing network. (z) The processed output, subsequent t (i) (i>1) is based on the previous layer t (i-1) The calculation results, w (i) (i = 1, 2, ..., n) The weight matrix of the i-th fully connected layer, whose dimension determines the mapping relationship between the input and output, is used to perform weighted transformation on the input data. (i) The bias vector of the i-th fully connected layer, ReLU represents the activation function, which introduces non-linearity into the neural network;
[0083] Step S404: Train the constructed general-purpose model using the backpropagation algorithm and define the loss function. During model training, the difference between the predicted output and the true label is quantitatively evaluated using the loss function. Further, the mean squared error loss function is selected. This function measures the overall deviation between the model's prediction and the actual result by calculating the square of the difference between the predicted value and the true value for each training sample and averaging the results over all samples. The specific formula is shown below:
[0084]
[0085] In the above formula, δ represents the model's loss function, used to measure the difference between the model's prediction and the actual result. The smaller the loss function value, the better the model's prediction performance. N represents the number of samples, i.e., the total number of samples used to calculate the loss. I represents an index variable used to iterate through each sample from 1 to N, i∈{1,2,...,N}. Let y represent the model's prediction for the i-th sample. (i) This represents the true value of the i-th sample.
[0086] Further, step S5 specifically includes:
[0087] Step S501: To effectively integrate the inference graph model from step S3 with the general large model from step S4, a unified data interface and format specification system is constructed. For the inference graph model G = (V, B, W), its output data includes the set of node association scores S'(v) and the set of node degree information H(v). i The model consists of a set of candidate entities (CE) and a set of candidate entities (CE); the output data of the general large model is defined as the output feature vector t of the last layer of the model. (n) ;
[0088] Furthermore, a feature mapping strategy is used to perform data alignment, addressing the significant differences in dimensionality and semantics between the output data of the two models. A linear transformation maps features such as node association scores and node degrees from the inference graph model to the same dimensionality space as the output feature vector of the general large model. Setting the output data of the inference graph model to O, the aligned data can be obtained after the transformation. The specific formula is shown below.
[0089] O = w o ·O+b o
[0090] In the above formula, w o Let b represent the alignment weight matrix. o This represents the alignment bias vector, and the initial value of the parameter can be determined by minimizing the difference between the aligned data and the output data of the general large model.
[0091] Step S502: A hybrid fusion mechanism is adopted to organically integrate the inference graph model with the general large model to achieve synergistic enhancement of knowledge reasoning capabilities. Specifically, a gated fusion unit is used as the core architecture, and its inputs are the aligned inference graph model data O and the output feature vector t of the general large model at layer n. (n) The output is the fused feature vector Y. The fusion process is implemented using the following formula:
[0092] Y = σ(w) g [O;t (n) ]+b g )·O+(1-σ(w g [O;t (n) ]+b g ))·t (n)
[0093] Among them, [O;t (n) ] indicates a concatenation operation between two data points, w g and b g These are the trainable weight matrix and bias vector, respectively, and σ is the Sigmoid activation function, which dynamically adjusts the fusion weights of the two model outputs through nonlinear transformation. This mechanism is based on a task-driven adaptive fusion strategy, effectively combining the inference graph model's ability to deeply mine knowledge graph structure information with the advantages of general large models in semantic feature extraction, thereby achieving a significant improvement in inference performance.
[0094] Step S503: Construct a joint optimization objective function δ1 to quantify the deviation between the predicted output and the true value of the fusion model. Based on the cross-entropy principle in information theory, a cross-entropy loss function is used for modeling.
[0095]
[0096] In the above formula, N represents the total number of training samples, H is the total number of categories in the classification task, and y ih and These correspond to the true label and predicted probability of the i-th sample in the h-th class, respectively.
[0097] Furthermore, iterative optimization is performed using a gradient-based optimization strategy. During each training iteration, the parameter set related to edge weight calculation in the inference graph model, the neural network parameters of the general large model, and the weight parameters w of the gated fusion unit are updated synchronously. g and bias parameter b g Through multiple rounds of iterative training, the parameter configuration of the dual-model architecture and fusion mechanism is continuously adjusted so that the loss function value of the fusion model on the training dataset gradually decreases. When the loss function meets the preset convergence condition (such as the gradient norm being less than a threshold) or reaches the maximum number of iterations, the training process is terminated.
[0098] The beneficial effects of this invention include:
[0099] (1) This invention first performs data cleaning, standardization, deduplication and word segmentation preprocessing on interactive datasets, domain datasets and general datasets. On the one hand, it judges and removes noisy data by feature threshold to ensure data accuracy. On the other hand, it unifies the format and scale of multi-source data to eliminate the difference in units. At the same time, it performs deduplication to reduce data redundancy. Word segmentation process transforms text into structured word sequences, providing a high-quality data foundation for subsequent knowledge graph embedding.
[0100] (2) This invention uses a knowledge graph embedding model to construct a triple embedding matrix and performs L2 normalization. This method maps entities and relations to a low-dimensional vector space, enabling the semantic relations to be expressed quantitatively. The tensor product operation encodes the semantics of the triples into a multi-dimensional matrix, and the normalization process eliminates vector length bias, providing a structured and standardized knowledge representation for reasoning and improving the computational stability of subsequent models.
[0101] (3) Based on the reasoning graph model, this invention combines breadth-first search and pruning strategies to traverse the domain knowledge graph. It calculates the correlation degree of the relation vector through cosine similarity, dynamically updates the node score, and combines path length threshold and score threshold pruning to efficiently filter candidate entities. It introduces confidence calculation to integrate normalized score, node degree and occurrence frequency to ensure that the filtered entities are not only semantically related to the interactive content, but also have importance and universality in the knowledge graph.
[0102] (4) This invention constructs a multi-layer self-focusing network and a general feedforward neural network. The self-focusing layer automatically captures the local key features of the input data through the attention mechanism, while the feedforward network extracts the global dependency through the fully connected layer. The ReLU activation function enhances the nonlinear expression capability, and backpropagation combined with the mean squared error loss function is used for parameter optimization, so that the model can not only capture detailed features but also understand the overall semantics, thereby improving prediction accuracy and generalization ability.
[0103] (5) This invention integrates the advantages of both reasoning graph models and general large models. The gated fusion mechanism dynamically adjusts the weights through Sigmoid. When the reasoning graph model captures entity associations in the knowledge graph structure analysis, the general large model extracts semantic features simultaneously. The two complement each other to improve the stability of the algorithm. The fusion of different feature representation methods can cover more comprehensive data dimensions and effectively deal with noise and data distribution imbalance problems.
[0104] (6) This invention realizes a collaborative optimization mechanism for data alignment and gating fusion models. Feature mapping unifies the output dimension of heterogeneous models. The joint optimization objective function updates the parameters of the two models synchronously through cross-entropy and gradient descent, so that the knowledge graph structure information of the reasoning graph model is deeply integrated with the semantic features of the general large model. This not only enhances the multi-source knowledge processing capability, but also adapts to the accurate reasoning needs in multiple scenarios such as intelligent decision-making and risk prevention.
[0105] Other advantages, objectives, and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination, or may be learned from practice of the invention. The objectives and other advantages of the invention can be realized and obtained from the following description and the foregoing claims. Attached Figure Description
[0106] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will now be described in further detail with reference to the accompanying drawings, wherein:
[0107] Figure 1 This is a schematic diagram of a knowledge base-based reasoning comparison method according to the present invention.
[0108] Figure 2 This is a metadata graph representing the domain data of the present invention;
[0109] Figure 3 This is a diagram of the internal structure of the general-purpose large model of the present invention;
[0110] Figure 4 This is a schematic diagram illustrating the working principle of the reasoning and comparison method of the present invention.
[0111] Figure 5 This is a diagram showing the intelligent question-answering results of the Mingtu large model of the present invention. Detailed Implementation
[0112] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be understood that the preferred embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0113] like Figure 1 As shown, the reasoning comparison method based on a knowledge base of the present invention specifically includes the following steps:
[0114] Step S1: Establish interactive datasets, domain datasets, and general datasets, and perform data preprocessing on the three databases respectively;
[0115] Step S2: Embed the preprocessed data using a knowledge graph and construct a ternary embedding matrix;
[0116] Step S3: Use the inference graph model to traverse the domain knowledge graph and search for candidate entities related to the interactive content;
[0117] Step S4: Using a multi-layer self-focusing network and a feedforward neural network, capture the global dependencies and local feature representations of the input data to build a large model and pre-train the large model;
[0118] Step S5: Utilize the implicit knowledge within the pre-trained large model and the explicit knowledge within the domain knowledge graph, and gradually generate accurate prediction or analysis results through multi-dimensional feature fusion and optimization iteration.
[0119] In this embodiment, step S1 specifically includes the following steps:
[0120] Step S101: Establish an interactive dataset, a domain dataset, and a general dataset. The interactive dataset provides dynamic scene data directly related to the analysis task to clarify the task context; the domain dataset provides structured knowledge from a specific professional field to support the construction of a domain knowledge graph and professional reasoning; the general dataset provides cross-domain common sense and basic logic to ensure the model's basic understanding and generalization ability. These three datasets work together to provide the data foundation for subsequent processes. Perform data cleaning on the domain dataset D1, general dataset D2, and interactive dataset D3 to remove noisy and invalid data. Let the domain dataset be... The general dataset is Interactive dataset is Where d i d j and d k These represent data instances in the domain dataset, the general dataset, and the interactive dataset, respectively, with I, J, and K representing the total number of data instances in the three datasets.
[0121] Define a unified noise data judgment function f(d). When f(d) = 1, d is judged as noise data; when f(d) = 0, d is considered as valid data. f(d) is set based on a feature threshold. If x l If the value is less than θ1 or θ2 (θ1 and θ2 are pre-set feature thresholds), then f(d) = 1, and the numerical feature x... l (l is the feature dimension index) is a numerical indicator used to describe a certain aspect of an attribute of a data instance; for each data instance d, multiple numerical features x can be used. l To represent its different dimensions of characteristics;
[0122] After cleaning, the domain dataset and the general dataset were updated as follows: and θ1 and θ2 are used to define the reasonable range of values for data features. Data exceeding this range will be identified as noise and removed, thus ensuring the accuracy and effectiveness of data processing. The domain data metadata diagram is shown below. Figure 2 As shown;
[0123] In this embodiment, the domain dataset D1 consists of 1000 customer credit records from a bank, containing three numerical features: "monthly income," "debt ratio," and "credit score." The general dataset D2 consists of 500 publicly available industry credit data entries, and the interactive dataset D3 consists of 200 real-time credit consultation texts from users. The feature thresholds for the noise data judgment function f(d) are defined as θ1 = 0.2 and θ2 = 8000 (e.g., "debt ratio" less than 0.2 or greater than 0.8, and "monthly income" less than 2000 or greater than 80000 are considered noise). During cleaning, a record in D1 with "monthly income 1500 (<2000), debt ratio 0.7" is removed because it triggers f(d) = 1. A record in D2 with "monthly income 5000, debt ratio 0.3" is retained because its features are within the thresholds. The query in D3 with "monthly income 30000" retains its features effectively after parsing. After cleaning D1, 920 valid data points were obtained, D2 yielded 470, and D3 yielded 195, verifying that the noise data judgment function based on feature thresholds can effectively remove outliers.
[0124] Step S102: Process the cleaned domain dataset General Dataset and interactive datasets Standardization is performed to unify the data to the same format and metric, resulting in a standardized domain dataset. General Dataset and interactive datasets For the numerical features in all three datasets, a standardization formula was used, as shown below:
[0125]
[0126] Where μ represents the feature in the merged dataset (where μ is the sum of the features in the merged dataset). and The mean of the merged dataset (x) is used to characterize the central tendency of the data. i x j x kThese correspond to the numerical features of a certain attribute of a certain data instance in the three datasets. Through this standardization formula, data of different magnitudes and distribution characteristics can be transformed to the same scale, which is convenient for subsequent data analysis and model training. σ is the standard deviation of the feature in the merged dataset, which is used to measure the dispersion of the data. x' is a numerical feature, which represents the standardized numerical index used to describe a certain attribute of a certain data instance.
[0127] Step S103: Process the standardized domain dataset and general datasets A deduplication operation is performed to eliminate duplicate data records. A hash function h(d) is used to hash the data records in both datasets, mapping each data record d to a fixed-length hash value. If h(d1) = h(d2) (where d1 and d2 come from the domain dataset or the general dataset respectively, and d1 = d2), and further comparison confirms that the key features of d1 and d2 are completely identical, then these two data records are determined to be duplicates, and only one is retained while the other is deleted. After deduplication, the domain dataset and the general dataset are updated to... and
[0128] In this embodiment, let the standardized domain dataset D1 be... 2 The dataset D2 contains 1000 customer credit records from a bank (including customer ID, ID number, loan amount, and loan term). 2 It contains 500 publicly available industry credit data entries; using the MD5 hash function as h(d), the hash value is calculated by concatenating the key features (customer ID + ID card number) of each data entry. For example, D1 2 In the context of d1 = {"Customer ID":C001,"ID Number":110101199001011234,"Loan Amount":50000,"Term":12} and D2 2 In the data d2 = {"Customer ID":C001,"ID Number":110101199001011234,"Loan Amount":50000,"Term":12}, after hash calculation: h(d1) = h(d2) = a1b2c3d4, further comparison shows that the key features are completely consistent, so it is determined to be duplicate data, and only D1 is retained. 2 d1 in; and D1 2In the original text, d3 = {"Customer ID":C002,"ID Number":110101199102022345,"Loan Amount":50000,"Term":12} and d4 = {"Customer ID":C002,"ID Number":110101199102022345,"Loan Amount":50000,"Term":24}, although h(d3) = h(d4) = e5f6g7h8, the loan terms are different, so they are not considered duplicates; after deduplication, D1 3 980 records are retained (20 duplicates are removed), D2 3 485 records were retained (15 duplicates were removed).
[0129] Step S104: Process the deduplicated domain dataset General Dataset and the standardized interactive dataset Word segmentation is performed to convert the text data into a word sequence; a statistical word segmentation method is adopted, and the probability function of the word segmentation model is known to be P(w1,w2,…,w…). s |T), where T is the input text data, w1, w2, ..., w s The sequence of words obtained after word segmentation is s, where s is the length of the word sequence. The probability function is statistical information obtained from training on a large corpus, which helps the model determine the segmentation method of words in the text, transforming the text data into a computer-processable word sequence form, and is used to calculate the word sequence w1, w2, ..., w3 given a text T. s The probability is calculated, and the optimal word segmentation result is obtained by maximizing this probability function. After word segmentation, the domain dataset, general dataset, and interactive dataset are updated to D1', D2', and D3', respectively.
[0130] Step S2 specifically includes the following steps:
[0131] Step S201: Perform knowledge graph embedding processing on the deduplicated and segmented domain dataset D1', general dataset D2', and interactive dataset D3' from S104. First, extract the entity set E1 from the domain dataset D1' using part-of-speech tagging, ensuring that the entities in E1 cover the key concepts and objects within the domain. Then, use a knowledge graph embedding model to represent the triples (h,r,t) (where h is the head entity, r is the relation, and t is the tail entity) in the knowledge graph as relations in the vector space, i.e., h⊕r≈t. This yields the domain data triple set S1 and the general data triple set S2. Optimize the entity vectors by minimizing the objective function, and then use a scoring function. After multiple rounds of iterative training, obtain the entity vector representation E1' of the domain dataset.
[0132] For the general dataset D2' and the interactive dataset D3', after extracting the entity sets E2 and E3 using a similar method, the entities in the general dataset are embedded using the above embedding method and optimization strategy to obtain the entity vector representation E2' of the general dataset and the entity vector representation E3' of the interactive dataset.
[0133] Step S202: Using the same model as the entity embedding in S201, extract relation sets R1 and R2 from the domain dataset D1' and the general dataset D2' respectively. Map the relations to a vector space of the same dimension L as the entity vectors to obtain the relation vector r∈R L During the mapping process, the model learns the semantic features of the relationship, so that relationships with similar semantics are closer in the vector space;
[0134] For each triple (h, r, t) in the domain knowledge graph, the tensor product operation M is performed. h,r,t =e h ⊙r⊙e t Construct its embedding matrix representation, where e h and e t For the elements in the entity vector representation E1' of the domain dataset or the entity vector representation E2' of the general dataset, the head entity vector, relation vector, and tail entity vector are fused using the tensor product operation to form a multidimensional tensor. The semantic information of the triples is then encoded into a matrix, ultimately constructing the triple embedding matrix M1∈R of the domain dataset. L ×L×L Similarly, the same operation is performed on the triples in the general knowledge graph to generate the triple embedding matrix M2∈R of the general dataset. L×L×L ;
[0135] Step S203: Perform L2 normalization on the embedding matrix M1 generated from the domain dataset and the embedding matrix M2 generated from the general dataset to obtain the normalized embedding matrix M1' of the domain dataset and the embedding matrix M2' of the general dataset. This is done by dividing each element of the matrices M1 and M2 by their norm, so that the length (modulus) of each vector in the matrix becomes 1.
[0136] The calculation formula in step S201 is as follows:
[0137] f(h,r,t)=-||h+rt||2
[0138] ξ=∑ (h,r,t∈S) ∑ (h,r,t∈S') |γ+f(h,r,t)-f(h',r,t')|
[0139] In the above formula, f(h,r,t) represents the scoring function, which is used to evaluate the rationality of the triples. The smaller the value, the more reasonable the relationship of the triples in the vector space. ξ represents the minimization objective function, which is used to optimize the entity vectors. S represents the domain data triple set S1 or the general dataset triple set S2. S' represents the set of negative triples generated by randomly replacing the head or tail entity in the positive triples. h' is the replaced head entity vector in the negative triples. t' is the replaced tail entity vector in the negative triples. γ represents the marginal hyperparameter, which is used to control the interval between positive and negative examples in the vector space to avoid overfitting of the model.
[0140] The calculation formula in step S203 is as follows:
[0141]
[0142] In the above formula, ||M|| represents the L2 norm of the matrix, M represents the embedding matrix M1 generated from the domain dataset or the embedding matrix M2 generated from the general dataset, L represents the size of the ternary matrix M in a certain dimension, i.e., the order of the matrix, i, j, k represent the index variables for each triple summation, i.e., the element positions in the three dimensions of the ternary matrix M, M ijk M represents the element value of the ternary matrix M in the i-th row, j-th column, and k-th layer; M' represents the normalized domain dataset embedding matrix M1' or the general dataset embedding matrix M2'.
[0143] Step S3 specifically includes the following steps:
[0144] Step S301: Use the normalized domain knowledge embedding matrix M1'∈R from S203 L×L×L Based on this, a directed weighted inference graph model G = (V, B, W) is constructed, and the specific construction methods of each component are as follows:
[0145] (1) Node set V: Consists of entities in the domain knowledge graph, denoted as V = E1', where each node v i Each node in V corresponds one-to-one with an entity in the domain knowledge graph. If there are |E1'| entities in the domain knowledge graph, then the set of nodes can be represented as V = {v1, v2, ..., v...} |E1'|};
[0146] (2) Edge set B: Edges in the graph (v i ,v j )∈B represents entity v i With v j The semantic association between entities v is encoded in the normalized domain knowledge graph matrix M1; in the domain knowledge graph, if there exists a relationship between entity v and entity v... i Pointing to v jThe relationship path is defined in the graph model (v i ,v j )∈B;
[0147] (3) Weight matrix W: To accurately measure the semantic strength of an edge, a weight function f is introduced. For an edge (v i ,v j ), its weight element w ij via f(M1'[h i ,r ij ,t j ]) calculate, where h i and t j The head entity v i With tail entity v j The vector representation in M1', r ij The corresponding relation vector is used; the weight function f adopts a proportional weighting strategy, which determines the weight of the edge by the proportion of the relation vector's magnitude to the sum of all relation vectors' magnitudes. That is, the larger the magnitude of the relation vector, the higher the weight of the corresponding edge, and the more important the relation is in the knowledge graph.
[0148] Step S302: Locate the corresponding node v in the inference graph model G of the key entity set E3 extracted from the interactive dataset in S201. c Starting from this point, a breadth-first search algorithm is executed, as follows:
[0149] (1) Initialize an empty queue Q, enqueue the starting node, and mark v. c For items that have been visited, denoted as visited(v c =True;
[0150] (2) When queue Q is not empty, perform the following operations:
[0151] 1) Retrieve the head node of the queue, denoted as u;
[0152] 2) Traverse all unvisited adjacent nodes v of node u. For each adjacent node v:
[0153] i) Calculate the relation vector r uv Relationship vector r with interactive content c Cosine similarity;
[0154] ii) Update the association score of node v according to the association score formula;
[0155] iii) Mark v as visited, i.e., visited(v) c = True, and enqueue v;
[0156] To effectively improve computational efficiency, a pruning optimization strategy is adopted. Constraints are constructed by setting an association score threshold β and an upper limit on the search depth: when the association score s(v) of node v is less than the threshold β, or when the association score ... the association score of node v is less than the threshold β. c When the path length path_length(v) to node v exceeds the upper limit, the expansion operation on that node and its subsequent paths is terminated immediately.
[0157] The path length path_length(v) is precisely calculated using a breadth-first search algorithm. Specifically, during the node enqueue phase, its hierarchical information is recorded, and the initial node v... c The path length is defined as 0, and the path length is incremented by 1 for each level of nodes traversed.
[0158] Step S303: After completing the traversal operation of the domain knowledge graph G = (V, E1', W), obtain the set of association scores {S(v)}v of all nodes in the node set V. ∈V To eliminate the impact of differences on the comparison of results and enhance the comparability between scores at different nodes, a linear normalization method is used to standardize the score set.
[0159] Based on the known threshold N for the number of candidate entities, the top N nodes in terms of association score are selected, and their corresponding entities form the candidate entity set CE = {e...} i} N i=1 Where N represents the number of candidate entities to be selected, and e i Let i represent the i-th entity in the candidate entity set (i = 1, 2, ..., N). Then, evaluate the reliability of the candidate entities and introduce execution degree calculation. Combined with the frequency of the entity's appearance in the knowledge graph, use the sum confidence calculation formula to sort the candidate entities by confidence degree and perform a second screening to ensure that the final output candidate entities are highly relevant to the interactive content.
[0160] In this embodiment, the preset threshold for the number of candidate entities is N=3. The interaction content is "Loan Risk Assessment for High-Income Customers," and the starting entity is "High-Income Customers." The entity association scores calculated by the inference graph model are ranked as follows: "Credit Score ≥ 700" (0.85), "No Bad Loan Records" (0.78), "Stable Occupation" (0.72), "Debt-to-Asset Ratio < 0.5" (0.68), and "Loan Term ≤ 3 Years" (0.65). After selecting the first three entities, weight parameters γ1 = 0.5, γ2 = 0.3, and γ3 = 0.2 are set, and the node degree of each entity is calculated ("Credit Score ≥ 700" degree = 10, "No Bad Loan Records" degree = 10). The confidence scores were calculated using the following parameters: "credit score ≥ 700" (200 occurrences), "no bad loan record" (150 occurrences), and "stable occupation" (120 occurrences, with a maximum frequency of 200 occurrences). After normalization, the overall confidence scores were calculated as follows: "credit score ≥ 700" = 0.5 × 0.85 + 0.3 × 1.0 + 0.2 × 1.0 = 0.925; "no bad loan record" = 0.5 × 0.78 + 0.3 × 0.8 + 0.2 × 0.75 = 0.78; and "stable occupation" = 0.5 × 0.72 + 0.3 × 0.6 + 0.2 × 0.6 = 0.66. The candidate entity set was then determined by sorting the confidence scores.
[0161] The calculation formula in step S301 is as follows:
[0162]
[0163] In the above formula, w ij Represents an edge (v) i ,v j The weights of |r ij | represents the relation vector r ij The larger the modulus, the higher the importance of the relation in the semantic space; This represents the sum of the magnitudes of all relation vectors, used as the denominator for normalization, so that the weight of each edge is between 0 and 1;
[0164] The calculation formula in step S302 is as follows:
[0165]
[0166] S'(v)=S(v)+α*sim(r uv ,r c )*S(u)*w uv
[0167] In the above formula, sim(r) uv ,r c ) represents the relation vector r uv With relation content vector r cThe cosine similarity; S(v) represents a certain score or weight value of node v, which is dynamically updated based on subsequent calculation results. S(v) is the value before the update, and S'(v) is the value after the update; α is a hyperparameter used to control the degree of influence of the entire update item; S(u) represents the score or weight value of node u, which is used to influence the update of node v, reflecting the ability of node u to influence node v; w uv It is the connection weight from node u to node v, reflecting the strength or importance of the relationship between nodes u and v. The larger the weight, the closer the connection between the two.
[0168] The calculation formula in step S303 is as follows:
[0169]
[0170] In the above formula, S'(v) represents the normalized association score. The normalized score is in the interval [0,1], which is a difference distribution of the original scores that are too large, so that the association scores of different nodes are comparable. S(v) represents the original association score of node v mentioned above. min({S(v)}) represents the minimum value among all node association scores in the set S(v), and max({S(v)}) represents the maximum value among all node association scores in the set S(v).
[0171]
[0172] In the above formula, CF(e i ) represents entity e i The overall confidence score is quantified by weighted fusion of three different evaluation indicators to determine the reliability and importance of entities in knowledge base reasoning scenarios; γ1, γ2, γ3∈[0,1] and γ1+γ2+γ3=1, which are weight parameters used to balance the influence of normalized association score, node degree, and occurrence frequency on confidence score; S'(v) represents the confidence score of node v. i The normalized correlation score, H(v) i ) is node v i In a reasoning graph, the degree reflects the richness of the connections between entities, f(e i ) represents entity e i The frequency of an entity's appearance in a domain knowledge graph indicates its importance and prevalence within the domain knowledge.
[0173] Step S4 specifically includes the following steps:
[0174] Step S401: The normalized general knowledge embedding matrix M2' from S203 is directly used as input data to construct the input layer of the neural network. This directly affects the subsequent network's data processing efficiency and feature extraction capability. The dimension of M2' is R. L×L×LThis three-dimensional structure utilizes the multi-relation embedding characteristics of knowledge graphs, where the three dimensions correspond to different semantic association dimensions of knowledge. To match the design principles of neural networks, the number of neurons in the input layer must match the dimensions of the input data; therefore, the number of neurons in this input layer is L. 3 To meet the input data format requirements of subsequent network layers, the tensor flattening operation function is used to transform the dimension of M2', reshaping it from a three-dimensional matrix into a one-dimensional vector X∈R in sequence. L3 That is, the operation X = Flatten(M2') is performed, and the generated X will be used as the input data for subsequent network layers;
[0175] Step S402: After the input layer, a multi-layer self-focusing network consisting of Z cascaded self-focusing layers is constructed. This network adaptively focuses on key regions of the input data, effectively extracting and representing local features. For the z-th self-focusing layer (z∈{1,2,…,Z}), let the input be the output X of the previous layer. (z-1) The output is X (z) Its calculation process follows these steps:
[0176] First, the input feature matrix X is transformed through three independent linear transformation operations. (z-1) Mapped to query vector Q respectively (z) Key vector K (z) Sum vector U (z) The specific formula is as follows:
[0177]
[0178] In the above formula, These are all trainable weight matrices, whose function is to perform a linear transformation on the input data. The bias vector is used to adjust the result after linear transformation, thereby improving the model's expressive power; {·} represents matrix multiplication.
[0179] Then, the attention score matrix A is calculated using the scaled dot product attention mechanism. (z) The specific formula is as follows:
[0180]
[0181] In the above formula, d k For the key vector K (z) In terms of dimensions, the Softmax function normalizes the scores, making A... (z) The sum of each row of elements is 1, thus quantifying the importance weight of features at different positions;
[0182] Finally, the value vectors are weighted and summed using attention weights to obtain the output feature matrix of this layer, as shown in the following formula:
[0183] X (z) =A (z) ·U (z)
[0184] In the above formula, X (z) This represents the output of the self-focusing layer at layer z, which is the final vector obtained after weighted summation. It will be used as input for subsequent calculations or other layers of the model; A (z) U is the attention weight matrix of the z-th layer. Each element in this matrix represents the importance of the corresponding feature in the weighted summation; the larger the value, the greater the contribution of the corresponding feature to the final output. (z) Let A be the feature vector matrix of the z-th layer, containing various feature information extracted from the input data at this layer. This is then combined with the attention weight matrix A. (z) Multiplication enables weighted combinations of different features;
[0185] Step S403: After the multi-layer self-focusing network, a Mingtu general feedforward neural network is connected. This network architecture consists of several fully connected layers and aims to capture the global dependencies of the input data, thus controlling the output X of the self-focusing network. (z) The data is sequentially fed into n fully connected layers for processing.
[0186] t (1) =ReLU(w (1) ·X (z) +b (1) )
[0187] t (2) =ReLU(w (2) ·X (z) +b (2) ...
[0188] t (n) =ReLU(w (n) ·X (z) +b (n) )
[0189] In the above formula, t (i) (i = 1, 2, ..., n) represents the output of the i-th fully connected layer, t (1) It is the first fully connected layer to output X of the self-focusing network. (z) The processed output, subsequent t (i) (i>1) is based on the previous layer t (i-1) The calculation results, w (i)(i = 1, 2, ..., n) The weight matrix of the i-th fully connected layer, whose dimension determines the mapping relationship between the input and output, is used to perform weighted transformation on the input data. (i) The bias vector of the i-th fully connected layer, ReLU represents the activation function, which introduces non-linearity into the neural network;
[0190] Step S404: Train the constructed Mingtu general large model using the backpropagation algorithm and define the loss function. During model training, the difference between the predicted output and the true label is quantitatively evaluated using the loss function. Further, the mean squared error loss function is selected. This function measures the overall deviation between the model's prediction results and the actual results by calculating the square of the difference between the predicted value and the true value for each training sample and averaging the results across all samples. The specific formula is shown below:
[0191]
[0192] In the above formula, δ represents the model's loss function, used to measure the difference between the model's prediction and the actual result. The smaller the loss function value, the better the model's prediction performance. N represents the number of samples, i.e., the total number of samples used to calculate the loss. I represents an index variable used to iterate through each sample from 1 to N, i∈{1,2,...,N}. Let y represent the model's prediction for the i-th sample. (i) This represents the true value of the i-th sample. The Mingtu general large model is as follows: Figure 3 As shown.
[0193] Step S5 specifically includes the following steps:
[0194] Step S501: To effectively integrate the inference graph model from step S3 with the Mingtu general large model from step S4, a unified data interface and format specification system is constructed. For the inference graph model G = (V, B, W), its output data includes the set of node association scores S'(v) and the set of node degree information H(v). i The model consists of a set of candidate entities (CE) and a set of candidate entities (CE); the output data of the general large model is defined as the output feature vector t of the last layer of the model. (n) ;
[0195] Furthermore, a feature mapping strategy is used to perform data alignment, addressing the significant differences in dimensionality and semantics between the output data of the two models. A linear transformation maps features such as node association scores and node degrees from the inference graph model to the same dimensional space as the output feature vector of the Mingtu general-purpose model. Setting the output data of the inference graph model to O, the aligned data can be obtained after the transformation. The specific formula is shown below.
[0196] O = w o ·O+bo
[0197] In the above formula, w o Let b represent the alignment weight matrix. o This represents the alignment bias vector. The initial value of the parameter can be determined by minimizing the difference between the aligned data and the output data of the Mingtu general large model.
[0198] Step S502: A hybrid fusion mechanism is adopted to organically integrate the inference graph model with the general large model to achieve synergistic enhancement of knowledge reasoning capabilities. Specifically, a gated fusion unit is used as the core architecture, and its inputs are the aligned inference graph model data O and the output feature vector t of the general large model at layer n. (n) The output is the fused feature vector Y. The fusion process is implemented using the following formula:
[0199] Y = σ(w) g [O;t (n) ]+b g )·O+(1-σ(w g [O;t (n) ]+b g ))·t (n)
[0200] Among them, [O;t (n) ] indicates a concatenation operation between two data points, w g and b g These are the trainable weight matrix and bias vector, respectively, with σ being the Sigmoid activation function. This function dynamically adjusts the fusion weights of the two models' outputs through nonlinear transformation. This mechanism, based on a task-driven adaptive fusion strategy, effectively combines the deep mining capabilities of inference graph models for knowledge graph structure information with the advantages of general-purpose large models in semantic feature extraction, thereby achieving a significant improvement in inference performance. The system diagram of the inference comparison method is shown below. Figure 4 As shown;
[0201] Step S503: Construct a joint optimization objective function δ1 to quantify the deviation between the predicted output and the true value of the fusion model. Based on the cross-entropy principle in information theory, a cross-entropy loss function is used for modeling.
[0202]
[0203] In the above formula, N represents the total number of training samples, H is the total number of categories in the classification task, and y ih and These correspond to the true label and predicted probability of the i-th sample in the h-th class, respectively.
[0204] Furthermore, iterative optimization is performed using a gradient-based optimization strategy. During each training iteration, the parameter set related to edge weight calculation in the inference graph model, the neural network parameters of the Mingtu general large model, and the weight parameters w of the gated fusion unit are updated synchronously. g and bias parameter b g Through multiple rounds of iterative training, the parameter configurations of the dual-model architecture and fusion mechanism are continuously adjusted, causing the loss function value of the fusion model on the training dataset to gradually decrease. The training process terminates when the loss function meets the preset convergence condition (e.g., the gradient norm is less than a threshold) or the maximum number of iterations is reached. The Mingtu fusion large model question-answering effect is as follows: Figure 5 As shown.
[0205] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 The computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the functions specified in one or more boxes. Specific embodiments have been used to illustrate the principles and implementation of this invention. The above description of the embodiments is only for the purpose of helping to understand the method and core ideas of this invention; at the same time, for those skilled in the art, based on the ideas of this invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this invention.
[0206] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.
Claims
1. A reasoning comparison method based on a knowledge base, characterized in that: Includes the following steps: Step S1: Establish interactive dataset, domain dataset, and general dataset, and perform data preprocessing on the three datasets respectively. The preprocessing includes: data cleaning, standardization, deduplication, and word segmentation. Step S2: Embed the preprocessed data using a knowledge graph and construct a ternary embedding matrix; specifically including: Step S201: Perform knowledge graph embedding processing on the domain dataset D1', general dataset D2', and interactive dataset D3' after deduplication and word segmentation. Extract entity set E1 from D1'. Use the embedding model to represent the triples in the knowledge graph as relations in the vector space, thereby obtaining the domain data triple set S1. By minimizing the objective function and iterative training, obtain the domain entity vector E1'. Similarly process D2' and D3' to obtain E2 and E3, as well as the general data triple set S2, the entity vector representation of the general dataset E2', and the entity vector representation of the interactive dataset E3'. Step S202: Extract relation sets R1 and R2 from D1' and D2', map them to an L-dimensional vector space, and obtain the relation vector r∈R L For each triplet, an embedding matrix is constructed through tensor product operation, generating triplet embedding matrix M1 for the domain dataset and triplet embedding matrix M2 for the general dataset, respectively. Step S203: Perform L2 normalization on the ternary embedding matrix M1 of the domain dataset and the ternary embedding matrix M2 of the general dataset to obtain the normalized domain dataset embedding matrix M1' and general dataset embedding matrix M2'. That is, divide each element of matrix M1 and M2 by its norm to make the length of each vector in the matrix become 1. Step S3: Use the inference graph model to traverse the domain knowledge graph and search for candidate entities related to the interactive content; specifically including: Step S301: Based on matrix M1', construct a directed weighted inference graph G=(V, B, W), where the node set V corresponds to the domain entity E1'; the edge set B represents the relationship path between entities; the weight matrix W is calculated by the weight function f', and the edge weight is determined according to the proportion of the relation vector magnitude. Step S302: Locate the corresponding node v in the inference graph model G of the entity set E3 extracted from the interaction dataset in S201. c Starting from this point, a breadth-first search algorithm is executed; Step S303: After completing the traversal operation of the domain knowledge graph G=(E1',B,W), obtain the association score set of all nodes in the node set V and perform linear normalization processing; based on the known threshold N of the number of candidate entities, select the top N nodes in the association score to form a candidate entity set, evaluate the reliability of the candidate entities, introduce execution degree calculation, combine the entity occurrence frequency to perform comprehensive confidence calculation and ranking, perform secondary screening of candidate entities, and output candidate entities that are highly related to the interaction content; Step S4: Utilize multi-layer self-focusing networks and feedforward neural networks to capture the global dependencies and local feature representations of the input data, construct a large model, and pre-train the large model; specifically including: Step S401: Use the normalized general dataset embedding matrix M2' from S203 as input, and set the number of neurons in the input layer to L. 3 The tensor flattening operation function is used to transform the dimension of M2' into a one-dimensional vector for use by subsequent networks; Step S402: After the input layer, construct a multi-layer self-focusing network consisting of Z cascaded self-focusing layers. For the z-th self-focusing layer, z∈{1, 2,…, Z}, input the feature matrix X... (z-1) Mapped to query vector Q respectively (z) Key vector K (z) Sum vector U (z) The attention score matrix A is calculated using the scaled dot product attention mechanism. (z) The value vector is weighted and summed using attention weights to obtain the output feature matrix of this layer; Step S403: After the multi-layer self-focusing network, n fully connected layers are connected, and ReLU activation is performed layer by layer to capture global dependencies; Step S404: Train the constructed general large model using the backpropagation algorithm and define the loss function; During the model training process, the difference between the predicted output and the true label is quantitatively evaluated through the loss function. Further select the mean squared error loss function, which measures the overall deviation between the model prediction result and the actual result by calculating the square of the difference between the predicted value and the true value of each training sample and averaging it over all samples. Step S5: Utilizing the implicit knowledge within the pre-trained large model and the explicit knowledge within the domain knowledge graph, through multi-dimensional feature fusion and optimization iteration, accurate prediction or analysis results are gradually generated, specifically including: Step S501: Construct a unified data interface and format specification system. For the inference graph model G=(V, B, W), the features, including node association scores and node degree, are mapped to the same dimension space as the output feature vector of the general large model through linear transformation. The output data of the inference graph model is O. After transformation, the aligned data can be obtained. Step S502: A hybrid fusion mechanism is adopted to organically integrate the inference graph model with the general large model to achieve synergistic enhancement of knowledge reasoning capabilities. Specifically, a gated fusion unit is used as the core architecture, and its inputs are the aligned inference graph model data O and the output feature vector t of the general large model at layer n. (n) The output is the fused feature vector Y; Step S503: Construct the cross-entropy loss function as the joint optimization objective; perform iterative optimization using a gradient-based optimization strategy, and synchronously update the parameter set related to edge weight calculation in the inference graph model, the neural network parameters of the general large model, and the weight parameters w of the gated fusion unit during each training iteration. g and bias parameter b g Through multiple rounds of iterative training, the parameter configuration of the dual-model architecture and fusion mechanism is continuously adjusted so that the loss function value of the fusion model on the training dataset gradually decreases; when the loss function meets the preset convergence condition or reaches the maximum number of iterations, the training process is terminated.
2. The reasoning comparison method based on a knowledge base according to claim 1, characterized in that: Step S1 specifically includes: Step S101: Establish the interactive dataset, domain dataset, and general dataset. Perform data cleaning on the domain dataset D1, general dataset D2, and interactive dataset D3 to remove noisy and invalid data. Let the domain dataset be D1 = The general dataset is D2= The interactive dataset is D3= , where d i d j and d k These represent data instances in the domain dataset, the general dataset, and the interactive dataset, respectively, with I, J, and K representing the total number of data instances in the three datasets. Define a unified noise data judgment function f(d). When f(d) = 1, d is judged as noise data; when f(d) = 0, d is considered as valid data. f(d) is set based on a feature threshold. If x l <θ1 or x l If θ2, θ1, and θ2 are pre-set feature thresholds, then f(d) = 1, and the numerical feature x... l A numerical indicator used to describe a specific attribute of a data instance, where l is the feature dimension index; for each data instance d, multiple numerical features x can be used. l To represent its different dimensions of characteristics; After cleaning, the domain dataset, general dataset, and interactive dataset were updated as follows: =D1-{d∈D1|f(d=1}、 =D2-{d∈D2|f(d)=1} and =D3-{d∈D3|f(d)=1}, where θ1 and θ2 are used to define the reasonable range of values for data features. Data that exceeds this range will be judged as noise data and removed, thereby ensuring the accuracy and effectiveness of data processing; Step S102: Process the cleaned domain dataset General dataset and interactive datasets Standardization is performed to unify the data to the same format and metric, resulting in a standardized domain dataset. General dataset and interactive datasets For numerical features in all three datasets, a standardization formula was used, as shown below: x'= Where μ is the mean of this feature in the merged dataset, used to characterize the central tendency of the data, and the merged dataset is a collection of data from the merged dataset. , and The merged dataset; , , Each of the three datasets represents a numerical feature of a specific attribute of one of the data instances. This standardization formula transforms data of different magnitudes and distribution characteristics to the same scale, facilitating subsequent data analysis and model training. σ represents the standard deviation of this feature in the merged dataset, used to measure the dispersion of the data. x' is a numerical feature, representing a standardized numerical index used to describe a specific attribute of a data instance. Step S103: Process the standardized domain dataset and general datasets A deduplication operation is performed to eliminate duplicate data records. A hash function h(d) is used to hash the data records in both datasets, mapping each data record d to a fixed-length hash value. If h(d1) = h(d2), where d1 and d2 come from the domain dataset or the general dataset respectively, and d1 = d2, and further comparison confirms that the key features of d1 and d2 are completely identical, then these two data records are determined to be duplicates, and only one is retained while the other is deleted. After deduplication, the domain dataset and the general dataset are updated to... and ; Step S104: Process the deduplicated domain dataset General dataset and the standardized interactive dataset Word segmentation is performed to convert the text data into a word sequence; a statistical word segmentation method is adopted, and the probability function of the word segmentation model is known to be P(w1, w2,…, w…). s |T), where T is the input text data, w1, w2, ..., w s The word sequence is obtained after word segmentation, where s is the length of the word sequence. The probability function is statistical information obtained from training on a large corpus, which helps the model determine the segmentation method of words in the text, transforming text data into a computer-processable word sequence form. It is used to calculate the word sequence w1, w2, ..., w1 given a text T. s The probability is calculated, and the optimal word segmentation result is obtained by maximizing this probability function. After word segmentation, the domain dataset, general dataset, and interactive dataset are updated to D1', D2', and D3', respectively.
3. The reasoning comparison method based on a knowledge base according to claim 1, characterized in that: Step S201 specifically includes: performing knowledge graph embedding processing on the domain dataset D1', general dataset D2', and interactive dataset D3' after deduplication and word segmentation in step S1; firstly, extracting entity set E1 from the domain dataset D1' using part-of-speech tagging, so that the entities in E1 cover the key concepts and objects in the domain; using a knowledge graph embedding model, representing the triples (h, r, t) in the knowledge graph as relations in a vector space, where h is the head entity, r is the relation, and t is the tail entity, i.e., h⊕r≈t, thereby obtaining the domain data triple set S1; optimizing the entity vectors by minimizing the objective function; then using a scoring function and undergoing multiple rounds of iterative training to obtain the entity vector representation E1' of the domain dataset; For the general dataset D2' and the interactive dataset D3', after extracting the entity sets E2 and E3 using a similar method, the entities in the general dataset and the interactive dataset are embedded using the above embedding method and optimization strategy to obtain the general dataset triple set S2, the entity vector representation of the general dataset E2' and the entity vector representation of the interactive dataset E3'. Step S202 specifically includes: using the same model as entity embedding in S201, extracting relation sets R1 and R2 from the domain dataset D1' and the general dataset D2' respectively, mapping the relations to a vector space of the same dimension L as the entity vectors, and obtaining the relation vector r∈R L During the mapping process, the model learns the semantic features of the relationship, so that relationships with similar semantics are closer in the vector space; For each triple (h, r, t) in the domain knowledge graph, the tensor product operation M is performed. h, r, t =e h ⊙r⊙e t Construct its embedding matrix representation, where e h and e t For each element in the entity vector representation E1' of the domain dataset, the head entity vector, relation vector, and tail entity vector are fused using the tensor product operation to form a multidimensional tensor. The semantic information of the triples is then encoded into a matrix, ultimately constructing the triple embedding matrix M1∈R of the domain dataset. L×L×L Similarly, the same operation is performed on the triples in the general knowledge graph to generate the triple embedding matrix M2∈R of the general dataset. L×L×L .
4. The reasoning comparison method based on a knowledge base according to claim 3, characterized in that: The calculation formula in step S201 is as follows: In the above formula, f(h, r, t) represents the scoring function, which is used to evaluate the rationality of the triples. The smaller the value, the more reasonable the relationship of the triples in the vector space. ξ represents the minimization objective function, which optimizes the entity vectors. S represents the domain data triple set S1 or the general dataset triple set S2. S' represents the set of negative triples generated by randomly replacing the head or tail entity in the positive triples. h' is the replaced head entity vector in the negative triples, t' is the replaced tail entity vector in the negative triples, and γ represents the marginal hyperparameter, which is used to control the interval between positive and negative examples in the vector space to avoid overfitting of the model.
5. The reasoning comparison method based on a knowledge base according to claim 3, characterized in that: The calculation formula in step S203 is as follows: M'= In the above formula, ||M|| represents the L2 norm of the matrix, M represents the embedding matrix M1 generated from the domain dataset or the embedding matrix M2 generated from the general dataset, L represents the size of the ternary matrix M in a certain dimension, i.e., the order of the matrix, i, j, k represent the index variables for each triple summation, i.e., the element positions in the three dimensions of the ternary matrix M, M ijk This represents the element value of the ternary matrix M in the i-th row, j-th column, and k-th layer; M' represents the normalized domain dataset embedding matrix M1' or the general dataset embedding matrix M2'.
6. The reasoning comparison method based on a knowledge base according to claim 1, characterized in that: Step S301 specifically includes: using the normalized domain dataset embedding matrix M1'∈R from S203. L×L×L Based on this, a directed weighted inference graph model G=(V, B, W) is constructed, and the specific construction methods of each component are as follows: (1) Node set V: Consists of entities in the domain knowledge graph, denoted as V=E1', where each node v i Each node in V corresponds one-to-one with an entity in the domain knowledge graph. If there are |E1'| entities in the domain knowledge graph, then the node set can be represented as V={v1, v2,..., v...} |E1'| }; (2) Edge set B: Edges in the graph (v i , v j )∈B represents entity v i With v j Semantic relationships between entities; in a domain knowledge graph, if there exists a relationship between entity v i Pointing to v j The relationship path is defined in the graph model (v i , v j )∈B; (3) Weight matrix W: To accurately measure the semantic strength of an edge, a weight function f' is introduced. For an edge (v i , v j ), its weight element w ij via f'(M1'[h i , r ij , t j ]) calculate, where h i and t j The head entity v i With tail entity v j The vector representation in M1', r ij For the corresponding relation vector; the weight function f' adopts the proportional weighting strategy, which determines the weight of the edge by the proportion of the relation vector's magnitude to the sum of all relation vectors' magnitudes. That is, the larger the magnitude of the relation vector, the higher the weight of the corresponding edge, and the more important the relation is in the knowledge graph. Step S303 specifically includes: after completing the traversal operation of the domain knowledge graph G=(E1',B,W), obtaining the set of association scores {S(v)} of all nodes in the node set V. v∈V To eliminate the impact of differences on the comparison of results and enhance the comparability between scores at different nodes, a linear normalization method is used to standardize the score set. Based on the known threshold N for the number of candidate entities, the top N nodes by association score are selected, and their corresponding entities form the candidate entity set CE={e i } N i=1 Where N represents the number of candidate entities to be selected, and e i Let i represent the i-th entity in the candidate entity set, i=1, 2, ..., N; then, evaluate the reliability of the candidate entities and introduce execution degree calculation. Combined with the frequency of the entity's appearance in the knowledge graph, use the comprehensive confidence degree calculation formula to perform a second screening of the candidate entities through confidence degree ranking, so as to ensure that the final output candidate entities are highly relevant to the interaction content.
7. The reasoning comparison method based on a knowledge base according to claim 6, characterized in that: The specific steps of the breadth-first search algorithm in step S302 are as follows: (1) Initialize an empty queue Q, enqueue the starting node, and mark v. c For items that have been visited, denoted as visited(v c =True; (2) When queue Q is not empty, perform the following operations: 1) Retrieve the head node of the queue, denoted as u; 2) Traverse all unvisited adjacent nodes v of node u. For each adjacent node v: i) Calculate the relation vector r uv Relationship vector r with interactive content c Cosine similarity; ii) Update the association score of node v according to the association score formula; iii) Mark v as visited, i.e., visited(v) c = True, and enqueue v; To effectively improve computational efficiency, a pruning optimization strategy is adopted. Constraints are constructed by setting an association score threshold β and an upper limit on the search depth: when the association score s(v) of node v is less than the threshold β, or when the association score ... the association score of node v is less than the threshold β. c When the path length path_length(v) to node v exceeds the upper limit, the expansion operation on that node and its subsequent paths is terminated immediately. The path length path_length(v) is precisely calculated using a breadth-first search algorithm. Specifically, during the node enqueue phase, its hierarchical information is recorded, and the initial node v... c The path length is defined as 0, and the path length is incremented by 1 for each level of nodes traversed.
8. The reasoning comparison method based on a knowledge base according to claim 6, characterized in that: In step S301, the weight element w ij The calculation formula is as follows: In the above formula, w ij Represents an edge (v) i , v j The weights of |r ij | represents the relation vector r ij The larger the modulus, the higher the importance of the relation in the semantic space; This represents the sum of the magnitudes of all relation vectors, used as the denominator for normalization, so that the weight of each edge is between 0 and 1; The calculation formula in step S302 is as follows: sim(r uv , r c )= S'(v)=S(v)+α*sim(r uv , r c )*S(u)*w uv In the above formula, sim(r) uv , r c ) represents the relation vector r uv Relationship vector r with interactive content c The cosine similarity; S(v) represents a certain score or weight value of node v, which is dynamically updated based on subsequent calculation results. S(v) is the value before the update, and S'(v) is the value after the update; α is a hyperparameter used to control the degree of influence of the entire update item; S(u) represents the score or weight value of node u, which is used to influence the update of node v, reflecting the ability of node u to influence node v; w uv It is the connection weight from node u to node v, reflecting the strength or importance of the relationship between nodes u and v. The larger the weight, the closer the connection between the two. The calculation formula in step S303 is as follows: In the above formula, S''(v) represents the normalized association score. The normalized score is in the interval [0, 1], which is a difference distribution of the original scores that are too large, so that the association scores of different nodes are comparable. S'(v) represents the original association score of node v mentioned above. min({S'(v)}) represents the minimum value among all node association scores in the set S'(v), and max({S'(v)}) represents the maximum value among all node association scores in the set S'(v). In the above formula, CF(e i ) represents entity e i The overall confidence level is quantified by weighted fusion of three different evaluation indicators to determine the reliability and importance of entities in knowledge base reasoning scenarios; γ1, γ2, γ3∈[0, 1] and γ1+γ2+γ3=1, which are weight parameters used to balance the impact of normalized association score, node degree, and occurrence frequency on confidence level; S''(v) represents the confidence level of node v. i The normalized correlation score, H(v) i ) is node v i In a reasoning graph, the degree reflects the richness of the connections between entities, f(e i ) represents entity e i The frequency of an entity's appearance in the domain knowledge graph indicates its importance and prevalence within the domain knowledge.
9. The reasoning comparison method based on a knowledge base according to claim 1, characterized in that: Step S401 specifically includes: using the normalized general dataset embedding matrix M2' from S203 directly as input data to construct the input layer of the neural network, which directly affects the subsequent network's data processing efficiency and feature extraction capability. The dimension of M2' is R. L ×L×L This dimensional structure leverages the multi-relation embedding characteristics of knowledge graphs, where the three dimensions correspond to different semantic association dimensions of knowledge. To match neural network design principles, the number of neurons in the input layer must match the dimensions of the input data; therefore, the number of neurons in this input layer is L. 3 To meet the input data format requirements of subsequent network layers, a tensor flattening operation function is used to transform the dimension of M2', reshaping it from a three-dimensional matrix into a one-dimensional vector in sequence. That is, the operation X=Flatten(M2') is performed, and the generated X will be used as the input data for subsequent network layers; Step S402 specifically includes: after the input layer, constructing a multi-layer self-focusing network consisting of Z cascaded self-focusing layers, and using an adaptive approach to focus on key regions of the input data to achieve effective extraction and representation of local data features; for the z-th self-focusing layer, z∈{1, 2, ⋯, Z}, let the input be the output X of the previous layer. (z-1) The output is X (z) The calculation process follows these steps: First, the input feature matrix X is transformed through three independent linear transformation operations. (z-1) Mapped to query vector Q respectively (z) Key vector K (z) Sum vector U (z) The specific formula is shown below: Q (z) = ·X (z-1) + K (z) = ·X (z-1) + U (z) = ·X (z-1) + In the above formula, , , These are all trainable weight matrices, whose function is to perform a linear transformation on the input data. , , The bias vector is used to adjust the result after linear transformation, thereby improving the model's expressive power; {·} represents matrix multiplication. Then, the attention score matrix A is calculated using the scaled dot product attention mechanism. (z) The specific formula is as follows: A (z) =Softmax( ) In the above formula, d k For the key vector K (z) In terms of dimensions, the Softmax function normalizes the scores, making A... (z) The sum of each row of elements is 1, thus quantifying the importance weight of features at different positions; Finally, the value vectors are weighted and summed using attention weights to obtain the output feature matrix of this layer, as shown in the following formula: X (z) =A (z) ·U (z) In the above formula, X (z) This represents the output of the self-focusing layer at layer z, which is the final vector obtained after weighted summation. It will be used as input for subsequent calculations or other layers of the model; A (z) U is the attention weight matrix of the z-th layer. Each element in this matrix represents the importance of the corresponding feature in the weighted summation; the larger the value, the greater the contribution of the corresponding feature to the final output. (z) Let A be the feature vector matrix of the z-th layer, containing various feature information extracted from the input data at this layer. This is then combined with the attention weight matrix A. (z) Multiplication enables weighted combinations of different features; Step S403 specifically includes: after the multi-layer self-focusing network, a general feedforward neural network is connected. This neural network consists of several fully connected layers and aims to capture the global dependencies of the input data, and to process the output X of the self-focusing network. (z) The data is sequentially fed into n fully connected layers for processing. t (1) =ReLU(w (1) ·X (z) +b (1) ) t (2) =ReLU(w (2) ·X (z) +b (2) ) ... t (n) =ReLU(w (n) ·X (z) +b (n) ) In the above formula, t (i) t represents the output of the i-th fully connected layer, i=1, 2, …, n; (1) It is the first fully connected layer to output X of the self-focusing network. (z) The processed output, subsequent t (i) Based on the previous layer t (i-1) The calculation results are: i>1; w (i) Let be the weight matrix of the i-th fully connected layer, i=1, 2, …, n. Its dimension determines the mapping relationship between the input and output, and it is used to perform weighted transformation on the input data. (i) Let be the bias vector of the i-th fully connected layer, and ReLU represent the activation function, which introduces non-linearity into the neural network. In step S404, the specific formula for the mean squared error loss function is as follows: In the above formula, δ represents the model's loss function, used to measure the difference between the model's prediction and the actual result. The smaller the loss function value, the better the model's prediction performance. N represents the number of samples, i.e., the total number of samples used to calculate the loss. i represents an index variable used to iterate through each sample from 1 to N, i∈{1, 2, ..., N}. This represents the model's predicted value for the i-th sample. This represents the true value of the i-th sample.
10. The reasoning comparison method based on a knowledge base according to claim 1, characterized in that: Step S501 specifically includes: to achieve effective integration of the inference graph model in step S3 and the general large model in step S4, a unified data interface and format specification system is constructed. For the inference graph model G=(V, B, W), its output data includes the set of node association scores S'(v) and the set of node degree information H(v). i The model consists of a set of candidate entities (CE) and a set of candidate entities (CE); the output data of the general large model is defined as the output feature vector t of the last layer of the model. (n) This study utilizes a feature mapping strategy to align data, addressing the significant differences in dimensionality and semantics between the output data of the two models. A linear transformation maps features in the inference graph model, including node association scores and node degrees, to the same dimensional space as the output feature vector of the general large model. Setting the output data of the inference graph model to O, the aligned data is obtained after the transformation. The specific formula is shown below: O=w o O+b o In the above formula, w o Let b represent the alignment weight matrix. o This represents the alignment bias vector, and the initial value of the parameter can be determined by minimizing the difference between the aligned data and the output data of the general large model. In step S502, the inputs to the gated fusion unit are the aligned inference graph model data O and the output feature vector t of the general large model at layer n, respectively. (n) The output is the fused feature vector Y. The fusion process is implemented using the following formula: Y=σ(w g [O; t (n) ]+b g )·O +(1-σ(w g [O; t (n) ]+ b g ))·t (n) Among them, [O; t (n) ] indicates a concatenation operation between two data points, w g and b g These are the trainable weight matrix and bias vector, respectively, and σ is the Sigmoid activation function, which dynamically adjusts the fusion weights of the two model outputs through nonlinear transformation. In step S503, a joint optimization objective function δ1 is constructed to quantify the deviation between the predicted output of the fusion model and the true value. Based on the cross-entropy principle in information theory, the cross-entropy loss function is used for modeling. In the above formula, N represents the total number of training samples, and H represents the total number of categories in the classification task. and Corresponding to the i-th sample in the i-th... The true label and predicted probability of the class.
Citation Information
Patent Citations
Intelligent manufacturing management platform for source-known brain data in aeronautical manufacturing industry
CN116485576A
Expert re-learning reasoning question-answering method combined with gating network
CN116795958A