Knowledge fusion method based on dynamic word meaning representation and large language model

Through the method of dynamic word meaning representation and large language model, the fragmentation problem of threat intelligence knowledge base in the field of network security is solved, effective fusion and entity alignment of multi-source intelligence are achieved, and the accuracy and correlation analysis capabilities of intelligence are improved.

CN120471152APending Publication Date: 2025-08-12SICHUAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510523136.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the field of network security, the threat intelligence knowledge base has problems such as fragmentation, redundant repetition, diverse semantics and uneven quality, making it difficult to achieve effective fusion of multi-source heterogeneous threat intelligence, resulting in insufficient accuracy and correlation analysis capabilities of attack intelligence.

Method used

Using a method based on dynamic word meaning representation and large language model, the data preprocessing of the attack intelligence corpus, the entity relationship projection and graph neural network feature extraction are used using the TransR model, and the entity alignment and fusion are combined with the large language model, the intelligence entity fusion problem under complex confrontation conditions is solved.

Benefits of technology

The alignment and fusion of attack intelligence entities in complex confrontation scenarios is achieved, the quality and correlation analysis capabilities of the intelligence knowledge base are improved, and the accuracy of attack detection and traceability is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471152A_ABST
    Figure CN120471152A_ABST
Patent Text Reader

Abstract

The invention relates to a knowledge fusion method based on dynamic word meaning representation and a large language model, which belongs to the field of network security, and comprises the following steps: firstly, carrying out data preprocessing on corpus data in an attack intelligence duration corpus, labeling context meaning items by using a dynamic word meaning representation model, and then, carrying out dynamic word meaning representation on the context meaning items; using a TransR model to project the entity relationship triad to a low-dimensional vector space to realize semantic representation, completing entity relationship aggregation, and on this basis, performing multi-view feature extraction on three views of a graph structure, entity information and attribute description in an entity based on a graph neural network model; and finally, the correlation between the entity to be fused and the candidate entity is inferred in combination with the large language model, and entity alignment and integration are completed. Compared with the prior art, the method has the advantages that the intelligence entity fusion problem under the complex confrontation condition is converted into the entity alignment relation measurement problem from the multi-view angle, and the key technology of attack intelligence knowledge fusion is broken through.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network security, and in particular relates to a knowledge fusion method based on dynamic word meaning representation and a large language model. Background Art

[0002] With the widespread adoption of technologies such as 5G networks and the Internet of Things, cybersecurity threats are becoming increasingly severe. Hacker attack methods are becoming more complex, covert, and intelligent, demonstrating cognitive adversarial characteristics. This has severely impacted people's lives and production, and posed new challenges to national cyberspace governance. Currently, threat intelligence-driven information security defense has become a key research area in cyberspace security. Attack detection and tracing based on threat intelligence have been widely used in national and corporate security development and strategic development planning. However, traditional rule-based and behavior-based threat identification methods face challenges in detecting advanced persistent threat attacks, such as delayed timeliness and poor accuracy.

[0003] Threat intelligence-driven attack detection and attribution have alleviated these issues to a certain extent. However, the heterogeneous threat intelligence knowledge from multiple sources suffers from redundancy, semantic diversity, and uneven quality, hindering information integration and sharing. Therefore, it is necessary to integrate knowledge from different sources to form a unified, high-quality cybersecurity knowledge base. Currently, several knowledge bases based on threat intelligence exist. For example, Kaiser et al. integrated multiple open-source threat intelligence sources to construct a comprehensive multi-level threat knowledge base. Zhang et al. combined cybersecurity threat intelligence with management security requirements data for critical infrastructure to construct a heterogeneous cybersecurity requirements knowledge graph, effectively identifying management vulnerabilities in cybersecurity incidents. However, these knowledge bases still face challenges such as fragmented attack intelligence, difficulty extracting effective information, complex and changing cognitive confrontations, and difficulty integrating isolated intelligence. Therefore, addressing information fusion theory under the conditions of concept drift and behavioral evolution, addressing the increasing complexity, scenario-specificity, and time-shifting nature of attack techniques and threat entities, and enhancing the accuracy of attack intelligence is a current research hotspot.

[0004] Currently, research on knowledge fusion in general domains is relatively mature, examining entities, attributes, and relationships from various perspectives, and applying advanced methods such as distributed word embedding and graph neural networks. However, research on knowledge graph fusion in the security field is still in its infancy. Research approaches based on entity-text similarity are relatively narrow, often focusing only on superficial features, lacking a comprehensive perspective and framework, and failing to fully utilize the large data volume and randomness of attack intelligence. Therefore, it is necessary to deeply analyze the differences in domain knowledge and propose knowledge fusion methods suitable for the cybersecurity field, effectively integrate various isolated intelligence types, and improve the overall quality of intelligence knowledge bases and their correlation analysis capabilities.

[0005] Based on the above ideas, the present invention proposes a knowledge fusion method based on dynamic word meaning representation and large language model, which converts the intelligence entity fusion problem under complex adversarial conditions into an entity alignment relationship measurement problem from a multi-view perspective. It breaks through the key technology of attack intelligence knowledge fusion and proposes a theoretically innovative and technologically advanced multi-view word meaning dynamic embedding and large language model knowledge fusion method to solve the key scientific problems faced by knowledge-driven attack detection and tracing. Summary of the Invention

[0006] In view of this, the present invention provides a knowledge fusion method based on dynamic word meaning representation and large language model, aiming to solve the problems of word meaning, behavior evolution and concept drift in attack intelligence corpus when facing complex confrontation scenarios, accurately identify attack intelligence data, and realize entity alignment and fusion in the field of multi-semantics attack intelligence.

[0007] A knowledge fusion method based on dynamic word meaning representation and a large language model, the method comprising:

[0008] Step 1: Preprocess the corpus data in the attack intelligence corpus and annotate contextual meanings using a dynamic word meaning representation model.

[0009] Step 2: Use the TransR model to project entity-relationship triples into a low-dimensional vector space to achieve semantic representation and complete entity-relationship aggregation;

[0010] Step 3: Perform multi-view feature extraction on the three views of the entity graph structure, entity information, and attribute description based on the graph neural network model;

[0011] Step 4: Combine the large language model to infer the correlation between the entity to be fused and the candidate entity and complete the entity alignment integration.

[0012] Preferably, the dynamic word meaning representation process includes:

[0013] According to the rules, a time series sliding window of appropriate size, in years, is set. Each time, the word meaning representation algorithm processes the attack intelligence corpus data within this window.

[0014] The attack intelligence data set systematically collected from security forums and security vendors constitutes the diachronic corpus in this application. New data is selected to enter the window while the original and equal amount of data are discarded to ensure that the window size remains unchanged.

[0015] Segment the text in the attack intelligence corpus into token sequences, replacing key words with masked tokens, such as jargon and inflected words.

[0016] By predicting the original tokens of the masked part through self-supervised learning, the model can understand the specific semantics of key words in attack intelligence;

[0017] After training the dynamic word meaning representation model based on a time series window, we use this model to obtain the meaning vectors of polysemous words. Each meaning of a polysemous word is formally represented as a vector. The T5 model based on the encoder-decoder structure is used to input the example sentence corresponding to the polysemous word meaning to obtain the context vector of the polysemous word.

[0018] After obtaining each sense of a polysemous word, the cosine distance is used to measure the similarity between the context vector of the polysemous word and the vector of each sense of the polysemous word. The sense with the highest similarity is the semantics of the polysemous word in the text to be annotated, thereby achieving automatic annotation of the context sense and completing the dynamic word meaning representation of the required corpus data. The specific calculation method of the cosine distance similarity is designed as follows:

[0019]

[0020] Where S is the similarity between the two meanings of the polysemous word M and N, c m is the vector representation of the meaning M, c n is the vector representation of the sense N.

[0021] Preferably, the entity relationship aggregation process includes:

[0022] For each relation triple (h, r, t), where h is the head entity, r is the entity's relation, and t is the tail entity, after obtaining the dynamic semantic representation of each element using the semantic dynamic representation model in step 1, the entities in the entity space are transformed through the matrix M based on the TransR model. r Projected into the r relation space, they are h r and t r , there are h r +r≈t r , h r =hM r ,t r =tM r , a relation-specific projection can bring head or tail entities that actually hold the same relation closer together, while those holding different relations are further away from each other. The proximity of the head and tail entities is measured by the triple distance separation, and its specific calculation formula is designed as follows:

[0023]

[0024] Among them, the calculated f r(h, t) is the distance of the correct triplet, which is optimized to a lower range in this application, with a value between 1.0 and 4.0; L is the distance separation that the negative samples need to meet after the correct optimization of the triplet to verify its mathematical correctness.

[0025] Preferably, the multi-view feature extraction process for the three views of the graph structure, entity information and attribute description in the entity includes:

[0026] Retrieve candidate entities with semantic similarity to the triples in step 2 from the knowledge graph;

[0027] The entity information view embeds the entity name using the dynamic representation model described above, concatenates it with the embedding vector obtained from the relationship view, and uses a graph neural network to understand the semantics of the concatenated vector.

[0028] A graph neural network with an attention mechanism is used to mine the structural features between entities;

[0029] Concatenate all attributes in the attribute description view and use the Transformer model to understand the semantics of the concatenated attributes to generate an attribute feature matrix.

[0030] By designing a transformation matrix, the graph structure features and attribute features are combined to form a new joint entity representation, and then the cosine similarity between the joint entity representation of the candidate entity and the triple representation of the entity to be aligned is used to characterize the correlation between the two.

[0031] Preferably, the entity alignment integration process includes:

[0032] Retrieve a set of entity pairs that are semantically similar to the entity to be fused and aligned from the knowledge graph as sample construction prompts;

[0033] Continue to build 3 to 5 sets of high-quality positive and negative entity pairs using the method in step 3, where positive examples represent entity pairs that should be fused, and negative examples represent irrelevant entity pairs. Clearly mark the relevance, and then use a large language model to infer the relevance between the entity to be fused and the candidate entity through a few-shot contextual learning approach.

[0034] Design prompt words for the selected large language model DeepSeek, so that the large language model can build matching rules based on context examples, thereby judging the relevance between the fusion entity and the candidate entity based on the input, and classifying them into five levels from high to low according to the relevance: very relevant, highly relevant, partially relevant, less relevant, and irrelevant. The corresponding output relevance scores are 4, 3, 2, 1, and 0 points.

[0035] The cosine similarity between entities calculated in step 3 is integrated with the entity relevance score given by the large language model to obtain the final entity relevance score. The specific formula for the final entity relevance score is designed as follows:

[0036] cor_score i =α·cos_sim(e,e i )+β·cor_score llm (e,e i )

[0037] Among them, cos_sim represents cosine similarity calculation, cor_score llm represents the entity relevance score given by the large language model, e represents the semantic vector of the entity to be fused, and e i represents the semantic vector of the i-th candidate entity, α and β are weight parameters;

[0038] Finally, the entity relevance is ranked, and the top-ranked entity pairs are considered to be able to be aligned.

[0039] The present application provides a knowledge fusion method based on dynamic word meaning representation and a large language model. Compared with the existing technology, the beneficial effects of the present invention are: combining emerging technologies in the field of network security in recent years to solve the problem of intelligent processing of attack intelligence with low information fusion degree caused by the evolution of attack behavior and concept drift in edge scenarios, and laying a theoretical foundation for research such as automatic intelligence analysis and forensic tracing; proposing a knowledge fusion method based on dynamic embedding of multi-view word meanings, solving the problems of inaccurate word meaning representation and poor fusion effect in the application of current general domain entity alignment models in the security field, and realizing the effective fusion of multi-source attack intelligence. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in this embodiment or the prior art, the following briefly introduces the drawings required for use in the embodiment or the prior art description. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0041] Figure 1 A flowchart of a knowledge fusion method based on dynamic word meaning representation and a large language model provided in an embodiment of the present application.

[0042] Figure 2 A flowchart of the entity alignment and fusion method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0043] The following embodiments of the present invention are described in further detail in conjunction with the accompanying drawings and specific embodiments. The following examples or drawings are used to illustrate the present invention, but are not intended to limit the scope of the present invention.

[0044] See also Figure 1 , Figure 1 This is a flow chart of a method for knowledge fusion based on dynamic word meaning representation and a large language model, provided in an embodiment of the present application. This method extracts dynamic word meaning representation from attack intelligence corpus as its core, combines it with a large language model to perform multi-view feature extraction and entity alignment, achieves attack intelligence knowledge fusion, and builds a high-quality network security knowledge base, including:

[0045] Step 1: Preprocess the corpus data in the attack intelligence corpus and annotate contextual meanings using a dynamic word meaning representation model.

[0046] Step 2: Use the TransR model to project entity-relationship triples into a low-dimensional vector space to achieve semantic representation and complete entity-relationship aggregation;

[0047] Step 3: Perform multi-view feature extraction on the three views of the entity graph structure, entity information, and attribute description based on the graph neural network model;

[0048] Step 4: Combine the large language model to infer the correlation between the entity to be fused and the candidate entity and complete the entity alignment integration.

[0049] The specific steps for preprocessing the corpus data in the attack intelligence corpus, automatically annotating contextual meanings, and completing dynamic word meaning representation include:

[0050] Step 1a: Set a time series sliding window of appropriate size in years according to the rules. Each time, the word meaning representation algorithm processes the attack intelligence corpus data within this window.

[0051] Optionally, when setting the temporal sliding window, the size of the temporal sliding window is directly related to the recognition performance of the new word concept. If the window size is too large, the concept description is vague and the sensitivity to concept drift is weak. On the contrary, if the window size is too small, the constructed semantic representation model is more sensitive to concept drift, but the small amount of data included will also lead to incomplete concept description and low recognition rate. Therefore, this application defines the maximum window size w max and the minimum window size w min , set the current window size to w current The sizes of the first and second sliding windows are set to w min If no concept drift is detected from window 1 to window 2, the size of the next window is set to min{w max ,2×w current}; Otherwise, resize the next window to w min .

[0052] Step 1b: The attack intelligence data set systematically collected from security forums and security vendors constitutes the diachronic corpus in this application. New data input windows are selected while discarding the original, equal amount of data to ensure that the window size remains unchanged.

[0053] Step 1c: Segment the text in the attack intelligence corpus into token sequences, replacing key words with masked tokens, such as jargon and inflected words.

[0054] Optionally, word segmentation is the process of breaking down a continuous text sequence into smaller units. These units can be words, phrases, or subwords. Currently, the commonly used word segmentation methods are divided into three types: rule-based, statistics-based, and subword-based. This application chooses to use a subword-based word segmentation method and use WordPiece for word segmentation, which can also process words outside the vocabulary.

[0055] Optionally, masking techniques are often used to randomly mask some tokens during model training, allowing the model to predict these masked tokens, thereby improving the model's language understanding capabilities. Different models will select different proportions of tokens for masking. This application will mask key words such as blackened words and inflected words, and then randomly select 15% of the tokens for masking.

[0056] Step 1d: Predict the original word units of the masked part through self-supervised learning, so that the model can understand the specific semantics of key words in the attack intelligence.

[0057] Optionally, at the input layer of the self-supervised learning model, the masked text is converted into word embedding vectors and processed through a multi-layer Transformer encoder. For each masked token, the model outputs a probability distribution representing the probability of each possible token appearing at that location. A cross-entropy loss function is used to measure the difference between the model's prediction and the true token, and backpropagation and parameter updates are performed.

[0058] Preferably, the present application establishes the following calculation formula for the definition of the cross entropy loss function:

[0059]

[0060] The output of the model is a probability distribution p = [p1, p2, ..., p C ], where p i Indicates the probability that the sample belongs to the i-th category. The true label is usually represented as a one-hot vector y = [y1, y2, ..., yC ], where y i =1 means the true category of the sample is category i, and the rest of the positions are 0. When the model predicts a higher probability for the true category, the loss value is lower; otherwise, the loss value is higher.

[0061] Step 1e: After training the dynamic word meaning representation model based on a time series window, use this model to obtain the meaning vectors of polysemous words. Each meaning of a polysemous word is formally represented as a vector. The T5 model based on the encoder-decoder structure is used to input the example sentence corresponding to the polysemous word meaning to obtain the context vector of the polysemous word.

[0062] Alternatively, among Transformer-based language models, the T5 model, based on the Encoder-Decoder architecture, has demonstrated excellent performance on numerous NLP tasks. Unlike traditional static word embedding models, T5 utilizes large amounts of text data for self-supervised learning during pre-training, enabling dynamic integration of contextual information to accurately represent lexical semantics. Therefore, this application proposes to use T5 as a word meaning representation model to achieve accurate representation of entity relationships. First, the input text is dynamically contextualized using T5's multi-layer Transformer encoder, and position-sensitive hidden states are generated through a self-attention mechanism to simultaneously capture the local grammatical features of vocabulary (such as subject-verb agreement) and global semantic dependencies (such as cross-sentence reference). Secondly, a cross-layer feature fusion strategy is designed to perform weighted aggregation on hidden states of different depths in the encoder to enhance the contextual representation ability of polysemous words. Furthermore, the target entity is located through a structured input template, and the entity vector (based on entity span pooling) and context vector (based on the [CLS] flag) are extracted respectively by combining a two-stream representation mechanism. A bilinear transformation is used to model the relationship matrix between entities. Finally, a multi-task fine-tuning framework is used to jointly optimize the relationship classification and text generation goals.

[0063] Step 1f: After obtaining each sense of a polysemous word, use the cosine distance to measure the similarity between the context vector of the polysemous word and the vector of each sense of the polysemous word. The sense with the highest similarity is the semantics of the polysemous word in the text to be annotated, thereby achieving automatic annotation of contextual senses and completing the dynamic word meaning representation of the required corpus data.

[0064] Optionally, the cosine distance similarity calculation formula used in this application is Where S is the similarity between the two meanings of the polysemous word M and N, c m is the vector representation of the meaning M, c n is the vector representation of the sense N.

[0065] The specific steps for projecting entity relationship triples into a low-dimensional vector space to realize semantic representation and entity relationship aggregation include:

[0066] Step 2a: For each relation triple (h, r, t), where h is the head entity, r is the entity's relation, and t is the tail entity, after obtaining the dynamic semantic representation of each element, the entities in the entity space are projected into the relation space based on the TransR model. The relation-specific projection can make the head entities or tail entities that actually hold the same relation close to each other, while those holding different relations are kept away from each other.

[0067] Alternatively, in view of the complexity of the attack intelligence knowledge graph, the same entity often has multiple relationships. By establishing separate relationship spaces for different relationships, the entity is first mapped to the relationship space for calculation, so that the vector representation of the same node in different relationship planes is different. For each relationship triple (h, r, t), after obtaining the dynamic semantic representation of each element, the entity in the entity space is mapped through the matrix M based on the TransR model. r Projected into the r relation space, they are h r and t r , there are h r +r≈t r , h r =hM r ,t r =tM r .

[0068] Preferably, in order to calculate the relationship between the head and tail entities, the following calculation formula can be established:

[0069]

[0070] Among them, the calculated f r (h, t) is the distance of the correct triplet, which is optimized to a lower range in this application, with a value between 1.0 and 4.0; L is the distance separation that the negative samples need to meet after the correct optimization of the triplet to verify its mathematical correctness.

[0071] See also Figure 2 , Figure 2 The following is a flow chart of the entity alignment and fusion method according to an embodiment of the present application. Entity alignment and fusion are divided into two key steps: multi-view feature extraction and entity alignment integration. For multi-view feature extraction, the specific steps include:

[0072] Step 3a: Retrieve candidate entities with similar semantics to the triple from the knowledge graph.

[0073] Preferably, the knowledge graph representation learning (KGE) method is used to map the entities and relationships in the knowledge graph to a low-dimensional vector space to capture their semantic information; then, a graph neural network is used to perform representation learning on the nodes in the knowledge graph, and the information of neighboring nodes is aggregated through a message passing mechanism to further improve the vector representation accuracy of entities and relationships; then, for a given triple, by calculating its cosine similarity score in the vector space, candidate entities with similar semantics to the triple are screened out.

[0074] Optionally, the path reasoning method is combined to further optimize the retrieval results of candidate entities by generating and representing paths in the knowledge graph; the similarity scores of multiple methods are combined to select several entities with the highest scores as the final candidate entities.

[0075] Step 3b: The entity information view embeds the entity name through the above dynamic representation model, concatenates it with the embedding vector obtained from the relationship view, and uses a graph neural network to understand the semantics of the concatenated vector.

[0076] Optionally, to obtain embedding vectors for the relational view, we first generate a relational embedding vector for each entity based on the relational path or graph structure in the knowledge graph. This vector reflects the entity's association information and structural features in the knowledge graph. We then use a relational graph convolutional network approach to aggregate the entity's features along different relational paths to generate a comprehensive relational embedding vector.

[0077] Optionally, for the dynamic representation vectors and relationship embedding vectors of the spliced entity names, the vectors from different sources need to be dimensionally aligned or reduced before splicing to ensure that the spliced vectors have a consistent structure. This application uses the linear discriminant analysis (LDA) method to adjust the vector dimensions, which finds a set of projection directions so that the projected data has the maximum inter-class divergence and the minimum intra-class divergence between different categories.

[0078] Preferably, the maximum inter-class divergence S used in this application is B and the minimum intra-class divergence S W The specific calculation formula is as follows:

[0079]

[0080] Where C is the number of categories, N i is the number of samples in the i-th class, μ i is the mean vector of the i-th class, μ is the mean vector of all samples, X i is the sample set of the i-th class, and x is the sample point.

[0081] Preferably, the projection direction ω found by the linear discriminant analysis method is obtained by solving the following generalized eigenvalue:

[0082] S B ω=λS W ω

[0083] Where λ is the eigenvalue and ω is the corresponding eigenvector.

[0084] Alternatively, to use a graph neural network to understand the semantics of the concatenated vectors, the concatenated vector is first fed into the graph neural network as the initial representation of the node. In the graph neural network, nodes aggregate information from their neighbors through message passing and update their own representations. Finally, by stacking multiple layers of graph neural networks, the abstraction level of the node representation and the semantic understanding capabilities are gradually improved.

[0085] Step 3c: Use a graph neural network with an attention mechanism to mine the structural features between entities.

[0086] Preferably, a graph neural network based on the attention mechanism scores multiple neighbors of an entity, assigns a weight to each neighbor information, and aggregates neighbor entities to embed. In addition, when aggregating neighbor information, it is necessary to normalize the attention of all neighbors of each node.

[0087] Optionally, multi-hop sampling is performed on each layer of the entity's neighbor information. A K-hop subgraph is derived from the knowledge graph with the target entity as the node. The node vectors and adjacency matrix of the subgraph serve as the input to the graph neural network. A graph attention layer is then added to learn the directedness between entities. The attention network a needs to consider the influence of both nodes simultaneously.

[0088] Step 3d: Concatenate all attributes in the attribute description view and use the Transformer model to understand the semantics of the concatenated attributes to generate an attribute feature matrix.

[0089] Optionally, to accommodate the input constraints of the Transformer model, a maximum length is set for the concatenated text. Any excess length is truncated, and any less than this length is padded. The attribute text is fed into the Transformer model, which leverages its multi-layer self-attention mechanism to capture the semantic relationships and contextual information between attributes. Finally, the model's last hidden state is selected as the generated attribute feature matrix.

[0090] Step 3e: By designing a transformation matrix, the graph structure features and attribute features are combined to form a new joint entity representation. Then, the cosine similarity between the joint entity representation of the candidate entity and the triple representation of the entity to be aligned is used to characterize the correlation between the two.

[0091] Alternatively, for the design of the transformation matrix, a linear transformation is first performed to design a learnable transformation matrix W to map the graph structure features and attribute features to the same feature space. The graph structure features are designed as The characteristic attributes are The latitude of the transformation matrix W is (d1+d2)×d joint , where d joint It is the latitude of the joint feature space. After splicing the graph structure features and attribute features, a comprehensive feature vector is obtained Then, a linear transformation is performed through the transformation matrix W, and the design calculation formula is as follows:

[0092] f joint =Wf concat

[0093] in, is the final joint entity representation. After obtaining the joint entity representation, in order to facilitate similarity calculation, the joint entity representation is normalized using L2 normalization. The calculation formula is designed as follows:

[0094]

[0095] For entity alignment and integration, the specific steps include:

[0096] Step 4a: Retrieve a set of entity pairs that are semantically similar to the entity to be fused and aligned from the knowledge graph as sample construction prompts.

[0097] Step 4b: Continue to construct 3 to 5 sets of high-quality positive and negative entity pairs using the method in step 3, where positive examples represent entity pairs that should be fused, and negative examples represent irrelevant entity pairs. Clearly mark the relevance, and then use the large language model to infer the relevance between the entity to be fused and the candidate entity through a few-shot contextual learning approach.

[0098] Optionally, design prompt words for the selected large language model DeepSeek so that the large language model can build matching rules based on context examples, thereby judging the relevance between the fusion entity and the candidate entity based on the input, and classifying the relevance into five levels from high to low: very relevant, highly relevant, partially relevant, less relevant, and irrelevant, with the corresponding output relevance scores being 4, 3, 2, 1, and 0 points;

[0099] Optionally, the prompt word design of the DeepSeek model is as follows:

[0100]

[0101] Step 4c: Integrate the previously calculated inter-entity cosine similarity with the entity relevance score given by the large language model to obtain the final entity relevance score.

[0102] Preferably, since the relevance score calculated by cosine similarity and the relevance score given by the large language model are not performed under the same standard framework, an integration function can be designed to normalize and integrate the scores to obtain the final relevance score. The specific calculation formula is as follows:

[0103] cor_score i =α·cos_sim(e,e i )+β·cor_score llm (e,e i )

[0104] Among them, cos_sim represents cosine similarity calculation, cor_score llm represents the entity relevance score given by the large language model, e represents the semantic vector of the entity to be fused, and e i represents the semantic vector of the i-th candidate entity, and α and β are weight parameters.

[0105] Step 4d: Finally, the entity relevance is ranked, and the top-ranked entity pairs are considered to be able to be aligned, thus completing the entire entity fusion process.

[0106] Optionally, the tolerance for entity alignment in this application will change according to the changes in the actual application scenarios of the business. Generally speaking, entities whose relevance is ranked from high to low and accounts for the top 30% of the total participating entities and whose relevance score is greater than 0.5 can be aligned.

[0107] It should be noted that for the above method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and processes involved are not necessarily required by this application.

[0108] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0109] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent replacements for some of the technical features therein.

[0110] Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A knowledge fusion method based on dynamic word meaning representation and large language model, characterized by: The method comprises: Step 1: Preprocess the corpus data in the attack intelligence corpus and annotate contextual meanings using a dynamic word meaning representation model. Step 2: Use the TransR model to project entity-relationship triples into a low-dimensional vector space to achieve semantic representation and complete entity-relationship aggregation; Step 3: Perform multi-view feature extraction on the three views of the entity graph structure, entity information, and attribute description based on the graph neural network model; Step 4: Combine the large language model to infer the correlation between the entity to be fused and the candidate entity and complete the entity alignment integration.

2. The knowledge fusion method based on dynamic word meaning representation and large language model according to claim 1 is characterized in that: In step 1: According to the rules, a time series sliding window of appropriate size, in years, is set. Each time, the word meaning representation algorithm processes the attack intelligence corpus data within this window. The attack intelligence data set systematically collected from security forums and security vendors constitutes the diachronic corpus in this application. New data is selected to enter the window while the original and equal amount of data are discarded to ensure that the window size remains unchanged. Segment the text in the attack intelligence corpus into token sequences, replacing key words with masked tokens, such as jargon and inflected words. By predicting the original tokens of the masked part through self-supervised learning, the model can understand the specific semantics of key words in attack intelligence; After training the dynamic word meaning representation model based on a time series window, we use this model to obtain the meaning vectors of polysemous words. Each meaning of a polysemous word is formally represented as a vector. The T5 model based on the encoder-decoder structure is used to input the example sentence corresponding to the polysemous word meaning to obtain the context vector of the polysemous word. After obtaining each sense of a polysemous word, the cosine distance similarity is used to measure the similarity between the context vector of the polysemous word and the vector of each sense of the polysemous word. The sense with the highest similarity is the semantics of the polysemous word in the text to be annotated, thereby achieving automatic annotation of the context sense and completing the dynamic word meaning representation of the required corpus data. The specific calculation method of the cosine distance similarity is designed as follows: Where S is the similarity between the two meanings of the polysemous word M and N, c m is the vector representation of the meaning M, c n is the vector representation of the sense N.

3. The knowledge fusion method based on dynamic word meaning representation and large language model according to claim 2 is characterized in that: In step 2: For each relation triple (h, r, t), where h is the head entity, r is the entity's relation, and t is the tail entity, after obtaining the dynamic semantic representation of each element using the semantic dynamic representation model in claim 2, the entities in the entity space are transformed through the matrix M based on the TransR model. r Projected into the r relation space, they are h r and t r , there are h r +r≈t r , h r =hM r ,t r =tM r , a relation-specific projection can bring head or tail entities that actually hold the same relation closer together, while those holding different relations are further away from each other. The proximity of the head and tail entities is measured by the triple distance separation, and its specific calculation formula is designed as follows: Among them, the calculated f r (h, t) is the distance of the correct triplet, which is optimized to a lower range in this application, with a value between 1.0 and 4.0; L is the distance separation that the negative samples need to meet after the correct optimization of the triplet to verify its mathematical correctness.

4. The knowledge fusion method based on dynamic word meaning representation and large language model according to claim 3 is characterized in that: In step 3: Retrieve candidate entities from the knowledge graph that are semantically similar to the aggregated triples; The entity information view embeds the entity name using the dynamic representation model described above, concatenates it with the embedding vector obtained from the relationship view, and uses a graph neural network to understand the semantics of the concatenated vector. A graph neural network with an attention mechanism is used to mine the structural features between entities; Concatenate all attributes in the attribute description view and use the Transformer model to understand the semantics of the concatenated attributes to generate an attribute feature matrix. By designing a transformation matrix, the graph structure features and attribute features are combined to form a new joint entity representation. Then, the cosine distance similarity between the joint entity representation of the candidate entity and the triple representation of the entity to be aligned is used to characterize the correlation between the two.

5. The knowledge fusion method based on dynamic word meaning representation and large language model according to claim 4 is characterized in that: In step 4: Retrieve a set of entity pairs that are semantically similar to the entity to be fused and aligned from the knowledge graph as sample construction prompts; Continue to build 3 to 5 sets of high-quality positive and negative entity pairs, where positive examples represent entity pairs that should be fused, and negative examples represent irrelevant entity pairs. Clearly mark the relevance, and then use a large language model to infer the relevance between the entity to be fused and the candidate entity through a few-shot contextual learning approach. Design prompt words for the selected large language model DeepSeek, so that the large language model can build matching rules based on context examples, thereby judging the relevance between the fusion entity and the candidate entity based on the input, and classifying them into five levels from high to low according to the relevance: very relevant, highly relevant, partially relevant, less relevant, and irrelevant. The corresponding output relevance scores are 4, 3, 2, 1, and 0 points. The calculated cosine similarity between entities is integrated with the entity relevance score given by the large language model to obtain the final entity relevance score. The specific formula for the final entity relevance score is designed as follows: cor_score i =α·cos_sim(e,e i )+β·cor_score llm (and, and i ) Among them, cos_sim represents cosine similarity calculation, cor_score llm represents the entity relevance score given by the large language model, e represents the semantic vector of the entity to be fused, and e i represents the semantic vector of the i-th candidate entity, α and β are weight parameters; Finally, the entity relevance is ranked from high to low according to the final entity relevance score, and the top-ranked entity pairs are considered to be able to be aligned.