Knowledge graph generation method based on large language model
By acquiring real-time radar and text data, frequency domain feature extraction and text semantic feature vector extraction, combined with large language models for spatiotemporal alignment and feature fusion, triplets are generated and data completion are formed, triplet chains are formed, knowledge graphs are constructed and adversarial strategy generation lags are solved, and efficient multimodal information fusion and dynamic tactical decision-making support are achieved.
Patent Information
- Application Number
- CN202510759711.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-09
AI Technical Summary
The existing radar interference technology analysis methods have poor adaptability in complex environments, and it is difficult to identify enemy interference patterns in real time. They lack effective fusion and in-depth analysis of multimodal information, resulting in low recognition accuracy and lag in the generation of adversarial strategies. The static knowledge graph lacks real-time adaptive update capabilities.
By obtaining real-time radar and text data, frequency domain feature extraction and text semantic feature vector extraction, combined with large language models for spatiotemporal alignment and feature fusion, generate triplets and complete data, form triplet chains, build knowledge graphs and update in real time.
It improves the accuracy of enemy radar jamming pattern recognition, enhances the intelligent generation ability of adversarial strategies, solves the problem of knowledge imbalance under sparse data distribution, and dynamically generates adversarial strategies, improving the real-time adaptability of the radar system and the accuracy of tactical decision-making.
Smart Images

Figure CN120258117A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge graph generation, and specifically provides a method for generating a knowledge graph based on a large language model. Background Technique
[0002] In modern radar countermeasures, radar technology and jamming technology have become important components of the confrontation. The enemy often uses advanced radar jamming technologies, such as range gate pull-off (RGPO) and other means, to effectively confuse or deceive our own radar systems, thereby obtaining an information advantage in the confrontation. This jamming technology changes the characteristics of radar signals, interferes with the signal reception of radar systems, reduces detection accuracy, and affects the tactical decisions of both sides. Therefore, how to accurately identify the radar jamming patterns of the enemy in a complex environment and effectively counter them through reasonable strategies has become an urgent technical challenge.
[0003] Currently, the analysis of radar jamming technology mostly relies on traditional signal processing methods, such as frequency domain analysis and time domain analysis. These methods have certain limitations in dealing with high-frequency and complex jamming patterns and are difficult to respond to the changing battlefield environment in real time. At the same time, the tactical strategies and the relationship between the enemy and us contained in text data are often ignored or not fully utilized. Therefore, the lack of effective fusion and in-depth analysis of multi-modal information leads to poor adaptability of existing systems in complex environments, resulting in low accuracy in jamming pattern recognition, lag in the generation of countermeasure strategies, and the inability to dynamically model the spatio-temporal correlation of jamming technology and generate accurate countermeasure strategies in real time. Existing knowledge graph construction technologies mostly focus on static knowledge bases and lack the ability to adaptively update real-time battlefield data. There are significant deficiencies especially in cross-modal feature alignment and sparse data completion, which limit the timeliness and accuracy of tactical decisions.
[0004] The above information disclosed in the background technical part is only used to strengthen the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for generating a knowledge graph based on a large language model to solve the problems raised in the above background technique.
[0006] To achieve the above purpose, the present invention provides the following technical solutions: A method for generating a knowledge graph based on a large language model, the specific steps include: Step 1: Obtain the real-time radar data and text data of both sides of the enemy and us during radar detection and counter-jamming. Extract the frequency domain features of the radar data to identify the jamming pattern, and extract the text semantic feature vectors of the text data to identify the relationship between the subjects; Step 2: Perform spatio-temporal alignment on the extracted radar frequency-domain features and text semantic feature vectors. Based on the aligned radar frequency-domain features and text semantic feature vectors, perform feature fusion through the constructed fusion model according to the cross-modal attention mechanism, and output the fused feature vectors; Step 3: Generate triples by performing relation extraction on the fused feature vectors through a large language model. Judge the distribution of the triple data according to the similarity of the triples, and complete the triple data for the sparse distribution locations to ensure the balance of the triple data distribution; Step 4: Form triple chains based on the continuity of radar countermeasures for the triples. According to the triple chains, output a knowledge graph by constructing a knowledge graph generation model, and update the knowledge graph based on real-time radar data and text data.
[0007] Further, the radar data includes the operating frequency, waveform, and transmit power of the radar; The specific process of extracting frequency-domain features from the radar data is as follows: Obtain radar data at the same time interval Perform sampling. After preprocessing through maximum normalization, remove the DC component in the signal. Through a time window composed of time intervals, divide the radar data, and extract the frequency-domain features from the radar data within the time window; The calculation formula for removing the DC component in the signal is: where is the signal after removing the DC component, is the preprocessed radar signal at time, is the total number of signal samples; The calculation formula for extracting the frequency-domain features from the radar data within the time window is: where is the th frequency component in the frequency-domain signal of the radar signal, is the th sample in the radar signal, is the total number of signal samples, is the imaginary unit, represents the th frequency component in the frequency domain, .
[0008] Further, the text data includes the working principle, technical features, tactical applications, countermeasure strategies, and work logs of radar jamming technology; The specific process of extracting text semantic feature vectors from the text data is as follows: The text data is split into word units by adopting a vocabulary, and the [CLS] token is added at the beginning of each input text, and the [SEP] token is added at the end of the text to define the text boundary. The RoBERTa is used to extract the text semantic feature vector of the [CLS] token.
[0009] Furthermore, the method for constructing the fusion model is as follows: Take the radar feature as the query, the text feature as the key and the value; Through the query and the key, calculate the similarity to obtain the attention weight. After obtaining the attention weight, apply it to the value vector to obtain the weighted output feature vector: Among them, is the query vector, is the key vector, is the transpose of the key vector, is the value vector, is the dimension of the key vector, is the output fusion feature vector, is the function; Among them, the query vector, the key vector, and the value vector are respectively: Among them, is the weight matrix of the query, is the radar frequency domain feature vector, is the weight matrix of the key, is the text semantic feature vector.
[0010] Furthermore, the specific method for spatio-temporal alignment is as follows: Among them, is the loss function for alignment in time, is the alignment mask matrix of the radar feature and the text feature, is the acquisition time length of the radar feature, is the number of text description words, is the radar signal at is the feature vector at the moment, is the text description at the position is the text semantic feature vector at the place.
[0011] Furthermore, the large language model is a joint model of RoBERTa-GCN-EGPLinker, and the specific structure includes RoBERTa semantic encoding, GCN, and EGPLinker models: Through RoBERTa semantic encoding, the fusion feature vector Extract it into context vector representations, and update the node of the context vector representations through GCN. Use the representations of the context and entities in the updated text as the input of the EGPLinker model to perform entity graph relationship connection; Entity graph relationship connection: Among them, represents the input sequence data respectively represent the start and end projection matrices of the th entity, is the bias of the th entity, respectively represent the extracted entities , represents the representation of the entity in the context of the text, is the entity relationship score between, is the transpose operation, is the global pointer model; Determine the triple relationship between entities through the relationship score threshold of entities and output the triple.
[0012] Furthermore, the calculation method for judging the distribution of triple data according to the similarity of triples is: Among them, is the probability that the th relationship type appears in the triple, is the th relationship type in the triple, is the number of relationship types in the triple; When, judge that the distribution position of the triple is sparse; Among them, is the triple distribution position sparse judgment threshold; The specific method for complementing triple data at sparse distribution positions is: Among them, are the head entity, tail entity and relationship type in the sparse triple of the distribution position respectively, are the embedding vectors of the head entity and the tail entity respectively, is the output candidate relationship type, is a random noise vector, following the standard normal distribution, is the relationship type in the sparse triple of the distribution position, is the generated learning weight matrix, is the discriminative learning weight matrix, is the complemented triple.
[0013] Further, the method of forming a triple chain from triples according to the continuity of radar countermeasure is as follows: Divide the process of radar countermeasure into stages, assign stage labels to each triple, and form a sequence: where, is the stage label sequence, is the stage, is the number of divided stages; Statistically analyze the transition frequencies between stages and construct a probability matrix: where, is the transition probability matrix, is the number of occurrences of stage , is the number of times that stage transitions to stage ; Generate a triple chain through dynamic programming: where, is the cumulative probability that the th triple is in stage , is the maximum cumulative probability in stage , is the relationship score of the th triple.
[0014] Further, the knowledge graph is constructed based on a graph neural network. The specific process is as follows: Use the entities in each triple chain as the nodes of the graph and the relationships in the triple chain as the edges: where, is the representation of node in the th layer of the graph neural network, is the representation of node in the th layer of the graph neural network, is the set of neighbor nodes of node , is the relationship between nodes , , is the weight matrix of the th layer, is the aggregation function.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention obtains real-time radar data and text data, extracts frequency-domain features from the radar data, extracts text semantic feature vectors from the text data, performs spatio-temporal alignment, performs feature fusion through a constructed fusion model, outputs a fused feature vector, generates triples through relation extraction of the fused feature vector by establishing a large language model, judges the distribution of triple data according to the similarity of the triples, completes the triple data in places with sparse distribution positions, ensures the balance of the data distribution of the triples, forms a triple chain according to the continuity of radar countermeasures for the triples, generates a knowledge graph through a constructed knowledge graph generation model according to the triple chain, and updates the knowledge graph according to real-time radar data and text data; This solution solves the data island problem existing in the analysis of traditional radar jamming technology and the generation of countermeasure strategies by introducing a large language model and multi-modal fusion technology. By spatio-temporally aligning radar frequency-domain data and text semantic features, deep fusion of multi-dimensional data is achieved based on a cross-modal attention mechanism, which not only improves the recognition accuracy of enemy radar jamming patterns but also enhances the intelligent generation ability of countermeasure strategies. Based on the relation extraction and data completion mechanism of the large language model, the problem of knowledge imbalance under sparse data distribution is solved. In addition, by constructing a triple chain and updating the knowledge graph, a continuous event stream can be formed, so that when constructing the knowledge graph, not only static relation information is included, but also the change process of time and actions is introduced, making the generated knowledge graph more in line with actual dynamic characteristics. This solution can dynamically generate countermeasure strategies based on real-time data, provide more accurate support for radar countermeasure decision-making, significantly improve the real-time adaptability of the radar system, enable more efficient identification of enemy jamming and implementation of effective countermeasures in a complex electronic warfare environment, and provide a more comprehensive and accurate basis for countermeasure decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 It is a schematic diagram of the overall method flow of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to specific embodiments.
[0018] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those with ordinary skills in the field to which the present invention pertains. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The terms such as "connected" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0019] Embodiment
[0020] Please refer to Figure 1 , the present invention provides a technical solution: A method for generating a knowledge graph based on a large language model, the specific steps including: Step 1: Obtain the real-time radar data and text data of both the friendly and enemy sides during the radar detection and countermeasure interference process, extract the frequency domain features of the radar data to identify the interference pattern, and extract the text semantic feature vectors of the text data to identify the relationships between the entities.
[0021] The radar data includes information such as the operating frequency, waveform and transmit power of the radar. By extracting the frequency domain features of the radar data, the interference pattern and the characteristics of the radar waveform can be identified. Using the fast Fourier transform to perform frequency domain analysis on the radar data can convert the radar signal from the time domain to the frequency domain, facilitating the identification of the intensity, periodicity and their changes of different frequency components, thereby effectively detecting the interference pattern in the radar signal and the characteristics of the friendly and enemy radar systems. This process can not only help identify the nature of the interference source, but also reflect the working state of the radar, thus providing data support for constructing countermeasure strategies.
[0022] The text data includes the working principles, technical features, tactical applications, countermeasure strategies, work logs, etc. of the radar jamming technology. By extracting the text semantic feature vectors of these text data, various interference technologies, tactical applications and the behavioral relationships between the friendly and enemy sides in different situations can be identified. The processing of the text data can not only extract the direct countermeasure strategies and technical details, but also reveal the interactions and information flows between different entities, forming a comprehensive knowledge graph.
[0023] In this embodiment, the radar data includes the operating frequency, waveform and transmit power of the radar; The specific process of extracting frequency-domain features from radar data is as follows: Obtain radar data at the same time interval Perform sampling. After preprocessing through maximum normalization, remove the DC component in the signal. Through The time window composed of time intervals divides the radar data, and extracts the frequency-domain features of the radar data within the time window; The calculation formula for removing the DC component in the signal is: where, is the signal after removing the DC component, is the radar signal after preprocessing at moment, is the total number of signal samples; The calculation formula for extracting the frequency-domain features of the radar data within the time window is: where, is the th frequency component in the frequency-domain signal of the radar signal, is the th sample in the radar signal, is the total number of signal samples, is the imaginary unit, represents the th frequency component in the frequency domain, .
[0024] In this embodiment, the text data includes the working principle, technical features, tactical applications, countermeasure strategies, and work logs of radar jamming technology; The specific process of extracting text semantic feature vectors from the text data is as follows: By using a vocabulary, the text data is split into word units, and a [CLS] token is added at the beginning of each input text, and a [SEP] token is added at the end of the text to define the text boundary. RoBERTa is used to extract the text semantic feature vector of the [CLS] token.
[0025] Step 2: Align the extracted radar frequency-domain features and text semantic feature vectors in space and time. According to the aligned radar frequency-domain features and text semantic feature vectors, through the constructed fusion model, feature fusion is performed according to the cross-modal attention mechanism, and a fused feature vector is output.
[0026] Through the cross-modal attention mechanism, the correlation between these two types of information can be dynamically captured and weighted, and they can be organically combined, so that the final fused features contain more information, avoiding the isolated existence of the two information sources. The cross-modal attention mechanism can assign a weight to each feature, thus highlighting its importance in a specific context. For example, at a certain moment, radar data may be more important, while at another moment, the semantic information of the text description may be more crucial. Through this mechanism, the contribution degrees of the two types of features can be flexibly adjusted, improving the model's expression ability for complex situations. Traditional feature fusion methods often adopt simple methods such as splicing and weighted averaging, while the cross-modal attention mechanism can automatically adjust the importance of each modality in different situations by learning the mutual relationship between modalities, thus more accurately extracting key information.
[0027] Traditional feature fusion methods (such as splicing, weighted summation, etc.) often ignore the possible mutual dependence relationships between different modalities. The cross-modal attention mechanism effectively avoids this problem by adaptively adjusting the attention degrees of different modalities, thus better combining the two types of features. In complex tasks, the features of a single modality may not be sufficient to comprehensively describe the complexity of the task. By fusing radar frequency domain features and text semantic features, the implicit semantic information in the text can be used to supplement the high-level knowledge that may be lacking in radar data, thus providing the model with a more multi-dimensional understanding.
[0028] In this embodiment, the method for constructing the fusion model is as follows: Regarding the radar features as queries, and the text features as keys and values; Through the queries and keys, calculate the similarity to obtain the attention weights. After obtaining the attention weights, apply them to the value vectors to obtain the weighted output feature vectors: Wherein, is the query vector, is the key vector, is the transpose of the key vector, is the value vector, is the dimension of the key vector, is the output fused feature vector, is the function; Among them, the query vector, key vector, and value vector are respectively: Wherein, is the weight matrix of the query, is the radar frequency domain feature vector, is the weight matrix of the key, is the text semantic feature vector.
[0029] The radar frequency-domain features are based on the spectral analysis results of radar signals and reflect the changes of radar signals in time and frequency. These data have clear spatio-temporal characteristics. For example, within a specific time window, the frequency, waveform, and transmission power of the radar may change. The text semantic feature vector is the result of vectorizing text data through natural language processing and reflects the semantic information in the text, such as the principle and tactical application of radar jamming technology. Due to the different sources of these two types of data (one is based on signal processing and the other is based on text understanding), their spatio-temporal alignment becomes very important. The purpose of spatio-temporal alignment is to ensure that radar data is associated with the corresponding text data within the same time period. For example, the change of radar data at a certain moment may be directly related to the technical description or countermeasure strategy within a certain time period. Through spatio-temporal alignment, the two types of features can be synchronized in time sequence, enabling them to be associated within the same time dimension, thus more accurately reflecting the dynamic changes and confrontation process of the system.
[0030] In this embodiment, the specific method of spatio-temporal alignment is as follows: Among them, is the loss function for temporal alignment, is the alignment mask matrix of radar features and text features, is the acquisition time length of radar features, is the number of text description words, is the radar signal at the feature vector at time, is the text description at position the semantic feature vector here.
[0031] Step 3: Generate triples by establishing a large language model to extract relationships from the fused feature vectors, judge the distribution of triple data according to the similarity of triples, and complete the triple data in sparse distribution positions to ensure the balance of the triple data distribution.
[0032] The RoBERTa-GCN-EGPLinker joint model is a composite deep learning model that combines a pre-trained language model, a graph convolutional network (GCN), and a graph spectrum reasoning module, aiming to handle tasks related to knowledge graph generation, relationship extraction, reasoning, and data completion. By combining three different technologies, the model can perform triple extraction, relationship reasoning, and knowledge graph optimization more efficiently and accurately in large-scale data, thereby improving the integrity, accuracy, and balance of the knowledge graph.
[0033] RoBERTa is an improved model based on BERT, which is optimized by increasing the amount of training data, increasing the training time, changing the training strategy, etc., thus significantly improving BERT's performance in natural language processing tasks. RoBERTa is widely used in relation extraction tasks because it can understand and capture fine-grained semantic relationships in text, identify entities and their mutual relationships in the fused feature vectors, and effectively process complex language features and implicit syntactic information, thereby providing basic data for subsequent graph convolutional networks and graph reasoning modules. GCN is a deep learning model used to process graph data, especially performing well in tasks with graph structures. GCN can effectively capture the dependency relationships between nodes in a graph and enhance the representation of nodes through information transmission between adjacent nodes. In knowledge graph generation, GCN can encode the relationships between entities as features of graph nodes based on the structural information of the graph, helping to better understand the connections between entities. EGPLinker is an efficient global pointer model for named entity recognition and relation extraction tasks, which solves some limitations in traditional sequence labeling models through a pointer network based on inner product calculation. Its core idea is to transform the entity recognition and relation extraction tasks into a prediction task for each pair of word positions, thereby being able to directly detect the start and end positions of entities, as well as identify different types of entities and relationships.
[0034] In this embodiment, the large language model is a joint model of RoBERTa-GCN-EGPLinker, and its specific structure includes RoBERTa semantic encoding, GCN, and the EGPLinker model: Through RoBERTa semantic encoding, the fused feature vector is extracted into a context vector representation, and GCN is used to update the nodes of the context vector representation. Based on the representation of the context and entities in the updated text as the input of the EGPLinker model, entity graph relationships are connected; RoBERTa semantic encoding: Among them, is the context vector representation, is the output fused feature vector; Based on GCN, the relationships between text nodes are updated: Among them, is the node representation matrix of the th layer, and the initial node representation matrix is the context vector representation, is the node representation matrix of the th layer, is the symmetric normalization factor, is the adjacency matrix, is the degree matrix, is the node update weight matrix, is the activation function; Entity graph relationship connection: in, Represents input sequence data Respectively represent The starting and ending projection matrices of each entity, For the The bias of an entity, Represent the extracted entities , Represents the context of the entity in the text, For Entity The relationship score between is the transpose operation, It is a global pointer model; The triple relationship between entities is determined by the relationship score threshold between entities, and the triple is output.
[0035] Triple completion is a key task in knowledge graph construction, especially when certain fields or relations in the graph are scarce. Completion through triple similarity analysis can effectively fill the blank areas in the knowledge graph and avoid information loss.
[0036] By judging the distribution based on similarity, we can identify which fields have scarce or uneven triple data and make targeted supplements. This helps to fill the problem of excessive concentration of low-frequency triples or high-frequency triples and avoid the situation where the knowledge graph is overly biased towards certain fields or relationships.
[0037] In traditional knowledge graph generation methods, due to the imbalance of data sources, there may be too many triples of certain relationship types in the generated knowledge graph, while triples of other relationship types are relatively scarce. By analyzing and completing the graph based on the similarity of triples, the distribution of different relationships can be balanced to ensure that all fields and relationship types in the graph are fully expressed.
[0038] This balance not only improves the information coverage of the knowledge graph, but also enhances the diversity and practicality of the graph. For example, after completion, some less popular entity types or relationships will receive similar attention as mainstream relationships, which enhances the breadth and depth of the graph.
[0039] In this embodiment, the calculation method for determining the distribution of triple data according to the similarity of triples is: in, is the first The probability of occurrence of a relationship type, is the first Types of relationships, is the number of relationship types in the triple; When it is determined that the distribution position of the triple is sparse; where is the threshold for judging the sparsity of the distribution position of the triple; The specific method for completing the triple data in the sparse distribution position is to complete it through an adversarial neural network. The specific formula is: where are the head entity, tail entity, and relationship type in the triple with a sparse distribution position respectively, are the embedding vectors of the head entity and tail entity respectively, is the output candidate relationship type, is a random noise vector that follows a standard normal distribution, is the relationship type in the triple with a sparse distribution position, is the generation learning weight matrix, is the discriminative learning weight matrix, is the completed triple.
[0040] Through the adversarial training of the generator and discriminator of the adversarial neural network, synthetic data similar to the real data distribution can be generated. In the construction of the knowledge graph, the adversarial neural network is used to complete the triple data in the sparse area. Through the method of adversarial training, the generated data can not only fill the sparse area, but also ensure that the generated data is consistent with the data distribution in the existing knowledge graph. The generator will generate completion content similar to the real triple data, while the discriminator will judge the authenticity of the generated data. Through this process, the generated triples are more in line with the existing data pattern in the graph, reducing the situation of low-quality completion or inconsistency with the existing data.
[0041] Step 4: Form triple chains based on the continuity of radar countermeasures for the triples. According to the triple chains, construct a knowledge graph generation model, output the knowledge graph, and update the knowledge graph based on real-time radar data and text data.
[0042] Traditional knowledge graphs usually consist of discrete triples, and the relationships between these triples may be isolated, lacking clear coherence. By organizing the triples in a chain based on the "continuity of radar countermeasures", more logical triple chains can be formed. This enables the information in the graph to be not only interconnected, but also capable of reasoning and analysis through continuous relationship chains. For example, the relationships between certain entities or events can be traced through continuous triple chains to obtain a higher level of understanding and reasoning ability.
[0043] The continuity of radar countermeasure refers to associating the continuous observation results in radar data with the triples in the knowledge graph to form a time - serialized triple chain. With the real - time input of new radar data and text data, the system can update the triple chain in a timely manner, seamlessly integrating the new information into the existing knowledge graph. This method based on continuity and adversarial relationships can quickly reflect changes in a dynamic environment, being more efficient and intelligent than traditional static update methods. In some application scenarios, the knowledge graph not only needs to express the semantic relationships between entities but also needs to handle the dimensions of time and space. By organizing triples into a triple chain with time and space continuity, it can help the system better understand the relationships between entities over time. For example, continuous observations based on radar data can help track the change trajectory or dynamic state of a target object, which is particularly important in fields such as transportation and meteorology. Traditional knowledge graph generation methods often rely on static triples and lack the processing of dynamic updates and continuity. This makes the knowledge graph update slowly when facing real - time data and changing data sources (such as radar data and text data), and may not be able to reflect the temporal relationships between data in a timely manner. For example, when new radar data comes in, traditional methods may simply add new triples without forming an effective continuous relationship chain, resulting in loose information integration.
[0044] In this embodiment, the method for forming a triple chain from triples according to the continuity of radar countermeasure is as follows: Divide the stages according to the process of radar countermeasure, assign stage labels to each triple, and form a sequence: where, is the stage label sequence, is the stage, is the number of divided stages; Statistically analyze the transition frequencies between stages and construct a probability matrix: where, is the transition probability matrix, is the number of times stage appears, is the number of times stage transfers to stage ; Generate a triple chain through dynamic programming: where, is the cumulative probability that the -th triple is in stage , is the maximum cumulative probability in stage , is the relationship score of the -th triple.
[0045] In this embodiment, the knowledge graph is constructed based on a graph neural network. The specific process is as follows: Taking the entities in each triple chain as the graph nodes and the relationships in the triple chain as the edges: Among them, is the representation of node in the -th layer of the graph neural network, is the representation of node in the -th layer of the graph neural network, is the set of neighbor nodes of node , is the relationship between node and , is the weight matrix of the -th layer, is the aggregation function.
[0046] Graph neural networks can learn and represent complex graph structures through the relationships between nodes and edges. In a knowledge graph, entities are usually represented as nodes, and the relationships between entities are represented as edges. Traditional knowledge graph methods often rely only on simple adjacency matrices or manually constructed rules to represent these relationships. In contrast, graph neural networks can better capture the complex and multi-level relationships between entities by leveraging the information of adjacent nodes (such as the features of nodes and the influence of neighboring nodes), thereby enhancing the expressive power and reasoning ability of the knowledge graph.
[0047] Graph neural networks (GNNs) can automatically learn the features of nodes and edges and propagate information through methods such as graph convolution. Through this information propagation, GNNs can capture the potential associations between entities, enhancing the context understanding and reasoning ability of the knowledge graph. Especially when dealing with data with complex correlations and dynamic changes, GNNs can adaptively adjust the node representations, making the knowledge graph more adaptable in dynamic scenarios. Graph neural networks are particularly suitable for processing dynamically changing graph data. When new triple chains are input in real time, GNNs can adaptively adjust the graph structure by iteratively updating the features of nodes and edges. Compared with traditional static knowledge graphs, GNNs can effectively perform incremental learning and update the entity relationships in the knowledge graph in real time. Especially when facing continuously changing data, they can ensure the accuracy and timeliness of the graph.
[0048] All the above formulas are dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula that is closest to the real situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0049] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will appreciate that the units and algorithm steps of the examples described in connection with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed by hardware or software methods depends on the specific application and design constraints of the technical solution.
[0050] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. They may be located in one place or distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0051] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.
Claims
1. A method for generating a knowledge graph based on a large language model, characterized in that, The specific steps include: Step 1: Obtain the real-time radar data and text data of both friendly and enemy sides during radar detection and countermeasure jamming. Extract the frequency-domain features of the radar data to identify the interference pattern, and extract the text semantic feature vectors of the text data to identify the relationships between the entities. Step 2: Align the extracted radar frequency-domain features and text semantic feature vectors in space and time. According to the aligned radar frequency-domain features and text semantic feature vectors, perform feature fusion through the constructed fusion model based on the cross-modal attention mechanism, and output the fused feature vectors. Step 3: Use a large language model to perform relation extraction on the fused feature vectors to generate triples. Judge the distribution of the triple data according to the similarity of the triples, and complete the triple data in the sparse areas of the distribution position to ensure the balance of the triple data distribution. Step 4: Form a triple chain from the triples according to the continuity of radar countermeasures. According to the triple chain, generate a knowledge graph through the constructed knowledge graph generation model, and update the knowledge graph according to the real-time radar data and text data.
2. The method for generating a knowledge graph based on a large language model according to claim 1, wherein: The radar data includes the operating frequency, waveform, and transmit power of the radar. The specific process of extracting the frequency-domain features of the radar data is as follows: Obtain radar data at the same time interval Perform sampling. After preprocessing by maximum normalization, remove the DC component in the signal. Through a time interval compose a time window to divide the radar data, and extract the frequency domain features of the radar data within the time window; The calculation formula for removing the DC component from the signal is as follows: where is the signal with the DC component removed, is the radar signal preprocessed at time, is the total number of signal samples; The calculation formula for extracting the frequency-domain features of radar data within a time window is as follows: Among them, is the th frequency component in the frequency-domain signal of the radar signal, is the th sample in the radar signal, is the total number of signal samplings, is the imaginary unit, represents the th frequency component in the frequency domain, .
3. A method for generating a knowledge graph based on a large language model according to claim 1, characterized in that: The text data includes the working principle, technical features, tactical applications, countermeasure strategies, and work logs of radar jamming technologies. The specific process of extracting the text semantic feature vectors of the text data is as follows: The text data is split into word units by using a vocabulary table, and a [CLS] token is added at the beginning of each input text, and a [SEP] token is added at the end of the text to define the text boundary. The RoBERTa is used to extract the text semantic feature vectors of the [CLS] token.
4. A method for generating a knowledge graph based on a large language model according to claim 1, characterized in that: The method for constructing the fusion model is as follows: Use the radar features as queries, and the text features as keys and values. Calculate the similarity through the query and the key to obtain the attention weights. After obtaining the attention weights, apply them to the value vectors to get the weighted output feature vectors: Among them, is the query vector, is the key vector, is the transpose of the key vector, is the value vector, is the dimension of the key vector, is the output fused feature vector, is the function; Among them, the query vector, the key vector, and the value vector are respectively: Among them, is the weight matrix of the query, is the radar frequency domain feature vector, is the weight matrix of the key, is the text semantic feature vector.
5. A method for generating a knowledge graph based on a large language model according to claim 1, characterized in that: The specific method of spatio-temporal alignment is as follows: Among them, is the loss function for temporal alignment, is the alignment mask matrix between radar features and text features, is the acquisition time length of radar features, is the number of text description words, is the radar signal at the feature vector at time, is the semantic feature vector of the text description at position at.
6. A method for generating a knowledge graph based on a large language model according to claim 1, characterized in that: The large language model is a joint model of RoBERTa-GCN-EGPLinker, and its specific structure includes RoBERTa semantic encoding, GCN, and EGPLinker models: Through RoBERTa semantic encoding, the fused feature vectors are extracted into context vector representations, and the context vector representations are updated by nodes through GCN. Based on the representations of the context and entities in the updated text, they are used as the input of the EGPLinker model to perform entity graph relationship connection; Entity graph relationship connection: Among them, represents the input sequence data respectively represent the start and end projection matrices of the th entity, is the bias of the th entity, respectively represent the extracted entities , represents the representation of the entity in the text context, is the entity relationship score between, is the transpose operation, is the global pointer model; Determine the triple relationship between entities through the relationship score threshold between entities, and output the triples.
7. A method for generating a knowledge graph based on a large language model according to claim 1, characterized in that: The calculation method for judging the distribution of triple data based on the similarity of triples is as follows: Among them, is the probability that the th type of relationship appears in the triple, is the th type of relationship in the triple, is the number of relationship types in the triple; When, it is determined that the distribution position of the triple is sparse; Among them, is the sparse judgment threshold for the distribution position of the triple; The specific method for triple data completion in sparse distribution locations is as follows: Among them, are the head entity, tail entity, and relationship type in the triple with sparse distribution locations respectively, are the embedding vectors of the head entity and tail entity respectively, is the output candidate relationship type, is a random noise vector, following the standard normal distribution, is the relationship type in the triple with sparse distribution locations, is the generated learning weight matrix, is the discriminative learning weight matrix, is the completed triple.
8. A method for generating a knowledge graph based on a large language model according to claim 1, characterized in that: The method for forming a triple chain from the triples according to the continuity of radar countermeasures is as follows: Divide the process of radar countermeasure into stages, assign stage labels to each triple, and form a sequence: Among them, is the stage label sequence, is the stage, is the number of divided stages; Statistically analyze the transition frequencies between stages and construct a probability matrix: Among them, is the transition probability matrix, is the number of occurrences of stage ; is the number of times that stage transitions to stage . Dynamic programming to generate triple chains: where is the cumulative probability that the -th triple is in stage , is the maximum cumulative probability in stage , is the relationship score of the -th triple.
9. A method for generating a knowledge graph based on a large language model according to claim 1, characterized in that: The knowledge graph is constructed based on a graph neural network, and the specific process is as follows: Taking the entities in each triple chain as graph nodes and the relationships in the triple chain as edges: Among them, is the representation of node at the th layer in the graph neural network, is the representation of node at the th layer in the graph neural network, is the set of neighbor nodes of node , is the relationship between nodes , , is the weight matrix at the th layer, is the aggregation function.
Citation Information
Patent Citations
Knowledge fusion method for industrial chain knowledge graph
CN113157940A
Knowledge graph completion method and system based on unstructured information
CN113934847A
Electromagnetic target multi-mode sensing method and device based on knowledge graph
CN117763388A
Radar active interference visual language combined identification method, system, medium and equipment
CN118519110A
Radar text processing method and related device
CN119128574A
Cited By
Military simulation knowledge graph generation method based on large language model
CN122021839A
Military simulation knowledge graph generation method based on large language model
CN122021839B